DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Remarks
General Note:
In the interest of compact prosecution and timely examination, Examiner respectfully requests that Applicant point out the support for any claim amendments which are considered to have changed the scope of a claim in accordance with the guidance of MPEP 2163. Lack thereof may be considered as an indication that Applicant does not find the claim scope to have changed or may establish a prima facie case for a rejection under 35 USC § 112(a). MPEP 2163 relates.
Examiner notes that a significant number of amendments appear to have changed the scope of the claim(s) (e.g. specifying a particular structure as the particular structure having a function where no structure was previously recited).
Claim Objections
The objections to the claims provided in the previous Office Action dated 12/18/2025 (hereafter typically referred to as the previous or prior Office Action) are withdrawn in light of Applicant’s amendments.
Claim Rejections - 35 USC § 112(b)
Those rejections properly addressed in Applicant’s amendments filed 4/20/2026 are withdrawn. Those only partially, incompletely, or not addressed remain. Examiner notes that as indicated in the prior Office Action, the list of issues was exemplary and non-exhaustive due to the extensive quantity and inter-related nature of the issues. There remains an extensive list of remaining issues which at a minimum were categorically and exemplarily raised in the prior Office Action. Applicant is requested to carefully review all of the claims and rectify all potential issues, not only those which Examiner was able to identify in this Office Action.
Furthermore, in the prior Office Action Examiner made clear statements where possible of interpretations of particular phrasing(s) in the rejection. As Applicant has now had the opportunity to amend and clarify the positively recited limitations of the claims and otherwise correct issues of clarity in light of Examiner’s comments, the claims will be interpreted as Applicant intending a particular meaning based thereon. See the Claim Interpretations section below.
Claim Rejections - 35 USC § 103
Applicant's arguments filed 04/20/2026 have been fully considered but they are not persuasive.
Applicant first begins with what appears to be providing a summary of their understanding of the disclosure of Tremblay (e.g. page 14). It is evident from the statements made that Applicant either misunderstands or is otherwise mischaracterizing the disclosure of Tremblay. For example, Applicant states “Tremblay relates to a robotic system using three different trained neural networks (KNN)”. Tremblay relates to more than just “three different neural networks”. Applicant provides no supporting evidence for this statement and [0029] which was the closest support Examiner could find states “Approaches in accordance with various embodiments can utilize a set of three learning modules” in the context of also stating “although other types, numbers, and arrangements of modules can be used as well within the scope of the various embodiments” as well as in the context that any given “network” may be one or more networks. Thus, Applicant’s statement appears to be a gross oversimplification of the disclosure of Tremblay in its entirety. Furthermore, the phrase “KNN” does not appear within Tremblay. Finally, Tremblay is specifically designed around a non-finite number of neural networks. See the model repository 134 and general disclosure of Tremblay which is clearly designed around retrieving/implementing/using an appropriate network for a given task (see e.g. [0024]).
As another example, Applicant states “the task is not further specified”. Tremblay explicitly provides examples such as a “block stacking task” (See e.g. [0027] or [0040]) illustrated in Figures 2A – 2D, and thus while the disclosure of Tremblay is designed to be general, Tremblay does in fact provide specific task examples.
As another example, Applicant states that “Tremblay does not disclose a distributed system for controlling a robot during a grasping task involving the grasping of objects of different types”. Applicant does not qualify this statement and appears to be referring to the preamble of the claims. Applicant’s claim preambles, however, make no statement as to a different object type, and furthermore are of the preamble wherein the patentable weight of a given recitation is not guaranteed (MPEP 2111.02 relates). Additionally, Applicant explicitly agrees prior to this that “Tremblay describes a local computing instance and a cloud-based instance”. Consequently, Applicant appears to at a minimum indicate that Tremblay discloses a distributed system. As demonstrated above, gripping is clearly and explicitly disclosed, [0026] discloses “a robotic arm, gripper assembly”, and “object type” is not claimed with any particularity. Thus, it is further clear that Tremblay discloses this stated feature, let alone any less specific feature of the preamble.
In summary, it is clear that Applicant first relies on features which are not actually claimed with the level of detail, specificity, narrowness, etc. that the arguments rely upon. Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
Referring now to the numbered arguments which are addressed below, the above general issue holds true, as well as the following:
Applicant frequently states that “the Action alleges that paragraph … discloses this feature”. However, Applicant in all of these circumstances has amended the claims such that Examiner has not made any such allegation previously. Furthermore, the reference as a whole must be considered (MPEP 2141.02(VI)). Instead, Applicant appears to focus only on recitations provided by the Examiner with respect to the previous versions of the claims, which as demonstrated in the previous Office Action were replete with issues of clarity and rejected under 35 U.S.C. 112(b).
Applicant’s claims refer to “the neural network”, however, the metes and bounds of the neural network are not ever claimed. For example, whether it is a single distinct and separate model and furthermore, wherein how it may be understood to be single, distinct, and separate is not claimed. At present, the neural network is effectively a “black box” wherein the claim only defines the contents by what is done to it, will be done to it, has been done to it, what it does, or similar. Similar issues exist with respect to other terms, phrases, structures, etc. Applicant appears to attempt to argue distinctions where none exist in the claims as presently constructed which mostly just use labelling to associate functions without otherwise actually distinguishing structures.
With respect to (I):
Tremblay does not need to use the word “pre-training” or “post-training” to read on these limitations. Tremblay instead discloses iterative and continuous training, and therefore there are an abundance of instances which may be considered “pre” or “post” “training”, what makes something “pre” or “post” being an undefined, arbitrary selected point in time from which to consider (i.e. any “capturing image data” qualifies for selection). See e.g. [0051] “Deep learning is a technique that models the neural learning process of the human brain, continually learning, continually getting smarter”, [0061] “In some embodiments, the training manager can make multiple passes or iterations over the training data”, [0065] “If the trained model does not satisfy at least a minimum performance criterion, or other such accuracy threshold, then the training manager 504 can be instructed to perform further training”, [0068] “In one embodiment building a machine learning application is an iterative process”, [0069] “In some embodiments the model will be continually trained as new data is available”. In summary, there is no requirement in the claims that Tremblay be interpreted from the perspective of starting with no existing trained models. Applicant appears to be imposing particular exclusive timing constraints not present within the claims.
Additionally, while this claim limitation recites that there is at least some form of pre-training and post-training having a particular timing, none of the following limitations appear to refer back to these limitations, reiterate this timing, or similar, and thus any “pre” or “post” training meeting this timing will satisfy the claim. See e.g. wherein Applicant recites after this limitation “perform a pre-training” or “perform a post-training”. There is no requirement that any of these “training” likewise satisfy this alleged timing requirement.
With respect to (II):
Applicant states “there is no mention of a neutral network being trained in this passage”. There is no requirement of a neural network being trained in the recited limitation. The limitation states “is trained” rather than a function (Claim 1) or step (Claim 13) of training the neural network.
Furthermore, as demonstrated clearly above and referenced by Applicant, various models/neural networks/etc. are involved at all stages. See also at least Figure 3 wherein demonstration data is input into a perception network and [0030] “In some embodiments, the perception network 302 can include or utilize two neural networks. A first network is a deep neural network (DNN) trained for object detection”. Thus, it is clear a trained neural network of the variety claimed is disclosed.
Applicant then states that “the Office Action has not demonstrated that the “plan” is stored at Model Repository 134”. This argument is wholly unclear. The limitation in question does not make any reference to a storage location requirement for anything. Furthermore, the relationship of the “plan” of Tremblay to the limitation is not clear. Finally, Applicant appears to be arguing against the level of recitation of the Office Action rather than the disclosure of Tremblay. See again above discussion wherein Applicant should consider the reference as a whole, not simply that which was referenced/recited in a prior Office Action and which may be with respect to a different limitation than is presently argued (being amended). Finally, the recitations provided by the Examiner are made in the context of the preceding recitations already provided, especially as they relate to related limitations.
Finally, Applicant begins a form of argument that seems wholly unsupported. Specifically, Applicant begins to make arguments that “disparate embodiments” are disclosed throughout Tremblay, in particular with respect to Figures 1 and 5. It is abundantly clear that this is not the case from the brief descriptions of the drawings alone:
[0004] “FIG. 1 an example system that can be utilized to implement aspects in accordance with various embodiments” (emphasis added)
[0008] “FIG. 5 illustrates an example system for training an image synthesis network that can be utilized in accordance with various embodiments” (emphasis added)
While Tremblay uses the phrase “various embodiments”, Tremblay does not appear to ever clearly delineate between what any given embodiment is, and furthermore at the very beginning of the Detailed Description states ([0015]) “In the following description, various embodiments will be described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will also be apparent to one skilled in the art that the embodiments may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified in order not to obscure the embodiment being described”.
Thus, Applicant appears to be arbitrarily defining distinct embodiments based on their own characterizations.
With respect to (III):
Applicant’s arguments appear to rely on ignoring the disclosure of Tremblay in general. There is no indication that the models referred to with respect to the training manager are exclusive of the perception networks of [0037]. Instead, there clearly are teachings, suggestions, and especially motivation that the clear reference to synthetic images is implemented by a training manager. Furthermore, [0037] clearly teaches all but “pre-training” of the limitation, and [0067] stands to show that there may be training after, thus making [0037] clearly “pre” and Applicant’s arguments with respect to [0065] largely moot.
Additionally, and separately, the limitation recites “a pre-training”. The limitation does not specify what it is a pre-training of.
With respect to (IV):
The phrase “statistical model” is not recited with respect to “object data”. See the entire preceding portion of [0056] (“The classified data can include instances of at least one type of object”). As this limitation is with respect to/in the context of training, the following “for which a statistical model is to be trained” was/is also recited.
Applicant does not appear to argue any of the other recitations or the rest of [0056] which states “type of object” which reads on “object-type-specific”.
With respect to (V):
As an initial matter, the Response, let alone the claims, do not set forth or otherwise define the metes and bounds of critical argued claim elements such as “central training computer”, “pre-training parameters”, “local processing unit”, etc. Consequently, there does not appear to be any merit to stating that “the “data” of Tremblay is not “pre-training parameters”, but instead represents “video and position data, representative of the objects 120” ”. It is further abundantly clear that “for processing” of the cited passage still describes “a training module 110” and thus any data related thereto is training related data, and as the parameters are wholly undescribed and undefined by the claims, may readily be considered as “parameters”. Examiner has not set forth with particularity any “pre-training parameters” as effectively any data, value, information, etc. “generated” as a part or related to the “pre-training” appears to read on the limitation as broadly claimed.
With respect to Applicant’s argument with respect to Figure 5, it appears to be a reiteration of the above wherein Applicant still does not articulate how the nature of their claimed features in the limitations are claimed such that they are narrow to the point of excluding the disclosure of Tremblay, and furthermore appear to rely on the “disparate embodiments” argument addressed above.
With respect to (VI):
Applicant does not clearly define or describe what “continuously and cyclically” means within the claims. Thus, [0065] which discloses continuing to train for at least one cycle reads on this claim. Furthermore, and alternatively, see again also [0051], [0061], [0068], and [0069] which both individually and especially collectively make it clear that training of a given network may never be completely halted and at a minimum will undergo a large number of iterations.
With respect to (VII):
This argument appears to rely on previous arguments (I) and (VI) already addressed above and to not be a stand-alone argument.
Furthermore, and alternatively, the phrasing of the limitation in question is particularly broad and does not appear to be in alignment with Applicant’s arguments. The limitation does not actually recite a function of doing so (a particular process), but instead of a neural network that can undergo such a particular process which are very different things. Theoretically any neural network might be capable of being continuously and cyclically replaced.
In other words, the claim does not recite, “the at least one local processing unit being configured to: continuously and cyclically replace a pre-trained neural network with a subsequent post-trained neural network until a convergence criterion is fulfilled” or similar. As was emphasized heavily in the previous Office Action, Applicant should clearly recite those features which are to be positively claimed. At present, the claim limitation focuses on the neural network, not any training, replacement, etc. process which Applicant’s arguments indicate maybe should have been claimed.
As a similar, unargued limitation, Claim 1 recites an ICP algorithm, however, the claim focuses on the ICP algorithm being “configured to receive” particular data, rather than on a function of inputting said data into the ICP algorithm. Again, this holds a particular distinction as an algorithm may generally be considered as “configured to” take any number of different inputs, but in practice a may be operated with different inputs, for example by a processor configured to input different inputs.
In summary, Applicant appears to rely on incorporating features into the claims not actually recited at such detail, or not even actually claimed (the focus of the claim being wholly different than possibly argued as constructed), as well as mischaracterizing the disclosure of Tremblay and/or selectively focusing only on portions thereof rather than the whole disclosure.
Claim Objections
Claim 1 is objected to because of the following informalities:
Claim 1 recites “to transmit these to the robot controller for execution”. For clarity, it should read “to transmit the gripping instructions to the robot controller for execution”.
Appropriate correction is required.
Claim Interpretation
With respect to recitations of “designed to”, “serve” or “serves” or “serves as” and “serves to”, and “intended for” and “intended to”: These been interpreted as reciting a non-functional description, and therefore are non-limiting (not all phrases are necessarily presently used within the claims).
With respect to recitations of transitional phrases such as “to”, “for”, “as a result of”, “in order to”, “so that”, “used to”, and “for the purpose of”: The broadest reasonable interpretation of the prepositions of “to” and “for” includes meanings of indicating a purpose, intention, tendency, or result, while the other phrases indicate similar meanings with less nuance towards other interpretations. This is regardless of if the claim is an apparatus or method claim. Thus, these are typically interpreted as non-functional, non-limiting phrases merely indicating intended results, purpose, preferences, etc. (not all phrases are necessarily presently used within the claims).
In the interest of compact prosecution, an effort has been made to provide recitations to features where disclosed by a reference even if the feature is not considered positively recited, such as a mere indication of intended use, purpose, or similar. Examiner has also made an effort to point out features interpreted under the above, at least wherein the feature is generally first provided.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1 – 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The claims, while improved (in particular the independent claim, Claim 1), continue to exhibit issues of being narrative and indefinite, failing to conform with current U.S. practice. They appear to be a literal translation into English from a foreign document and are replete with errors. Applicant is respectfully requested to review the following rejections and amend the claims such that they conform with current U.S. practice. The rejections found under 35 USC § 103 are made in light of these amendments.
Furthermore, due to the extensive quantity and inter-related nature of the issues identified below, the rejections below are considered to be non-exhaustive.
Regarding Claim 1, the claim recites the limitation “the at least one local processing unit configured for interacting with the robot controller”. There is insufficient antecedent basis for this limitation in the claim. The claim does not previously recite that the local processing unit is configured as such. It is believed that Applicant is intending to recite define or describe the local processing unit as appears done in the following clause.
In the interest of compact prosecution, the limitation has instead been interpreted as reading: “wherein the at least one local processing unit is configured to interact with the robot controller”
Examiner additionally notes that Applicant uses “configured for” and “configured to”. Examiner finds the phrasing “configured to” to more clearly indicate a function rather than a purpose, particularly when coupled with especially broad phrases.
Claim 1 and many of its dependent claims recite, and are directed to, an apparatus. “Features of an apparatus may be recited either structurally or functionally. In re Schreiber, 128 F.3d 1473, 1478, 44 USPQ2d 1429, 1432 (Fed. Cir. 1997)” (MPEP 2114(I)). Furthermore, “ “[A]pparatus claims cover what a device is, not what a device does.” Hewlett-Packard Co.v.Bausch & Lomb Inc., 909 F.2d 1464, 1469, 15 USPQ2d 1525, 1528 (Fed. Cir. 1990) (emphasis in original). A claim containing a “recitation with respect to the manner in which a claimed apparatus is intended to be employed does not differentiate the claimed apparatus from a prior art apparatus” if the prior art apparatus teaches all the structural limitations of the claim. Ex parte Masham, 2 USPQ2d 1647 (Bd. Pat. App. & Inter. 1987)” (MPEP 2114(II)). A positively recited apparatus claim limitation should not be a recitation of how the apparatus is designed or intended to operate or an intended or achieved result or effect, particularly wherein there is no clear and definite structural distinction required by said design or intent.
The phrasing of “is [verb]ed” is not a clear positive recitation of a functional or structural limitation, especially as no particular timing is claimed. If pertinent to the apparatus, it should be recited as a clear functional recitation such as Structure [Y] configured to [verb] possibly by steps or acts of [Z]. For example, if training the neural network is pertinent, then the claim might recite “wherein the central training computer is configured to train the neural network to perform object recognition and position detection …” or similar.
Regarding Claim 1, the claim recites the limitation “wherein the neural network is trained” twice. Claim 19 also relates, however the issue depends on if the intent is to refer back by antecedent basis or to separately claim. Presently Examiner believes the former.
As the nature of the what performed the training is not presently claimed, the claim limitations have been interpreted as instead reading “wherein the neural network is configured to be trained such that it is configured for”. Examiner, however, notes that this leaves the level of positive recitation to “such that” somewhat unclear, and Applicant is advised to use phrasing more in line with the above wherein a particular structure is recited as performing training to do a certain function or similar.
Relatedly, Claim 1 also recites the limitation “a set of local resources that interact via a local network”.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
a set of local resources configured to interact via a local network
Relatedly, Claim 1 also recites the limitation “the robot having …”.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“the robot, wherein the robot comprises: …”
Relatedly, Claim 3 recites the limitation “which is rendered based on …”.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“based on”.
Relatedly, Claim 11 recites the limitation “wherein the post-training of the neural network is configured to be performed iteratively and cyclically”. It is unclear how this is a functional limitation, regardless of the use of the typical phrasing configured to, especially as it is actually “configured to be”
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“wherein the central training computer is configured to perform post-training of the neural network iteratively and cyclically”.
Relatedly, Claim 11 recites the limitation “refined result data sets … which are automatically annotated and which have been transmitted from the at least one local processing unit to the central training computer”.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“wherein the system is configured to annotate refined result data sets which have been transmitted from the at least one local processing unit to the central training computer”.
Relatedly, Claim 12 recites the limitation “wherein a post-training data set for post-training the neural network is gradually and continuously expanded by image data acquired by sensors of the optical acquisition device in the vicinity of the robot”.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“wherein the system is configured to gradually and continuously expand a post-training data set with image data acquired by sensors of the optical acquisition device in the vicinity of the robot”.
Regarding Claim 1, the claim recites the limitation “the image data in the working area of the robot”. There is insufficient antecedent basis for this limitation in the claim. Previously the claim only recites “image data of the plurality of objects”, not that it is specifically in the working area. Furthermore, what is previously recited with respect the plurality of objects is “a respective object, of a plurality of objects, having a respective object type of different object types which are arranged in a working area of the robot”. In other words, the “a working area” is only in relation to “different object types”, not the plurality of objects.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“the image data”.
Regarding Claim 1, the claim recites the limitation “the set of local resources comprising: … the at least one local processing unit configured for interacting with the robot controller”. There is insufficient antecedent basis for this limitation in the claim. The claim only recites “at least one local processing unit via a network interface” previously.
In the interest of compact prosecution, the limitation has instead been interpreted as reading: “the set of local resources comprising: … the at least one local processing”.
Regarding Claim 1, the claim recites the limitation “the set of local resources comprising: …” wherein the clauses following are formatted with an indentation indicating inclusion in the local resources. However, many of the clauses appear to be directed towards further describing structures rather than specifying further structures that are a part of the local resources.
Regarding Claim 1, the claim recites the limitation “the gripping instructions for the end effector unit for gripping the respective object”. There is insufficient antecedent basis for this limitation in the claim. The claim never previously recites such instructions.
In the interest of compact prosecution, the limitation has instead been interpreted as reading: “gripping instructions for the end effector unit for gripping the respective object”.
Regarding Claim 1, the claim recites the limitation “the result data set determined by the pretrained or post-trained neural network”. There is insufficient antecedent basis for this limitation in the claim. The claim never previously recites such a result data set. The claim previously recites “the at least one local processing unit is configured to apply, in an inference phase, the pretrained or post-trained neural network for determining a result data set”. In other words, the local processing unit applies the pretrained or post-trained neural network in an inference phase for the purpose of determining a result data set, but neither neural network is stated as actually determining the result data set. The nature of applying, the inference phase, and determining are not claimed with any particularity and are a part of a comprising claim and therefore any other structures, functions, etc. may be involved as a part of the recited limitations.
In the interest of compact prosecution, the limitation has instead been interpreted as reading:
“the result data set”
Regarding Claim 1, the claim recites the limitation “receive as input data the image data captured by the optical acquisition device which has been fed to the implemented neural network for application”. There is insufficient antecedent basis for the limitation in the claim. The claim does not previously recite that the data “has been fed to the implemented neural network for application” or a function or similar thereof.
In the interest of compact prosecution, the limitation has been interpreted as instead reading:
“receive as input data the image data”.
Regarding Claim 2, the claim recites the limitation “the refined result data set generated on the at least one local processing unit”. There is insufficient antecedent basis for the limitation. Claim 1 previously recites “the at least one local processing unit is configured to execute a modified Iterative Closest Point, ICP, algorithm, the ICP algorithm being configured to … receive as input data reference image data and compare the image data with the reference image data in order to minimize errors and to generate a refined result data set”. In other words, the refined result data set is only an intended result and not a positively recited occurrence, and furthermore, is not actually stated as being generated by anything in specific.
In the interest of compact prosecution, the limitation has been interpreted as instead reading:
“the refined result data set”
Regarding Claims 1, 4 and 19, the claim recites the limitation “the pre-training”. There is more than one “pre-training” to be referred back to for antecedent basis. In the case of Claims 1 and 4, there is first, “pre-training of the neural network prior to capturing image data of the plurality of objects” and second, “pre-training exclusively with synthetically generated object data …”. These two “pre-training” have not presently been claimed as being the same. In the case of Claim 19, there are potentially even more.
For the purposes of compact prosecution, the limitation has been interpreted as simply reading “(a) pre-training”.
Regarding Claims 1, 11, 13 and 19, the claim recites the limitation “the post-training”. There is more than one “post-training” to be referred back to for antecedent basis. In the case of Claims 1 and 11, there is first, “post-training of the neural network during or after capturing the image data” and second, “a post-training of the neural network”. These two “post-training” have not presently been claimed as being the same. In the case of Claims 13 and 19, there are potentially even more.
For the purposes of compact prosecution, the limitation has been interpreted as simply reading “(a) post-training”.
Regarding Claims 13 and 14, the claim recites the limitation “the pre-training parameters”. There is more than one “pre-training parameters” to be referred back to for antecedent basis. In the case of Claims 13 and 14, there is first, “pre-training parameters of a pre-trained neural network” and second, “pre-training parameters … from the central training computer”. These two “pre-training parameters” have not presently been claimed as being the same.
For the purposes of compact prosecution, the limitation has been interpreted as simply reading “(a) pre-training parameters”.
Regarding Claims 13, 14, and 15, the claim recites the limitation “the pre-training parameters”. There is more than one “post-training parameters” to be referred back to for antecedent basis. In the case of Claims 13, 14, and 15, there is first, “post-training parameters of a post-trained neural network” and second, “post-training parameters … from the central training computer”. These two “post-training parameters” have not presently been claimed as being the same.
For the purposes of compact prosecution, the limitation has been interpreted as simply reading “(a) pre-training parameters”.
Regarding Claim 3, the claim recites the limitation “the image data acquired with the optical acquisition device and fed to the neural network”. There is insufficient antecedent basis for this limitation in the claim. Previously recited image data is not recited as being “fed” to anything.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“the image data”
Regarding Claims 3 and 13, the claims recite the limitation “the result data set determined by the neural network and the 3D model”. There is insufficient antecedent basis for this limitation in the claim. Applicant specifically amended the limitation in Claim 1 to instead recite “the result data set determined by the pretrained or post-trained neural network”, which itself has antecedent basis issues.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“the result data set”
Regarding Claim 4, the claim recites the limitation “the respective object to be grasped”. There is insufficient antecedent basis for the limitation in the claim.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“the respective object”
Regarding Claim 4, the claim recites the limitation “the central training computer, in response to the determined object type, being configured to load the object-type-specific 3D model from a model storage”. This clause is incomprehensible. The phrase “in response to the determined object type” is an incomplete contingent condition. It is unknown what the condition is, or even should be. If it is mere existence, there is no reason to even phrase it as “in response to”. No determining is previously recited. Additionally, this appears to indicate configuring the central training computer based on some condition having no recited function actually in existence that it corresponds to.
In the interest of compact prosecution, the claim is interpreted as ending prior to this limitation.
Regarding Claim 5, the claim recites the limitations “the individual feature vectors” and “the accumulation”. There is insufficient antecedent basis for these limitations in the claim. No “individual feature vectors” are previously recited, and the claim previously recites “evaluating and/or accumulating” such that first, no accumulation is specifically recited, and second, any potential accumulation from accumulating is explicitly recited as optional, and yet the claim appears to remove the “or” portion of “and/or” if said accumulation is therefrom.
In the interest of compact prosecution, the limitations have been interpreted as reading:
“individual feature vectors” and “an accumulation”
Regarding Claim 10, the claim recites the limitation “the gripper”. There is insufficient antecedent basis for the limitation in the claim. No gripper is ever previously recited.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“a gripper”
Regarding Claim 10, the claim recites the limitation “the calculated visualization of the grasping instructions”. There is insufficient antecedent basis for the limitation in the claim. No calculated or calculating of visualization of the grasping instructions is ever previously recited.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“a calculated visualization of the grasping instructions”
Regarding Claim 13, the claim is replete with errors of improperly claimed steps and antecedent basis, of the same general form as the rejections preceding. A non-exhaustive list of errors are as follows:
“the 3D model, comprising a CAD model, assigned to the selected respective object type” and “the recorded retraining data” (three issues of antecedent basis)
“and use it” (no clear active claiming of a step, e.g. “using”)
“Provisioning of pre-training parameters” (no clear active claiming of a step, e.g. just reciting “provisioning”, no “of recited)
“executing a modified ICP algorithm which i) receives … evaluates and compares … ii) receives … and compares …” (no clear active claiming of steps, e.g. “receiving”, “evaluating”, “comparing”)
“which is rendered” (no clear active claiming of a step, e.g. “rendering”)
“is transmitted” (no clear active claiming of a step, e.g. “transmitting”)
“Continuous and cyclical retraining of the neural network” (no clear active claiming of a step, e.g. “retraining continuously and cyclically”)
Regarding Claim 14, the claim is replete with errors of improperly claimed steps and antecedent basis, of the same general form as the rejections preceding. A non-exhaustive list of errors are as follows:
“the model storage” and “the acquired retraining data” (two issues of antecedent basis)
“which is annotated in an automatic process” (no clear active claiming of steps, e.g. “annotating”)
“continuous and cyclical retraining” (no clear active claiming of a step, e.g. “retraining continuously and cyclically”)
Regarding Claim 16, the claim is replete with errors of improperly claimed steps and antecedent basis, of the same general form as the rejections preceding. A non-exhaustive list of errors are as follows:
“capturing of the image data” (no clear active claiming of steps, e.g. “capturing the image data”)
“the image data with the optical capture device of the objects in the working area of the robot” (antecedent basis issue)
“which evaluates and compares” (no clear active claiming of steps, e.g. “evaluating and comparing”)
“refined result data set” (two antecedent options, no clarity as to which refined data set is referred to of the two)
Regarding Claim 17, the claim recites “wherein the acquisition of the image data for the respective object is triggered before the instructions for gripping the respective object are executed”. This is not of the form of a step, no “instructions for gripping the respective object” are previously recited, no execution of such is previously recited so as to confer significance even if previously recited, no step of acquiring image data or triggering thereof is previously recited, and no “image data for the respective object” is previously recited.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“wherein the method further comprises acquiring image data for the respective object”
Regarding Claim 18, the claim as a whole does not particularly make sense. The claim recites “when applying the … neural network .. in a pre-training phase”, however no “pre-training phase” is previously recited, this being the first use of such a phrase in the claims. The claim also recites “the plurality of objects are arranged in a certain configuration comprising on a plane and disjointly in the working area” and “the plurality of objects are arranged in the working area without adhering to the certain configuration”. These are vague and unclear phrasings. Furthermore, this is a contingent limitation (see “when …”) in a method claim, and may therefore not be required. MPEP 2111.04 relates.
The claim has been interpreted as merely indicating that image data input in an initial training phase is synthetic data and that image data input in a later or further training phase is real.
Regarding Claim 19, the claim recites the limitation “the neural network”. There are two potential neural networks to be referred to under antecedent basis. It is unclear which is referred to.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“a neural network”
Regarding Claim 19, the claim recites the limitation “wherein, as a result of the pre-training, pre-training parameters of a pre-trained neural network are transmitted”. This is not a functional or structural limitation.
In the interest of compact prosecution, the limitation has been interpreted as reading:
“wherein the system is further configured to transmit parameters of a pre-trained neural network” (the “as a result of” being broad to the point of little to no narrowing scope)
Regarding Claim 20, the claim is replete with errors categorically address and identified already above. A non-exhaustive list of errors are as follows:
“wherein the pretrained or post-trained neural network is applied” (non-functional or structural limitation)
“wherein a modified Iterative Closest Point, ICP, algorithm is executed” (non-functional or structural limitation)
“which evaluates and compares” (non-functional or structural limitation)
“transmitting these” (non-functional or structural limitation)
“is rendered” (non-functional or structural limitation)
“the image data of the optical acquisition device supplied to the implemented neural network for evaluation” (antecedent basis).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Examiner notes that recitations not necessarily required to be disclosed by the prior art have still been addressed where expedient in the interest of compact prosecution. See Claim Interpretation Section above. Furthermore, the rejections provided below are a best effort in light of the 112(b) rejections above.
Claims 1 – 4 and 6 – 21 are rejected under 35 U.S.C. 103 as being unpatentable over Tremblay et al. (US 20190228495 A1) in light of Kehoe et al. (B. Kehoe, A. Matsukawa, S. Candido, J. Kuffner and K. Goldberg, "Cloud-based robot grasping with the google object recognition engine," 2013 IEEE International Conference on Robotics and Automation, Karlsruhe, Germany, 2013, pp. 4263-4270) and Shanley (US 10133696 B1).
Regarding Claim 1, Tremblay teaches:
A distributed system (See at least example environment 100 and Figure 1) for controlling at least one robot (See at least robot 102) in a gripping task for gripping a respective object, of a plurality of objects, having a respective object type of different object types which are arranged in a working area of the robot (See at least Figures 2A – 2D and [0026] “The robotics can be any appropriate automated, or at least partially automated, mechanism, as may include a robotic arm, gripper assembly, multi-link manipulator, end effector, motion control system, or other such physical hardware component, module, or sub-system that may be contained within, or connected in some way to, the robot 102 to perform one or more tasks as instructed”), comprising:
a central training computer (See at least Provider Environment 125), having a memory (See at least memory 704) on which an instance of a neural network ANN is configured to be stored (See at least model repository 134), wherein the central training computer is configured for i) pre-training of the neural network prior to capturing image data of the plurality of objects (Examiner notes that this indicates a timing and does not claim the process/function of “capturing…”. Furthermore, the timing itself is not particularly claimed and the capacity/capability is sufficient), and ii) post-training of the neural network during or after capturing the image data (See at least [0024] “The communication, or information from the communication, can be directed to a training manager 130, which can select an appropriate model or network and then train the model using relevant training data 132” as well as [0051] “Deep learning is a technique that models the neural learning process of the human brain, continually learning, continually getting smarter”, [0061] “In some embodiments, the training manager can make multiple passes or iterations over the training data”, [0065] “If the trained model does not satisfy at least a minimum performance criterion, or other such accuracy threshold, then the training manager 504 can be instructed to perform further training”, [0068] “In one embodiment building a machine learning application is an iterative process”, [0069] “In some embodiments the model will be continually trained as new data is available” which make clear the timing of training may be effectively any time), wherein the neural network is trained and configured for object recognition and position detection, including detection of an orientation of the respective object (See at least [0027] “FIGS. 2A through 2C illustrate portions of a basic task that can be learned … A robot capturing image data representative of these actions could analyze the image data to determine orientation, location, relationship, and other information about the objects), and the neural network is trained and configured to calculate grasping instructions for an end effector unit (See at least [0026] “end effector”) of the robot for grasping the object (See at least [0026] “gripper assembly” and [0027] “The plan can be a program, file, database, or set of actions or instructions, which could include steps such as “Place Block B on Block A” followed by “Place Block C to the right of Block A” ”);
wherein the central training computer is configured to receive the respective object type (See at least [0027] “identifiable by their respective colors or other such aspects … orientation, location, relationship, and other information about the objects”.
Alternatively, the claim does not recite the nature of “receive” with any particularity. Simply observing the object with a sensor would also read on this limitation); and
wherein the central training computer is configured to perform a pre-training (See at least [0065] “the training manager 504 can be instructed to perform further training” (meaning there is prior initial training).
Examiner this “pre-training” holds none of the timing limitations of previous “i) pre-training” and the “pre-training” is not defined with particularity what it is of or even what it entails. Presently the claim appears to merely describe what is used, but not how or in what manner) exclusively with synthetically generated object data (See at least [0037] “Leveraging convolutional pose machines, object cuboids can be reliably detected in images even when severely occluded, after training only on synthetic images” (emphasis added)) which is generated by means of a geometric, object-type-specific (See at least [0056] “The classified data can include instances of at least one type of object for which a statistical model is to be trained”) 3D model (See at least [0041] “Each object of interest can be modeled, such as by a bounding cuboid” and [0083] “our system operates in 3D”) of the respective object, and wherein, the central training computer is configured to transmit pre-training parameters of a pre-trained neural network generated from the pre-training to at least one local processing unit via a network interface (See at least [0023] “This can involve, for example, using a training module 110 on the robot itself, or sending the data across the at least one network 122 for processing … At least some functionality may also operate on a remote device, networked device, or in “the cloud” in some embodiments”), and
wherein the central training computer is further configured to continuously and cyclically perform a post-training of the neural network (See at least [0065] “the training manager 504 can be instructed to perform further training, or in some instances try training a new or different model”) and to transmit post-training parameters of a post-trained neural network, generated from the post-training, to the at least one local processing unit via the network interface (See at least [0026] “The execution neural network can perform the inference on the robot 102, on the client device 138, or using an inference 136 in the provider environment 124, among other such options. Once the instructions are generated, the instructions can be provided to the control system 104 of the robot, either directly or upon execution by the processor 112, etc.”);
a set of local resources that interact via a local network (See at least Figure 1), the set of local resources comprising:
the robot having a robot controller (See at least controller 104), a manipulator (See at least [0026] “multi-link manipulator”), and the end effector unit, wherein the robot controller is configured for controlling the robot and the end effector unit for (Recitation of intended purpose of the “controlling” …) executing the gripping task for a respective object of the respective object type (See again at least Figures 2A – 2D);
an optical acquisition device configured for capturing the image data in the working area of the robot (See at least [0022] “sensors 108 … for example, one or more cameras to capture images or video of the performance in the environment within a field of view 118 of the respective sensors … the sensors 108 can capture information, such as video and position data, representative of the objects 120 in the task environment”); and
the at least one local processing unit configured for interacting with the robot controller (See at least processor 112), the at least one local processing unit being configured to: store different instances of the neural network (See at least training program 110 and/or memory 114), receive pre-training parameters and post-training parameters from the central training computer (See at least [0026] “The execution neural network can perform the inference on the robot 102, on the client device 138, or using an inference 136 in the provider environment 124, among other such options. Once the instructions are generated, the instructions can be provided to the control system 104 of the robot, either directly or upon execution by the processor 112, etc.”), implement a pre-trained neural network that is configured to be continuously and cyclically replaced by a subsequent post-trained neural network until a convergence criterion is fulfilled (See at least [0061] “In some embodiments the training manager can monitor the quality of patterns (i.e., the model convergence) during training, and can automatically stop the training when there are no more data points or patterns to discover”), and
wherein the at least one local processing unit is configured to apply, in an inference phase, the pretrained or post-trained neural network for (Indicates the following is just the purpose or intent of the previous function rather than a positively recited limitation) determining a result data set from the image data captured by the optical acquisition device, the at least one local processing unit being configured to calculate, based upon the result data set, the gripping instructions for the end effector unit for (Indication of intended purpose) gripping the respective object and to transmit these to the robot controller for execution (See at least [0041] “a camera can acquire a live video feed of a scene, from which a pair of networks can infer the positions and relationships of objects in the scene in real time. The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts”);
…
and whereby the image data captured with the optical acquisition device and the refined result data set comprise a post-training data set, the local computing unit being configured to transmit the post-training data set to the central training computer for (Clear statement of intended purpose) post-training by the central computer (See again [0054]);
wherein the network interface is configured for data exchange
…
between the central training computer and the at least one local processing unit (See at least network 122),
…
Tremblay does not teach, but Kehoe teaches:
…
wherein the at least one local processing unit is configured to execute a modified Iterative Closest Point, ICP, algorithm, the ICP algorithm being configured to i) receive as input data the image data captured by the optical acquisition device which has been fed to the implemented neural network for application (See Section IV, E, “First, estimating the pose of the object using a least-squares fit between the detected 3D point cloud and the reference point set using the iterative closest point method (ICP) [36] [38]. We use the ICP implementation from PCL. The ICP algorithm performs a local optimization and therefore requires a reasonable initial pose estimate to find the correct alignment.”), and,
Tremblay further teaches, and therefore in combination with Kehoe teaches:
ii) receive as input data reference image data and compare the image data with the reference image data in order to (Explicit statement of intended result or purpose of the prior recited function) minimize errors and to generate a refined result data set, the reference image data being a synthesized, rendered image which is rendered based on the result data set determined by the pretrained or post-trained neural network (See at least [0054] “During training, data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. If the neural network does not correctly label the input, then errors between the correct label and the predicted label are analyzed, and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels the input and other inputs in a training dataset”);
…
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to utilize a well-known pose estimation technique such as that using an ICP method as disclosed in Kehoe in the system of Tremblay with a reasonable expectation of success. ICP is a well known method with particular advantages which would be obvious to utilize in Tremblay, which is not particular as to how ground truth and other data for comparison is generated.
Tremblay does not teach, but Shanley teaches:
…
via an asynchronous protocol (See at least Column 4, Lines 29 – 33 “An example system having a bridge, an asynchronous channel based bus, and a message broker to provide asynchronous communication is shown in accordance with various embodiments, are then described”).
…
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to utilize an asynchronous protocol as disclosed in Shanley in the system of Tremblay with a reasonable expectation of success. It is common for different components within a system to operate with different timings such that an asynchronous protocol is required. Asynchronous protocols are well known and routine in computer systems including networking and that disclosed by Shanley would merely be one of many different solutions to a typical problem.
Regarding Claim 2, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the network interface is configured to transmit parameters for (Indicates intended purpose of the preceding function) instantiating the pre-trained or post-trained neural network from the central training computer to the at least one local processing unit (See at least [0026] “The execution neural network can perform the inference on the robot 102, on the client device 138, or using an inference 136 in the provider environment 124, among other such options. Once the instructions are generated, the instructions can be provided to the control system 104 of the robot, either directly or upon execution by the processor 112, etc.” and/or Figure 1), and/or wherein the network interface is configured to transmit the image data captured with the optical acquisition device and the refined result data set generated on the at least one local processing unit to the central training computer for (indicates purpose of preceding) post-training (See at least Figure 1 and [0069] “the now classified data instances can be stored to the classified data repository, which can be used for further training of the trained model 508 by the training manager”) and/or wherein the network interface is configured to load the geometric, object-type-specific 3D model on the local processing unit (See again at least [0023], [0041], and [0056]).
Regarding Claim 3, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the at least one local processing unit is configured to generate i) annotated post-training data from the image data acquired with the optical acquisition device and fed to the neural network (See at least [0056] “For example, the classified data might include a set of images that each includes a representation of a type of object, where each image also includes, or is associated with, a label, metadata, classification, or other piece of information identifying the type of object represented in the respective image”) and ii) synthesized reference image data via an annotation algorithm, the at least one local processing unit being configured to transmit the annotated post-training data to the central training computer for (Indicates purpose) post-training by the central training computer, the synthesized reference image data being a synthesized, rendered image which is rendered based on the result data set determined by the neural network and the 3D model (See again preceding).
Regarding Claim 4, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the system comprises a user interface (See at least [0024] “The interface layer 126 can include application programming interfaces (APIs) or other exposed interfaces enabling a user, client device, or other such source to submit requests or other communications to the provider environment”) configured to provide one selection field (The claim does not particularly define what this term means. See above and at least [0026] “The plan can be at least partially human-readable, and can be sent to the client device 138, provided through a UI of the training program 110 executing on the robot, or otherwise provided” and [0063] “When creating a machine learning model, the training manager in some embodiments can enable a user to specify settings or apply custom options”) in order to (Intended purpose of the preceding) determine the respective object type of the respective object to be grasped and wherein the user interface is configured to transmit the determined respective object type to the central training computer, the central training computer, in response to the determined object type, being configured to load the object-type-specific 3D model from a model storage in order to (Intended purpose of the preceding) synthesize object-type-specific images in all physically plausible positions and/or orientations via of a synthesis algorithm, the central training computer being configured to perform the pre-training of the neural network based upon the object-type specific images (Dependent upon non-positively recited limitation).
Regarding Claim 6, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay does not teach, but Shanley has already been shown to teach in combination with Tremblay:
wherein the network interface is configured to facilitate synchronization using a message broker implemented as a microservice (See at least Column 10, Lines 11 – 14, “The function of bridge 510 is to extend specific channel(s) 640 on bus 410 out to an application's message broker 555 (broker 555 may be a broker, platform, designated microservice or the like)”).
Regarding Claim 7, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the at least one local processing unit is configured as a gateway and the data exchange between the local resources and the central training computer takes place exclusively via the local processing unit (See at least Figure 1).
Regarding Claim 8, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the grasping instructions comprise an identification data set (See at least [0041] “The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts. Finally, an execution network reads the plan and generates actions for the robot”) used to (Indicates the intended purpose of the preceding limitation) identify at least one end effector suitable for the respective object (Does not specify how it is suitable, likely a 112(b) rejection as this is broad to the point of subjectivity if this was positively recited) from a set of end effectors of the end effector unit.
Regarding Claim 9, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the optical acquisition device is a device for (Indicates the intended use or purpose of the optical acquisition device) capturing depth images (See at least [0022] “Other sensors or mechanisms can be utilized as well, as may include depth sensors”) and for (Indicates the intended use or purpose of the optical acquisition device) capturing intensity images in the visible or infrared spectrum.
Regarding Claim 10, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 9,
Tremblay further teaches:
wherein the computed grasping instructions can be (This indicates capability or capacity for which is especially broad. There is no indication that any plan comprising “grasping instructions” in Tremblay is unable to meet the following requirements, and furthermore such “visualization” is standard practice for robotic interfaces. See also [0083] “It also generates human-readable plans, unlike those of the recent work” and [0016] “In embodiments where the plan is human readable, a human can view the plan and make any corrections, either manually or through another demonstration of the task”) visualized by showing a virtual scene of the gripper grasping the respective object, the user interface being configured to output the calculated visualization of the grasping instructions (This appears to merely be describing display of grasping instructions. See again at least [0083]).
Regarding Claim 11, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein the post-training of the neural network is configured to be performed iteratively and cyclically (See at least [0068] “In one embodiment building a machine learning application is an iterative process that involves a sequence of steps”) following a transmission of post-training data in the form of refined result data sets comprising the image data acquired by the optical acquisition device, which are automatically annotated and which have been transmitted from the at least one local processing unit to the central training computer (See at least [0069] “the now classified data instances can be stored to the classified data repository, which can be used for further training of the trained model 508 by the training manager. In some embodiments the model will be continually trained as new data is available” and [0049] “During a training process, performance data is captured 402 or otherwise obtained or received that is representative of a task to be performed at least partially in the physical world. As mentioned, this can include image data captured by at least one camera, among other such options”).
Regarding Claim 12, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay further teaches:
wherein a post-training data set for post-training the neural network is gradually and continuously expanded by image data acquired by sensors of the optical acquisition device in the vicinity of the robot (See at least [0049] “During a training process, performance data is captured 402 or otherwise obtained or received that is representative of a task to be performed at least partially in the physical world. As mentioned, this can include image data captured by at least one camera, among other such options”).
Regarding Claim 13, the combination of Tremblay, Kehoe, and Shanley teaches:
a system according claim 1,
Tremblay further teaches or has already been shown to teach:
An operating method for operating a system according claim 1, comprising the following method steps:
on the central training computer: reading in the respective object (See at least [0027] “identifiable by their respective colors or other such aspects … orientation, location, relationship, and other information about the objects”);
on the central training computer: accessing a model storage (See at least model repository 134) to (Indicates intended purpose of proceeding) load the 3D model, comprising a CAD model, assigned to the selected respective object type and generating synthetic object data from it (See at least [0047] “synthetic data generated by randomly sampling”) and use it (See at least [0032] “The object detection network can be a convolutional neural network that is trained on a set of training images, using domain randomization to overcome any reality gap resulting from the use of synthetic data”) for the purpose of (Intended purpose) pre-training;
on the central training computer: Pre-training a neural network with the generated synthetic object data (See again at least [0032]);
on the central training computer: Provisioning of pre-training parameters (See again at least [0032]);
on the central training computer: Transmitting the pre-training parameters via the network interface to at least one local processing unit (See at least [0023] “This can involve, for example, using a training module 110 on the robot itself, or sending the data across the at least one network 122 for processing … At least some functionality may also operate on a remote device, networked device, or in “the cloud” in some embodiments”);
on the at least one local processing unit: reading pre-training parameters or post-training parameters of a pre-trained or post-trained neural network via the network interface (See at least [0026] “The execution neural network can perform the inference on the robot 102, on the client device 138, or using an inference 136 in the provider environment 124, among other such options. Once the instructions are generated, the instructions can be provided to the control system 104 of the robot, either directly or upon execution by the processor 112, etc.”) in order to (Intended purpose) implement the pre-trained or post-trained neural network;
on the at least one local processing unit: Acquiring the image data (See at least Figure 1 and field of view 118);
on the at least one local processing unit: applying the pre-trained or post-trained neural network with the acquired image data to determine the result dataset (See at least [0041] “a camera can acquire a live video feed of a scene, from which a pair of networks can infer the positions and relationships of objects in the scene in real time. The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts”);
on the at least one local processing unit:
…
ii) receives as input data reference image data and compares the image data with the reference image data to (Intended purpose of preceding) minimize alignment errors and to generate a refined result data set, wherein the reference image data is a synthesized, rendered image which is rendered based on the result data set determined by the neural network and the 3D model (See at least [0054] “During training, data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. If the neural network does not correctly label the input, then errors between the correct label and the predicted label are analyzed, and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels the input and other inputs in a training dataset”);
on the at least one local processing unit: calculating gripping instructions for the end effector unit of the robot based on the generated refined result data set (See at least [0041] “The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts. Finally, an execution network reads the plan and generates actions for the robot”);
on the at least one local processing unit: exchanging data with the robot controller (See at least Figure 1) for (Intended purpose) controlling the end effector unit of the robot with the generated gripping instructions;
on the at least one local processing unit: generating post-training data, wherein the refined result data set serves as the post-training data set and the refined result data set is transmitted to the central training computer (CTC) for the purpose of post-training (See at least [0069] “the now classified data instances can be stored to the classified data repository, which can be used for further training of the trained model 508 by the training manager. In some embodiments the model will be continually trained as new data is available” and [0049] “During a training process, performance data is captured 402 or otherwise obtained or received that is representative of a task to be performed at least partially in the physical world. As mentioned, this can include image data captured by at least one camera, among other such options”);
on the central training computer: acquiring the post-training data via the network interface (See at least Figure 1), the post-training data comprising the labeled real image data acquired with the optical acquisition device (See again at least [0049] and [0069]);
on the central training computer: Continuous and cyclical retraining of the neural network with the recorded retraining data until a convergence criterion is fulfilled for the provision of post-training parameters (See at least [0061] “In some embodiments the training manager can monitor the quality of patterns (i.e., the model convergence) during training, and can automatically stop the training when there are no more data points or patterns to discover”);
on the central training computer: transmitting the post-training parameters via the network interface to at least one local processing unit (See at least Figure 1).
Kehoe has already been shown to teach in combination with Tremblay:
…
executing a modified ICP algorithm which, as input data, firstly evaluates and compares the image data of the optical acquisition device which have been supplied to the implemented neural network for application and
…
Regarding Claim 14, the combination of Tremblay, Kehoe, and Shanley teaches:
a system according claim 1,
Tremblay further teaches or has already been shown to teach:
A method for operating a central training computer in a system according to claim 1, comprising the following method steps:
reading in the respective object type (See at least [0027] “identifiable by their respective colors or other such aspects … orientation, location, relationship, and other information about the objects”);
accessing the model storage (See at least model repository 134) in order to load the 3D model, in particular the CAD model, assigned to the detected object type and to generate synthetic object data from it and use it for the purpose of pre-training;
pre-training a neural network (See again at least [0032]) with the generated synthetic object data, which serve as pre-training data, to provide pre-training parameters;
transmitting the pre-training parameters via the network interface to the at least one local processing unit (See at least [0023] “This can involve, for example, using a training module 110 on the robot itself, or sending the data across the at least one network 122 for processing … At least some functionality may also operate on a remote device, networked device, or in “the cloud” in some embodiments”);
retrieving post-training data via the network interface, wherein the post-training data comprises a refined result data set based on image data acquired with the optical acquisition device, which is annotated in an automatic process (See at least [0056] “For example, the classified data might include a set of images that each includes a representation of a type of object, where each image also includes, or is associated with, a label, metadata, classification, or other piece of information identifying the type of object represented in the respective image”);
continuous and cyclical retraining of the neural network with the acquired retraining data until a convergence criterion is fulfilled (See at least [0061] “In some embodiments the training manager can monitor the quality of patterns (i.e., the model convergence) during training, and can automatically stop the training when there are no more data points or patterns to discover”) to provide retraining parameters
transmitting the post-training parameters via the network interface to the at least one local processing unit (See at least Figure 1).
Regarding Claim 15, the combination of Tremblay, Kehoe, and Shanley teaches:
The central operating method according to claim 14,
Tremblay further teaches:
in which the steps of acquiring post-training data, post-training, providing and transmitting the post-training parameters are carried out iteratively on the basis of newly acquired post-training data (See at least [0049] “During a training process, performance data is captured 402 or otherwise obtained or received that is representative of a task to be performed at least partially in the physical world. As mentioned, this can include image data captured by at least one camera, among other such options” and [0068] “building a machine learning application is an iterative process”).
Regarding Claim 16, the combination of Tremblay, Kehoe, and Shanley teaches:
a system according claim 1,
Tremblay further teaches or has already been shown to teach:
A local operating method for operating a local processing unit in a system according to claim 1, comprising:
reading of pre-training parameters or post-training parameters of a pre-trained or post-trained ANN via the network interface (See at least [0026] “The execution neural network can perform the inference on the robot 102, on the client device 138, or using an inference 136 in the provider environment 124, among other such options. Once the instructions are generated, the instructions can be provided to the control system 104 of the robot, either directly or upon execution by the processor 112, etc.”) in order to implement the pre-trained or post-trained neural network;
capturing of the image data with the optical capture device of the objects in the working area of the robot (See at least [0049] “During a training process, performance data is captured 402 or otherwise obtained or received that is representative of a task to be performed at least partially in the physical world. As mentioned, this can include image data captured by at least one camera, among other such options”);
applying the pre-trained or post-trained neural network to the captured image data (See at least [0069] “the now classified data instances can be stored to the classified data repository, which can be used for further training of the trained model 508 by the training manager. In some embodiments the model will be continually trained as new data is available”) to determine the respective result data set for the plurality of objects depicted in the image data;
…
and, secondly, reference image data, to minimize alignment errors and to generate a refined result data set, wherein the reference image data is a synthesized image which is rendered based on the result data set determined by the neural network and the 3D model (See at least [0054] “During training, data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. If the neural network does not correctly label the input, then errors between the correct label and the predicted label are analyzed, and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels the input and other inputs in a training dataset”);
calculating gripping instructions for application on the end effector unit of the robot based on the generated refined result data set See at least [0041] “The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts. Finally, an execution network reads the plan and generates actions for the robot”);
exchanging data with the robot controller to control the end effector unit of the robot with the generated gripping instructions (See at least Figure 1);
generating post-training data, wherein the generated refined result data set serves as a post-training data set and the refined result data set is transmitted to the central training computer for the purpose of post-training (See at least [0069] “the now classified data instances can be stored to the classified data repository, which can be used for further training of the trained model 508 by the training manager. In some embodiments the model will be continually trained as new data is available”).
Kehoe has already been shown to teach in combination with Tremblay:
…
executing a modified Iterative Closest Point algorithm which evaluates and compares as input data, firstly, the image data of the optical acquisition device supplied to the implemented neural network for evaluation
…
Regarding Claim 17, the combination of Tremblay, Kehoe, and Shanley teaches:
The local operating method according to claim 16
Tremblay further teaches:
The local operating method according to claim 16, wherein the acquisition of the image data for the respective object is triggered before the instructions for gripping the respective object are executed (See at least [0041] “a camera can acquire a live video feed of a scene, from which a pair of networks can infer the positions and relationships of objects in the scene in real time. The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts”).
Regarding Claim 18, the combination of Tremblay, Kehoe, and Shanley teaches:
The local operating method according to claim 16
Tremblay further teaches:
in which, when using the pre-trained neural network in a pre-training phase, the objects are arranged under certain simplifying assumptions, in particular on a plane and disjointly in the working area (See at least [0032] “using domain randomization to overcome any reality gap resulting from the use of synthetic data”), and in which, when using the post-trained neural network, the objects are arranged in the working area without adhering to any simplifying assumptions (These are already demonstrated as being in the real world which adheres to no assumptions).
Regarding Claim 19, the combination of Tremblay, Kehoe, and Shanley teaches:
a distributed system according claim 1,
Tremblay further teaches or has already been shown to teach:
A central training computer in a distributed system according to claim 1, comprising a storage in which an instance of a neural network stored (Redundant limitation to Claim 1), wherein the central training computer is configured for pre-training and for post-training of the neural network (Redundant recitation to Claim 1), which is trained for object recognition and for position detection, including detection of an orientation of the object (Redundant limitation to Claim 1), in order to calculate gripping instructions for an end effector unit of the robot for gripping the respective object (Redundant recitation to Claim 1);
wherein the central training computer is configured to read in the respective object type (Redundant recitation to Claim 1), and
wherein the central training computer has an interface to a model storage (See various memory and Figure 1), in which a geometric 3D model of objects of the respective object type is stored for a respective object type (See at least [0041] “Each object of interest can be modeled, such as by a bounding cuboid” and [0083] “our system operates in 3D”), and
(The remaining items appear to again be redundant recitations and limitations to Claim 1)
wherein the central training computer is configured to perform a pre-training exclusively with synthetically generated pre-training data, which are generated by means of the geometric, object-type-specific 3D model of the objects of the respective object type, and wherein, as a result of the pre-training, pre-training parameters of a pre-trained neural network are transmitted via a network interface to at the least one local processing unit, and
wherein the central training computer is configured to continuously and cyclically perform a post-training of the pre-trained neural network on the basis of post- training data and to transmit post-training parameters of a post-trained neural network via the network interface to the at least one local processing unit as a result of the post-training.
Regarding Claim 20, the combination of Tremblay, Kehoe, and Shanley teaches:
a distributed system according claim 1,
Tremblay further teaches or has already been shown to teach:
A local computing unit in a distributed system according to claim 1, wherein the local computing unit is configured for data exchange with a controller of the robot for controlling the robot and its end effector unit for executing the gripping task for one object, of the plurality of objects, at a time (Redundant limitations to Claim 1), and
wherein the at least one local processing unit is configured to store different instances of the neural network, in that the at least one local processing unit is configured to receive pre-training parameters and post-training parameters from the central training computer to (Intended purpose) implement a pre-trained neural network which is continuously and cyclically replaced by a post-trained neural network until a convergence criterion is satisfied (Redundant limitations to Claim 1), and
wherein the pretrained or post-trained neural network is applied in an inference phase by determining a result data set for the image data captured by the optical capture device (Redundant limitation to Claim 1),
…
secondly, reference image data (Redundant limitation to Claim 1) to minimize alignment errors and to generate a refined result data set, wherein the reference image data is a synthesized image which is rendered based on the result data set determined by the neural network and the 3D model (Redundant limitation to Claim 1), and wherein the refined result data set serves as a basis for calculating the gripping instructions for the end effector unit for gripping the object and transmitting these to the robot controller of the robot for executing the gripping task (Close to redundant, see again at least [0041] “a camera can acquire a live video feed of a scene, from which a pair of networks can infer the positions and relationships of objects in the scene in real time. The resulting percepts can be fed to another network that generates a plan to explain how to recreate those percepts”).
Kehoe has already been shown to teach in combination with Tremblay:
…
and wherein a modified Iterative Closest Point, ICP, algorithm is executed which evaluates and compares as input data, firstly, the image data of the optical acquisition device supplied to the implemented neural network for evaluation and,
…
Regarding Claim 21, the combination of Tremblay, Kehoe, and Shanley teaches:
The local processing computing unit according to claim 20,
Tremblay further teaches:
wherein the at least one local processing unit comprises a graphics processing unit used to (Intended use limitation) evaluate the neural network (See at least [0025] “In various embodiments the processor 112 (or a processor of the training manager 130 or inferencer 136) will be a central processing unit (CPU). As mentioned, however, resources in such environments can utilize GPUs to process data for at least certain types of requests”).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Tremblay et al. in light of Kehoe et al., Shanley, and Qi et al. (Qi, Charles R., et al. "Deep hough voting for 3d object detection in point clouds." proceedings of the IEEE/CVF International Conference on Computer Vision. 2019).
Regarding Claim 5, the combination of Tremblay, Kehoe, and Shanley teaches:
The system according to claim 1,
Tremblay does not teach, but in combination with Qi teaches:
wherein the neural network has a Votenet architecture comprising three modules comprising: a backbone configured for learning local features; an evaluation module configured for evaluating and/or accumulating the individual feature vectors; and a conversion module configured to convert a result of the accumulation into object detections (See at least Figure 2. See also Page 11 of Applicant’s originally filed specification which refers to this same reference and appears to merely summarize or paraphrase rather than modify, as well as Figure 10 of the originally filed drawings which is identical except for a slight change to the input, which is not the architecture itself).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to utilize the Votenet architecture as disclosed in Qi in one or more neural networks of Tremblay with a reasonable expectation of success. The Votenet architecture is a well-known architecture for object processing, including when using depth data.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW C GAMMON whose telephone number is (571)272-4919. The examiner can normally be reached M - F 10:00 - 6:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ADAM MOTT can be reached on (571) 270-5376. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW C GAMMON/Examiner, Art Unit 3657 /ADAM R MOTT/Supervisory Patent Examiner, Art Unit 3657