DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Remarks
Claim Objections
The objections to the claims are withdrawn in light of Applicant’s amendments. Examiner notes that Applicant’s amendments of these claims, in combination with those of the independent claims, have raised the issue to that of a rejection under Claim Rejections - 35 USC § 112(b).
Claim Rejections - 35 USC § 101
Applicant's arguments filed 08/13/2026 have been fully considered but they are not persuasive.
First, Applicant’s arguments rely on a limitation exhibiting issues under 35 USC 112(b). Any final determination with respect to 35 USC 101 cannot be made until Applicant’s own language rather than Examiner’s interpretation is provided.
Second, Applicant’s arguments appear to rely on features not found in the present claims. Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). The present claims do not ever claim a robot or equivalent. The present claims do not ever claim “a wider variety of tasks”, especially inasmuch as there is no baseline even for “wider” in the claims. The present claims do not claim “additional training data”, only “training data”.
Third, Applicant’s arguments appear to be that there is a “practical application” wherein “the effectiveness of training the robot may be increased”. As disclosed, the “hand-held manipulation device” is a human-used substitute for a robotic manipulator in collection of training data for use in training said robotic manipulator. The only training, rather than merely of data gathering or outputting, that appears claimed is of a “vision language model” which is not for, or of, a robot, and as disclosed and likewise claimed appears to be merely related to suggesting tasks to a user for collection of training data for further unclaimed training. In other words, training of a model for better training data for training a robot not trained itself by said model. Thus, if Applicant’s argument is that practical application occurs in training the robot, the claims must have limitations directed as such. MPEP 2106.04(d)(1) relates:
“In short, first the specification should be evaluated to determine if the disclosure provides sufficient details such that one of ordinary skill in the art would recognize the claimed invention as providing an improvement. The specification need not explicitly set forth the improvement, but it must describe the invention such that the improvement would be apparent to one of ordinary skill in the art. Conversely, if the specification explicitly sets forth an improvement but in a conclusory manner (i.e., a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology. Second, if the specification sets forth an improvement in technology, the claim must be evaluated to ensure that the claim itself reflects the disclosed improvement. That is, the claim includes the components or steps of the invention that provide the improvement described in the specification” (emphasis added).
Examiner finally notes that the only claims directed to any training are of Claims 3 and 12, and not the independent claims, or in other words all claims, and thus even if Applicant’s argument was considered to hold merit, it would only be with respect to two out of twenty claims.
Claim Rejections - 35 USC § 102
Applicant's arguments filed 08/13/2026 have been fully considered but they are not persuasive.
First, Applicant’s arguments appear to rely on features not found in the present claims. Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Applicant argues on Page 9 of the Remarks that Wang does not disclose “suggesting new tasks to be performed”. The claims as presently constructed do not require that the suggested tasks be “new”, and furthermore even if they recited “new”, the claims would need to further define the term as modifications to an existing task are still “new” otherwise they would not be modifications.
Second, Applicant states that “Wang does not disclose analyzing the image of the scene to determine one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device” which appears wholly unsupported by the disclosure of Wang. However, Applicant then further states “That is, Wang merely discloses instructing a user to modify a task based on feedback data rather than suggesting new tasks to be performed based on analyzing an image of a scene”, thus it is unclear if Applicant is indicating that they find Wang not to teach “analyzing the image of the scene to determine one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device” or, ‘analyzing the image of the scene to determine one or more suggested new tasks that may be performed by the user utilizing the hand-held manipulation device’.
Regardless, Wang discloses either of these. See the following recitations from Wang as necessary which may help walk Applicant through the relevant portions of Figure 4 (emphasis provided via underlining):
[0056] A robot control instruction 425 is generated based on the robot control module 420. The robot control instruction 425 indicates how a human-driven robot task, such as those tasks previously described in relation to FIGS. 2A-2B and 3A-3C, is to be performed when training or testing the robot control module 420. … In such embodiments, the robot control instruction 425 could be considered a first robot control instruction, a subsequent robot control instruction could be considered a second robot control instruction, and so on.
While not as clearly explicit as subsequent robot control instructions, it appears that robot control instruction is based on image analysis of the scene. See [0057] and Figure 4. Thus, the first robot control instruction likely reads on the claims.
However, Wang further discloses subsequent instructions with very explicitly described relation to the input data.
[0060] In addition to the feedback data 472, performance data and environmental data are also collected. This is referred to as collected data 474 in FIG. 4. … The collected performance data may be collected by the data collection device 450 and/or one or more visual sensing devices such as the visual sensing devices 120.
Thus, collected data 474 is clearly of images of the scene.
[0063] In response to receiving the feedback data 472 and the collected data 474 that indicates that the human-driven robot task was successfully performed (i.e., the cup 201 was successfully picked up), a robot control instruction 426 is generated based on the module 420.
Thus, the “Yes” path of Figure 4 is described, and specifically states that a new task is provided to the user, as illustrated by the loop of Figure 4, referred to as a robot control instruction 426
Furthermore, the “No” path of Figure 4 is similarly described, wherein robot control instruction 427 is generated (See e.g. [0068]), based on Collected Data 478.
[0067] The collected data 478 may include the sensor data, the camera data, and raw visual data as discussed previously for collected data 474.
In summary, Applicant does not specify the “suggested tasks” in any manner such that they must be “new”, a “first”, etc. or furthermore how those terms may even be defined such that Wang still does not read on them.
Claim Rejections - 35 USC § 103
Applicant’s arguments appear wholly dependent on those previously presented with respect to 35 USC § 102 already addressed above.
Conclusion
Examiner identified prior art very clearly significantly overlapping with Applicant’s disclosure having two authors in common. Applicant was respectfully requested to address the nature of this reference in 52. of the prior Office Action. Applicant appears to remain wholly silent with respect to this reference and request in the Remarks filed 08/13/2026.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1 – 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 10 and 19 recite the limitation “the one or more suggested tasks performed”. There is insufficient antecedent basis for this limitation in the claim. The claims only ever previously recite “one or more suggested tasks that may be performed”. There is never any positive recitation of the suggested tasks actually being performed.
In the interest of compact prosecution, the overall limitation of:
“receiving training data based on the one or more suggested tasks performed by the user utilizing the hand-held manipulation device” (Claim 1, considered representative for all independent claims)
is instead interpreted as reading:
“receiving training data of the one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device after or during performance of the one or more suggested tasks”, or equivalent.
Regarding Claim 6 and 15, the claims recite the limitation “the one or more suggested tasks”. The claims previously recite “the one or more suggested tasks performed” and “one or more suggested tasks that may be performed”. Therefore, there are two items both having antecedent basis for these limitations.
However, no further correction beyond that addressed with respect to Claims 1, 10, and 19 appears required, as the interpretation provided therein removes the “performed” noun phrase.
Regarding Claims 2 – 5, 7 – 9, 11 – 14, 16 – 18, and 20, the claims depend from claim(s) rejected above and inherit the deficiencies of said claim(s) as described above. Therefore, Claims 2 – 5, 7 – 9, 11 – 14, 16 – 18, and 20 are rejected under the same logic presented above.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1 – 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 10 will be used to illustrate the rejection with respect to all independent claims (Claims 1 and 19), which are consequently rejected under the same logic presented below.
Claim 10 recites:
A computing device comprising one or more processors configured to:
receive an image of a scene; receive a request to identify tasks that may be performed by a user utilizing a hand-held manipulation device, based on the image of the scene;
in response to receiving the request to identify the tasks that may be performed by the user utilizing the hand-held manipulation device, analyze the image of the scene to determine one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device;
output the one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device; and
receive training data based on the one or more suggested tasks performed by the user utilizing the hand-held manipulation device.
101 Analysis – Step 1: Statutory Category – Yes
The claim recites a machine. The claim falls within one of the four statutory categories. MPEP 2106.03 relates.
101 Analysis – Step 2A Prong One Evaluation: Judicial Exception – Yes – Mental Processes
In Step 2A, Prong one of the 2019 Patent Eligibility Guidance (PEG), a claim is to be analyzed to determine whether it recites subject matter that falls within one of the following groups of abstract ideas: a) mathematical concepts, b) mental processes, and/or c) certain methods of organizing human activity.
The Office submits that the foregoing bolded limitation(s) constitutes judicial exceptions in terms of “mental processes”. MPEP 2106.04(a)(2)(III) relates.
The claim recites limitations of analyzing information and determining (or possibly due to phrasing only for the purpose of determining) tasks from an image. These limitations, as drafted, are the performance of mental processes but for the recitation of “a computing device comprising one or more processors”. That is, other than reciting “a computing device comprising one or more processors” nothing in the claim elements precludes the limitations from being performed in the human mind, with or without the aid of pen and paper. The mere nominal recitation of “a computing device comprising one or more processors” does not take the claim limitations out of the mental processes grouping.
For example, a person might readily look at an image and think of possible tasks to be performed, particularly as it is clear that one may do so without even the need for an image but using mere imagination.
Thus, the claim recites mental processes.
101 Analysis – Step 2A Prong Two Evaluation: Practical Application – No
In Step 2A, Prong two of the 2019 PEG, a claim is to be evaluated whether, as a whole, it integrates the recited judicial exception into a practical application. As noted in MPEP 2106.04(d), it must be determined whether any additional elements in the claim beyond the abstract idea integrate the exception into a practical application in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize the judicial exception. The courts have indicated that additional elements such as: merely using a computer to implement an abstract idea, adding insignificant extra solution activity, or generally linking use of a judicial exception to a particular technological environment or field of use do not integrate a judicial exception into a “practical application.”
The Office submits that the foregoing underlined limitation(s) recite additional elements that do not integrate the recited judicial exception into a practical application.
The claim recites additional elements or steps of the activities being functions of “a computing device comprising one or more processors”. The use of a computing device merely describes how to generally perform the computations using a generic device, i.e. a computer and is recited at a high level of generality and is merely automating or performing the functions.
With respect to the remaining underlined limitations reciting functions of “receive” and “output”, these are insignificant extra-solution activity of mere data gathering and outputting. MPEP 2106.05(g)(3) relates.
Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
101 Analysis – Step 2B Evaluation: Inventive Concept – No
In Step 2B of the 2019 PEG, a claim is to be evaluated as to whether the claim, as a whole, amounts to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. MPEP 2106.05 relates.
As discussed with respect to Step 2A Prong Two, the additional elements in the claim amount to no more than mere instructions to apply the exception using a generic computer component(s). The same analysis applies here in 2B, i.e., mere instructions to apply an exception on a generic computer or computer component(s) cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Under the 2019 PEG, a conclusion that an additional element is insignificant extra-solution activity in Step 2A should be re-evaluated in Step 2B. Again, the functions of “obtain …” and “distribute …” are insignificant extra-solution activity of mere data gathering and outputting which is insufficient for both Step 2A Prong Two and Step 2B considerations. MPEP 2106.05(g)(3) relates.
The specification does not provide any indication that the computing device is anything other than a conventional computer. MPEP 2106.05(d)(II), and the cases cited therein, including Intellectual Ventures I, LLC v. Symantec Corp., 838 F.3d 1307, 1321 (Fed. Cir. 2016), TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610 (Fed. Cir. 2016), and OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363 (Fed. Cir. 2015), indicate that mere receipt or transmission of data over a network is a well‐understood, routine, and conventional function when it is claimed in a merely generic manner (as it is here). Accordingly, a conclusion that the providing step is well-understood, routine, conventional activity is supported under Berkheimer.
Furthermore, while Applicant may clearly indicate in their disclosure that the technical improvement is of automated task selection for a hand-held manipulation device identified for a user to manipulate in a particular manner (See e.g. [0003] – [0004]), the present claim language merely recites “outputting” with no specifics that actually require that the user be actually informed. For example, outputting may simply be retaining the information in storage where the user is ignorant of any determinations. Furthermore, as argued by Applicant in the Remarks filed on 8/13/2026, the technical improvement presented in the argument was of the training of a robot, rather than in the collection of training data for use thereof.
Thus, the claim is ineligible.
With respect to the dependent claims, the claims merely recite additional details to the mental processes already recited or circumstantial items to the devices involved. In other words, the additional limitations merely recite details which are categorically already addressed above. For example, Claims 2 and 3 do not meaningfully limit the model and merely recite it at a generic and high level. Claim 5 recites “a task” and does not refer back to any particular task, especially not a determined task. The audio of claim 6 is merely claimed as “indicating” which is a broad phrasing and leaves it such that it encompasses audio taken, or audio emitted, etc.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 4 – 8, 10, 13 – 17, and 19 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Wang et al. (US 20240189993 A1).
Regarding Claim 1, Wang teaches:
A method comprising:
receiving an image of a scene (See at least [0060] “In addition to the feedback data 472, performance data and environmental data are also collected. This is referred to as collected data 474 in FIG. 4. The collected performance data is data that is related to the performance of the human-driven robot task. For example, when the human-driven robot task is to pick up the cup 201, the collected performance data would be data related to all the actions that the human data collector 430 performed, such as gripping the cup and lifting the cup, that are related to picking up the cup. The collected performance data may be collected by the data collection device 450 and/or one or more visual sensing devices such as the visual sensing devices 120” and [0062] “Thus, in the embodiment, the collected data 474 includes sensor data collected from one or more sensors associated with the data collection device 450 and/or camera data from one or more cameras associated with the data collection device 450”);
receiving a request to identify tasks that may be performed by a user utilizing a hand-held manipulation device (See at least [0009] “In some embodiments, the data collection device is a forearm-mounted human-machine operation interface that is used to operate one or more robotic grippers or robotic hands in the execution of complex grasping and manipulation robot control tasks. In other embodiments, the data collection device is a palm-mounted human-machine operation interface that is used to operate one or more robotic grippers or robotic hands in the execution of complex grasping and manipulation robot control tasks” and Figure 1C), based on the image of the scene (See at least [0068] “In response to receiving the feedback data 476 and the collected data 478 that indicates that the human-driven robot task specified by the robot control instruction 425 was not successfully performed (i.e., the cup 201 was not picked up), a robot control instruction 427 is generated based on the module 420”);
in response to receiving the request to identify the tasks that may be performed by the user utilizing the hand-held manipulation device, analyzing the image of the scene to determine one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device (See again at least [0068]. See also [0040] wherein it is clear that all rendered instructions are provided to the user in a mixed reality which inherently requires image/scene analysis to appropriately locate the instructions such as the disclosed bounding box ([0042]), trajectory/path ([0040]), etc.);
outputting the one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device (See at least [0040] “In some embodiments, the the mixed reality device 140 renders the robot control instructions as data visualization instructions that show the human data collector 105 how to perform the human driven robot tasks” and [0057] “The mixed reality device 440 renders the robot control instruction 425 instruction 425 into a rendered instruction 445 that shows the human data collector 430 how to perform the human-driven robot task”); and
receiving training data based on the one or more suggested tasks performed by the user utilizing the hand-held manipulation device (See at least [0066] If the cup 201 is successfully moved to the dishwasher, then the human data collector 430 will provide feedback data 472 that has been marked to indicate successful completion of the human-driven robot task to the computation device 410 and the robot control module 420 for use in retraining or updating the robot control module 420. Collected data 474 that includes the sensor data, camera data and/or raw visual data as discussed previously may also be provided to the computation device 410 and the robot control module 420 for use in retraining or updating the robot control module 420” and [0071] “The process described in FIG. 4 may be performed during a training mode and/or a testing mode. During the training mode, the robot control module 420 may not have prior data collected that can be used for controlling a robot”).
Regarding Claim 4, Wang teaches:
The method of claim 1, further comprising:
receiving preferences associated with the tasks that may be performed by a user utilizing the hand-held manipulation device; and
determining the one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device based at least in part on the preferences (Examiner notes that this claim at present is extremely broad. The claim does not define or further claim what the “preferences” are, how they are “associated”, or how the determining is “based on the preferences”.
See at least [0072] “During the testing mode, the robot control module 420 will typically have an overall robot control mission that is to be tested based on a series of related human-driven robot tasks. For example, the overall robot control mission may be to move a cup from a counter to a dishwasher as discussed previously in relation to FIGS. 3A-3C. In the testing mode, the robot control module 420 will provide each of the related human-driven robot tasks one by one to the human data collector until the overall robot control mission is performed and thus successfully tested or feedback data is received indicating that the overall robot control mission is not able to be performed, in which case a new human-driven robot task may be provided as previously described that is intended to help in achieving the overall robot control mission”).
Regarding Claim 5, Wang teaches:
The method of claim 1, wherein the training data comprises a video of the user utilizing the hand-held manipulation device to perform a task (See at least [0030] “In operation, the various cameras included as the visual sensing devices 120 provide raw sensing data to the computation device 110. The raw sensing camera data may include position (RGB and depth) data of the data collection device 130 when the human data collector 105 moves the data collection device 130 while performing one or more robot control tasks related to training and/or testing the machine learning models 126 and/or other software 127. This data may be collected as video data and collected on a frame-by-frame basis”).
Regarding Claim 6, Wang teaches:
The method of claim 5, further comprising:
receiving audio indicating one or more suggested tasks (See at least [0039] “For example, in one embodiment the mixed reality device 140 renders the instructions as an audio or voice instruction that allows the human human data collector 105 to hear voice instructions from the computation device 110”); and
storing the video in association with the audio as the training data (See again [0039] above. Presently the nature of the association is not claimed, and video captured with respect to instructions is already shown as being disclosed).
Regarding Claim 7, Wang teaches:
The method of claim 1, further comprising:
receiving, from a first hand-held manipulation device, a first video of the first hand-held manipulation device and a second hand-held manipulation device performing a task;
receiving, from the second hand-held manipulation device, a second video of the first hand- held manipulation device and the second hand-held manipulation device performing the task (See at least [0007] “In some embodiments, the visual sensing devices include multiple cameras, such as depth cameras, which are located on the data collection device”, Figure 1D, and [0009] “In some embodiments, the data collection device is a forearm-mounted human-machine operation interface that is used to operate one or more robotic grippers or robotic hands in the execution of complex grasping and manipulation robot control tasks. In other embodiments, the data collection device is a palm-mounted human-machine operation interface that is used to operate one or more robotic grippers or robotic hands in the execution of complex grasping and manipulation robot control tasks. In still other embodiments, the hands and/or arms of the human data collector can be considered as the data collection device … In further embodiments, the data collection device can be sensing gloves or hand pose tracking devices, such as motion capture gloves. In still further embodiments, the data collection device can be considered any combination of the human-machine operation interface, the human data collector hands and/or arms, the sensing gloves, or the hand pose tracking devices”); and
storing the first video and the second video as the training data (Wang already discloses recording manipulation data or even effectively all task related sensor data for training. See previous recitations as necessary).
Regarding Claim 8, Wang teaches:
The method of claim 1, further comprising:
receiving, from a first hand-held manipulation device, a first video of the first hand-held manipulation device and a second hand-held manipulation device performing a task;
receiving, from the second hand-held manipulation device, a second video of the first hand- held manipulation device and the second hand-held manipulation device performing the task;
receiving, from a head-mounted camera, a third video of the first hand-held manipulation device and the second hand-held manipulation device performing the task (While not explicit, it is clear the VR/AR/Mixed Reality device inherently includes a camera taking video, both from naturally understood structure of such a device (Examiner’s takes official notice of such common knowledge), and the disclosure. See e.g. [0038] “As illustrated, the testing and/or training system 100 also includes the mixed reality device 140, which may be a virtual reality/augmented reality (VR/AR) device or other human usable interface such as a screen a human can view or a voice interface that can receive audio messages from”, [0048] “The upper left of FIG. 3A shows an example embodiment of what is seen by the human data collector 105 through the mixed reality device 140. As shown, the human data collector 105 sees the object or cup 150 along with a portion of the data collection device 130. In addition, the mixed reality device 140 renders a visualization instruction that shows the human data collector 105 a visual path 310, which is an example of a visual marker, to take to pick up the cup 150 and move the cup to the dishwasher machine 330”, and Figures 3A – 3C of Wang. It is clear that in order to render a properly located visualization as disclosed, a video image of the operator’s view is required
Furthermore, and alternatively, the claim does not describe or otherwise claim with any particularity the meaning of “head-mounted”, for example if it is specifically mounted to the head of the user rather than some other device or object which may arbitrarily or conventionally be considered or labeled a “head”. Under this broad construction, any of the other cameras disclosed such as camera 120 may read on the limitation); and
storing the first video, the second video, and the third video as the training data (Wang already discloses recording manipulation data or even effectively all task related sensor data for training. See previous recitations as necessary. Furthermore, Applicant at this point has still not particularly defined “training data”).
Regarding Claims 10, 13 – 17, and 19, the claims are directed to effectively the same subject matter as Claims 1 and 4 – 8 with respect to the application of prior art. The claims are therefore rejected under the same logic as Claims 1, 4 – 8 above. The only distinction appears to be the recitation of generic computing components, which are still disclosed by Wang. See e.g. [0076].
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2 – 3, 11 – 12, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Nie et al. (US 20250308082 A1).
Regarding Claim 2, Wang teaches:
The method of claim 1, further comprising:
Wang does not explicitly teach, but in combination with Nie teaches:
inputting the image of the scene into a vision language model, and
determining the one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device based on an output of the vision language model (See at least [0029] “The general overview includes a process 100 for generating the blob representations … For instance, based on an input image 102, the process 100 includes using an open vocabulary segmentation 104 and vision language model 108 to generate the blob representations including the blob parameters 106 and the blob descriptions 110” and [0068] “Among other benefits and advantages, embodiments of the present disclosure provide a process 100 to decompose an input image 102 into blob parameters 106 to describe the location and size of objects within the image 102 and blob descriptions 110 to describe the visual appearance of the objects”).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to utilize a vision language model as disclosed in Nie to identify object parameters such as object location, size, and descriptions as taught by Nie within the processing for determination of instructions provided to the user of Wang (See e.g. Figure 4 and [0068]) with a reasonable expectation of success. Wang discloses providing the user with various refined instructions which are context aware such that the person might accurately perform the desired actions for training recordation. Nie discloses a clear means for accomplishing facilitating these features which do not appear explicitly disclosed in Wang. Furthermore, the inclusion of clear descriptions will significantly aid in identifying a unique feature of a target item out of a group of like-featured items.
Regarding Claim 3, the combination of Wang and Nie teaches:
The method of claim 2, further comprising:
Wang does not explicitly teach, but the combination with Nie further teaches (Nie recited below):
training the vision language model to receive the image of the scene as input, and output the one or more suggested tasks that may be performed by the user utilizing the hand-held manipulation device based on the image of the scene (See at least [0076] “In an embodiment, prior to performing steps 310 and 320, the method 300 further comprises training the blob-grounded text-to-image diffusion model using training data comprising a training image and one or more training models”).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to train any models within the system of Wang or Wang in combination with Nie with a reasonable expectation of success. It is well understood and routine to train a model to input and output the items for which it is to be used. Therefore, this is clearly what the combination of Wang and Nie contemplates in the combination provided in the rejection of Claim 2 above.
Regarding Claims 11 – 12 and 20, the claims are directed to effectively the same subject matter as Claims 2 – 3 with respect to the application of prior art. The claims are therefore rejected under the same logic as Claims 2 – 3 above. The only distinction appears to be the recitation of generic computing components, which are still disclosed by Wang. See e.g. [0076].
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Fuchs et al. (US 20020177967 A1).
Regarding Claim 9, Wang teaches:
The method of claim 8, further comprising:
synchronizing … the first hand-held manipulation device, … the second hand-held manipulation device, and … the head-mounted camera … first video, the second video and the third video … (Examiner notes that at present “are recorded” is non-specific and not necessarily exclusive of storage of the information rather than active sensor data capture.
See at least [0010] “In some embodiments, the computation device oversees real-time synchronization of multiple data resources, data processing, and data visualization by providing commands to and receiving collected data from the other elements of the testing and/or training system”, [0026] “The computation device 110 provides timestamps for incoming data. The timestamps are useful for future time synchronization”, and [0062] “The collected raw visual data and collected sensor data and camera data may then be synchronized into a data set that is used to retrain or otherwise update the robot control module 420”)
Wang does not explicitly teach (though the above may be considered as teaching these limitations), but in combination with Fuchs very clearly teaches:
…
a first clock of [the first hand-held manipulation device] …
… a second clock of [the second hand-held manipulation device] …
… a third clock of [the head-mounted camera] …
… before the [first video, the second video and the third video] are recorded …
(See at least [0044] “it is important to synchronize the internal clocks for the video recorder and each of the DAT recorders prior to the recording process”).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to synchronize the clocks of all data collection devices prior to data recording including any video recording devices as taught by Fuchs in the system and method of Wang with a reasonable expectation of success. As disclosed by Fuchs in the following sentence of [0044], “Once this is done, the video and DAT recorders can be set to automatically provide the recordings with synchronized time/date stamps”.
Regarding Claim 18, the claim is directed to effectively the same subject matter as Claim 9 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 9 above. The only distinction appears to be the recitation of generic computing components, which are still disclosed by Wang. See e.g. [0076].
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chi et al. (Chi, Cheng, et al. "Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots." arXiv preprint arXiv:2402.10329 (2024)) recited in the IDS dated 04/17/2025 which appears to disclose most of the same disclosure except for that of a model suggesting the task(s) to be performed.
Examiner notes that this reference has two authors which common to the inventorship of the instant Applicant. However, this reference also has six other authors not in common with the instant application, two of which are explicitly listed as having “equal contribution”, wherein the two inventors in common with the authorship of this reference are not indicated as such. Therefore, at a minimum the inventorship is “by another” and it is additionally at minimum implied that the majority of the contributions to the reference are from the two indicated authors which are not in common with the instant application. MPEP 2109(V) and MPEP 2153.01(a) relate.
Applicant is respectfully requested to address the nature of this reference prior to its potential use as prior art in a future Office Action in the interest of compact prosecution if Applicant believes that such a rejection might be overcome.
Such a rejection, for example under 35 U.S.C. 103, might be overcome by: (1) a showing under 37 CFR 1.130(a) that the subject matter disclosed in the reference was obtained directly or indirectly from the inventor or a joint inventor of this application and is thus not prior art in accordance with 35 U.S.C.102(b)(2)(A); (2) a showing under 37 CFR 1.130(b) of a prior public disclosure under 35 U.S.C. 102(b)(2)(B); or (3) a statement pursuant to 35 U.S.C. 102(b)(2)(C) establishing that, not later than the effective filing date of the claimed invention, the subject matter disclosed and the claimed invention were either owned by the same person or subject to an obligation of assignment to the same person or subject to a joint research agreement. See generally MPEP § 717.02.
Turkelson et al. (US 20200210768 A1) which discloses a training collection method for computer vision involving instructing a user/operator in the collection of training data.
Jarvis et al. (US 20240085974 A1) which appears to be from the same family of disclosure as Wang and which appears to more specifically focus on the potential structures of a system such as is disclosed in Wang. For example, Jarvis very clearly discloses cameras mounted to two user manipulated robotic grippers for robotic training (features of Claim 8)
Florence et al. (US 20250144795 A1) which discloses controlling a robot using multi-modal language models wherein the modality includes images/video.
Reher (US 12611767 B2) which discloses the use of vision-language-action models for processing data and generating decisions.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW C GAMMON whose telephone number is (571)272-4919. The examiner can normally be reached M - F 10:00 - 6:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ADAM MOTT can be reached on (571) 270-5376. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW C GAMMON/Examiner, Art Unit 3657
/ADAM R MOTT/Supervisory Patent Examiner, Art Unit 3657