Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This Office action is in response to the amendment filed on 05/28/2026. Claims 1-20 are currently pending with claims 1 and 13-19 being amended.
Response to Amendment
The amendments to the claims submitted on 05/28/2026 overcome the claim objections set forth in the previous Office action except for those set forth in the claim objection section.
Response to Arguments
Examiner notes wherein Applicant argues the newly amended limitations, which have not been addressed by the prior art of record. As such, Examiner has augmented the below rejection(s) in view of the prior art of record to address the newly amended limitations.
Applicant’s amendments/arguments, see communications, filed 05/28/2026, with respect to the rejection(s) of claim(s) 1-11, 13-17, and 19-20 under U.S.C. 103 in view of Song et al. (US 20220096197 A1) have been fully considered and are persuasive. However, upon further consideration, a new ground(s) of rejection is made in view of Wang et al. (Explainable Human-Robot Training and Cooperation with Augmented Reality).
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 1 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “human data collector” is introduced multiple times without clarification as to whether these are meant to indicate a plurality of different humans or if this is the same human.
Claim 1 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “mixed reality device” is introduced multiple times without clarification as to whether these are meant to indicate a plurality of different devices or if this is the same device.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-11, 13-17, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (Explainable Human-Robot Training and Cooperation with Augmented Reality), hereinafter Wang in view of Rose et al. (US 20240253224 A1), hereinafter Rose, Bank et al. (US 20200030979 A1), hereinafter Bank and Brown et al. (US 12455896 B2), hereinafter Brown.
Regarding claim 1, Wang teaches:
1. (Currently Amended) A method for testing and/or improving a robot control system, the method comprising:
…
converting the high-level instructions into human-operated robot tasks (Page 4, "Learned skills contain both high-level knowledge about the physical effects of different devices and low-level knowledge on how to operate these devices and manipulate objects. We use a standard symbolic STRIPS planner to combine the skills to solve novel and more complex tasks considering present objects, both real and virtual ones. For example, when asked to ’Prepare an ice tea’ the planner might consider the kettle or the microwave to heat some water before putting a tea bag inside and the fridge or some ice cubes to cool it later. Such a high-level solution is extended with the required low-level manipulations, like placing objects, opening doors and pressing buttons. By learning the skills and applying them to new tasks and environments, the system generalizes from previously observed episodes. As a consequence, the generated plans may not be feasible or desired and need validation by a human. Here, we propose an interactive system which can show the plan of the robot to the human via AR glasses before it is executed (see Fig. 3). The user can give a command to the robot via speech. Then the robot will generate a plan to solve the query according to its knowledge at different levels. Before execution, the robot will ask the user to validate the plan in AR. The virtual "avatar" of the robot appears overlaid on the physical body of the robot and real object "shadows" (holographic twins) are displayed in the AR glasses. Then the virtual robot will execute the plan with the virtual objects. In this way, the human can understand the robot reasoning and provide feedback to its plan, approving or correcting it." This demonstrates the ability of the system to break down a high level task into distinct steps which may be performed in order to accomplish the goal.) that are performable by a human data collector wearing a mixed reality device, in lieu of executing the high-level instructions on the robot; (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels). The robot follows the teacher’s gaze or looks back at him to show its attention. The teacher is further continuously informed about the robot’s perception: Object and action labels (XAI cues) are popping up whenever the teacher gazes at some object or a manual action is recognized (cfr. Fig. 2 and [12]). The state of the environment, including the objects and agents, is recorded before and after the demonstrated skill by the episodic memory. These observations are parsed into predefined symbolic representations of the environment, consisting of logical predicates. Together with the skill label, these observations are used to learn the symbolic skill, capturing in which situations it can be applied and what is changed in the environment by executing it." This demonstrates that the system may learn each skill by observing a human performing the skill while wearing AR glasses.)
providing the human-operated robot tasks to a mixed reality device worn by a human data collector, the mixed reality device rendering the human-operated robot tasks Page 3, Paragraphs 2-4, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. The HoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig. 1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.
The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input. More importantly, as the HoloLens can track the user’s hand and fingers, then the manual actions (e.g., "pick" or "drop") are also detected. The manipulation of objects by a human hand is implemented via Microsoft MRTK SDK4, which enables the corresponding virtual object to stick to the human’s hand while the "picking/holding" gesture is applied, and release from the hand after a "drop" gesture is detected. Moreover, colliders are attached to the user’s fingers, enabling the teacher, for instance, to press the toaster lever, turn on the power button of the microwave and close the microwave door. Finally, the user behavior and related manipulated object information are sent to the back-end system via ROS (see Fig. 1, "INPUT"). The system can also display and animate a holographic virtual robot, which looks almost identical as the physical robot. The HoloLens receives the robot behavior data (see Fig. 1 "OUTPUT"), including the pose of the virtual robot and the speech commands, and visualizes / speaks them. Furthermore, the back-end triggers the display of the XAI cues which are shown in the AR environment. The next sections introduce three use cases for enhancing human-robot interaction via AR based on this system architecture." in a manner that shows the human data collector how to perform the human-operated robot task; (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels). The robot follows the teacher’s gaze or looks back at him to show its attention. The teacher is further continuously informed about the robot’s perception: Object and action labels (XAI cues) are popping up whenever the teacher gazes at some object or a manual action is recognized (cfr. Fig. 2 and [12]). The state of the environment, including the objects and agents, is recorded before and after the demonstrated skill by the episodic memory. These observations are parsed into predefined symbolic representations of the environment, consisting of logical predicates. Together with the skill label, these observations are used to learn the symbolic skill, capturing in which situations it can be applied and what is changed in the environment by executing it." This demonstrates that the system may provide cues indicating how it is desired to interact with different objects within the environment.)
…
Wang does not specifically discuss using a library of actions which may be used in sequence to perform operations, using machine learning to generate instructions, receiving feedback from an operator in response to an attempt/simulation and updating the control system accordingly. However, Rose, in the same field of endeavor of robotics, teaches:
providing at a computation device a library of prompt templates that define one or more steps of one or more robot control tasks; (Paragraph 0032, "In some implementations, a robot system or control module may employ a finite Instruction Set comprising generalized reusable work primitives that can be combined (in various combinations and/or permutations) to execute a task. For example, a robot control system may store a library of reusable work primitives each corresponding to a respective basic sub-task or sub-action that the robot is operative to autonomously perform (hereafter referred to as an Instruction Set). A work objective may be analyzed to determine a sequence (i.e., a combination and/or permutation) of reusable work primitives that, when executed by the robot, will complete the work objective. The robot may execute the sequence of reusable work primitives to complete the work objective. In this way, a finite Instruction Set may be used to execute a wide range of different types of tasks and work objectives across a wide range of industries.")
providing a prompt to one or more generative Al(artificial intelligence) models, the one or more generative Al models generating high-level instructions, the high-level instructions configured to trigger, when executed, a robot to perform the one or more robot control tasks comprising one or more steps defined by the library of prompt templates; (Paragraph 0049, "In some implementations of the present systems, methods, control modules, and computer program products, an LLM is used to assist in determining a sequence of reusable work primitives (hereafter “Instructions”), selected from a finite library of reusable work primitives (hereafter “Instruction Set”), that when executed by a robot will cause or enable the robot to complete a task. In some implementations, an LLM is used to assist in determining a “workflow”. For example, a robot control system may take a Natural Language (NL) command as input and return a Task Plan formed of a sequence of allowed Instructions drawn from an Instruction Set whose completion achieves the intent of the NL input. Throughout this specification and the appended claims, unless the specific context requires otherwise a Task Plan may comprise, or consist of, a workflow depending on the specific implementation. Take as an exemplary application the task of “kitting” a chess set comprising sixteen white chess pieces and sixteen black chess pieces. A person could say, or type, to the robot, e.g., “Put all the white pieces in the right hand bin and all the black pieces in the left hand bin” and an LLM could support a fully autonomous system that converts this input into a sequence of allowed Instructions that successfully performs the task. In this case, the LLM may help to allow the robot to perform general tasks specified in NL. General tasks include but are not limited to all work in the current economy.")
However, Bank, in the same field of endeavor of robotics, teaches:
… receiving feedback data in response to the human data collector attempting to perform the human-operated robot tasks, (Paragraph 0032, "The MR simulation may be rerun to test the adjusted application, and with additional iterations as necessary, until the simulated operation of the robotic device is successful. The above simulation provides an example of programming the robotic device to learn various possible paths, such as a robotic arm with gripper, for which instructions are to be executed for motion control in conjunction with feedback from various sensor inputs.") the feedback data comprising at least one of (i) an overwrite, by the human data collector, of one or more of the human-operated robot tasks performed in real time without restarting the method, (Paragraphs 0018-0019, "FIG. 1 shows an example of a system of computer based tools according to embodiments of this disclosure. In an embodiment, a user 101 may wear a mixed reality (MR) device 115, such as Microsoft HoloLens, while operating a graphical user interface (GUI) device 105 to train a skill engine 110 of a robotic device. The skill engine 110 may have multiple application programs 120 stored in a memory, each program for performing a particular skill of a skill set or task. For each application 120, one or more modules 122 may be developed based on learning assisted by MR tools, such as MR device 115 and MR system data 130. As the application modules 122 are programmed to learn tasks for robotic devices, task parameters may be stored as local skill data 112, or cloud based skill data 114. The MR device 115 may be a wearable viewing device that can display a digital representation of a simulated object on a display screen superimposed and aligned with real objects in a work environment. The MR device 115 may be configured to display the MR work environment as the user 101 monitors execution of the application programs 120. The real environment may be viewable in the MR device 115 as a direct view in a wearable headset (e.g., through a transparent or semi-transparent material) or by a video image rendered on a display screen, generated by an internal or external camera. For example, the user 101 may operate the MR device 115, which may launch a locally stored application to load the virtual aspects of the MR environment and sync with the real aspects through a sync server 166. The sync server 166 may receive real and virtual information captured by the MR device 115, real information from sensors 162, and virtual information generated by simulator 164. The MR device 115 may be configured to accept hand gesture inputs from user 101 for editing, tuning, updating, or interrupting the applications 120. The GUI device 105 may be implemented as a computer device such as a tablet, keypad, or touchscreen to enable entry of initial settings and parameters, and editing, tuning or updating of the applications 120 based on user 101 inputs. The GUI device 105 may work in tandem with the MR device 115 to program the applications 120.") … and updating the robot control system using the feedback data. (Paragraphs 0031-0032, "For the example simulation, the real object 214 may act as an obstacle for the virtual workpiece 222. For an initial programming of a spatial-related application, the MR device 115 may be used to observe the path of the virtual workpiece as it travels along a designed path 223 to a target 225. The user 101 may initially setup the application program using initial parameters to allow programming of the application. For example, the initial spatial parameters and constraints may be estimated with the knowledge that adjustments can be made in subsequent trials until a spatial tolerance threshold is met. Entry of the initial parameters may be input via the GUI device 105 or by an interface of the MR device 115. Adjustments to spatial parameters of the virtual robotic unit 231 and object 222 may be implemented using an interface tool displayed by the MR device 115. For example, spatial and orientation coordinates of the virtual gripper 224 may be set using a visual interface application running on the MR device 115. As the simulated operation is executed, one or more application modules 122 may receive inputs from the sensors, compute motion of virtual robotic unit 231 and/or gripper 224, monitor for obstacles based on additional sensor inputs, receive coordinates of obstacle object 214 from the vision system 212, and compute a new path if necessary. Should the virtual workpiece fail to follow the path 223 around the obstacle object 214, the user may interrupt the simulation, such as by using a hand gesture with the MR device 115, then modify the application using GUI 105 to make necessary adjustments.
The MR simulation may be rerun to test the adjusted application, and with additional iterations as necessary, until operation of the simulated robotic unit 231 is successful according to constraints, such as a spatial tolerance threshold. For example, the path 223 may be required to remain within a set of spatial boundaries to avoid collision with surrounding structures. As another example, the placement of object 222 may be constrained by a spatial range surrounding target location 225 based on coordination with subsequent tasks to be executed upon the object 222.")
However, Brown, in the same field of endeavor of robotics, teaches:
… or (ii) a score assigned by the human data collector to one or more of the human-operated robot tasks; … (Column 78, Lines 1-20, "In some embodiments, the model updater 108 uses train the additional data generated by the other models in combination with the unstructured data of the unstructured service reports to configure the trained second model 116. The additional data generated by the other models can also or alternatively be used by the applications 120 in combination with an output of the second model 116 to select an action to perform. For example, the output of the trained second model 116 (e.g., a recommended action to perform) can be provided as an input to the other models to predict a consequence of the recommended action on energy consumption, occupant comfort, air quality, sustainability, infection risk, or any other variable state or condition predicted or modeled by the other models. The output of the other models can then be used by the system 100 to evaluate the consequences of the recommended action (e.g., score the recommended action relative to other recommended actions based on the consequences) and/or provide a user interface that informs the user of the consequences when presenting the recommended actions for user consideration.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic teaching system and methods as taught by Wang with the ability to utilize minimanipulations which are combined into a sequence to achieve an operation with machine learning as taught by Rose as well as with the ability to utilize feedback data to update the control of the robot as taught by Bank and Brown. This would provide a system capable of efficiently learning minimanipulations as well as effective sequences of minimanipulations to achieve an operation from a human operator.
Regarding claim 2, where all the limitations of claim 1 are discussed above, Wang further teaches:
2. (Original) The method of claim 1, wherein the rendering of the human- operated robot tasks at the mixed reality device includes one or more of (1) input/output variables such as location of objects, (2) visual/voice instructions for display on the mixed reality device, and (3) checkable events or metrics as a result of tasks that can be examined/scored by the mixed reality device or by external sensors/systems. (Page 3, Paragraphs 2-4, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. The HoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig. 1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.
The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input. More importantly, as the HoloLens can track the user’s hand and fingers, then the manual actions (e.g., "pick" or "drop") are also detected. The manipulation of objects by a human hand is implemented via Microsoft MRTK SDK4, which enables the corresponding virtual object to stick to the human’s hand while the "picking/holding" gesture is applied, and release from the hand after a "drop" gesture is detected. Moreover, colliders are attached to the user’s fingers, enabling the teacher, for instance, to press the toaster lever, turn on the power button of the microwave and close the microwave door. Finally, the user behavior and related manipulated object information are sent to the back-end system via ROS (see Fig. 1, "INPUT"). The system can also display and animate a holographic virtual robot, which looks almost identical as the physical robot. The HoloLens receives the robot behavior data (see Fig. 1 "OUTPUT"), including the pose of the virtual robot and the speech commands, and visualizes / speaks them. Furthermore, the back-end triggers the display of the XAI cues which are shown in the AR environment. The next sections introduce three use cases for enhancing human-robot interaction via AR based on this system architecture.")
Regarding claim 3, where all the limitations of claim 1 are discussed above, Wang does not specifically teach generating a prompt from a combination of a template form the library and a user input. However, Rose, in the same field of endeavor of robotic control, teaches:
3. (Original) The method of claim 1, further comprising generating the prompt from a combination of a prompt template from the library of prompt templates (Paragraph 0037, "In accordance with the present robots, systems, control modules, computer program products, and methods, a catalog of reusable work primitives may be defined, identified, developed, or constructed such that any given work objective across multiple different work objectives may be completed by executing a corresponding workflow comprising a particular combination and/or permutation of reusable work primitives selected from the catalog of reusable work primitives. Once such a catalog of reusable work primitives has been established, one or more robot(s) may be trained to autonomously or automatically perform each individual reusable work primitive in the catalog of reusable work primitives without necessarily including the context of: i) a particular workflow of which the particular reusable work primitive being trained is a part, and/or ii) any other reusable work primitive that may, in a particular workflow, precede or succeed the particular reusable work primitive being trained. In this way, a semi-autonomous robot may be operative to autonomously or automatically perform each individual reusable work primitive in a catalog of reusable work primitives and only require instruction, direction, or guidance from another party (e.g., from an operator, user, or pilot) when it comes to deciding which reusable work primitive(s) to perform and/or in what order. In other words, an operator, user, pilot, or LLM module may provide a workflow consisting of reusable work primitives to a semi-autonomous robot system and the semi-autonomous robot system may autonomously or automatically execute the reusable work primitives according to the workflow to complete a work objective. For example, a semi-autonomous humanoid robot may be operative to autonomously look left when directed to look left, autonomously open its right end effector when directed to open its right end effector, and so on, without relying upon detailed low-level control of such functions by a third party. Such a semi-autonomous humanoid robot may autonomously complete a work objective once given instructions regarding a workflow detailing which reusable work primitives it must perform, and in what order, in order to complete the work objective. Furthermore, in accordance with the present robots, systems, methods, control modules and computer program products, a robot system may operate fully autonomously if it is trained or otherwise configured to (e.g. via consultation with an LLM module, which can be included in the robot system) analyze a work objective and independently define a corresponding workflow itself by deconstructing the work objective into a set of reusable work primitives from a library of reusable work primitives that the robot system is operative to autonomously perform.") and a user prompt. (Paragraph 00146, "Various implementations of the present systems, methods, control modules, and computer program products involve using NL expressions (descriptions) (e.g., via a NL prompt, which may be entered directly in text by a user or may be spoken vocally by a user and converted to text by an intervening voice-to-text system) to control functions and operations of a robot, where an LLM module may provide an interface between the NL expressions and the robot control system. This framework can be particularly advantageous when certain elements of the robot control architecture employ programming and/or instructions that can be expressed in NL. A suitable, but non-limiting, example of this is the aforementioned Instruction Set. For example, as mentioned earlier, a task plan output of an LLM module can be parsed (e.g., autonomously by the robot control system) by looking for a word match to Instruction Set commands, and the arguments of the Instruction Set can be found by string matching within the input NL prompt (e.g. by a text-string matching module as discussed earlier). In some implementations, a 1-1 map may be generated between the arguments used in the robot control system and NL variants, in order to increase the chance of the LLM module processing the text properly. For example, even though an object is represented in the robot control system (e.g., in a world model environment portion of the robot control system) as chess_pawn_54677, it may be referred to in the NL prompt as “chess pawn 1”. In this case, if the returned task plan contains the phrase “grasp chess pawn 1”, this may be matched to Instruction Set “grasp” and the object “chess pawn 1” so the phrase may be mapped to grasp(chess_pawn_54677). Such parsing and/or word matching (e.g. the text-string matching module) can be employed in any of the situations discussed herein where robot language is converted to natural language or vice-versa.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods as taught by Wang with the ability to control the system and generate instructions based on the library of operations as well as a user input as taught by Rose. This would allow a user to control the system using natural language inputs necessitating less training before they are able to operate the system effectively.
Regarding claim 4, where all the limitations of claim 1 are discussed above, Wang does not specifically teach using high level instructions which correspond to pre-defined low level libraries. However, Rose, in the same field of endeavor of robotics, teaches:
4. (Original) The method of claim 1, wherein the high-level instructions call predefined low-level libraries. (Paragraph 0037, "In accordance with the present robots, systems, control modules, computer program products, and methods, a catalog of reusable work primitives may be defined, identified, developed, or constructed such that any given work objective across multiple different work objectives may be completed by executing a corresponding workflow comprising a particular combination and/or permutation of reusable work primitives selected from the catalog of reusable work primitives. Once such a catalog of reusable work primitives has been established, one or more robot(s) may be trained to autonomously or automatically perform each individual reusable work primitive in the catalog of reusable work primitives without necessarily including the context of: i) a particular workflow of which the particular reusable work primitive being trained is a part, and/or ii) any other reusable work primitive that may, in a particular workflow, precede or succeed the particular reusable work primitive being trained. In this way, a semi-autonomous robot may be operative to autonomously or automatically perform each individual reusable work primitive in a catalog of reusable work primitives and only require instruction, direction, or guidance from another party (e.g., from an operator, user, or pilot) when it comes to deciding which reusable work primitive(s) to perform and/or in what order. In other words, an operator, user, pilot, or LLM module may provide a workflow consisting of reusable work primitives to a semi-autonomous robot system and the semi-autonomous robot system may autonomously or automatically execute the reusable work primitives according to the workflow to complete a work objective. For example, a semi-autonomous humanoid robot may be operative to autonomously look left when directed to look left, autonomously open its right end effector when directed to open its right end effector, and so on, without relying upon detailed low-level control of such functions by a third party. Such a semi-autonomous humanoid robot may autonomously complete a work objective once given instructions regarding a workflow detailing which reusable work primitives it must perform, and in what order, in order to complete the work objective. Furthermore, in accordance with the present robots, systems, methods, control modules and computer program products, a robot system may operate fully autonomously if it is trained or otherwise configured to (e.g. via consultation with an LLM module, which can be included in the robot system) analyze a work objective and independently define a corresponding workflow itself by deconstructing the work objective into a set of reusable work primitives from a library of reusable work primitives that the robot system is operative to autonomously perform.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods of operating as taught by Wang with the low level operation libraries which may be used to accomplish high level tasks as taught by Rose. This allows the system to re-use may low level operations and efficiently program new tasks/actions in order for the system to complete high level operations.
Regarding claim 5, where all the limitations of claim 4 are discussed above, Wang does not specifically teach the libraries including computer vision, motion planning, or motion execution. However, Rose, in the same field of endeavor of robotics, teaches:
5. (Original) The method of claim 4, wherein the predefined low-level libraries include one or more of computer vision libraries, motion planning libraries, (Paragraph 0152, "The various implementations described herein include systems, methods, control modules, and computer program products for leveraging one or more LLM(s) in a robot control system, including for example establishing an NL interface between the LLM(s) and the robot control system and calling the LLM(s) to help autonomously instruct the robot what to do. Example applications of this approach include task planning, motion planning, reasoning about the robot's environment (e.g., “what could I do now?”), and so on. Such implementations are particularly well-suited in robot control systems for which at least some control parameters and/or instructions (e.g., the Instruction Set described previously) are amenable to being specified in NL. Thus, some implementations may include converting or translating robot control instructions and/or parameters into NL for communicating such with the LLM(s) via the NL interface.") and motion execution libraries. (Paragraph 0078, "FIG. 2 is a flowchart diagram which illustrates an exemplary method 200 of operation of a robot system. Method 200 in FIG. 2 is similar in at least some respects to the method 100 of FIG. 1. In general, method 200 in FIG. 2 describes detailed implementations by which method 100 in FIG. 1 can be achieved. Method 200 is a method of operation of a robot system (such as robot system 700 discussed with reference to FIG. 7). In general, throughout this specification and the appended claims, a method of operation of a robot system is a method in which at least some, if not all, of the various acts are performed by the robot system. For example, certain acts of a method of operation of a robot system may be performed by at least one processor or processing unit (hereafter “processor”) of the robot system communicatively coupled to a non-transitory processor-readable storage medium of the robot system (collectively a robot controller of the robot system) and, in some implementations, certain acts of a method of operation of a robot system may be performed by peripheral components of the robot system that are communicatively coupled to the at least one processor, such as one or more physically actuatable components (e.g., arms, legs, end effectors, grippers, hands), one or more sensors (e.g., optical sensors, audio sensors, tactile sensors, haptic sensors), mobility systems (e.g., wheels, legs), communications and networking hardware (e.g., receivers, transmitters, transceivers), and so on. The non-transitory processor-readable storage medium of the robot system may store data (including, e.g., at least one library of reusable work primitives and at least one library of associated percepts) and/or processor-executable instructions that, when executed by the at least one processor, cause the robot system to perform the method and/or cause the at least one processor to perform those acts of the method that are performed by the at least one processor. The robot system may communicate, via communications and networking hardware communicatively coupled to the robot system's at least one processor, with remote systems and/or remote non-transitory processor-readable storage media. Thus, unless the specific context requires otherwise, references to a robot system's non-transitory processor-readable storage medium, as well as data and/or processor-executable instructions stored in a non-transitory processor-readable storage medium, are not intended to be limiting as to the physical location of the non-transitory processor-readable storage medium in relation to the at least one processor of the robot system and the rest of the robot hardware. In other words, a robot system's non-transitory processor-readable storage medium may include non-transitory processor-readable storage media located on-board a robot body of the robot system and/or non-transitory processor-readable storage media located remotely from the robot body, unless the specific context requires otherwise. Further, a method of operation of a robot system such as method 200 (or any of the other methods discussed herein) can be implemented as a robot control module or computer program product. Such a control module or computer program product comprises processor-executable instructions or data that, when the control module or computer program product is stored on a non-transitory processor-readable storage medium of the robot system, and the control module or computer program product is executed by at least one processor of the robot system, the control module or computer program product (or the processor-executable instructions or data thereof) cause the robot system to perform acts of the method.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods of operating as taught by Wang with the low level operation libraries defining motion planning and motion execution which may be used to accomplish high level tasks as taught by Rose. This allows the system to re-use may low level operations and efficiently program new tasks/actions in order for the system to complete high level operations.
Regarding claim 6, where all the limitations of claim 1 are discussed above, Wang does not specifically discuss the feedback overwriting the task. However, Bank, in the same field of endeavor of robotics, teaches:
6. (Original) The method of claim 1, wherein the feedback data comprises an overwrite of one or more of the human-operated robot tasks by the human data collector. (Paragraphs 0018-0019, "FIG. 1 shows an example of a system of computer based tools according to embodiments of this disclosure. In an embodiment, a user 101 may wear a mixed reality (MR) device 115, such as Microsoft HoloLens, while operating a graphical user interface (GUI) device 105 to train a skill engine 110 of a robotic device. The skill engine 110 may have multiple application programs 120 stored in a memory, each program for performing a particular skill of a skill set or task. For each application 120, one or more modules 122 may be developed based on learning assisted by MR tools, such as MR device 115 and MR system data 130. As the application modules 122 are programmed to learn tasks for robotic devices, task parameters may be stored as local skill data 112, or cloud based skill data 114. The MR device 115 may be a wearable viewing device that can display a digital representation of a simulated object on a display screen superimposed and aligned with real objects in a work environment. The MR device 115 may be configured to display the MR work environment as the user 101 monitors execution of the application programs 120. The real environment may be viewable in the MR device 115 as a direct view in a wearable headset (e.g., through a transparent or semi-transparent material) or by a video image rendered on a display screen, generated by an internal or external camera. For example, the user 101 may operate the MR device 115, which may launch a locally stored application to load the virtual aspects of the MR environment and sync with the real aspects through a sync server 166. The sync server 166 may receive real and virtual information captured by the MR device 115, real information from sensors 162, and virtual information generated by simulator 164. The MR device 115 may be configured to accept hand gesture inputs from user 101 for editing, tuning, updating, or interrupting the applications 120. The GUI device 105 may be implemented as a computer device such as a tablet, keypad, or touchscreen to enable entry of initial settings and parameters, and editing, tuning or updating of the applications 120 based on user 101 inputs. The GUI device 105 may work in tandem with the MR device 115 to program the applications 120.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods as taught by Wang with the ability to replace a task based on feedback data as taught by Bank. This would allow the system to fine-tune the operation of the robot for more effective control.
Regarding claim 17, where all the limitations of claim 13 are discussed above, Wang does not specifically discuss the feedback scoring the task. However, Brown, in the same field of endeavor of robotics, teaches:
7. (Original) The method of claim 1, wherein the feedback data comprises a score for one or more of the human-operated robot tasks provided by the human data collector. (Column 78, Lines 1-20, "In some embodiments, the model updater 108 uses train the additional data generated by the other models in combination with the unstructured data of the unstructured service reports to configure the trained second model 116. The additional data generated by the other models can also or alternatively be used by the applications 120 in combination with an output of the second model 116 to select an action to perform. For example, the output of the trained second model 116 (e.g., a recommended action to perform) can be provided as an input to the other models to predict a consequence of the recommended action on energy consumption, occupant comfort, air quality, sustainability, infection risk, or any other variable state or condition predicted or modeled by the other models. The output of the other models can then be used by the system 100 to evaluate the consequences of the recommended action (e.g., score the recommended action relative to other recommended actions based on the consequences) and/or provide a user interface that informs the user of the consequences when presenting the recommended actions for user consideration.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods as taught by Wang with the ability to utilize a score when determining how to perform a task as taught by Brown. This would allow the system to fine-tune the operation of the robot for more effective control.
Regarding claim 8, where all the limitations of claim 1 are discussed above, Wang further teaches:
8. (Original) The method of claim 1, wherein the human data collector performs the human-operated robot tasks (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels). The robot follows the teacher’s gaze or looks back at him to show its attention. The teacher is further continuously informed about the robot’s perception: Object and action labels (XAI cues) are popping up whenever the teacher gazes at some object or a manual action is recognized (cfr. Fig. 2 and [12]). The state of the environment, including the objects and agents, is recorded before and after the demonstrated skill by the episodic memory. These observations are parsed into predefined symbolic representations of the environment, consisting of logical predicates. Together with the skill label, these observations are used to learn the symbolic skill, capturing in which situations it can be applied and what is changed in the environment by executing it." This demonstrates that the system may learn each skill by observing a human performing the skill while wearing AR glasses.) using a human-machine interface. (Page 3, Paragraph 2, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. TheHoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig.1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.")
Regarding claim 9, where all the limitations of claim 8 are discussed above, Wang further teaches:
9. (Original) The method of claim 8, wherein the human-machine interface (Page 3, Paragraph 2, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. TheHoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig.1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.") comprises at least one input interface that receives user input (Page 3, Paragraph 4, "The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input.") and provides corresponding control commands to a microcontroller unit to control the operation of one or more robotic components. (Page 3, Paragraph 1, “Finally, the action generation module (Fig. 1) receives action commands from other modules (e.g., where to look, what to grasp, etc.) and executes the corresponding motor behavior.”)
Regarding claim 10, where all the limitations of claim 9 are discussed above, Wang further teaches:
10. (Original) The method of claim 9, wherein the one or more robotic components comprise one or more of a robotic gripper, a robotic limb, and a robotic wheel. (Page 2, Paragraph 4, "The physical robot is a torso with two Kinova arms1 with 7 DOF each and a pan-tilt unit, part of the mobile platform "Johnny" [13].")
Regarding claim 11, where all the limitations of claim 9 are discussed above, Wang further teaches:
11. (Original) The method of claim 9, wherein the human-machine interface is a human-mounted human-machine interface. (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels).")
Regarding claim 13, Wang further teaches:
13. (Currently Amended) A system for training and/or testing a robot control system, the system comprising:
a computation device comprising at least one processor and a non-transitory processor- readable storage medium configured to train and/or test the robot control system; (Page 2, Paragraph 4 – Page 3, Paragraph 1, “The robot’s cognitive skills are realized in multiple ROS2 nodes, communicating with each other and with the front-end interface. The behavior engine orchestrates the social behavior of the robot. It controls the robot gaze, while concurrently issuing XAI cues to be shown in the AR environment (cfr. [12]). It also regulates the speech interaction, e.g. acknowledging the user commands or asking curiosity-driven questions (see Sec. 3). The episodic memory collects world state observations during demonstrations by the human teachers. The learning module realizes symbolic skill learning which integrates demonstrations from the episodic memory into a knowledge graph, and generates new hypotheses to be queried to the user (Sec. 3). The planning module uses the learned semantic skills to generate high-level plans to solve new tasks. Such plans are yet to be checked with the human tutor (see Sec. 4). The human action prediction module operates when assisting a user. The robot can predict what the user is intending to achieve, and plan a supportive action, e.g., moving the next required object closer to the user. The ergonomics module generates interventions in a ergonomically optimal way for the user (see Sec. 5). Finally, the action generation module (Fig. 1) receives action commands from other modules (e.g., where to look, what to grasp, etc.) and executes the corresponding motor behavior.”)
a data collection device comprising one or more cameras and one or more sensors configured to collect feedback data related to one or more human-operated robot tasks that are to be performed by a human data collector;
one or more sensing devices comprising one or more optical sensors located in a location of the human data collector; and
a mixed reality device comprising an augmented reality or virtual reality headset that is worn by the human data collector, (Page 3, Paragraph 4, “The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input. More importantly, as the HoloLens can track the user’s hand and fingers, then the manual actions (e.g., "pick" or "drop") are also detected. The manipulation of objects by a human hand is implemented via Microsoft MRTK SDK4, which enables the corresponding virtual object to stick to the human’s hand while the "picking/holding" gesture is applied, and release from the hand after a "drop" gesture is detected. Moreover, colliders are attached to the user’s fingers, enabling the teacher, for instance, to press the toaster lever, turn on the power button of the microwave and close the microwave door. Finally, the user behavior and related manipulated object information are sent to the back-end system via ROS (see Fig. 1, "INPUT"). The system can also display and animate a holographic virtual robot, which looks almost identical as the physical robot. The HoloLens receives the robot behavior data (see Fig. 1 "OUTPUT"), including the pose of the virtual robot and the speech commands, and visualizes / speaks them. Furthermore, the back-end triggers the display of the XAI cues which are shown in the AR environment. The next sections introduce three use cases for enhancing human-robot interaction via AR based on this system architecture.”) the system configured to perform the following:
…
convert the high-level instructions into human-operated robot tasks (Page 4, "Learned skills contain both high-level knowledge about the physical effects of different devices and low-level knowledge
on how to operate these devices and manipulate objects. We use a standard symbolic STRIPS planner to combine the skills to solve novel and more complex tasks considering present objects, both real and virtual ones. For example, when asked to ’Prepare an ice tea’ the planner might consider the kettle or the microwave to heat some water before putting a tea bag inside and the fridge or some ice cubes to cool it later. Such a high-level solution is extended with the required low-level manipulations, like placing objects, opening doors and pressing buttons. By learning the skills and applying them to new tasks and environments, the system generalizes from previously observed episodes. As a consequence, the generated plans may not be feasible or desired and need validation by a human. Here, we propose an interactive system which can show the plan of the robot to the human via AR glasses before it is executed (see Fig. 3). The user can give a command to the robot via speech. Then the robot will generate a plan to solve the query according to its knowledge at different levels. Before execution, the robot will ask the user to validate the plan in AR. The virtual "avatar" of the robot appears overlaid on the physical body of the robot and real object "shadows" (holographic twins) are displayed in the AR glasses. Then the virtual robot will execute the plan with the virtual objects. In this way, the human can understand the robot reasoning and provide feedback to its plan, approving or correcting it." This demonstrates the ability of the system to break down a high level task into distinct steps which may be performed in order to accomplish the goal.) that are performable by the human data collector in lieu of executing the high-level instructions on the robot; (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels). The robot follows the teacher’s gaze or looks back at him to show its attention. The teacher is further continuously informed about the robot’s perception: Object and action labels (XAI cues) are popping up whenever the teacher gazes at some object or a manual action is recognized (cfr. Fig. 2 and [12]). The state of the environment, including the objects and agents, is recorded before and after the demonstrated skill by the episodic memory. These observations are parsed into predefined symbolic representations of the environment, consisting of logical predicates. Together with the skill label, these observations are used to learn the symbolic skill, capturing in which situations it can be applied and what is changed in the environment by executing it." This demonstrates that the system may learn each skill by observing a human performing the skill while wearing AR glasses.)
provide the human-operated robot tasks to the mixed reality device worn by the human data collector, the mixed reality device rendering the human-operated robot tasks (Page 3, Paragraphs 2-4, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. The HoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig. 1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.
The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input. More importantly, as the HoloLens can track the user’s hand and fingers, then the manual actions (e.g., "pick" or "drop") are also detected. The manipulation of objects by a human hand is implemented via Microsoft MRTK SDK4, which enables the corresponding virtual object to stick to the human’s hand while the "picking/holding" gesture is applied, and release from the hand after a "drop" gesture is detected. Moreover, colliders are attached to the user’s fingers, enabling the teacher, for instance, to press the toaster lever, turn on the power button of the microwave and close the microwave door. Finally, the user behavior and related manipulated object information are sent to the back-end system via ROS (see Fig. 1, "INPUT"). The system can also display and animate a holographic virtual robot, which looks almost identical as the physical robot. The HoloLens receives the robot behavior data (see Fig. 1 "OUTPUT"), including the pose of the virtual robot and the speech commands, and visualizes / speaks them. Furthermore, the back-end triggers the display of the XAI cues which are shown in the AR environment. The next sections introduce three use cases for enhancing human-robot interaction via AR based on this system architecture.") in a manner that shows the human data collector how to perform the human-operated robot task; (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels). The robot follows the teacher’s gaze or looks back at him to show its attention. The teacher is further continuously informed about the robot’s perception: Object and action labels (XAI cues) are popping up whenever the teacher gazes at some object or a manual action is recognized (cfr. Fig. 2 and [12]). The state of the environment, including the objects and agents, is recorded before and after the demonstrated skill by the episodic memory. These observations are parsed into predefined symbolic representations of the environment, consisting of logical predicates. Together with the skill label, these observations are used to learn the symbolic skill, capturing in which situations it can be applied and what is changed in the environment by executing it." This demonstrates that the system may provide cues indicating how it is desired to interact with different objects within the environment.)
…
Wang does not specifically discuss using a library of actions which may be used in sequence to perform operations, using machine learning to generate instructions, receiving feedback from an operator in response to an attempt/simulation and updating the control system accordingly. However, Rose, in the same field of endeavor of robotics, teaches:
… provide at the computation device a library of prompt templates that define one or more steps of one or more robot control tasks; (Paragraph 0032, "In some implementations, a robot system or control module may employ a finite Instruction Set comprising generalized reusable work primitives that can be combined (in various combinations and/or permutations) to execute a task. For example, a robot control system may store a library of reusable work primitives each corresponding to a respective basic sub-task or sub-action that the robot is operative to autonomously perform (hereafter referred to as an Instruction Set). A work objective may be analyzed to determine a sequence (i.e., a combination and/or permutation) of reusable work primitives that, when executed by the robot, will complete the work objective. The robot may execute the sequence of reusable work primitives to complete the work objective. In this way, a finite Instruction Set may be used to execute a wide range of different types of tasks and work objectives across a wide range of industries.")
provide a prompt to one or more generative AI(artificial intelligence) models, the one or more generative AI models generating high-level instructions, the high-level instructions configured to trigger, when executed, a robot to perform the one or more robot control tasks comprising said one or more steps defined by the library of prompt templates; (Paragraph 0049, "In some implementations of the present systems, methods, control modules, and computer program products, an LLM is used to assist in determining a sequence of reusable work primitives (hereafter “Instructions”), selected from a finite library of reusable work primitives (hereafter “Instruction Set”), that when executed by a robot will cause or enable the robot to complete a task. In some implementations, an LLM is used to assist in determining a “workflow”. For example, a robot control system may take a Natural Language (NL) command as input and return a Task Plan formed of a sequence of allowed Instructions drawn from an Instruction Set whose completion achieves the intent of the NL input. Throughout this specification and the appended claims, unless the specific context requires otherwise a Task Plan may comprise, or consist of, a workflow depending on the specific implementation. Take as an exemplary application the task of “kitting” a chess set comprising sixteen white chess pieces and sixteen black chess pieces. A person could say, or type, to the robot, e.g., “Put all the white pieces in the right hand bin and all the black pieces in the left hand bin” and an LLM could support a fully autonomous system that converts this input into a sequence of allowed Instructions that successfully performs the task. In this case, the LLM may help to allow the robot to perform general tasks specified in NL. General tasks include but are not limited to all work in the current economy.") …
However, Bank, in the same field of endeavor of robotics, teaches:
… receive feedback data in response to the human data collector attempting to perform the human-operated robot tasks, (Paragraph 0032, "The MR simulation may be rerun to test the adjusted application, and with additional iterations as necessary, until the simulated operation of the robotic device is successful. The above simulation provides an example of programming the robotic device to learn various possible paths, such as a robotic arm with gripper, for which instructions are to be executed for motion control in conjunction with feedback from various sensor inputs.") the feedback data comprising at least one of (i) an overwrite, by the human data collector, of one or more of the human-operated robot tasks performed in real time without restarting the system, (Paragraphs 0018-0019, "FIG. 1 shows an example of a system of computer based tools according to embodiments of this disclosure. In an embodiment, a user 101 may wear a mixed reality (MR) device 115, such as Microsoft HoloLens, while operating a graphical user interface (GUI) device 105 to train a skill engine 110 of a robotic device. The skill engine 110 may have multiple application programs 120 stored in a memory, each program for performing a particular skill of a skill set or task. For each application 120, one or more modules 122 may be developed based on learning assisted by MR tools, such as MR device 115 and MR system data 130. As the application modules 122 are programmed to learn tasks for robotic devices, task parameters may be stored as local skill data 112, or cloud based skill data 114.
The MR device 115 may be a wearable viewing device that can display a digital representation of a simulated object on a display screen superimposed and aligned with real objects in a work environment. The MR device 115 may be configured to display the MR work environment as the user 101 monitors execution of the application programs 120. The real environment may be viewable in the MR device 115 as a direct view in a wearable headset (e.g., through a transparent or semi-transparent material) or by a video image rendered on a display screen, generated by an internal or external camera. For example, the user 101 may operate the MR device 115, which may launch a locally stored application to load the virtual aspects of the MR environment and sync with the real aspects through a sync server 166. The sync server 166 may receive real and virtual information captured by the MR device 115, real information from sensors 162, and virtual information generated by simulator 164. The MR device 115 may be configured to accept hand gesture inputs from user 101 for editing, tuning, updating, or interrupting the applications 120. The GUI device 105 may be implemented as a computer device such as a tablet, keypad, or touchscreen to enable entry of initial settings and parameters, and editing, tuning or updating of the applications 120 based on user 101 inputs. The GUI device 105 may work in tandem with the MR device 115 to program the applications 120.") … and
update the robot control system using the feedback data. (Paragraphs 0031-0032, "For the example simulation, the real object 214 may act as an obstacle for the virtual workpiece 222. For an initial programming of a spatial-related application, the MR device 115 may be used to observe the path of the virtual workpiece as it travels along a designed path 223 to a target 225. The user 101 may initially setup the application program using initial parameters to allow programming of the application. For example, the initial spatial parameters and constraints may be estimated with the knowledge that adjustments can be made in subsequent trials until a spatial tolerance threshold is met. Entry of the initial parameters may be input via the GUI device 105 or by an interface of the MR device 115. Adjustments to spatial parameters of the virtual robotic unit 231 and object 222 may be implemented using an interface tool displayed by the MR device 115. For example, spatial and orientation coordinates of the virtual gripper 224 may be set using a visual interface application running on the MR device 115. As the simulated operation is executed, one or more application modules 122 may receive inputs from the sensors, compute motion of virtual robotic unit 231 and/or gripper 224, monitor for obstacles based on additional sensor inputs, receive coordinates of obstacle object 214 from the vision system 212, and compute a new path if necessary. Should the virtual workpiece fail to follow the path 223 around the obstacle object 214, the user may interrupt the simulation, such as by using a hand gesture with the MR device 115, then modify the application using GUI 105 to make necessary adjustments.
The MR simulation may be rerun to test the adjusted application, and with additional iterations as necessary, until operation of the simulated robotic unit 231 is successful according to constraints, such as a spatial tolerance threshold. For example, the path 223 may be required to remain within a set of spatial boundaries to avoid collision with surrounding structures. As another example, the placement of object 222 may be constrained by a spatial range surrounding target location 225 based on coordination with subsequent tasks to be executed upon the object 222.")
However, Brown, in the same field of endeavor of robotics, teaches:
… or (ii) a score assigned by the human data collector to one or more of the human-operated robot tasks; (Column 78, Lines 1-20, "In some embodiments, the model updater 108 uses train the additional data generated by the other models in combination with the unstructured data of the unstructured service reports to configure the trained second model 116. The additional data generated by the other models can also or alternatively be used by the applications 120 in combination with an output of the second model 116 to select an action to perform. For example, the output of the trained second model 116 (e.g., a recommended action to perform) can be provided as an input to the other models to predict a consequence of the recommended action on energy consumption, occupant comfort, air quality, sustainability, infection risk, or any other variable state or condition predicted or modeled by the other models. The output of the other models can then be used by the system 100 to evaluate the consequences of the recommended action (e.g., score the recommended action relative to other recommended actions based on the consequences) and/or provide a user interface that informs the user of the consequences when presenting the recommended actions for user consideration.") …
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic teaching system and methods as taught by Wang with the ability to utilize minimanipulations which are combined into a sequence to achieve an operation with machine learning as taught by Rose as well as with the ability to utilize feedback data to update the control of the robot as taught by Bank and Brown. This would provide a system capable of efficiently learning minimanipulations as well as effective sequences of minimanipulations to achieve an operation from a human operator.
Regarding claim 14, where all the limitations of claim 13 are discussed above, Wang further teaches:
14. (Currently Amended) The system of claim [[1]]13, wherein the rendering of the human-operated robot tasks at the mixed reality device includes one or more of (1) input/output variables such as location of objects, (2) visual/voice instructions for display on the mixed reality device, and (3) checkable events or metrics as a result of tasks that can be examined/scored by the mixed reality device or by external sensors/systems. (Page 3, Paragraphs 2-4, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. The HoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig. 1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.
The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input. More importantly, as the HoloLens can track the user’s hand and fingers, then the manual actions (e.g., "pick" or "drop") are also detected. The manipulation of objects by a human hand is implemented via Microsoft MRTK SDK4, which enables the corresponding virtual object to stick to the human’s hand while the "picking/holding" gesture is applied, and release from the hand after a "drop" gesture is detected. Moreover, colliders are attached to the user’s fingers, enabling the teacher, for instance, to press the toaster lever, turn on the power button of the microwave and close the microwave door. Finally, the user behavior and related manipulated object information are sent to the back-end system via ROS (see Fig. 1, "INPUT"). The system can also display and animate a holographic virtual robot, which looks almost identical as the physical robot. The HoloLens receives the robot behavior data (see Fig. 1 "OUTPUT"), including the pose of the virtual robot and the speech commands, and visualizes / speaks them. Furthermore, the back-end triggers the display of the XAI cues which are shown in the AR environment. The next sections introduce three use cases for enhancing human-robot interaction via AR based on this system architecture.")
Regarding claim 15, where all the limitations of claim 1 are discussed above, Wang does not specifically teach generating a prompt from a combination of a template form the library and a user input. However, Rose, in the same field of endeavor of robotic control, teaches:
15. (Currently Amended) The system of claim [[1]]13, wherein the system is further configured for generating the prompt from a combination of a prompt template from the library of prompt templates (Paragraph 0037, "In accordance with the present robots, systems, control modules, computer program products, and methods, a catalog of reusable work primitives may be defined, identified, developed, or constructed such that any given work objective across multiple different work objectives may be completed by executing a corresponding workflow comprising a particular combination and/or permutation of reusable work primitives selected from the catalog of reusable work primitives. Once such a catalog of reusable work primitives has been established, one or more robot(s) may be trained to autonomously or automatically perform each individual reusable work primitive in the catalog of reusable work primitives without necessarily including the context of: i) a particular workflow of which the particular reusable work primitive being trained is a part, and/or ii) any other reusable work primitive that may, in a particular workflow, precede or succeed the particular reusable work primitive being trained. In this way, a semi-autonomous robot may be operative to autonomously or automatically perform each individual reusable work primitive in a catalog of reusable work primitives and only require instruction, direction, or guidance from another party (e.g., from an operator, user, or pilot) when it comes to deciding which reusable work primitive(s) to perform and/or in what order. In other words, an operator, user, pilot, or LLM module may provide a workflow consisting of reusable work primitives to a semi-autonomous robot system and the semi-autonomous robot system may autonomously or automatically execute the reusable work primitives according to the workflow to complete a work objective. For example, a semi-autonomous humanoid robot may be operative to autonomously look left when directed to look left, autonomously open its right end effector when directed to open its right end effector, and so on, without relying upon detailed low-level control of such functions by a third party. Such a semi-autonomous humanoid robot may autonomously complete a work objective once given instructions regarding a workflow detailing which reusable work primitives it must perform, and in what order, in order to complete the work objective. Furthermore, in accordance with the present robots, systems, methods, control modules and computer program products, a robot system may operate fully autonomously if it is trained or otherwise configured to (e.g. via consultation with an LLM module, which can be included in the robot system) analyze a work objective and independently define a corresponding workflow itself by deconstructing the work objective into a set of reusable work primitives from a library of reusable work primitives that the robot system is operative to autonomously perform.") and a user prompt. (Paragraph 00146, "Various implementations of the present systems, methods, control modules, and computer program products involve using NL expressions (descriptions) (e.g., via a NL prompt, which may be entered directly in text by a user or may be spoken vocally by a user and converted to text by an intervening voice-to-text system) to control functions and operations of a robot, where an LLM module may provide an interface between the NL expressions and the robot control system. This framework can be particularly advantageous when certain elements of the robot control architecture employ programming and/or instructions that can be expressed in NL. A suitable, but non-limiting, example of this is the aforementioned Instruction Set. For example, as mentioned earlier, a task plan output of an LLM module can be parsed (e.g., autonomously by the robot control system) by looking for a word match to Instruction Set commands, and the arguments of the Instruction Set can be found by string matching within the input NL prompt (e.g. by a text-string matching module as discussed earlier). In some implementations, a 1-1 map may be generated between the arguments used in the robot control system and NL variants, in order to increase the chance of the LLM module processing the text properly. For example, even though an object is represented in the robot control system (e.g., in a world model environment portion of the robot control system) as chess_pawn_54677, it may be referred to in the NL prompt as “chess pawn 1”. In this case, if the returned task plan contains the phrase “grasp chess pawn 1”, this may be matched to Instruction Set “grasp” and the object “chess pawn 1” so the phrase may be mapped to grasp(chess_pawn_54677). Such parsing and/or word matching (e.g. the text-string matching module) can be employed in any of the situations discussed herein where robot language is converted to natural language or vice-versa.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods as taught by Wang with the ability to control the system and generate instructions based on the library of operations as well as a user input as taught by Rose. This would allow a user to control the system using natural language inputs necessitating less training before they are able to operate the system effectively.
Regarding claim 16, where all the limitations of claim 13 are discussed above, Wang does not specifically teach using high level instructions which correspond to pre-defined low level libraries. However, Rose, in the same field of endeavor of robotics, teaches:
16. (Currently Amended) The system of claim [[1]]13, wherein the high-level instructions call predefined low-level libraries, (Paragraph 0037, "In accordance with the present robots, systems, control modules, computer program products, and methods, a catalog of reusable work primitives may be defined, identified, developed, or constructed such that any given work objective across multiple different work objectives may be completed by executing a corresponding workflow comprising a particular combination and/or permutation of reusable work primitives selected from the catalog of reusable work primitives. Once such a catalog of reusable work primitives has been established, one or more robot(s) may be trained to autonomously or automatically perform each individual reusable work primitive in the catalog of reusable work primitives without necessarily including the context of: i) a particular workflow of which the particular reusable work primitive being trained is a part, and/or ii) any other reusable work primitive that may, in a particular workflow, precede or succeed the particular reusable work primitive being trained. In this way, a semi-autonomous robot may be operative to autonomously or automatically perform each individual reusable work primitive in a catalog of reusable work primitives and only require instruction, direction, or guidance from another party (e.g., from an operator, user, or pilot) when it comes to deciding which reusable work primitive(s) to perform and/or in what order. In other words, an operator, user, pilot, or LLM module may provide a workflow consisting of reusable work primitives to a semi-autonomous robot system and the semi-autonomous robot system may autonomously or automatically execute the reusable work primitives according to the workflow to complete a work objective. For example, a semi-autonomous humanoid robot may be operative to autonomously look left when directed to look left, autonomously open its right end effector when directed to open its right end effector, and so on, without relying upon detailed low-level control of such functions by a third party. Such a semi-autonomous humanoid robot may autonomously complete a work objective once given instructions regarding a workflow detailing which reusable work primitives it must perform, and in what order, in order to complete the work objective. Furthermore, in accordance with the present robots, systems, methods, control modules and computer program products, a robot system may operate fully autonomously if it is trained or otherwise configured to (e.g. via consultation with an LLM module, which can be included in the robot system) analyze a work objective and independently define a corresponding workflow itself by deconstructing the work objective into a set of reusable work primitives from a library of reusable work primitives that the robot system is operative to autonomously perform.") the predefined low-level libraries including one or more of computer vision libraries, motion planning libraries, (Paragraph 0152, "The various implementations described herein include systems, methods, control modules, and computer program products for leveraging one or more LLM(s) in a robot control system, including for example establishing an NL interface between the LLM(s) and the robot control system and calling the LLM(s) to help autonomously instruct the robot what to do. Example applications of this approach include task planning, motion planning, reasoning about the robot's environment (e.g., “what could I do now?”), and so on. Such implementations are particularly well-suited in robot control systems for which at least some control parameters and/or instructions (e.g., the Instruction Set described previously) are amenable to being specified in NL. Thus, some implementations may include converting or translating robot control instructions and/or parameters into NL for communicating such with the LLM(s) via the NL interface.") and motion execution libraries. (Paragraph 0078, "FIG. 2 is a flowchart diagram which illustrates an exemplary method 200 of operation of a robot system. Method 200 in FIG. 2 is similar in at least some respects to the method 100 of FIG. 1. In general, method 200 in FIG. 2 describes detailed implementations by which method 100 in FIG. 1 can be achieved. Method 200 is a method of operation of a robot system (such as robot system 700 discussed with reference to FIG. 7). In general, throughout this specification and the appended claims, a method of operation of a robot system is a method in which at least some, if not all, of the various acts are performed by the robot system. For example, certain acts of a method of operation of a robot system may be performed by at least one processor or processing unit (hereafter “processor”) of the robot system communicatively coupled to a non-transitory processor-readable storage medium of the robot system (collectively a robot controller of the robot system) and, in some implementations, certain acts of a method of operation of a robot system may be performed by peripheral components of the robot system that are communicatively coupled to the at least one processor, such as one or more physically actuatable components (e.g., arms, legs, end effectors, grippers, hands), one or more sensors (e.g., optical sensors, audio sensors, tactile sensors, haptic sensors), mobility systems (e.g., wheels, legs), communications and networking hardware (e.g., receivers, transmitters, transceivers), and so on. The non-transitory processor-readable storage medium of the robot system may store data (including, e.g., at least one library of reusable work primitives and at least one library of associated percepts) and/or processor-executable instructions that, when executed by the at least one processor, cause the robot system to perform the method and/or cause the at least one processor to perform those acts of the method that are performed by the at least one processor. The robot system may communicate, via communications and networking hardware communicatively coupled to the robot system's at least one processor, with remote systems and/or remote non-transitory processor-readable storage media. Thus, unless the specific context requires otherwise, references to a robot system's non-transitory processor-readable storage medium, as well as data and/or processor-executable instructions stored in a non-transitory processor-readable storage medium, are not intended to be limiting as to the physical location of the non-transitory processor-readable storage medium in relation to the at least one processor of the robot system and the rest of the robot hardware. In other words, a robot system's non-transitory processor-readable storage medium may include non-transitory processor-readable storage media located on-board a robot body of the robot system and/or non-transitory processor-readable storage media located remotely from the robot body, unless the specific context requires otherwise. Further, a method of operation of a robot system such as method 200 (or any of the other methods discussed herein) can be implemented as a robot control module or computer program product. Such a control module or computer program product comprises processor-executable instructions or data that, when the control module or computer program product is stored on a non-transitory processor-readable storage medium of the robot system, and the control module or computer program product is executed by at least one processor of the robot system, the control module or computer program product (or the processor-executable instructions or data thereof) cause the robot system to perform acts of the method.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods of operating as taught by Wang with the low level operation libraries which may be used to accomplish high level tasks as taught by Rose. This allows the system to re-use may low level operations and efficiently program new tasks/actions in order for the system to complete high level operations.
Regarding claim 17, where all the limitations of claim 13 are discussed above, Wang does not specifically discuss the feedback overwriting the task or scoring the task. However, Bank, in the same field of endeavor of robotics, teaches:
17. (Currently Amended) The system of claim [[1]]13, wherein the feedback data comprises an overwrite of one or more of the human-operated robot tasks by the human data collector, (Paragraphs 0018-0019, "FIG. 1 shows an example of a system of computer based tools according to embodiments of this disclosure. In an embodiment, a user 101 may wear a mixed reality (MR) device 115, such as Microsoft HoloLens, while operating a graphical user interface (GUI) device 105 to train a skill engine 110 of a robotic device. The skill engine 110 may have multiple application programs 120 stored in a memory, each program for performing a particular skill of a skill set or task. For each application 120, one or more modules 122 may be developed based on learning assisted by MR tools, such as MR device 115 and MR system data 130. As the application modules 122 are programmed to learn tasks for robotic devices, task parameters may be stored as local skill data 112, or cloud based skill data 114. The MR device 115 may be a wearable viewing device that can display a digital representation of a simulated object on a display screen superimposed and aligned with real objects in a work environment. The MR device 115 may be configured to display the MR work environment as the user 101 monitors execution of the application programs 120. The real environment may be viewable in the MR device 115 as a direct view in a wearable headset (e.g., through a transparent or semi-transparent material) or by a video image rendered on a display screen, generated by an internal or external camera. For example, the user 101 may operate the MR device 115, which may launch a locally stored application to load the virtual aspects of the MR environment and sync with the real aspects through a sync server 166. The sync server 166 may receive real and virtual information captured by the MR device 115, real information from sensors 162, and virtual information generated by simulator 164. The MR device 115 may be configured to accept hand gesture inputs from user 101 for editing, tuning, updating, or interrupting the applications 120. The GUI device 105 may be implemented as a computer device such as a tablet, keypad, or touchscreen to enable entry of initial settings and parameters, and editing, tuning or updating of the applications 120 based on user 101 inputs. The GUI device 105 may work in tandem with the MR device 115 to program the applications 120.") …
However, Brown, in the same field of endeavor of robotics, teaches:
… and/or wherein the feedback data comprises a score for one or more of the human- operated robot tasks provided by the human data collector. (Column 78, Lines 1-20, "In some embodiments, the model updater 108 uses train the additional data generated by the other models in combination with the unstructured data of the unstructured service reports to configure the trained second model 116. The additional data generated by the other models can also or alternatively be used by the applications 120 in combination with an output of the second model 116 to select an action to perform. For example, the output of the trained second model 116 (e.g., a recommended action to perform) can be provided as an input to the other models to predict a consequence of the recommended action on energy consumption, occupant comfort, air quality, sustainability, infection risk, or any other variable state or condition predicted or modeled by the other models. The output of the other models can then be used by the system 100 to evaluate the consequences of the recommended action (e.g., score the recommended action relative to other recommended actions based on the consequences) and/or provide a user interface that informs the user of the consequences when presenting the recommended actions for user consideration.")
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the robotic control system and methods as taught by Wang with the ability to replace a task based on feedback data as taught by Bank and with the ability to utilize a score when determining how to perform a task as taught by Brown. This would allow the system to fine-tune the operation of the robot for more effective control.
Regarding claim 19, wherein all the limitations of claim 13 are discussed above, Wang further teaches:
19. (Currently Amended) The system of claim [[1]]13, further comprising a human-machine interface (Page 3, Paragraph 2, "Virtual objects, a virtual robot, and XAI cues (AR graphical elements) are displayed in the mixed-reality environment via the HoloLens3. TheHoloLens can scan the surroundings, build up 3D meshes of the environment objects and locate itself in the room, which enables it to stably overlay graphics in the environment considering occlusions with real objects. The shared environment, accessible to the user through the HoloLens (see Fig.1, right), includes multiple real and virtual objects. The object poses are continuously sent to the back-end, by the HoloLens for virtual objects and by a static camera for real objects (endowed with fiducial markers). Users can interact with the virtual objects in a similar way as with real ones; they can, for instance, pick up the bread, put it into a slot of the toaster, and push the button.") configured to be operated by the human data collector in performing the human-operated robot tasks, (Page 4, Paragraph 2, "Specifically, the user (wearing AR glasses) demonstrates a skill to the robot, e.g., by interacting with the (virtual) objects and giving language explanations (skill labels). The robot follows the teacher’s gaze or looks back at him to show its attention. The teacher is further continuously informed about the robot’s perception: Object and action labels (XAI cues) are popping up whenever the teacher gazes at some object or a manual action is recognized (cfr. Fig. 2 and [12]). The state of the environment, including the objects and agents, is recorded before and after the demonstrated skill by the episodic memory. These observations are parsed into predefined symbolic representations of the environment, consisting of logical predicates. Together with the skill label, these observations are used to learn the symbolic skill, capturing in which situations it can be applied and what is changed in the environment by executing it." This demonstrates that the system may learn each skill by observing a human performing the skill while wearing AR glasses.) wherein the human-machine interface comprises at least one input interface that receives user input (Page 3, Paragraph 4, "The HoloLens is also detecting the user’s behavior and communicates it to the back-end system. This includes the user’s head position, orientation and speech input.") and provides corresponding control commands to a microcontroller unit to control operation of one or more robotic components. (Page 3, Paragraph 1, “Finally, the action generation module (Fig. 1) receives action commands from other modules (e.g., where to look, what to grasp, etc.) and executes the corresponding motor behavior.”)
Regarding claim 20, where all the limitations of claim 19 are discussed above, Wang further teaches:
20. (Original) The system of claim 19, wherein the one or more robotic components comprise one or more of a robotic gripper, a robotic limb, and a robotic wheel. (Page 2, Paragraph 4, "The physical robot is a torso with two Kinova arms1 with 7 DOF each and a pan-tilt unit, part of the mobile platform "Johnny" [13].")
Allowable Subject Matter
Claim 12 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
Claim 18 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The Examiner has cited particular paragraphs or columns and line numbers in the referencesapplied to the claims above for the convenience of the Applicant. Although the specified citations arerepresentative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested of the Applicant in preparing responses, to fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. See MPEP 2141.02 [R-07.2015] VI. A prior art reference must be considered in its entirety, i.e., as a whole, including portions that would lead away from the claimed Invention. W.L. Gore & Associates, Inc. v. Garlock, Inc., 721 F.2d 1540, 220 USPQ 303 (Fed. Cir. 1983), cert, denied, 469 U.S. 851 (1984). See also MPEP §2123.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
/H.J.K./Examiner, Art Unit 3657
/JONATHAN L SAMPLE/Primary Examiner, Art Unit 3657