Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent claims 1, 18, and 20 are directed toward method. Therefore, each of the independent claims 1, 18, and 20 along with the corresponding dependent claims 2-17 and 19 are directed to a statutory category of invention under Step 1.
Under Step 2A, Prong 1, the claims are analyzed to determine whether one or more of the claims recites subject matter that falls within one of the following groups of abstract idea: (1) mental processes, (2) certain methods of organizing human activity, and/or (3) mathematical concepts. In this case, the independent claims 1, 18, and 20 are directed to an abstract idea without significantly more. Specifically, the claims, under their broadest reasonable interpretation cover certain mental processes/organizing human activity/mathematical concepts. The language of independent claim * is used for illustration:
assembling, as a first input prompt, representations of an observed initial state of a robot and a goal state of the robot;
processing the first input prompt using a goal-conditioned trajectory model to generate first output indicative of a sequence of predicted states to be reached by the robot between the observed initial and goal states;
assembling, as a second input prompt, representations of the sequence of predicted states; and
processing the second input prompt using an action prediction model to generate second output indicative of a sequence of predicted actions to be performed by the robot to reach the sequence of predicted states.
Independent claim 1 presents the mental processes of assembling and processing which can be directed towards the action of detailing and narrowing an idea with more detail.
As explained above, independent claim 1 recites at least one abstract idea. The other independent claims 18 and 20, which are of similar scope to claim 1, likewise recite at least one abstract Idea under Step 2A, Prong 1.
Under Step 2A, Prong 2, the claims are analyzed to determine whether the claim, as a whole, integrates the abstract idea into a practical application. As noted in the 2019 PEG, it must be determined whether any additional elements in the claim beyond the abstract idea integrate the exception into a practical application in a manner that imposes a meaningful limit on the judicial exception. The courts have indicated that additional elements such as merely using a computer to implement an abstract idea, adding insignificant extra solutions activity, or generally linking use of a judicial exception to a particular technological environment or a field of use do not integrate a judicial exception into a “practical application”; see at least MPEP 2106.04(d).
This judicial exception is not integrated into a practical application. In particular, the claim only recites the elements of — using a processor to perform the listed steps. The processor in all steps is recited at a high-level of generality (i.e., as a generic processor performing a generic computer functions) such that it amounts to no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Therefore, taken alone, the additional elements do not integrate the abstract idea into a practical application. Furthermore, looking at the additional limitations as an ordered combination or as a whole, the limitations add nothing significant that is not already present when looking at the elements taken individually. Because the additional elements, do not integrate the abstract idea into a practical application by imposing meaningful limits on practicing the abstract idea, independent claims 1, 11, and 20 are directed to an abstract idea.
Under Step 2B, the claims do not include any additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application in Step 2A, Prong Two, the additional element of limiting the use of the idea to one particular environment employs generic computer functions to execute an abstract idea and, therefore, does not add significantly more. Limiting the use of the abstract idea to a particular environment or field of use cannot provide an inventive concept.
Dependent claims 2-10, 14, 16-17, and 19 are also rejected as they do not provide significantly more.
Dependent claims 2-10 and 19 provides more detail regarding the training policy without providing any advantageous detail of how it improves the current technology.
Dependent claims 14 and 16-17 provides another step similar to the process of independent claim 1 but does not provide details on the control.
Examiner encourages Applicant to set an interview to discuss potential amendments for overcoming the above rejections under 35 U.S.C. § 101.
Examiner notes that dependent claims 11, 13, and 15 with dependent claim 12, comprise matter that would assist in overcoming the current rejection under 35 U.S.C. 101. Specifically, these claims appear to pertain to control of the robot based off of the training, where generating the control signal and outputting the signal would provide significantly more than the mental processes. This limitation amounts to a practical application and is not rejectable under 35 U.S.C. 101. Furthermore, examiner also directs applicant to the memo regarding changes to the MPEP in light of Ex Parte Desjardins within MPEP 2106.04(d) and MPEP 2106.05(a).
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 3, 7-18, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US20230311335A1 (Hausman).
Regarding claim 1, Hausman discloses a method implemented using one or more processors and comprising:
assembling, as a first input prompt, representations of an observed initial state of a robot and a goal state of the robot;
See at least Fig. 2A, where the figure illustrates the process flow of outputting the robot control. Hausman discloses assembling, as a first input prompt (LLM output 206A), representations of an observed initial state (see at least [0062]) and a goal state of the robot ([0063], “The explanation 204A can be an explanation generated based on processing a prior LLM prompt, also based on the FF NL input, in a prior pass and utilizing the LLM 150. For example, the prior LLM prompt could be “explain how you would bring me a snack from the table”, and the explanation 204A can be generated based on the highest probability decoding from the prior LLM output. For instance, the explanation 204A can be “I would find the table, then find a snack on the table, then bring it to you”. The explanation 204A can be prepended to the LLM prompt 205A, replace term(s) of the FF NL input 105 in the LLM prompt 205A, or otherwise incorporated into the LLM prompt 205A.”)
processing the first input prompt using a goal-conditioned trajectory model to generate first output indicative of a sequence of predicted states to be reached by the robot between the observed initial and goal states;
See at least Fig. 2A, where the figure discloses further processing the first input prompt (LLM prompt 205A) with a goal-conditioned trajectory model (task-grounding engine 132) to generate first output indicative of a sequence of predicted states to be reached by the robot between the observed initial and goal states (see at least [0065]).
assembling, as a second input prompt, representations of the sequence of predicted states; and
See Fig. 2A of Hausman, where the representations of the sequence of predicted states (as overall measures 212A in [0071]) are assembled.
processing the second input prompt using an action prediction model to generate second output indicative of a sequence of predicted actions to be performed by the robot to reach the sequence of predicted states.
See Fig. 2A of Hausman, where the figure shows processing the second input prompt using an action prediction model (selection engine 136) to generate output indicative of a sequence of predicted actions to be performed by the robot to reach the sequence of predicted states.
Regarding claim 3, with all of the limitations of claim 1, the method further comprises:
wherein the goal-conditioned trajectory model is trained using reference sequences of previously observed robot states, wherein for each reference sequence of previously observed robot states, each previously observed robot state is annotated based on a final previously observed robot state of the reference sequence.
Hausman discloses the goal-conditioned trajectory model is trained using reference sequences of previously observed robot states, wherein for each reference sequence of previously observed robot states, each previously observed robot state is annotated based on a final previously observed robot state of the reference sequence ([0065], “For example, “go to the table” can be descriptive of a “navigate to table” skill that the robot can perform by utilizing a trained navigation policy with a navigation target of “table” (or of a location corresponding to a “table”).”, where the reference sequence “go to the table” is an annotation of “navigate to table”).
Regarding claim 7, with all of the limitations of claim 1, the method further comprises:
wherein the representation of the goal state of the robot comprises a set of reference points on the robot that collectively represent the goal state.
Hausman discloses the representation of the goal state of the robot (see at least [0063]) comprise a set of reference points on the robot that collectively represent the goal state (see at least [0062-0063] regarding the goal state and [0081] regarding the grasping operations)
Regarding claim 8, with all of the limitations of claim 1, the method further comprises:
wherein the goal state of the robot comprises a representation of the robot itself in the goal state.
See at least [0063], where the goal state comprises a representation of the robot itself in the goal state (“For instance, the explanation 204A can be “I would find the table, then find a snack on the table, then bring it to you”, where the robot is represented as giving the snack to the human)
Regarding claim 9, with all of the limitations of claim 1, the method further comprises:
wherein the representation of the goal state of the robot comprises
a representation of a human or
different robot
in a pose that corresponds to the goal state of the robot.
See at least [0063], where the “you” within the representation of the goal state is the representation of a human in a pose that corresponds to the goal state of the robot.
Regarding claim 10, with all of the limitations of claim 1, the method further comprises:
prior to assembling the first input prompt, assembling, as a third input prompt, representations of the observed initial state of a robot and a task to be performed by the robot;
See Fig. 2A of Hausman, where Hausman discloses, prior to assembling the first input prompt, assembling, as a third input prompt, representations of the observed initial state of a robot (scene descriptors 202A) and a task to be performed by the robot (FF NL input 105).
processing the third input prompt using a generative model to generate the representation of the goal state.
See Fig. 2A, where the third input prompt is provided further into LLM engine 130 to generate the representation of the goal state.
Regarding claim 11, with all of the limitations of claim 1, the method further comprises:
generating a control signal for controlling one or more robots based on one or more of the sequence of predicted actions.
See Fig. 2A, where robotic skill indication 213A is used to generate a control signal for controlling one or more robots based on one or more of the sequence of predicted actions ([0071]).
Regarding claim 12, with all of the limitations of claim 1, the method further comprises:
operating one or more robots based on one or more of the sequence of predicted actions.
See Fig. 2A, where Hausman discloses operating one or more robots based on one or more of the sequence of predict actions ([0071], “In response, the implementation engine 136 controls the robot 110 based on the selected robotic skill A.”).
Regarding claim 13, with all of the limitations of claim 1, the method further comprises: generating a first control signal for controlling the robot based on a subset of one or more predicted actions selected from the sequence of predicted actions.
See the citation of claim 11.
Regarding claim 14, with all of the limitations of claim 13, the method further comprises:
assembling, as a third input prompt, a representation of a subsequent observed state of the robot upon the robot being controlled based on the first control signal;
See the citations to claim 1 regarding “assembling, as a first input …” and Fig. 2B of Hausman, where Fig. 2B continues from the subsequent action and state of Fig. 2A.
processing the third input prompt using the goal-conditioned trajectory model to generate third output indicative of a subsequent sequence of predicted states to be reached by the robot after the subsequent observed state;
See the citations to claim 1 regarding “processing the first input …” and Fig. 2B of Hausman, where Fig. 2B continues from the subsequent action and state of Fig. 2A.
assembling, as a fourth input prompt, representations of the subsequent sequence of predicted states to be reached by the robot after the subsequent observed state; and
See the citations to claim 1 regarding “assembling, as a second input …” and Fig. 2B of Hausman, where Fig. 2B continues from the subsequent action and state of Fig. 2A.
processing the fourth input prompt using the action prediction model to generate fourth output indicative of a subsequent sequence of predicted actions to be performed by the robot to reach the subsequent sequence of predicted states.
See the citations to claim 1 regarding “processing the second input …”and Fig. 2B of Hausman, where Fig. 2B continues from the subsequent action and state of Fig. 2A.
Regarding claim 15, with all of the limitations of claim 14, the method further comprises:
generating a control signal for controlling the robot based on a new subset of predicted actions selected from the subsequent sequence of predicted actions.
See the citation of claim 11 in view of the sequence of Fig. 2A and Fig 2B.
Regarding claim 16, with all of the limitations of claim 14, the method further comprises:
wherein the fourth input prompt is further assembled to include a representation of the subsequent observed state of the robot upon being controlled based on the first control signal.
See Fig. 3 of Hausman. The figure discloses an exemplary method of implementing robotic skills. Hausman discloses that the fourth input prompt (control based off of indication 213B) is further assembled to include a presentation of the subsequent observed state of the robot upon being controlled based on the first control signal (indication 213A) by processing the current environmental state data (box 358A, current environmental state data), where the current environmental state data would be the subsequent observed state of the robot upon being controlled based on the first control signal.
Regarding claim 17, with all of the limitations of claim 1, the method further comprises:
wherein the second input prompt is further assembled to include a representation of the observed initial state of a robot.
See Fig. 3 of Hausman, where the second input prompt is further assembled include a representation of the observed initial state of a robot (box 358A, current robot state data).
Regarding claim 18, Hausman discloses a method implemented using one or more processors and comprising:
assembling, as a first input prompt, representations of an observed initial state of a robot and a goal state of the robot;
See the citation of claim 1 regarding “assembling, as a first input …”.
processing the first input prompt using a goal-conditioned trajectory model to generate first output indicative of an interpolated sequence of predicted states to be reached by the robot between the observed initial and goal states;
See the citation of claim 1 regarding “processing the first input …”, where the sequence of predicted states was interpolated by the language model reasoning out the steps to accomplish the goal from an input prompt (see at least [0063]).
assembling, as a second input prompt, representations of the interpolated sequence of predicted states; and
See the citation of claim 1 regarding “assembling, as a second input …”.
processing the second input prompt using an action prediction model to generate second output indicative of an interpolated sequence of predicted actions to be performed by the robot to reach the interpolated sequence of predicted states.
See the citation of claim 1 regarding “processing, the second input …”.
Regarding claim 20, Hausman discloses a method implemented using one or more processors and comprising:
collecting a reference sequence of previously observed states of a kinematic entity;
Hausman discloses training (see at least [0029]) based off of a collection of a reference sequence of previously observed states of a kinematic entity (see at least [0031]).
annotating each previously observed state of the kinematic entity based on a selected previously observed state of the kinematic entity in the reference sequence;
Hausman discloses annotating each previously observed state of the kinematic entity based on a selected previously observed state of the kinematic entity in the reference sequence (see at least [0032], where the skills have a textual label indicating the skill was successfully completed from the state).
assembling, as an input prompt, representations of an observed initial state of the kinematic entity and the selected previously observed state of the kinematic entity in the reference sequence;
See [0035] of Hausman, where the process described iteratively appends new instructions to the assembly of representations of an observed initial states (“Practically, this can be viewed as a dialog between a user and a robot, in which a user provides the high level-instruction (e.g., “How would you bring me a coke can?”) and the language model responds with an explicit sequence (“I would: 1.,”, e.g., “I would: 1. find a coke can, 2. pick up the coke can, 3. bring it to you”).”)
processing the input prompt using a trajectory model to generate output indicative of an interpolated sequence of predicted states to be reached by the kinematic entity between the observed initial state and the selected previously observed state of the kinematic entity in the reference sequence;
See [0035] of Hausman cited above, where the input prompt (“How would you bring me a coke can?”) was processed using a trajectory model to generate output indicative of an interpolated sequence of predicted states (the steps to bring the user a coke can) to be reached by the kinematic entity (a robot) between the observed initial state and the selected previously observed state of the kinematic entity in the reference sequence.
comparing the sequence of predicted states to the reference sequence of previously observed states; and
See algorithm 1 of Hausman and [0032] of Hausman. Hausman discloses an algorithm that takes an instruction (A, state, skills, and descriptions to improve the nth skill (where pi_n is the skill that maximizes C) based off of subsequent descriptions (see lines 4-9 of algorithm 1) and skills (see line 11 of algorithm 1). [0032] of Hausman disclose that the probabilities involved are calculated to indicate success or failure of approaching the goal state with 1 or 0, where the probabilities represent the comparison of predicted states to observed states.
training the trajectory model based on the comparing.
See algorithm 1 of Hausman.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2, 4, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over US20230311335A1 (Hausman) further in view of “Planning with Diffusion for Flexible Behavior Synthesis” (Janner) from the IDS.
Regarding claim 2, with all of the limitations of claim 1, the method further comprises:
training a diffusion policy for controlling one or more robots based on the sequences of predicted states and predicted actions.
While Hausman discloses that the robotic control policy may be based on a reinforcement learning and/or imitation learning (Hausman, [0119]), Hausman does not explicitly disclose the use of a diffusion policy for controlling one or more robots based on the sequence of predicted states and predicted actions.
From a similar field of endeavor, Janner discloses a diffusion-based policy approach to control a robot to improve on the flaws of model-based reinforcement learning (Janner, Abstract). Janner specifically discloses training a diffusion policy for controlling one or more robots based on the sequences of predicted states and predicted actions (see at least Fig. 2)
One of ordinary skill in the art would find it obvious, prior to the applicant’s effective filing date, to implement the diffusion-based model of Janner to the system of Hausman as Janner results improvements over model-based reinforcement learning such as temporal compositionality producing out-of-distribution trajectories from available subsequences (Janner, Conclusion).
Regarding claim 4, with all of the limitations of claim 1, the method further comprises:
wherein the goal-conditioned trajectory model comprises a diffusion model.
While Hausman discloses that the goal-condition trajectory model comprises skill descriptions that are descriptive of a skill performed by the robot (Hausman, [0065]), where the skills performed by the robot controlled by reinforcement learning (Hausman, [0119]), Hausman does not explicitly disclose that the goal-conditioned trajectory model comprises a diffusion model.
In view of the rationale of claim 2, one of ordinary skill in the art would find it obvious, prior to the applicant’s effective filing date, to implement the diffusion-based model of Janner to the system of Hausman for at least the improvements mentioned in claim 2.
Regarding claim 19, with all of the limitations of claim 18, the method further comprises:
training a diffusion policy for controlling one or more robots based on the interpolated sequences of predicted states and predicted actions.
While Hausman discloses that the robotic control policy may be based on a reinforcement learning and/or imitation learning (Hausman, [0119]) to control interpolated sequences of predicted states and predicted actions (Hausman, [0063]), Hausman does not explicitly disclose the use of a diffusion policy for controlling one or more robots based on the interpolated sequence of predicted states and predicted actions.
From a similar field of endeavor, Janner discloses a diffusion-based policy approach to control a robot to improve on the flaws of model-based reinforcement learning (Janner, Abstract). Janner specifically discloses training a diffusion policy for controlling one or more robots based on the sequences of predicted states and predicted actions (see at least Fig. 2)
One of ordinary skill in the art would find it obvious, prior to the applicant’s effective filing date, to implement the diffusion-based model of Janner to the system of Hausman as Janner results improvements over model-based reinforcement learning such as temporal compositionality producing out-of-distribution trajectories from available subsequences (Janner, Conclusion).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over US20230311335A1 (Hausman).
Regarding claim 5, with all of the limitations of claim 1, the method further comprises:
wherein the goal-conditioned trajectory model comprises a flow model.
While Hausman does not explicitly disclose that the goal-conditioned trajectory model comprises a flow model, one of ordinary skill in the art would find it obvious that the goal-conditioned trajectory model comprises a flow model as Hausman does disclose that the steps performed in Fig. 2A iterates flows to each successive action, where Fig. 2B shows the next step of the goal provided to the robot. As a result of the goal-conditioned trajectory model proceeding to the next step, the “task grounding measures 208A” of Fig. 2A and “task grounding measures 208B” of Fig. 2B have different task-grounding measures as a result of the inputted prompt and skill descriptions.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over US20230311335A1 (Hausman) further in view of “Asymmetric Actor Critic for Image-Based Robot Learning” (Pinto).
Regarding claim 6, with all of the limitations of claim 1, the method further comprises:
wherein the representation of the goal state of the robot comprises one or more synthetic digital images depicting the goal state.
While Hausman does not explicitly disclose that the representation of the goal state of the robot comprises one or more synthetic images depicting the goal state, from a similar field of endeavor, Pinto discloses a simulator training reinforcement learning system that trains actions based on rendered simulated images (see at least Fig. 3 and page 4, paragraph 1).
One of ordinary skill in the art would find it obvious, prior to the applicant’s effective filing date, to combine the synthetic digital image training system of Pinto to the system of Hausman as Pinto discloses advantages such as observability the full state of the robot and its environment (page 1, Introduction, paragraph 3), where the simulation-based training provides a superior training environment that is fully controllable.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAEWOOK JUNG whose telephone number is (571)272-5470. The examiner can normally be reached Monday - Friday, 9:00 AM - 5:00 PM..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Wade Miles can be reached on (571) 270-7777. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.J./Examiner, Art Unit 3656
/WADE MILES/Supervisory Patent Examiner, Art Unit 3656