DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
Claim Interpretation
Claims 2 and 3 recite the phrase “a different skill included in the one or more skills”. The term “different” has been interpreted as referring to the previously recited “each …” in each case. In other words, there is only a requirement for “different” between the potential plurality of “each first trained machine learning model” and separately, between “each second trained machine learning model”. There is no requirement that the “different skill” of “each first trained machine learning model” and “each second trained machine learning model” be different, or that a given model be trained to perform only a single skill or similar.
Some of the claims (at least Claim 17) recites the term “task and motion planner”. Examiner notes that this term has an ordinary and customary meaning to one of ordinary skill in the art, particularly under the plain dictionary definitions of the terms involved in the overall noun phrase, and furthermore that while Applicant appears to believe there is a significant distinction between this term and machine learning model-based activities, even Applicant’s disclosure does not appear to clearly establish the distinction, let alone clearly provide any special definition of the term. MPEP 2111 relates. Consequently, the broadest meaning of the term appears to simply be anything which coordinates, plans, strategizes, arranges, or similar, tasks, activities, goals, etc., and motions, trajectories, etc. A machine learning model which performs these functions thus reads on this term even if Applicant appears to treat them as separate.
Examiner notes that Applicant has made frequent use of phrases such as “for” or “to”. In some cases, these terms may not amount to an actual positive recitation of a claim limitation and appear to only be provided for context of the actual limitations within the claim or otherwise indicate an intended result, use, or purpose of a preceding limitation. In the interest of compact prosecution Examiner has provided prior art where possible which Examiner believes teaches these recitations as positively recited limitations, regardless of if said interpretation are considered appropriate.
The adjectives “first”, “second”, “third”, etc. have been interpreted merely to indicate that two items might, but are not required to be, different. It is common claim construction practice to use these adjectives for ease of referring to items while maintaining a potential distinction which might be further claimed in further limitations and dependent claims. The adjectives, under the broadest reasonable interpretation do not require a particular chronological order, difference, etc. except where Applicant’s specification sets for such an explicit definition or the claim otherwise specifies as such.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2 – 3, 6 – 8, and 17 – 19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding Claims 2 – 3, the claims recite the phrase “is trained”. The claims are directed towards “a computer-implemented method”. This phrase does not clearly indicate if the training recited is a part of the process or method, or even the particular timing thereof, leaving the scope of the claim unclear.
In the interest of compact prosecution, and believing that it is Applicant’s intent that these recitations be positively recited steps of the method/process, the claims are instead interpreted as being written:
Claim 2:
The computer-implemented method of claim 1, wherein the method further comprises:
training each first trained machine learning model included in the one or more first trained machine learning models to control the robot to perform a different skill included in the one or more skills; and
training each second trained machine learning model included in the one or more second trained machine learning models to control the robot to perform a different skill included in the one or more skills.
Claim 3:
The computer-implemented method of claim 1, wherein the method further comprises:
training each first trained machine learning model included in the one or more first trained machine learning models to generate a base action to control the robot; and
training each second trained machine learning model included in the one or more second trained machine learning models to generate a delta action that modifies the base action generated by a corresponding first trained machine learning model included in the one or more first trained machine learning models.
Regarding Claim 6, the claim recites the limitation, “the one or more parameters of the untrained machine learning model”. There is insufficient antecedent basis for this limitation in the claim. Only Claim 5, which Claim 6 does not presently depend from, previously recites a “one or more parameters of the untrained machine learning model”.
In the interest of compact prosecution, the limitation is instead interpreted as reading:
“one or more parameters of the untrained machine learning model” (no “the”).
Regarding Claim 7, the claim depends from Claim 6 rejected above and inherits the deficiencies of said claim(s) as described above. Therefore, Claim 7 is rejected under the same logic presented above.
Regarding Claim 8, the claim recites the limitation “a Kullback-Leibler (KL) divergence term that penalizes…”. The claims are directed towards “a computer-implemented method”. The phrase “that penalizes” does not clearly indicate if penalizing is a part of the process or method, or even the particular timing thereof, leaving the scope of the claim unclear. Furthermore, if the phrase is intended as a functional limitation, it is not properly constructed in a manner by which it is a functional limitation, and furthermore based on the meaning of the term is unsupported as such. As evidenced by [0077] and Equation 4 of Applicant’s specification, the term itself does nothing on its own, and instead it is how it is used, or in other words its mathematical relationship in the mathematical equation, that dictates any further function it has. The plain meaning of a Kullback-Leibler divergence term, or relative entropy and I-divergence term, is simply a measure of how two probability distributions differ. It does not inherently penalize, reward, etc. anything, but must be further used in some manner.
In the interest of compact prosecution, and as the claim was constructed such that an actual step of penalizing was not claimed, the limitation is instead interpreted as reading “a Kullback-Leibler (KL) divergence term measuring …”.
Regarding Claims 17 – 19, the claims recite the limitation “the step of” or “the steps of”. There is insufficient antecedent basis for this limitation in the claim. It is unclear if there are omitted steps which should have been recited previously in the claims and are to be referred back to for further limitation.
In the interest of compact prosecution, the limitation is instead interpreted as reading:
“a step of” or “steps of”, respectively.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 4, 8, 11 – 12, 16, and 20 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Zhao et al. (US 20250026009 A1).
Regarding Claim 1, Zhao teaches:
A computer-implemented method for training one or more robot control models (See at least [0006] “[0006] The following disclosure describes a method and system for robot skill learning applicable to high precision assembly tasks employing a force or compliance controller”), the method comprising:
performing, based on one or more demonstration trajectories of a robot performing one or more skills associated with a task, one or more training operations to generate one or more first trained machine learning models for controlling the robot (See at least [0006] “A reinforcement learning controller is first pre-trained in an offline mode using human demonstration data, where several repetitions of the human demonstration are performed while collecting state and action data for each demonstration repetition. The demonstration data is used to pre-train a neural network in the reinforcement learning controller, with no interaction of the reinforcement learning controller with the compliance controller/robot system during pre-training” and Figure 9); and
one or more reinforcement learning operations using the one or more first trained machine learning models to generate one or more second trained machine learning models for controlling the robot (See at least [0006] “Following initial pre-training, the reinforcement learning controller is moved to online production where it is coupled to the compliance controller/robot system in a self-learning mode. During self-learning, the neural network-based reinforcement learning controller uses action, state and reward data to continue learning correlations between states and effective actions” and Figure 9).
Regarding Claim 4, Zhao teaches:
The computer-implemented method of claim 1, further comprising generating the one or more demonstration trajectories based on one or more user inputs to control the robot via one or more input/output devices (See at least [0065] “In a first step of the process (see circled number 1), a human operator 910 demonstrates the assembly operation in cooperation with the robot 100. One technique for demonstrating the operation involves putting the robot 100 in a teach mode, where the human 910 either manually grasps the robot gripper and workpiece and moves the workpiece into the installed position in the second workpiece (while the robot and controller monitor robot and force states), or the human 910 uses a teach pendant to provide commands to the robot 100 to complete the workpiece installation. Another technique for demonstrating the operation is teleoperation. In one form of teleoperation, the human 910 manipulates a duplicate copy of the workpiece which the robot 100 is grasping, and the human 910 moves the duplicate workpiece (which is instrumented and provides motion commands to the robot 100) while watching the robot 100, using the visual feedback from the robotic assembly operation and the human's own tactile feel to guide the successful completion of the assembly operation by the robot 100. In another form of teleoperation, the human 910 uses a joystick-type input device to provide motion instructions (translations and rotations) to the robot 100. These or other human demonstration techniques may be used”).
Regarding Claim 8, Zhao teaches:
The computer-implemented method of claim 1, wherein performing one or more reinforcement learning operations comprises updating one or more parameters of an untrained machine learning model based on a Kullback-Leibler (KL) divergence term that penalizes differences between one or more first actions generated using a first trained machine learning model included in the one or more first trained machine learning models and one or more second actions generated using the first trained machine learning model and the untrained machine learning model (See at least [0090] “In the optimization problem above, the objective function includes a loss function computed by subtracting a divergence term from the reward computed by the Q function. In the presently disclosed technique, the loss function includes a Kullback-Leibler divergence calculation (D.sub.KL) which is a measure of how one probability distribution [the control policy of the actor, π(.Math.|s)] is different from a second, reference probability distribution [the training dataset, π.sub.demo(.Math.|s)]. In Equation (5), λ is a weighting constant. By subtracting the divergence term from the Q function reward term, the objective function penalizes behavior of the actor control policy π which deviates from the training dataset”).
Regarding Claim 11, the claim is directed to effectively the same subject matter as Claim 1 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 1 above.
Regarding Claim 12, the claim is directed to effectively the same subject matter as Claim 4 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 4 above.
Regarding Claim 16, the claim is directed to effectively the same subject matter as Claim 8 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 8 above.
Regarding Claim 20, the claim is directed to effectively the same subject matter as Claim 1 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 1 above.
Claims 1 – 3, 5 – 7, 10 – 11, 13 – 15, and 17 – 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Alakuijala et al. (Alakuijala, Minttu, et al. "Residual reinforcement learning from demonstrations." arXiv preprint arXiv:2106.08050 (2021).).
Regarding Claim 1, Alakuijala teaches:
A computer-implemented method for training one or more robot control models, the method comprising:
performing, based on one or more demonstration trajectories of a robot performing one or more skills associated with a task, one or more training operations to generate one or more first trained machine learning models for controlling the robot (See at least Section III, “RRLfD is trained in two stages (Fig. 1). A convolutional neural network (CNN) is first trained to predict demonstrated control actions given a short history of depth or RGB images as well as current robot proprioceptive state (Section III A). Using demonstrated trajectories, the network learns to capture visual features relevant to solving the control task”); and
one or more reinforcement learning operations using the one or more first trained machine learning models to generate one or more second trained machine learning models for controlling the robot (See at least Section III, “To improve the policy with autonomous environment interaction, a light-weight policy network on top of the learned CNN features is trained with RL to additively correct the base policy’s actions (Section III-B)”).
Regarding Claim 2, Alakuijala teaches:
The computer-implemented method of claim 1, wherein each first trained machine learning model included in the one or more first trained machine learning models is trained to control the robot to perform a different skill included in the one or more skills, and wherein each second trained machine learning model included in the one or more second trained machine learning models is trained to control the robot to perform a different skill included in the one or more skills (Examiner notes that wherein the minimum number of first and second trained machine learning models is each one, this claim appears broad to the point, or close thereto, of being inherent to Claim 1. In the interest of compact prosecution, see at least Section IV(G), “We train residual DMPO on top of BC policies of various strengths to investigate the effect of base success rate on final performance and the required training time”).
Regarding Claim 3, Alakuijala teaches:
The computer-implemented method of claim 1, wherein each first trained machine learning model included in the one or more first trained machine learning models is trained to generate a base action to control the robot, and each second trained machine learning model included in the one or more second trained machine learning models is trained to generate a delta action that modifies the base action generated by a corresponding first trained machine learning model included in the one or more first trained machine learning models (See again recitations made with respect to Claim 1 above and Figure 1 including the caption which reads: “a) We propose a way to leverage demonstration data to learn a control policy as well as task-relevant visual features through behavioral cloning on image and proprioceptive inputs. b) The policy is then improved through reinforcement learning by a superimposed residual policy, based on the learned visual features, allowing data-efficient learning of control policies in image space from sparse rewards”).
Regarding Claim 5, Alakuijala teaches:
The computer-implemented method of claim 1, wherein performing one or more training operations to generate the one or more first trained machine learning models comprises (Examiner notes that these limitations appear to merely describe Behavior Cloning using a loss function):
generating, using an untrained machine learning model, one or more robot actions (See at least Section III(A), “We learn a base policy using behavioral cloning (BC). First, N demonstrations are gathered for a task of interest to create a dataset consisting of the demonstrated trajectories’ states Si = [si 1,...,si Ti ] and actions Ai = [ai 1,...,ai Ti ] … In addition to predicting at”);
generating, based on the one or more robot actions and using a simulator, one or more state-action pairs (See again above, in particular predicting at, Section I, “We evaluate our method in seven experimental settings from two simulated task suites and present the results in Section IV”);
calculating, based on the one or more state-action pairs and at least one trajectory included in the one or more demonstration trajectories, a loss (See at least Section III, “In addition to predicting at, we include actions 10, 20, and 30 time steps ahead and minimize loss over each of these targets, as done by [27]”); and
updating, based on the loss, one or more parameters of the untrained machine learning model (See again above).
Regarding Claim 6, Alakuijala teaches:
The computer-implemented method of claim 1, wherein performing one or more reinforcement learning operations to generate the one or more second trained machine learning models comprises:
generating, using a first trained machine learning model included in the one or more first trained machine learning models and an untrained machine learning model, one or more actions (See at least Section III(B), “πrφ takes as input the base action u, the robot’s proprioceptive state, and the features of the bottleneck layer (i.e., the final hidden layer before the fully connected output layer) of the BC policy network fθ. The bottleneck layer features provide the residual policy with visual information learned during BC. πrφ then outputs a corrective action a by predicting the parameters of a Gaussian distribution with diagonal covariance”);
generating, based on the one or more actions and using a simulator, one or more state-action pairs (See at least Section III(B), “At evaluation time, the policy is made deterministic: u + µ is executed in the environment instead of drawing a sample from the predicted Gaussian” and Section I, “We evaluate our method in seven experimental settings from two simulated task suites and present the results in Section IV”);
calculating, based on the one or more state-action pairs, a reward (See at least Section IV, “For Adroit environments, we use the carefully shaped multi-component rewards defined by [15]. We also add sparse task completion to each dense reward, appropriately scaled to prioritize task success, as we found this to increase success rates”); and
updating, based on the reward, the one or more parameters of the untrained machine learning model to generate a second trained machine learning model included in the one or more second trained machine learning models (See again above).
Regarding Claim 7, Alakuijala teaches:
The computer-implemented method of claim 6, wherein the reward comprises at least one of:
a sparse reward of one upon successful completion of a first skill included in the one or more skills and zero otherwise;
a dense reward based on progress toward one or more goals associated with the first skill (See at least Section IV, “We also add sparse task completion to each dense reward, appropriately scaled to prioritize task success, as we found this to increase success rates”); or
one or more penalty terms for movements greater than a threshold and collisions.
Regarding Claim 10, Alakuijala teaches:
The computer-implemented method of claim 1, further comprising:
receiving sensor data from one or more sensors (See at least Section IV(C), “For Adroit, proprioception includes joint positions, joint velocities, palm position and the fingertips’ tactile sensor readings as defined by [31]”);
generating, based on the sensor data and using the one or more first trained machine learning models and the one or more second trained machine learning models, one or more actions (See at least Section III(B), “the resulting control action u+a is executed in the environment” and Figure 1); and
causing the robot to perform one or more first movements based on the one or more actions (See again above).
Regarding Claim 11, the claim is directed to effectively the same subject matter as Claim 1 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 1 above.
Regarding Claim 13, the claim is directed to effectively the same subject matter as Claim 5 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 5 above.
Regarding Claim 14, Alakuijala teaches:
The one or more non-transitory computer-readable media of claim 13, wherein the loss comprises a difference between one or more first robot actions generated using the untrained machine learning model and one or more second robot actions included in the one or more demonstration trajectories (See at least Equation (2) in Section III. As necessary, see also associated description which clearly delineates caret or “hat” terms as those generated using the untrained base policy and those without as those from demonstrated trajectories).
Regarding Claim 15, the claim is directed to effectively the same subject matter as Claim 6 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 6 above.
Regarding Claim 17, Alakuijala teaches:
The one or more non-transitory computer-readable media of claim 11, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to perform the step of controlling the robot to perform one or more other skills associated with the task using a task and motion planner (TAMP) (Examiner notes that the claim does not further define or describe “other skills”, “task”, or effectively any of the terms involved in the claim. See various skills specifically described as part of robotic assembly which are distinct, different, and therefore clearly “other” from each other. For example, grasping or picking as in [0024] “A gripper 202 grasps a part 210 ”, “a hole search, where the gripper 202 moves the part 210 back and forth” ([0024]), “a phase search, where the gripper 302 finely adjusts the rotational position of the part 310 about its pivot axis” [0025]), and insertion of a part in general (See [0024] and [0025])).
Regarding Claim 18, the claim is directed to effectively the same subject matter as Claim 10 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 10 above.
Regarding Claim 19, Alakuijala teaches:
The one or more non-transitory computer-readable media of claim 18, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to perform the step of causing the robot to perform one or more second movements based on one or more motions generated using a motion planning technique (See at least Section IV(A), “An inverse kinematics submodule allows control in task space by converting end-effector velocities to joint velocities at 10Hz”).
Regarding Claim 20, the claim is directed to effectively the same subject matter as Claim 1 with respect to the application of prior art. The claim is therefore rejected under the same logic as Claim 1 above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. further in view of Mandlekar et al. (Mandlekar, Ajay, et al. "Human-in-the-Loop Task and Motion Planning for Imitation Learning." arXiv preprint arXiv:2310.16014 (2023).).
Regarding Claim 9, Zhao teaches:
The computer-implemented method of claim 1,
Zhao does not teach, but Mandlekar teaches:
wherein performing one or more reinforcement learning operations comprises:
scheduling a plurality of workers based on a sampling strategy for sampling workers to execute and a queue that stores indications of workers that require scheduling (See at least Section 4, “Since the TAMP system only requires human assistance in small parts of an episode, a human operator has the opportunity to manage multiple robots and data collection sessions simultaneously. To this end, we propose a novel queue ing system (Fig. 3) allowing each operator to interact with a fleet of robots. We implement this by using several (Nrobot) robot processes, a single human process, and a queue”); and
executing the plurality of workers based on the scheduling to generate the one or more second trained machine learning models (See at least Section 4, “Each robot process runs asynchronously, and spends its time in 1 of 3 modes — (1) being controlled by the TAMP system, (2) waiting for human control, or (3) being controlled by the human”).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to incorporate the scaling data collection for learning system which includes multiple robots and a queuing system taught by Mandlekar in the system of Zhao with a reasonable expectation of success. Zhao discloses a “co-training” system during which only intermittent need for a human operator exists. The use of Mandlekar’s system would enable a human operator to readily maintain and operate as needed a plurality of “workers” (robots).
Furthermore, the system of Zhao may be further modified such that it operates as disclosed in Mandlekar, using a non-machine learning model-based system for portions of task completion, and only using and performing training of a machine learning model for particular portions of task completion.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Nair et al. (Nair, Ashvin, et al. "Overcoming Exploration in Reinforcement Learning with Demonstrations." arXiv preprint arXiv:1709.10089 (2017).) which discloses combining imitation and reinforcement learning techniques.
Hester et al. (Hester, Todd, et al. "Deep Q-learning from Demonstrations." arXiv preprint arXiv:1704.03732 (2017).) which discloses first pre-training via demonstrations, and then further training via reinforcement learning.
Le et al. (Le, Hoang M., et al. "Hierarchical Imitation and Reinforcement Learning." arXiv preprint arXiv:1803.00590 (2018).) which discloses combining IL and RL for hierarchical control.
Gupta et al. (Gupta, Abhishek, et al. "Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning." arXiv preprint arXiv:1910.11956 (2019).) which discloses first performing relay imitation learning and then relay reinforcement fine-tuning wherein both high and low level policies exist and each undergo both processes.
Staroverov et al. (Skrynnik, Alexey, et al. "Forgetful experience replay in hierarchical reinforcement learning from demonstrations." arXiv preprint arXiv:2006.09939 (2020).) which discloses what the title indicates; hierarchical reinforcement learning using demonstrations further utilizing a “forgetting mechanism to deal with a catastrophic drop in productivity due to poor-quality expert trajectories and extend the approach of using demonstrations to partially observed and hierarchical environments” (Section 1).
Riedmiller et al. (US 20250196347 A1) which discloses the following highly relevant paragraph to both combining imitation and reinforcement learning, and the nature of objective functions used for training: [0078] Training using such an imitation learning technique in general involves training the executor neural network system 104 such that actions selected according to the action selection output 116 match the actions of the demonstrating agent. Any imitation learning technique can be used, e.g., behavioral cloning, inverse reinforcement learning, or Generative Adversarial Imitation Learning (arXiv:1606.03476, Ho et al.). The training may be performed offline, i.e., based solely on the demonstration data, and/or online, e.g., to fine tune the actions using reinforcement learning. For example the executor neural network system 104 can be trained to optimize an objective function that depends on a difference between a distribution of actions 118 selected according to the action selection output 116 and a distribution of actions defined by the actions of the demonstrating agent.
Bennice et al. (US 20240253215 A1) which discloses training a first policy and then further training that same policy to create an updated policy and furthermore discloses simulation, evaluation for training, and hierarchical command structure.
Cella et al. (US 20220197306 A1) which discloses pre-training through simulation in a digatil twin system a reinforcement learning agent which is then further trained using reinforcement learning (See e.g. [1595])
Kalakrishnan et al. (US 20220105624 A1) which discloses training a meta-learning model using imitation learning and reinforcement learning.
Tunyasuvunakool et al. (US 20190126472 A1) which disclose a hybrid of both imitation and reinforcement learning wherein the hybridization is in the nature of the rewards and the use of demonstration data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW C GAMMON whose telephone number is (571)272-4919. The examiner can normally be reached M - F 10:00 - 6:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ADAM MOTT can be reached on (571) 270-5376. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW C GAMMON/Examiner, Art Unit 3657
/ADAM R MOTT/Supervisory Patent Examiner, Art Unit 3657