Prosecution Insights
Last updated: October 01, 2026
Application No. 18/061,965

SKILL COMPOSITION AND SKILL TRAINING METHOD FOR THE DESIGN OF AUTONOMOUS SYSTEMS

Non-Final OA §103
Filed
Dec 05, 2022
Priority
Aug 12, 2022 — provisional 63/371,308
Examiner
WOOD, BLAKE ANDREW
Art Unit
3658
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
Microsoft Technology Licensing, LLC
OA Round
5 (Non-Final)
71%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 71% — above average
71%
Career Allowance Rate
119 granted / 167 resolved
+19.3% vs TC avg
Moderate +14% lift
Without
With
+14.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
19 currently pending
Career history
193
Total Applications
across all art units

Statute-Specific Performance

§101
9.2%
-30.8% vs TC avg
§103
51.4%
+11.4% vs TC avg
§102
20.2%
-19.8% vs TC avg
§112
16.8%
-23.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 167 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks In view of the Appeal Brief filed on 05 May 2026, PROSECUTION IS HEREBY REOPENED. New grounds of rejection are set forth below. To avoid abandonment of the application, appellant must exercise one of the following two options: (1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or, (2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid. A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing below: /THOMAS E WORDEN/ Supervisory Patent Examiner, Art Unit 3658 Response to Amendment Claims 1-4, 6, 9-11, 13, 15, and 16 have been newly amended. No claims have been newly added nor canceled. Claims 1-20 remain pending application. Response to Arguments Applicant’s arguments, see Appeal Brief, filed 05 May 2026, with respect to the rejection(s) of claim(s) 1-20 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, new grounds of rejection have been made below. See the 35 U.S.C. § 103 rejections of claims 1-20 below for further details. Claim Objections Claims 4 and 8 are objected to because of the following informalities: Regarding claim 4, Applicant claims: “wherein the first subtask comprises a grasp sub-task, wherein the subsequent subtask comprises a lift sub-task, and wherein the termination signal indicates that the grasp sub-task is complete and the lift sub-task may begin.” The examiner recommends amending this limitation to recite: “wherein the first subtask comprises a grasp sub-task, wherein the subsequent sub-task comprises a lift sub-task, and wherein the termination signal indicates that the grasp sub-task is complete and the lift sub-task may begin.” Regarding claim 8, Applicant claims: “wherein the different criteria comprise different speeds, angles, locations of a robotic arm controlled by the first machine learning model.” The examiner recommends amending this limitation to recite: “wherein the different criteria comprise different speeds, angles, and/or locations of a robotic arm controlled by the first machine learning model…” or the like. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-3, 9-13, 15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri (US 20210362333 A1), hereafter Kolluri, in view of Kalashnikov (US 20210237266 A1), hereafter Kalashnikov. Regarding claim 1, Kolluri discloses a method for training a first machine learning model to perform a sub-task of a plurality of sub-tasks that collectively perform a long horizon task (0064, A skill template can also include, for each demonstration subtask, the software modules that are required to tune the demonstration subtask using local demonstration data. Each demonstration subtask can rely on a different type of machine learning model and can use different techniques for tuning. For example, a movement demonstration subtask can rely heavily on camera images of the local workcell environment in order to find a particular task goal. Thus, the tuning procedure for the movement demonstration subtask can more heavily tune the machine learning models to recognize features in camera images captured in the local demonstration data. In contrast, an insertion demonstration subtask can rely heavily on force feedback data for sensing the edges of a connection socket and using appropriately gentle forces to insert a connector into the socket. Therefore, the tuning procedure for the insertion demonstration subtask can more heavily tune the machine learning models that deal with force perception and corresponding feedback. In other words, even when the underlying models for subtasks in a skill template are the same, each subtask can have its own respective tuning procedures for incorporating local demonstration data in differing ways.), the method comprising: Providing an input to the first machine learning model (0127, The first subtask in the skill template 400 is a movement subtask 410. The movement subtask 410 is designed to locate a wire in the workcell, which requires moving a robot from an initial position to an expected location of the wire, e.g., as placed by a previous robot in an assembly line. Moving from one location to the next is typically not very dependent on the local characteristics of the robot, and thus the metadata of the movement subtask 410 specifies that the subtask is a nondemonstration subtask. The metadata of the movement subtask 410 also specifies that the camera stream is needed in order to locate the wire.); Receiving a termination signal inferred by the first machine learning model from the input, wherein the termination signal determines whether the first sub-task is complete (0127, The first subtask in the skill template 400 is a movement subtask 410. The movement subtask 410 is designed to locate a wire in the workcell, which requires moving a robot from an initial position to an expected location of the wire, e.g., as placed by a previous robot in an assembly line. Moving from one location to the next is typically not very dependent on the local characteristics of the robot, and thus the metadata of the movement subtask 410 specifies that the subtask is a nondemonstration subtask. The metadata of the movement subtask 410 also specifies that the camera stream is needed in order to locate the wire. 0128, The movement subtask 410 also specifies an “acquired wire visual” transition condition 405, which indicates when the robot should transition to the next subtask in the skill template.); Attempting to perform a subsequent sub-task of the plurality of sub-tasks with a second machine learning model (0129, The next subtask in the skill template 400 is a grasping subtask 420. The grasping subtask 420 is designed to grasp a wire in the workcell. This subtask is highly dependent on the characteristics of the wire and the characteristics of the robot, particularly the tool being used to grasp the wire. Therefore, the grasping subtask 420 is specified as a demonstration subtask that requires refinement with local demonstration data. The grasping subtask 420 is thus also associated with a base policy id that identifies a previously generated base control policy for grasping wires generally.); and Determining whether the subsequent sub-task was successfully performed (0131-0133, The grasping subtask 420 also includes three transition conditions. The first transition condition, the “lost wire visual” transition condition 415, is triggered when the robot loses visual contact with the wire. This can happen, for example, when the wire is unexpectedly moved in the workcell, e.g., by a human or another robot. In that case, the robot transitions back to the movement subtask 410. The second transition condition of the grasping subtask 420, the “grasp failure” transition condition 425, is triggered when the robot attempts to grasp the wire but fails. In that scenario, the robot can simply loop back and try the grasping subtask 420 again. The third transition condition of the grasping subtask 420, the “grasp success” transition condition 435, is triggered when the robot attempts to grasp the wire and succeeds.). Kolluri fails to explicitly disclose, however: Training the termination signal of the first machine learning model with a termination signal reward based on whether the subsequent sub-task was successfully performed. Kalashnikov, however, in an analogous field of endeavor, does teach: Training the termination signal of the first machine learning model with a termination signal reward based on whether the subsequent sub-task was successfully performed (0077-0078, At block 410, the system determines a reward based on the system executing the action using the current robot policy model. In some implementations, when the action is a non-terminal action, the reward can be, for example, “0” reward—or a small penalty (e.g., −0.05) to encourage faster robotic task completion. In some implementations, when the action is a terminal action, the reward can be a “0” if the robotic task was successful and a “1” if the robotic task was not successful. For example, for a grasping task the reward can be “1” if an object was successfully grasped, and a “0” otherwise. The system can utilize various techniques to determine whether a grasp or other robotic task is successful. For example, for a grasp, at termination of an episode the gripper can be moved out of the view of the camera and a first image captured when it is out of the view. Then the gripper can be returned to its prior position and “opened” (if closed at the end of the episode) to thereby drop any grasped object, and a second image captured. The first image and the second image can be compared, using background subtraction and/or other techniques, to determine whether the gripper was grasping an object (e.g., the object would be present in the second image, but not the first)—and an appropriate award assigned to the last time step. In some implementations, the height of the gripper and/or other metric(s) can also optionally be considered. For example, a grasp may only be considered if the height of the gripper is above a certain threshold.). Kolluri and Kalashnikov are analogous because they are in a similar field of endeavor, e.g., machine learning systems for robot control. It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the present invention, with a reasonable expectation of success, to have included the training of a first control policy based on the results of a subsequent control policy of Kalashnikov in order to provide a means of further refining the long-horizon task. The motivation to combine is to ensure that the robot system is able to properly perform a plurality of sub-tasks while working towards a unified goal. Claims 9 and 15 are similar in scope to claim 1, and are similarly rejected. Regarding claim 2, the combination of Kolluri and Kalashnikov teaches the method of claim 1, and Kolluri further teaches wherein the first sub-task and the subsequent sub-task are performed sequentially by an autonomous system performing the long-horizon task (0127, The first subtask in the skill template 400 is a movement subtask 410. The movement subtask 410 is designed to locate a wire in the workcell, which requires moving a robot from an initial position to an expected location of the wire, e.g., as placed by a previous robot in an assembly line. Moving from one location to the next is typically not very dependent on the local characteristics of the robot, and thus the metadata of the movement subtask 410 specifies that the subtask is a nondemonstration subtask. 0129, The next subtask in the skill template 400 is a grasping subtask 420. The grasping subtask 420 is designed to grasp a wire in the workcell. This subtask is highly dependent on the characteristics of the wire and the characteristics of the robot, particularly the tool being used to grasp the wire. Therefore, the grasping subtask 420 is specified as a demonstration subtask that requires refinement with local demonstration data. The grasping subtask 420 is thus also associated with a base policy id that identifies a previously generated base control policy for grasping wires generally.). Regarding claim 3, the combination of Kolluri and Kalashnikov teaches the method of claim 1, and Kolluri further teaches wherein the trained first machine learning model controls a robotic device performing the first sub-task (0127, The first subtask in the skill template 400 is a movement subtask 410. The movement subtask 410 is designed to locate a wire in the workcell, which requires moving a robot from an initial position to an expected location of the wire, e.g., as placed by a previous robot in an assembly line. Moving from one location to the next is typically not very dependent on the local characteristics of the robot, and thus the metadata of the movement subtask 410 specifies that the subtask is a nondemonstration subtask. The metadata of the movement subtask 410 also specifies that the camera stream is needed in order to locate the wire.) and wherein the termination signal of the first machine learning model indicates that the robotics device has completed the first sub-task when the termination signal is true (0128, The movement subtask 410 also specifies an “acquired wire visual” transition condition 405, which indicates when the robot should transition to the next subtask in the skill template.). Regarding claim 10, the combination of Kolluri and Kalashnikov teaches the non-transitory computer-readable storage medium of claim 9, and Kolluri further teaches wherein the first sub-task and the subsequent sub-task are simulated in a simulator (0241, FIG. 14 is a flowchart of an example process for training a skill template using a simulated working environment. Although using local demonstration data is fast relative to traditional methods, collecting the data still requires a nontrivial time investment in order to collect the amount of data required to make the models very precise. The process can be further sped up by parallel training on simulated workcell data. These techniques are particularly suited for demonstration subtasks that rely on perceptual streams of information. The process can be performed by a computer system having one or more computers in one or more locations, e.g., the system 100 of FIG. 1. The process will be described as being performed by a system of one or more computers.). Regarding claim 11, the combination on Kolluri and Kalashnikov teaches the non-transitory computer-readable storage medium of claim 10, and Kalashnikov further teaches wherein the first sub-task comprises lifting the object, and wherein the subsequent sub-task is determined not to be successfully performed with the object slips from the robotic arm while the subsequent sub-task is performed (0052, Each of the rewards can be assigned in view of a reward function that can assign a positive reward (e.g., “1”) or a negative reward (e.g., “0”) at the last time step of an episode of performing a task. The last time step is one where a termination action occurred, as a result of an action determined based on the policy model indicating termination, or based on a maximum number of time steps occurring. Various self-supervision techniques can be utilized to assign the reward. For example, for a grasping task, at the end of an episode the gripper can be moved out of the view of the camera and a first image captured when it is out of the view. Then the gripper can be returned to its prior position and “opened” (if closed at the end of the episode) to thereby drop any grasped object, and a second image captured. The first image and the second image can be compared, using background subtraction and/or other techniques, to determine whether the gripper was grasping an object (e.g., the object would be present in the second image, but not the first)—and an appropriate award assigned to the last time step. In some implementations, the reward function can assign a small penalty (e.g., −0.05) for all time steps where the termination action is not taken. The small penalty can encourage the robot to perform the task quickly.). Kolluri and Kalashnikov are analogous because they are in a similar field of endeavor, e.g., model-based robotic control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the subsequent task failure determination of Kalashnikov in order to provide further means of determining the success of a plurality of subtasks. The motivation to combine is to ensure that subtask status is able to be properly acquired. Regarding claim 12, the combination of Kolluri and Kalashnikov teaches the non-transitory computer-readable storage medium of claim 11, and Kalashnikov further teaches wherein negative reinforcement is provided to the termination condition of the first machine learning model in response to determining that the subsequent sub-task is not successfully performed (0052, Each of the rewards can be assigned in view of a reward function that can assign a positive reward (e.g., “1”) or a negative reward (e.g., “0”) at the last time step of an episode of performing a task. The last time step is one where a termination action occurred, as a result of an action determined based on the policy model indicating termination, or based on a maximum number of time steps occurring. Various self-supervision techniques can be utilized to assign the reward. For example, for a grasping task, at the end of an episode the gripper can be moved out of the view of the camera and a first image captured when it is out of the view. Then the gripper can be returned to its prior position and “opened” (if closed at the end of the episode) to thereby drop any grasped object, and a second image captured. The first image and the second image can be compared, using background subtraction and/or other techniques, to determine whether the gripper was grasping an object (e.g., the object would be present in the second image, but not the first)—and an appropriate award assigned to the last time step. In some implementations, the reward function can assign a small penalty (e.g., −0.05) for all time steps where the termination action is not taken. The small penalty can encourage the robot to perform the task quickly.). Kolluri and Kalashnikov are analogous because they are in an analogous field of endeavor, e.g., model-based robotic control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the negative reinforcement of Kalashnikov in order to provide a means of further training the machine learning models. The motivation to combine is to ensure that the machine learning models are properly trained to perform a plurality of tasks. Regarding claim 13, the combination of Kolluri and Kalashnikov teaches the non-transitory computer-readable storage medium of claim 10, and Kalashnikov further teaches wherein the first sub-task comprises grasping an object with a robotic arm, wherein the subsequent subtask comprises lifting the object, and wherein the subsequent sub-task is determined to be successfully performed with the robotic arm continues to grasp the object throughout the subsequent sub-task (0052, Each of the rewards can be assigned in view of a reward function that can assign a positive reward (e.g., “1”) or a negative reward (e.g., “0”) at the last time step of an episode of performing a task. The last time step is one where a termination action occurred, as a result of an action determined based on the policy model indicating termination, or based on a maximum number of time steps occurring. Various self-supervision techniques can be utilized to assign the reward. For example, for a grasping task, at the end of an episode the gripper can be moved out of the view of the camera and a first image captured when it is out of the view. Then the gripper can be returned to its prior position and “opened” (if closed at the end of the episode) to thereby drop any grasped object, and a second image captured. The first image and the second image can be compared, using background subtraction and/or other techniques, to determine whether the gripper was grasping an object (e.g., the object would be present in the second image, but not the first)—and an appropriate award assigned to the last time step. In some implementations, the reward function can assign a small penalty (e.g., −0.05) for all time steps where the termination action is not taken. The small penalty can encourage the robot to perform the task quickly.). Kolluri and Kalashnikov are analogous because they are in a similar field of endeavor, e.g., model-based robotics control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the grasping and lifting of Kalashnikov in order to provide a means of picking an object from a surface. The motivation to combine is to ensure that the robotic system is able to properly interact with its environment. Regarding claim 17, the combination of Kolluri and Kalashnikov teaches the computing device of claim 15, and Kolluri further teaches wherein the input to the first machine learning model comprises a video stream, an audio signal, force sensor data, or position sensor data (0087, Having multiple, separately tunable control policies can be advantageous in a sensor-rich environment that, for example, can use data from multiple sensors having different update rates. For example, the different control policies 210a-n can execute at different update rates, which allows the system to incorporate both simple and more sophisticated control algorithms into the same system. For example, one control policy can focus on robot commands using current force data, which can be updated at a much faster rate than image data. Meanwhile, another control policy can focus on robot commands using current image data, which may require more sophisticated image recognition algorithms, which may have nondeterministic run times. The result is a system that can both rapidly adapt to force data but also adapt to image data without slowing down its adaptations to force data. During training, the subtask hyperparameters can identify separate training procedures for each of the separately tunable control policies 210-an.). Regarding claim 18, the combination of Kolluri and Kalashnikov teaches the computing device of claim 15, and Kolluri further teaches wherein an output of the first machine learning model comprises a joint angle usable to control a robotic computing device in addition to the termination signal (0079, The sensors 260 also include one or more robot state sensors that generate robot state data streams 204 that represent physical characteristics of the robot or a component of the robot. For example, the robot state data streams 204 can represent force, torque, angles, positions, velocities, and accelerations, of the robot or respective components of the robot, to name just a few examples. Each of the robot state data streams 204 can be processed by a respective deep neural network 230a-m.). Regarding claim 19, the combination of Kolluri and Kalashnikov teaches the computing device of claim 15, and Kalashnikov further teaches wherein the termination signal reward provides positive reinforcement to the termination signal of the first machine learning model when the subsequent sub-task completes successfully (0052, Each of the rewards can be assigned in view of a reward function that can assign a positive reward (e.g., “1”) or a negative reward (e.g., “0”) at the last time step of an episode of performing a task. The last time step is one where a termination action occurred, as a result of an action determined based on the policy model indicating termination, or based on a maximum number of time steps occurring. Various self-supervision techniques can be utilized to assign the reward. For example, for a grasping task, at the end of an episode the gripper can be moved out of the view of the camera and a first image captured when it is out of the view. Then the gripper can be returned to its prior position and “opened” (if closed at the end of the episode) to thereby drop any grasped object, and a second image captured. The first image and the second image can be compared, using background subtraction and/or other techniques, to determine whether the gripper was grasping an object (e.g., the object would be present in the second image, but not the first)—and an appropriate award assigned to the last time step. In some implementations, the reward function can assign a small penalty (e.g., −0.05) for all time steps where the termination action is not taken. The small penalty can encourage the robot to perform the task quickly.). Kolluri and Kalashnikov are analogous because they are in a similar field of endeavor, e.g., model-based robotics control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the positive reinforcement of Kalashnikov in order to provide further means of training the overall control system. The motivation to combine is to positively reward the control system for successfully completing a requested sub-task. Regarding claim 20, the combination of Kolluri and Kalashnikov teaches the computing device of claim 15, and Kolluri further teaches wherein an input to the first machine learning model includes a state of a robotic computing device controlled by the first machine learning model, a state of an object being manipulated by the robotic computing device, force sensor data, or a state of an environment surrounding the robotic computing device and the object, and wherein an output of the first machine learning model includes a joint angle or a hand position of the robotic computing device in addition to the termination signal (0079, The sensors 260 also include one or more robot state sensors that generate robot state data streams 204 that represent physical characteristics of the robot or a component of the robot. For example, the robot state data streams 204 can represent force, torque, angles, positions, velocities, and accelerations, of the robot or respective components of the robot, to name just a few examples. Each of the robot state data streams 204 can be processed by a respective deep neural network 230a-m.). Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri in view of Kalashnikov, and further in view of Campos (US 20180293498 A1), hereafter Campos. Regarding claim 4, the combination of Kolluri and Kalashnikov teaches the method of claim 1, but fails to explicitly teach wherein the first subtask comprises a grasp sub-task, wherein the subsequent subtask comprises a lift sub-task, and wherein the termination signal indicates that the grasp sub-task is complete and the lift sub-task may begin. Campos, however, in an analogous field of endeavor, does teach wherein the first subtask comprises a grasp sub-task, wherein the subsequent subtask comprises a lift sub-task (0161, FIG. 4F illustrates a block diagram of an embodiment of the AI engine that solves the example “Grasp and Stack” complex task 400F with concept network reinforcement learning. In this example, the AI engine solves the example complex task of Grasping a rectangular prism and precisely Stacking it on top of a cube. The AI engine initially broke the overall task down into four concepts: 1) Reaching the working area (staging 1), 2) Grasping the prism, 3) Moving to the second working area (staging 2), and 4) Stacking the prism on top of the cube. The Grasp concept can further be decomposed into an Orient the hand concept and Lift concept. Thus, to simplify the learning problem by using a single policy for each individual sub-task the concept of Grasp, the AI engine broke the Grasping concept into two more concepts: Orienting the hand around the prism in preparation for grasping, as well as clasping the prism to Lift the prism, for a total of five actor concepts in the concept network. Three of these concepts—Orienting, Lifting, and Stacking used the TRPO algorithm to train, while the Reach concept (Staging-1) and the Moving concept to the working area (Staging-2) were handled with inverse kinematics), and wherein the termination signal indicates that the grasp sub-task is complete and the lift sub-task may begin (0181, A goal of a termination condition for a policy can be defined based on the state space. Orient for the orient concept—an episode would end early if the hand moved too far from the prism, if the prism tipped more than 15 degrees, or the goal was achieved by aligning the opposed fingers with the prism while the pinch point was 1.5 cm above the prism.). Kolluri, Kalashnikov, and Campos are analogous because they are in an analogous field of endeavor, e.g., robot control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the grasp and lift sub-tasks of Campos in order to provide a means of further manipulating an object. The motivation to combine is to ensure that a requested task is able to be decomposed into easily accomplishable sub-tasks. Regarding claim 14, the combination of Kolluri and Kalashnikov teaches the non-transitory computer-readable storage medium of claim 10, but fails to explicitly teach wherein the termination signal of the first machine learning model is trained multiple times with different simulated coefficients of friction or different simulated degrees of deformity of an object grasped by a robotic arm. Campos, however, in an analogous field of endeavor, does teach wherein the termination signal of the first machine learning model is trained multiple times with different simulated coefficients of friction or different simulated degrees of deformity of an object grasped by a robotic arm (0099, For example, in each iteration, the machine learning software makes a decision about the next set of parameters for friction compensation and the next set of parameters for motion. These decisions are made by the modules of the AI engine. It is anticipated that the many iterations involved will require that the optimization process be capable of running autonomously. To achieve this, a software layer is utilized to enable the AI engine software to configure the control with the next iteration's parameterization for friction compensation and its parameterization of the axis motion. The goal for deep reinforcement learning in this example user's case is to explore the potential of the AI engine to improve upon manual or current automatic calibration. Specifically, to eliminate the human expert and make the AI the expert in selecting parameter values, equal or improve upon the degree of precision, reduce the number of iterations of tests needed, and hence the overall time needed to complete the circularity test. The AI engine is coded to understand machine dynamics and develop initial model of machine's dynamics. Development of a simulation model is included based on initial measurements. The AI engine's ability to set friction and backlash compensation parameters occurs within the simulation model. After the initial model training occurs, then the training of the simulation model of friction and backlash compensation is extended with the advice from any experts in that field. The training of the simulation model moves from the simulation model world, after the deep reinforcement learning is complete, to a real world environment. The training of the concept takes the learning from the real machine and uses it to improve and tune the simulation model.). Kolluri, Kalashnikov, and Campos are analogous because they are in a similar field of endeavor, e.g., robotic control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the friction analysis of Campos in order to provide a means of compensating for an object’s friction. The motivation to combine is to allow the system to compensate for differing coefficients of friction of different objects. Claims 5-8 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri in view of Kalashnikov, and further in view of Bodnar (US 11571809 B1), hereafter Bodnar. Regarding claim 5, the combination of Kolluri and Kalashnikov teaches the method of claim 1, but fails to explicitly teach wherein attempting to perform the subsequent sub-task while training the first machine learning model comprises performing an operation similar to but different than the subsequent sub-task. Bodnar, however, in an analogous field of endeavor, does teach wherein attempting to perform the subsequent sub-task while training the first machine learning model comprises performing an operation similar to but different than the subsequent sub-task (Col. 11, Line 50 - Col. 12, Line 2, In some implementations, the desired measure of risk-seeking may be determined based on a measure of experience of the robot in performing the robot task or other robotic tasks that share one or more attributes with the robotic task. In some such implementations, features of the assigned robotic task may be extracted and used to determine an embedding in a latent space that includes embeddings of other robotic tasks. Similar tasks may be identified, for instance, based on their respective Euclidian distances from the assigned robotic task in latent space. Generally speaking, risk-averse robotic operation may be desirable in various scenarios, such as in highly dynamic environments (e.g., environments with large measures of entropy) in which there are multiple changing/moving objects that may increase the probability of failure beyond what is determined by critic network 152, or in uncertain environments in which the environment is not well known. In the latter case, the robot may be operated cautiously (e.g., slowly, with more impedance, etc.) to decrease the likelihood of failure caused by an unknown object.). Kolluri, Kalashnikov, and Bodnar are analogous because they are in a similar field of endeavor, e.g., machine learning model-based robot control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the similar operation of Bodnar in order to provide a means of allowing the system to further adapt to an operating environment. The motivation to combine is to ensure that the robot is operated in a safe and effective manner (see at least Col. 12, Lines 3-17 of Bodnar). Regarding claim 6, the combination of Kolluri, Kalashnikov, and Bodnar teaches the method of claim 5, and Kalashnikov further teaches wherein the first sub-task comprises a grasp sub-task that grasps an object laying on a surface (0077-0078, For example, for a grasping task the reward can be “1” if an object was successfully grasped, and a “0” otherwise. The system can utilize various techniques to determine whether a grasp or other robotic task is successful.). Kolluri, Kalashnikov, and Bodnar are analogous because they are in a similar field of endeavor, e.g., machine learning model-based robot control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the grasp sub-task of Kalashnikov in order to provide a means of picking an object from a surface. The motivation to combine is to ensure that the robotic system is able to properly interact with its environment. The combination of Kolluri, Kalashnikov, and Bodnar fails to explicitly teach, however, wherein the operation similar to the subsequent sub-task comprises dragging the object along the surface. The examiner asserts, however, that it would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, to have made the subsequent sub-task a driving operation, because to do so would have been an obvious matter of design choice. The examiner asserts that there is both a design need (i.e., to manipulate an object) as well as a finite number of predictable solutions (i.e., moving the object through the air, moving the object along a surface, or merely grasping-and-holding an object). Regarding claim 7, the combination of Kolluri and Kalashnikov teaches the method of claim 1, but fails to explicitly teach wherein attempting to perform the subsequent sub-task while training the first machine learning model comprises performing the subsequent sub-task multiple times with different criteria. Bodnar, however, in an analogous field of endeavor, does teach wherein attempting to perform the subsequent sub-task while training the first machine learning model comprises performing the subsequent sub-task multiple times with different criteria (Col. 12, Lines 3-17, Risk-averse robot operation may also be desirable, for instance, where the vision data applied by the critic network lacks important information. When processing vision data depicting a glass champagne flute, a critic network used to operate a robot may output a value distribution that indicates a low probability of failure. However, it might have been the case that, during training, the critic network was only trained on training data that included vision data of plastic champagne flutes, which are considerably less fragile than glass. Yet, another object recognition routine implemented by the robot may detect that the champagne flute is made of glass, and therefore is very fragile. Accordingly, the robot may be operated cautiously, in spite of the high confidence of critic network of success, to reduce the likelihood of the glass champagne flute being damaged or destroyed.). Kolluri, Kalashnikov, and Bodnar are analogous because they are in a similar field of endeavor, e.g., machine learning model-based robot control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the differing criteria of Bodnar in order to provide a means of allowing the system to further adapt to an operating environment. The motivation to combine is to ensure that the robot is operated in a safe and effective manner (see at least Col. 12, Lines 3-17 of Bodnar). Regarding claim 8, the combination of Kolluri, Kalashnikov, and Bodnar teaches the method of claim 7, and Kolluri further teaches wherein the different criteria comprise different speeds, angles, and/or locations of a robotic arm controlled by the first machine learning model (0272, At a next-highest level, the software stack can include Cartesian position controllers and Cartesian selection controllers. A Cartesian position controller can receive as input goals in Cartesian space and use inverse kinematics solvers to compute an output in joint position space. The Cartesian selection controller can then enforce limit policies on the results computed by the Cartesian position controllers before passing the computed results in joint position space to a joint position controller in the next lowest level of the stack. For example, a Cartesian position controller can be given three separate goal states in Cartesian coordinates x, y, and z. For some degrees, the goal state could be a position, while for other degrees, the goal state could be a desired velocity.). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Kolluri in view of Kalashnikov, and further in view of Berman (US 20160107396 A1), hereafter Berman. Regarding claim 16, the combination of Kolluri and Kalashnikov teaches the computing device of claim 15, but fails to explicitly teach wherein the first sub-task comprises creating a mold as part of a manufacturing process and the subsequent sub-task comprises installing a part in the mold. Berman, however, in an analogous field of endeavor, does teach wherein the first sub-task comprises creating a mold as part of a manufacturing process and the subsequent sub-task comprises installing a part in the mold (0059, In another embodiment, a computer implemented process for creating an architectural component, comprising the steps of: (a) providing a three-dimensional work piece; (b) supporting on a movable support table the three-dimensional work piece comprising an expanded polystyrene foam, movement of the movable support table being responsive to commands from a computer processor; (c) using a first end effector at a distal end of a multi-task robotic arm and in cooperation with the support table, sequentially removing material from the work piece to form a three-dimensional architectural mold, the end effector being responsive to commands of the computer processor; (d) using a second end effector at the distal end of the multi-task robotic arm and in cooperation with the support table, sequentially applying layers of concrete onto the three-dimensional architectural mold; and (e) using a third end effector at the distal end of the multi-task robotic arm and in cooperation with the support table, machining, milling, and/or drilling through a surface of the three-dimensional architectural mold to form an architectural component.). Kolluri, Kalashnikov, and Berman are analogous because they are in a similar field of endeavor, e.g., robot control systems. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the present invention, with a reasonable expectation of success, to have included the mold creation and machining of Berman in order to provide a means of expanding the capabilities of the task decomposition system. The motivation to combine is to allow the robotic system to perform a more varied array of tasks and sub-tasks. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Popov et al. ("Data-efficient Deep Reinforcement Learning for Dexterous Manipulation") teaches methods for decomposing a larger “task” into a plurality of smaller “subtasks”. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BLAKE A WOOD whose telephone number is (571)272-6830. The examiner can normally be reached M-F, 8:00 AM to 4:30 PM Eastern. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Thomas Worden can be reached at (571) 272-4876. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BLAKE A WOOD/ Primary Examiner, Art Unit 3658
Read full office action

Prosecution Timeline

Show 11 earlier events
Nov 07, 2025
Examiner Interview Summary
Nov 07, 2025
Applicant Interview (Telephonic)
Nov 25, 2025
Response Filed
Dec 29, 2025
Final Rejection mailed — §103
Feb 05, 2026
Notice of Allowance
May 05, 2026
Response after Non-Final Action
May 22, 2026
Response after Non-Final Action
Sep 16, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12741380
METHODS AND APPARATUSES FOR DROPPED OBJECT DETECTION
3y 10m to grant Granted Sep 22, 2026
Patent 12743104
METHOD FOR INTERACTIVELY PROVIDING WAYPOINTS TO A MOBILE ROBOT FOR USE IN THE MARKING OF A GEOMETRIC FIGURE ON A GROUND SURFACE
3y 2m to grant Granted Sep 22, 2026
Patent 12741383
SYSTEMS AND METHODS FOR MULTI-SECTIONAL SHOW ROBOT
2y 9m to grant Granted Sep 22, 2026
Patent 12743094
ROBOT AND ROBOT CONTROL METHOD
2y 5m to grant Granted Sep 22, 2026
Patent 12742906
WIND CONDITION LEARNING DEVICE, WIND CONDITION PREDICTING DEVICE, AND DRONE SYSTEM
2y 3m to grant Granted Sep 22, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
71%
Grant Probability
86%
With Interview (+14.4%)
2y 9m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 167 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month