DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Status of Claims
The amendments filed 05/26/2026 and 05/27/2026 have been entered. Claims 1-9 have been amended. Claims 1-9 are now pending. Applicant’s amendments to the claim language have overcome each and every 35 USC 101 rejection set forth in the Non-Final Rejection mailed 02/25/2026. Applicant’s arguments to the claim language have overcome the previously noted 35 USC 112(f) interpretation set forth in the Non-Final Rejection mailed 02/25/2026.
Priority
Acknowledgement is made of applicant’s claim for foreign priority under 35 USC 119 (a)-(d) to application JP2024-056408 filed 03/29/2024. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. As such, the effective filing date of the application is 03/29/2024.
Response to Arguments
Applicant’s arguments with regards to claims 1-9 have been fully considered, but are moot because amendments to the claim language have necessitated new grounds of rejection as set forth below.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-9 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri et al. (US 20210362330 A1), hereinafter Kolluri. in view of Horowitz et al. (US 20230191608 A1), hereinafter Horowitz.
Regarding claim 1, Kolluri discloses:
A robot motion learning device comprising:
a plurality of first learning models that receive motion information at a certain time and convert the motion information into motion features, and also receive external information at the time and convert the external information into external features, for a plurality of types of robots (see at least [0058]: “A skill template can also include, for each demonstration subtask, the software modules that are required to tune the demonstration subtask using local demonstration data. Each demonstration subtask can rely on a different type of machine learning model and can use different techniques for tuning… Therefore, the tuning procedure for the insertion demonstration subtask can more heavily tune the machine learning models that deal with force perception and corresponding feedback. In other words, even when the underlying models for subtasks in a skill template are the same, each subtask can have its own respective tuning procedures for incorporating local demonstration data in differing ways.”)
a shared learning model that converts the motion features and external features output by the first learning models into predicted motion features at a next time that are common to the plurality of types of robots (see at least [0044]: “In some implementations, the training system generates the base control policy from generalized training data 165. While the local demonstration data 115 collected by the online execution system 110 is typically specific to one particular robot or one particular robot model, the generalized training data 165 can in contrast be generated from one or more other robots, which need not be the same model, located at the same site, or built by the same manufacturer. For example, the generalized training data 165 can be generated offsite from tens or hundreds or thousands of different robots having different characteristics and being different models. In addition, the generalized training data 165 does not even need to be generated from physical robots. For example, the generalized training data can include data generated from simulations of physical robots.”)
a plurality of second learning models that convert the predicted motion features at the next time into predicted motion information, for the plurality of types of robots (see at least [0047]: “Adapting a base control policy using local demonstration data has the highly desirable effect that it is relatively fast compared to generating the base control policy, e.g., either by collecting system demonstration data or by training using the generalized training data 165. For example, the size of the generalized training data 165 for a particular task tends to be orders of magnitude larger than the local demonstration data 115 and thus training the base control policy is expected to take much longer than adapting it for a particular robot. For example, training the base control policy can require vast computing resources, in some instances, a datacenter having hundreds or thousands of machines working for days or weeks to train the base control policy from generalized training data. In contrast, adapting the base control policy using local demonstration data 115 can take just a few hours.”)
and a processor configured to use teaching data related to motion of each of the robots to train either the first learning model and the second learning model related to the robot or the shared learning model (see at least Figure 1, item 110, the Online Execution System.)
Kolluri does not explicitly disclose, but Horowitz, in an analogous field of endeavor teaches:
wherein the time is updated for each control cycle of the robot, and has a configuration in which input from the robot, the first learning model, the shared learning model, the second learning model, and output to the robot are connected in that order, so that the robot operates the prediction motion information as a control input (see at least [0053]: “ In some embodiments, model training logic 202 is configured to cross train machine learning models that correspond to different domains but share some overlapping materials. Examples of domains include single stream recyclables (SSR), construction and demolition (C&D), organics, and e-waste. In some embodiments, model training logic 202 is configured to train each of multiple machine learning models' known/sorted material of a corresponding domain. Then, input data can be fed into a particular machine learning model associated with a first domain to obtain that model's output of the first domain-specific classifications on that input data. Next, the input data with the labels of the first domain-specific classifications is then used as training data to train another, related machine learning model associated with a second domain that shares some overlapping materials with the first domain. In this way, the machine learning model corresponding to the second domain adds the recognition parameters associated with the first domain, without requiring the longer lead time and greater number of iterations necessary when starting from scratch and using only human annotation for a core machine learning model associated with the first domain.”)
It would have been prima facie obvious for one of ordinary skill in the art before the effective filing date of the claimed invention, with a reasonable expectation for success, to combine the invention of Kolluri with the method of training as taught by Horowitz. This is because as stated in [0053] of Horowitz’ disclosure: “By continuing this approach of cross-training a machine learning model associated with a first domain with training data that had been programmatically annotated by another related machine learning model with a different but related domain, ultimately a sophisticated, cross-domain machine learning model is achieved with much more efficient training times and diversity of data. Furthermore, existing machine learning models in one domain can be used to rapidly achieve state of the art sortation performance in entirely new domains, where overlapping materials are present.”
Regarding claim 2, Kolluri discloses:
The robot motion learning device according to claim 1, wherein the processor is configured to:
branch a process according to update conditions, and when an update condition for additional motion occurs: use teaching data related to motion unlearned by each of the robots, to train the shared learning model while keeping parameters of the first learning model and the second learning model fixed (See at least [0046]: “The base control policy can also be defined using system demonstration data that is collected during the process of developing the skill template. For example, a team of engineers associated with the entity that generates skill templates can perform demonstrations using one or more robots at a facility that is remote from and/or unassociated with the system 100. The robots used to generate the system demonstration data also need not be the same robots or the same robot models as the robots 170a-n in the workcell 170. In this case, the system demonstration data can be used to bootstrap the actions of the base control policy. The base control policy can then be adapted into a customized control policy using more computationally expensive and sophisticated learning methods.”)
Regarding claim 3, Kolluri discloses:
The robot motion learning device according to claim 1, wherein the processor is configured to:
branch a process according to update conditions, and when an update condition for adding a new robot occurs:
Newly add the first learning model and the second learning model corresponding to the robot to be added as an object to be controlled (see at least [0044]: “In some implementations, the training system generates the base control policy from generalized training data 165. While the local demonstration data 115 collected by the online execution system 110 is typically specific to one particular robot or one particular robot model, the generalized training data 165 can in contrast be generated from one or more other robots, which need not be the same model, located at the same site, or built by the same manufacturer. For example, the generalized training data 165 can be generated offsite from tens or hundreds or thousands of different robots having different characteristics and being different models. In addition, the generalized training data 165 does not even need to be generated from physical robots. For example, the generalized training data can include data generated from simulations of physical robots.”)
Regarding claim 4, Kolluri discloses:
The robot motion learning device according to claim 1, wherein the processor is configured to:
branch a process according to update conditions, and when an update condition for adding a new robot occurs: use teaching data for the robot to be added, related to motion already learned by the shared learning model, to train the first learning model and the second learning model while keeping parameters of the shared learning model fixed (see at least [0044-0045]: “In some implementations, the training system generates the base control policy from generalized training data 165. While the local demonstration data 115 collected by the online execution system 110 is typically specific to one particular robot or one particular robot model, the generalized training data 165 can in contrast be generated from one or more other robots, which need not be the same model, located at the same site, or built by the same manufacturer. For example, the generalized training data 165 can be generated offsite from tens or hundreds or thousands of different robots having different characteristics and being different models. In addition, the generalized training data 165 does not even need to be generated from physical robots. For example, the generalized training data can include data generated from simulations of physical robots.
Thus, the local demonstration data 115 is local in the sense that it is specific to a particular robot that a user can access and manipulate. The local demonstration data 115 thus represents data that is specific to a particular robot, but can also represent local variables, e.g., specific characteristics of the particular task as well as specific characteristics of the particular working environment.”)
Regarding claim 5, Kolluri discloses:
The robot motion learning device according to claim 1, wherein the processor is configured to:
branch a process according to update conditions, and when a tuning update condition for each robot occurs: upon newly obtaining teaching data for a robot, related to a shared motion already learned by the shared learning model, use the teaching data to perform either first training in which to only the first learning model and the second learning model corresponding to the robot are trained or second training in which only the shared learning model corresponding to the shared motion is trained, or alternate between the first training and the second training (see at least [0046]: “The base control policy can also be defined using system demonstration data that is collected during the process of developing the skill template. For example, a team of engineers associated with the entity that generates skill templates can perform demonstrations using one or more robots at a facility that is remote from and/or unassociated with the system 100. The robots used to generate the system demonstration data also need not be the same robots or the same robot models as the robots 170a-n in the workcell 170. In this case, the system demonstration data can be used to bootstrap the actions of the base control policy. The base control policy can then be adapted into a customized control policy using more computationally expensive and sophisticated learning methods.”)
Regarding claim 6, Kolluri discloses:
The robot motion learning device according to claim 1, wherein the shared learning model is provided for each of a plurality of shared motions, and the processor is configured to learn to have a predetermined value for a desired motion of the shared learning models, and select by the predetermined value and execute one of the shared learning models provided for each of the plurality of shared motions (see at least [0050]: “In execution mode, an execution engine 130 can use the customized control policy 125 to automatically perform the task without any user intervention. The online execution system 110 can use the customized control policy 125 to generate commands 155 to be provided to the robot interface subsystem 160, which drives one or more robots, e.g., robots 170a-n, in a workcell 170. The online execution system 110 can consume status messages 135 generated by the robots 170a-n and online observations 145 made by one or more sensors 171a-n making observations within the workcell 170. As illustrated in FIG. 1, each sensor 171 is coupled to a respective robot 170. However, the sensors need not have a one-to-one correspondence with robots and need not be coupled to the robots. In fact, each robot can have multiple sensors, and the sensors can be mounted on stationary or movable surfaces in the workcell 170.”)
Regarding claim 7, Kolluri discloses:
The robot motion learning device according to claim 1, wherein the shared learning model converts the external features output by the first learning models into predicted external features at the next time that are common to the plurality types of robots configured to have the same input/output structure regardless of the type of the robot or the motion (see at least [0046]: “The base control policy can also be defined using system demonstration data that is collected during the process of developing the skill template. For example, a team of engineers associated with the entity that generates skill templates can perform demonstrations using one or more robots at a facility that is remote from and/or unassociated with the system 100. The robots used to generate the system demonstration data also need not be the same robots or the same robot models as the robots 170a-n in the workcell 170. In this case, the system demonstration data can be used to bootstrap the actions of the base control policy. The base control policy can then be adapted into a customized control policy using more computationally expensive and sophisticated learning methods.”)
Regarding claim 8, Kolluri discloses:
A robot motion learning system comprising: the robot motion learning device according to claim 1; and a plurality of types of robots wherein each robot is configured to have a unique first learning model and a unique second learning model depending on the structure of the robot and the input/output features (see at least [0046]: “The base control policy can also be defined using system demonstration data that is collected during the process of developing the skill template. For example, a team of engineers associated with the entity that generates skill templates can perform demonstrations using one or more robots at a facility that is remote from and/or unassociated with the system 100. The robots used to generate the system demonstration data also need not be the same robots or the same robot models as the robots 170a-n in the workcell 170. In this case, the system demonstration data can be used to bootstrap the actions of the base control policy. The base control policy can then be adapted into a customized control policy using more computationally expensive and sophisticated learning methods.”)
Regarding claim 9, Kolluri discloses:
A robot motion learning method for learning motions of a plurality of types of robots, comprising the steps of:
causing a first learning model corresponding to a robot to learn processing for converting motion information and external information of the robot at a certain time into common motion features using teaching data related to the motion of the robot (see at least [0058]: “A skill template can also include, for each demonstration subtask, the software modules that are required to tune the demonstration subtask using local demonstration data. Each demonstration subtask can rely on a different type of machine learning model and can use different techniques for tuning… Therefore, the tuning procedure for the insertion demonstration subtask can more heavily tune the machine learning models that deal with force perception and corresponding feedback. In other words, even when the underlying models for subtasks in a skill template are the same, each subtask can have its own respective tuning procedures for incorporating local demonstration data in differing ways.”)
causing a shared learning model to learn a time-series relationship of the common motion features related to motions common to the plurality of types of robots using the teaching data related to the motions of the plurality of types of robots (see at least [0044]: “In some implementations, the training system generates the base control policy from generalized training data 165. While the local demonstration data 115 collected by the online execution system 110 is typically specific to one particular robot or one particular robot model, the generalized training data 165 can in contrast be generated from one or more other robots, which need not be the same model, located at the same site, or built by the same manufacturer. For example, the generalized training data 165 can be generated offsite from tens or hundreds or thousands of different robots having different characteristics and being different models. In addition, the generalized training data 165 does not even need to be generated from physical robots. For example, the generalized training data can include data generated from simulations of physical robots.”)
and causing a second learning model corresponding to the robot to learn processing for converting predicted values at a next time of the common motion features output by the shared learning model into predicted motion information of the robot at the next time using the teaching data related to the motion of the robot (see at least [0047]: “Adapting a base control policy using local demonstration data has the highly desirable effect that it is relatively fast compared to generating the base control policy, e.g., either by collecting system demonstration data or by training using the generalized training data 165. For example, the size of the generalized training data 165 for a particular task tends to be orders of magnitude larger than the local demonstration data 115 and thus training the base control policy is expected to take much longer than adapting it for a particular robot. For example, training the base control policy can require vast computing resources, in some instances, a datacenter having hundreds or thousands of machines working for days or weeks to train the base control policy from generalized training data. In contrast, adapting the base control policy using local demonstration data 115 can take just a few hours.”)
Kolluri does not explicitly disclose, but Horowitz, in an analogous field of endeavor teaches:
wherein the time is updated for each control cycle of the robot, and has a configuration in which input from the robot, the first learning model, the shared learning model, the second learning model, and output to the robot are processed in that order, so that the robot operates the prediction motion information as a control input (see at least [0053]: “ In some embodiments, model training logic 202 is configured to cross train machine learning models that correspond to different domains but share some overlapping materials. Examples of domains include single stream recyclables (SSR), construction and demolition (C&D), organics, and e-waste. In some embodiments, model training logic 202 is configured to train each of multiple machine learning models' known/sorted material of a corresponding domain. Then, input data can be fed into a particular machine learning model associated with a first domain to obtain that model's output of the first domain-specific classifications on that input data. Next, the input data with the labels of the first domain-specific classifications is then used as training data to train another, related machine learning model associated with a second domain that shares some overlapping materials with the first domain. In this way, the machine learning model corresponding to the second domain adds the recognition parameters associated with the first domain, without requiring the longer lead time and greater number of iterations necessary when starting from scratch and using only human annotation for a core machine learning model associated with the first domain.”)
It would have been prima facie obvious for one of ordinary skill in the art before the effective filing date of the claimed invention, with a reasonable expectation for success, to combine the invention of Kolluri with the method of training as taught by Horowitz. This is because as stated in [0053] of Horowitz’ disclosure: “By continuing this approach of cross-training a machine learning model associated with a first domain with training data that had been programmatically annotated by another related machine learning model with a different but related domain, ultimately a sophisticated, cross-domain machine learning model is achieved with much more efficient training times and diversity of data. Furthermore, existing machine learning models in one domain can be used to rapidly achieve state of the art sortation performance in entirely new domains, where overlapping materials are present.”
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ELIZABETH NELESKI whose telephone number is (571)272-6064. The examiner can normally be reached 10 - 6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, THOMAS WORDEN can be reached at (571) 272-4876. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/E.R.N./ Examiner, Art Unit 3658
/THOMAS E WORDEN/ Supervisory Patent Examiner, Art Unit 3658