Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This is a non-final Office Action on the merits. Claims 1-26 are currently pending and are addressed below.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 07/15/2025 is being considered by the examiner.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-26 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-24 of U.S. Patent No. 12384029. Although the claims at issue are not identical, they are not patentably distinct from each other because the patented claims contain substantially similar limitations, anticipating the pending claims.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 2, 5, 8, 15-16 and 20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kolluri US (2021/0362328)
Regarding claim 1:
Kolluri teaches A method for robotic skill learning, said method comprising:
providing a robot controlled by a robot controller (see at least Fig. 1), the robot being configured to perform an operation where the robot moves a first part into a goal position (multiple exemplary tasks meet this limitation, see at least ¶0009, ¶0011, ¶0171-0172);
providing a reinforcement learning controller in communication with the robot controller and the robot, said reinforcement learning controller having a neural network which defines a policy for determining an action in response to state data received as feedback from the robot (see at least Fig. 2C, base control policy generated using reinforcement learning techniques, see at least ¶0050, ¶0089-0097);
pre-training the reinforcement learning controller, including using a human demonstration dataset to train the neural network in the reinforcement learning controller to cause the policy to maximize reward data received as feedback from the robot (¶0050);
operating the robot to perform the operation with the reinforcement learning controller in a self-learning mode, where the reinforcement learning controller provides the action as input to the robot controller, and the state data and the reward data received as feedback from the robot are used to continuously train the neural network in the reinforcement learning controller (¶0084-0097).
Regarding claim 2:
Kolluri further teaches periodically operating the robot to perform the operation with the reinforcement learning controller in a co-training mode, where a human demonstrator provides supplemental input to the robot controller to control the robot to perform the operation (¶0095-0097).
Regarding claim 5:
Kolluri further teaches wherein the state data includes robot positions and velocities each including three translational and three rotational components, and contact forces and torques each in three directions (see at least ¶0080,¶0231).
Regarding claim 8 and 20:
Kolluri further teaches wherein the human demonstration dataset includes action, state and reward data captured for multiple human demonstrations of the operation, and pre-training the reinforcement learning controller includes using the human demonstration dataset for training with no interaction between the reinforcement learning controller and the robot (see at least ¶0051-0055).
Regarding claims 15-16, Kolluri teaches a robotic skill learning system comprising a robot and controller to perform the method as in claims 1-2 above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 3, 4 13-14, 17, 18, 25 and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri.
Regarding claims 3 and 17:
Kolluri further teaches3. The method according to Claim 2 wherein the co-training mode includes using human demonstration by teleoperation (see at least ¶0047), and action data from the teleoperation along with the state data and the reward data received as feedback from the robot in the co-training mode are used to further train the neural network in the reinforcement learning controller (¶0095-0097).
Kolluri does further teach automatically determining when a task or subtask requires further training/demonstration learning, but does not explicitly teach the demonstration being invoked when a robot success performance metric drops below a predefined threshold.
It would have been obvious to one of ordinary skill in the art at the time of filing of the invention that the robotic demonstration learning system and method would need to utilize some metric for determining when the demonstration learning is necessary, including some threshold metric in order for the system to make the determination as taught by Kolluri.
Regarding claims 4 and 18:
Kolluri teaches the limitations as in claim 1 above.
Kolluri further teaches a plurality of exemplary tasks, including insertion tasks (see at least ¶0064-0069) as well as the robot having multiple degrees of freedom (see at least ¶0255).
Kolluri is silent as to the specifics of the insertion task.
However, it would have been at least obvious if not inherent to one of ordinary skill in the art before the time of filing of the invention, that in implementing the exemplary tasks as taught by Kolluri, particular translations and rotations, as well as target positions and motion limits would be taken into account as is conventional in at least an insertion task.
Regarding claim 13 and 25:
Kolluri teaches the limitations as in claim 1 above.
Kolluri further teaches controlling the robot utilizing different controllers operating at different frequencies.
Kolluri does not explicitly teach the distribution of frequencies as claimed.
It would have been obvious to one of ordinary skill in the art before the time of filing of the invention to modify the robotic demonstration learning system and method as taught by Kolluri, including multiple controllers running at different frequencies, by implementing any configuration of clock relationships to ensure the data is processed in the various operation and training steps allowing the robot to learn a task as demonstration data becomes available.
Regarding claim 14 and 26:
Kolluri further teaches wherein the robot controller and the reinforcement learning controller are both executed on a controller device which provides joint motion commands to the robot and receives the state data as feedback from the robot (see at least ¶0262-0269).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 6 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri in view of De Magistris et al. (US 2019/0137954).
Regarding claims 6 and 19:
Kolluri teaches the limitations as in claim 1, 5, and 15 above.
Kolluri teaches reinforcement learning, which deals with rewards based on performance of a task.
Kolluri does not explicitly teach positive and negative rewards based on successful task completion.
De Magistris teaches a system and method of robotic reinforcement learning including wherein the reward data includes a positive value when the operation is successfully completed and a negative value when the operation is not successfully completed, and the positive value has a magnitude greater than the negative value (see at least abstract, ¶0020, ¶0058).
It would have been obvious to one of ordinary skill in the art at the time of filing of the invention to modify the robotic reinforcement learning system and method as taught by Kolluri with the technique of utilizing positive and negative rewards based on task success or failure as taught by De Magistris in order to allow the robotic system to learn from random actions, and slowly reduces exploration and increases exploitation (¶0058).
Claim Rejections - 35 USC § 103
Claims 7, 9-12, and 21-24 are rejected under 35 U.S.C. 103 as being unpatentable over Kolluri as applied to claim 1 above in view of Straele et al. (US 2023/0081738).
Regarding claim 7:
Kolluri teaches the limitations as in claim 1 above.
Kolluri is silent as to the policy being defined by a statistical distribution.
Straele teaches a system and method of training a control strategy for a robot manipulator including wherein the policy defined by the neural network in the reinforcement learning controller is a statistical distribution of actions relative to states, where the statistical distribution is defined by parameters including a mean and a standard deviation (see at least abstract, ¶0008-0020, Fig. 2).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to modify the robotic reinforcement learning system and method as taught by Kolluri with the technique of utilizing a statistical distribution in the policy as taught by Straele in order to enable efficient, non-adversarial learning from observations and enable the training of a successful control strategy with high data efficiency
Regarding claims 9 and 21:
Straele further teaches wherein pre-training the reinforcement learning controller includes an offline reinforcement learning technique using a Kullback–Leibler divergence calculation in a loss function which penalizes deviation of the policy learned by the neural network from a policy based on the human demonstration dataset (see at least ¶0054-0069, ¶0021).
Regarding claim 10 and 22:
Straele further teaches wherein the reinforcement learning controller has an actor module including the neural network and a critic module including a second neural network, where the actor module provides the action as input to the robot controller and the critic module updates parameters of the policy of the actor module based on a critic function (see at least ¶0063-0069).
Regarding claim 11 and 23:
Straele further teaches wherein training of the actor module by the critic module includes using an optimization computation which determines desired actions which maximize the reward calculated by the critic function, and adjusting the parameters of the policy in the actor module based on the desired actions (see at least ¶0063-0069).
Regarding claim 12 and 24, the Examiner notes that the recited limitations are conventional features of an actor critic learning algorithm, wherein a critic estimates a future reward therefore, it would have been at least obvious to one of ordinary skill in the art before the time of filing of the invention to optimize for expected future rewards in the actor-critic learning in order to accurately train the policy for the desired goal (see at least ¶0063-0069).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN J RINK whose telephone number is (571)272-4863. The examiner can normally be reached M-F 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Anna Momper can be reached on (571) 270-5788. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Ryan Rink/ Primary Examiner, Art Unit 3619