Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
2. This action is responsive to the Application filed on 7/19/2024. A filing date 7/19/2024 is acknowledged. Claims 1-20 are pending in this application. Claims 1, 17, 20 are independent claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
3. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Adrian Martin (US Publication 20200016748 A1, hereinafter Martin), and in view of Connor Basich et al (US Publication 20230382433).
As for independent claim 1, Martin discloses: A controller for controlling a collaboration of a set of agents jointly performing a task, wherein the set of agents includes at least one robot (Martin: Abstract, A distribution system (e.g., distributed robotic system) and method, may employ a plurality of agents, each agent associated with a respective set of processor-executable instructions, and a specification that defines routes amongst the agents; [0012], Each of shoulder servos 402 and 405, and servo in robot 400, work cooperatively with a respective joint, or joint and gearbox), wherein for at least some of different control steps, the set of agents include different combinations of active agents and inactive agents (Martin: [0028], a respective agent in the first plurality of agents can be in an active state or an inactive state) defined by a collaboration variable (Martin: [0025], The instance of the second agent may include a variable, and attempting to locate the instance of the host agent, may further include attempting to locate, by the at least one processor, the instance of the host agent that includes the variable set to a first value), the controller includes circuitry configured to: accept a feedback signal including observations of a state of execution of the task performed by the active agents from the set of agents specified by the collaboration variable (Martin: [0028], a respective agent in the first plurality of agents can be in an active state or an inactive state); process the observations with a neural network trained with machine learning (Martin: [0012], Artificial neural networks (ANNs) are a family of machine learning models inspired by biological neural networks. The models include methods for learning and associated networks) to determine actions for the active agents specified by the collaboration variable (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value), wherein the actions are selected from types of actions including activation actions calling for activating or deactivating a specific agent from the set of agents (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value); and update the collaboration variable (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value); and output with the neural network at least one activation action from the activation actions to update a combination of active agents and inactive agents and cause the active agents to execute the determined actions, wherein the active agents and the inactive agents belong to the set of agents, and wherein the combination of active agents and inactive agents is one of the different combinations of active agents and inactive agents defined by the collaboration variable (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value).
Martin discloses using machine learning technologies but does not clearly disclose process the observations with a neural network trained with machine learning to determine actions, in an analogous art of robot-huamn cooperation management using machine learning, Basich discloses: process the observations with a neural network trained with machine learning to determine actions for the active agents specified by the collaboration variable (Basich: [0003], identifying a current set of indiscriminate states; identifying a discriminator from the current set of indiscriminate states; training a feedback model for the discriminator; determining an autonomy level associated with the environment state and the action; and performing the action according to the autonomy level. The autonomy level can be selected based at least on an autonomy model and a feedback model. The feedback model may be trained with or without the presence of a discriminator and an iterative state space refinement approach);
Martin and Basich are analogous arts because they are in the same field of endeavor, robot-huamn cooperation management using machine learning. Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention, to modify the invention of Martin using the teachings of Basich to include process the observations of state data with a neural network trained with machine learning to determine actions for the active agents. It would provide Martin’s device with enhanced capabilities of improving the accuracy and confidence of the model.
As for claim 2, Martin-Basich discloses: wherein the neural network solves an open decentralized Markov decision process (oDec-MDP) model (Basich: [0117], Partially Observable Markov Decision Process (POMDP) models, Markov Decision Process (MDP) models, Classical Planning (CP) models, Partially Observable Stochastic Game (POSG) models, Decentralized Partially Observable Markov Decision Process (Dec-POMDP) models, Reinforcement Learning (RL) models, artificial neural networks, hardcoded expert logic, or any other suitable types of models).
As for claim 3, Martin-Basich discloses: the neural network is trained with reinforcement learning based on the oDec-MDP model (Basich: [0117], Partially Observable Markov Decision Process (POMDP) models, Markov Decision Process (MDP) models, Classical Planning (CP) models, Partially Observable Stochastic Game (POSG) models, Decentralized Partially Observable Markov Decision Process (Dec-POMDP) models, Reinforcement Learning (RL) models, artificial neural networks, hardcoded expert logic, or any other suitable types of models).
As for claim 4, Martin-Basich discloses: wherein the neural network is trained with inverse reinforcement learning (IRL) based on the oDec-MDP model (Basich: [0271], The models can be updated using model-free reinforcement learning, model-based reinforcement learning, or some other learning technique. In model-free reinforcement learning, probability values can be adjusted up or down. In model-based reinforcement learning, probability values can be calculated based on counts of successfully reaching a goal state in a scenario as compared to all the times that the scenario was encountered).
As for claim 5, Martin-Basich discloses: wherein the oDec-MDP model is solved using open decentralized adversarial inverse reinforcement learning (o-Dec-AIRL), the o-Dec-AIRL comprising learning a common reward function for the task (Basich: [0108], a reward function that outputs a positive or negative reward) and a corresponding vector of learned policies based on one or more expert trajectories (Basich: [0120], A MDP model may model a distinct vehicle operational scenario using a set of states, a set of actions, a set of state transition probabilities, a reward function, or a combination thereof. In some embodiments, modeling a distinct vehicle operational scenario may include using a discount factor, which may adjust, or discount, the output of the reward function applied to subsequent temporal periods; [0152], Given such descriptors and rewards, an optimal policy π*, which is a set of actions (or equivalently, a path through the operational environment), that maximizes reward, can be computed as a function of how the operational environment evolves over time).
As for claim 6, Martin-Basich discloses: the common reward function is learned using inverse reinforcement learning contingent of the collaboration variable, a state space, and an action space (Basich: [0037], Iterative state space refinement may enable a CAS to refine the granularity of its state representation online).
As for claim 7, Martin-Basich discloses: wherein the common reward function is used to learn the corresponding vector of learned policies, wherein the vector of learned policies includes one learned policy for each active agent involved in the task (Basich: [0120], A MDP model may model a distinct vehicle operational scenario using a set of states, a set of actions, a set of state transition probabilities, a reward function, or a combination thereof. In some embodiments, modeling a distinct vehicle operational scenario may include using a discount factor, which may adjust, or discount, the output of the reward function applied to subsequent temporal periods; [0152], Given such descriptors and rewards, an optimal policy π*, which is a set of actions (or equivalently, a path through the operational environment), that maximizes reward, can be computed as a function of how the operational environment evolves over time).
As for claim 8, Martin-Basich discloses: wherein the circuitry is configured to generate an activation signal to cause a currently active agent to activate a currently inactive agent (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value).
As for claim 9, Martin-Basich discloses: wherein the currently active agent is a currently active robot, and the currently inactive agent is a currently inactive robot, and wherein the currently active robot submits the activation signal to the currently inactive robot (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value).
As for claim 10, Martin-Basich discloses: wherein the currently active agent is a currently active robot, and the currently inactive agent is a currently inactive human, and wherein the currently active robot submits the activation signal to the currently inactive human (Martin: [0179], the controller could dynamically change one or more routes from active to inactive. That is, the routes could change depending on which agents have stopped or started, that is, are inactive or active. The controller could coordinate the operation of a plurality of instances of agents that include or is associated with a particular label value. The controller, for example, can update a plurality of instances of agents that include or is associated with a particular label value).
As for claim 11, Martin-Basich discloses: wherein the activation signal is at least one of: a radio signal, an audio signal, and a video signal (Martin: [0080], the human operator 105 observes representations of sensor data, for example, video, audio, or haptic data received from one or more environmental sensors or internal sensor. The human operator then acts, conditioned by a perception of the representation of the data, and creates information or executable instructions to direct the at least one of the one or more of robots 102).
As for claim 12, Martin-Basich discloses: wherein the collaboration variable is a binary vector of a size of the set of agents, wherein the state of execution of the task is formulated based on the binary vector before submission to the neural network (Martin: [0025], The instance of the second agent may include a variable, and attempting to locate the instance of the host agent, may further include attempting to locate, by the at least one processor, the instance of the host agent that includes the variable set to a first value.
As for claim 13, Martin-Basich discloses: wherein the collaboration variable is a unique identifier natural number for each team of agents in the set of agents (Martin: [0164], the controller updates a registry with information that identifies (e.g., name, unique identifier) the instance of the agent).
As for claim 14, Martin-Basich discloses: wherein the set of agents comprises at least: a robot agent and a human agent such that either of the robot and the human is able to exit and enter the task during execution of the task in an open human-robot collaboration environment (Martin: [0216], The routes 1708 can be dynamic, switching as agents are started or stopped, e.g., enter inactive state or active state. For example, in some instances plurality of agents 1700 could include the first sub-plurality of agents 1702 and the second sub-plurality of agents 1704. In some instances, the plurality of agents 1700 could include the third sub-plurality of agents 1706 and the second sub-plurality of agents 1704. One or more agents can be stopped or started by label; Basich: [0036], This model enables the system to operate more reliably in the open world, reduce improper reliance on the human and ultimately optimize the autonomous behavior of the system).
As for claim 15, Martin-Basich discloses: wherein a time of execution of the task associated with the human agent is minimized for the task in the open human-robot collaboration environment (Basich: [0036], This model enables the system to operate more reliably in the open world, reduce improper reliance on the human and ultimately optimize the autonomous behavior of the system; [0192], A component of the CAS model is the ability to adjust its autonomy profile over time using what the system has learned in order to optimize its autonomy by reducing unnecessary reliance on human assistance, regardless of how the autonomy profile is initialized).
As for claim 16, Martin-Basich discloses: wherein the circuitry is configured to generate a control command that causes active agents to execute the determined actions (Basich: [0264], a human operator can remotely control the vehicle 10004 or can send commands to the vehicle 10004 to perform actions).
As per claim 17, it recites features that are substantially same as those features claimed by claim 1, thus the rationales for rejecting claim 1 are incorporated herein.
As per claim 18, it recites features that are substantially same as those features claimed by claim 4, thus the rationales for rejecting claim 4 are incorporated herein.
As per claim 19, it recites features that are substantially same as those features claimed by claim 13, thus the rationales for rejecting claim 13 are incorporated herein.
As per claim 20, it recites features that are substantially same as those features claimed by claim 1, thus the rationales for rejecting claim 1 are incorporated herein.
Examiner’s Note
Examiner has cited particular columns/paragraph and line numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. This will assist in expediting compact prosecution. MPEP 714.02 recites: “Applicant should also specifically point out the support for any amendments made to the disclosure. See MPEP § 2163.06. An amendment which does not comply with the provisions of 37 CFR 1.121(b), (c), (d), and (h) may be held not fully responsive. See MPEP § 714.” Amendments not pointing to specific support in the disclosure may be deemed as not complying with provisions of 37 C.F.R. 1.131(b), (c), (d), and (h) and therefore held not fully responsive. Generic statements such as “Applicants believe no new matter has been introduced” may be deemed insufficient.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Applicants are required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action.
Sinyavskiy (US Publication 20180319015) APPARATUS AND METHODS FOR HIERARCHICAL TRAINING OF ROBOTS
Passot (US Publication 20150094850) APPARATUS AND METHODS FOR TRAINING OF ROBOTIC CONTROL ARBITRATION
Terra (US Publication 20250005450) DYNAMIC ML MODEL SELECTION
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hua Lu whose telephone number is 571-270-1410 and fax number is 571-270-2410. The examiner can normally be reached on Mon-Fri 9:00 am to 6:00 pm EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Scott Baderman can be reached on 571-272-3644. The fax phone number for the organization where this application or proceeding is assigned is 703-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Hua Lu/
Primary Examiner, Art Unit 2118