Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
2. This action is in response to the original filing on 05/17/2024. Claims 1-5 and 7-21 are pending and have been considered below.
Information Disclosure Statement
3. The information disclosure statement (IDS(s)) submitted on 05/17/2024 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
4. Claims 1, 3, 13, and 19 are objected to because of the following informalities:
Claim 1 recites “the state-action selector … a next joint states transition” where “the respective state-action selector … the next joint states transition” was apparently intended.
Claim 3 recites “the feel-environment” where “the environment” was apparently intended.
Claim 13 recites “The computer program product according to claim 13” where “The computer program product according to claim 12” was apparently intended.
Claim 19 recites “select the action based on the joint state of the agents … current actions of the agents in the swarm;” where “select the action based on the joint states of the agents … current actions of the agents in the swarm.” was apparently intended.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-5, 7-17, 20, and 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claims 1 and 12 recite “sending a control signal corresponding to the selected action to a said component associated with the agent.” Although the preamble introduces “components associated with a swarm of agents,” the claim does not establish a clear correspondence between each agent and a particular component. The expression “a said component” combines an indefinite article with the antecedent referring term “said” and leaves unclear whether each agent controls its own respective component, any one of the previously recited components, or a component shared by multiple agents.
Claims 1, 12, 20 recite “using the selected fictitious actions to update values in the state-action selector of each of the agents in the swarm.” It is unclear whether each agent uses only the fictitious actions selected by that agent to update that agent’s respective selector, whether all fictitious actions selected by all agents are used to update every selector, or whether fictitious actions selected by one agent can update the selector of another agent. Consequently, the relationship among the agents selected fictitious actions, and respective state-action selectors is uncertain.
Claims 2-5, 7-11, 13-17, 21 are dependent upon claims 1, 12, 20, incorporate the deficiencies of claims 1, 12, 20 and do not cure those deficiencies, they are likewise rejected.
Claim Rejections - 35 USC § 101
5. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 18 and 19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea without significantly more.
Step 1, the claim 18 is directed to a process and manufacture.
Step 2A Prong 1, Claim 18 recites, in part
select an action based on joint states of the agents in the swarm (Mental processes, evaluation and judgement).
selecting, using its trained state-action selector, an action; and (Mental processes, evaluation).
wherein the state-action selector is trained using reinforcement learning, and the reinforcement learning comprises a Markov decision process (MDP) (Mathematical concepts, mathematical relationships and operations).
Step 2A Prong 2, this judicial exception is not integrated into a practical application.
The additional elements:
a computer program product comprising one or more non-transitory machine readable mediums encoded with instructions which, when executed by one or more processors, cause a process to be carried out for controlling an agent to perform actions, the agent included in a swarm of agents and includes a state-action selector that is trained (mere instructions to apply the exception using a generic computer component).
performing, by the agent included in the swarm of agents (mere instructions to apply the exception using a generic computer component).
sending, by the agent included in the swarm of agents, a control signal corresponding to the selected action to a component associated with the agent included in the swarm of agents (mere data gathering and output recited at a high level of generality, and thus are insignificant extra-solution activity).
Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception, either alone or in combination.
The additional elements:
a computer program product comprising one or more non-transitory machine readable mediums encoded with instructions which, when executed by one or more processors, cause a process to be carried out for controlling an agent to perform actions, the agent included in a swarm of agents and includes a state-action selector that is trained (mere instructions to apply the exception using a generic computer component).
performing, by the agent included in the swarm of agents (mere instructions to apply the exception using a generic computer component).
sending, by the agent included in the swarm of agents, a control signal corresponding to the selected action to a component associated with the agent included in the swarm of agents (mere data gathering and output recited at a high level of generality, and thus are insignificant extra-solution activity).
Claim 19 provides further limitations to the abstract idea (Mathematical concepts and/or Mental processes) as rejected in claim 18, however, they do not disclose any additional elements that would amount to a practical application or significantly more than an abstract idea (data gathering/insignificant extra-solution activity and/or generic computer component).
Claim Rejections - 35 USC § 102
6. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
7. Claim 18 is rejected under 35 U.S.C. 102(a)(2) as being anticipated by Isele et al. (U.S. Patent Application Pub. No. US 20220230080 A1).
Claim 18: Isele teaches a computer program product comprising one or more non-transitory machine readable mediums encoded with instructions which, when executed by one or more processors, cause a process to be carried out (i.e. a non-transitory machine-readable storage medium, such as a volatile or non-volatile memory, which may be read and executed by at least one processor to perform the operations described in detail herein; para. [0109]) for controlling an agent to perform actions (i.e. the ego agent 102 and/or the target agent 104 to be autonomously operated to reach their respective goals 204, 206 while accounting for one another within the multi-agent environment 200; para. [0035, 0103-0105]), the agent included in a swarm of agents (i.e. the system 100 includes an ego agent 102 and one or more target agents 104. For purposes of simplicity, this disclosure will describe the embodiments of the system 100 with respect to a single ego agent 102 and a single target agent 104. However, it is appreciated that the system 100 may include more than one ego agent 102 and more than one target agent 104 and that the embodiments and processes discussed herein may be utilized in an environment that includes one or more ego agents 102 and one or more target agents 104; para. [0026, 0027, 0035]) and includes a state-action selector (i.e. the storage units 114 a, 114 b may respectively store the respectively learned agent-action polices associated with the ego agent 102 and/or the target agent 104. Accordingly, the storage units 114 a, 114 b may be accessed by the multi-agent application 106 to store the respective agent-action polices learned by the application 106 to be followed by the respective agents 102, 104. In some embodiments, the storage units 114 a, 114 b may be accessed by the application 106 to retrieve the respective agent-action polices to autonomously control the operation of the ego agent 102 and/or the target agent 104 to account for the presence of one another (e.g., other agents) within the multi-agent environment 200; para. [0042, 0061, 086, 0087]) that is trained to select an action based on joint states of the agents in the swarm (i.e. S is the state space containing the state for all agents 102 a, 104 a … the current state as well as the actions of all agents … for each agent i, Ai:S→Ai maps the state to its action … Additional terms like the policy entropy could also be added to Jπφ to improve the training. In one embodiment, the policy learning module 134 trains a central critic, Që i(s,ai,a−i), for each agent i, to estimate the return value of the state and the joint—action; para. [0084-0087]), the process comprising:
selecting and performing, by the agent included in the swarm of agents and using its trained state-action selector, an action (i.e. to select an optimum set of actions for the virtual ego agent 102 a and the virtual target agent 104 a … The output of the recursive reasoning graph will be the agent action policy for each of the virtual ego agent 102 a and the virtual target agent 104 a … implementing the agent action policy to operate the ego agent 102 and/or the target agent 104; para. [0085, 0087, 0100-0105); and
sending, by the agent included in the swarm of agents, a control signal corresponding to the selected action to a component associated with the agent included in the swarm of agents (i.e. the ECUs 110 a, 110 b may be configured to operably control the plurality of components of the respective agents 102, 104. The ECUs 110 a, 110 b may additionally provide one or more commands to one or more control units (not shown) of the agents 102, 104 including, but not limited to a respective engine control unit, a respective braking control unit, a respective transmission control unit, a respective steering control unit, and the like to control the ego agent 102 and/or target agent 104 to be autonomously operated; para. [0043]);
wherein the state-action selector is trained using reinforcement learning (i.e. the system 100 may include a multi-agent recursive reasoning reinforcement learning application (multi-agent application) 106 that may be configured to complete multi-agent reinforcement learning for simultaneously learning policies for multiple agents interacting amongst one another; para. [0030, 0032]), and the reinforcement learning comprises a Markov decision process (MDP) (i.e. using the simulated model 400, the Markov Game may be specified by (S,{Ai}i=1 n,T,{ri}i=1 n,Reject,s0), where n is the number of agents 102 a, 104 a; S is the state space containing the state for all agents 102 a, 104 a; A″ represent the action space for agent i (where agent i is the virtual ego agent 102 a when determining the reward for the virtual ego agent 102 a, where agent i is the virtual target agent 104 a when determining the reward for the virtual target agent 104 a); T:S×Πi=1 nAi×S→R represents the transition probability conditioned on the current state as well as the actions of all agents 102 a, 104 a; ri: S×Πi=1 nAi×S→R represents the reward for agent i; and s0:S→R represents the initial state distribution of all agents 102 a, 104 a; para. [0082-0087]).
Claim Rejections – 35 USC § 103
8. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
9. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Isele in view of Faust et al. (U.S. Patent Pub. No. US 11067988 B1).
Claim 19: Isele teaches the computer program product according to claim 18. Isele further teaches wherein data for performing the MDP includes:
the state-action selector (i.e. the policy learning module 134 may train the critic, Qθ (s, a), to estimate the return value of the state-action pair (s, a), the policy learning module 134 trains a central critic, Që i(s,ai,a−i), for each agent i, to estimate the return value of the state and the joint—action … Using centralized-training-decentralized execution, the centralized critic is defined as: Qi(s, ai, a−i) and the decentralized actor is defined as: πi(ai|s). The training may be defined as: Qi(s, ai, a−i)←r+γV(s′); πi(ai|s)←argmaxa i Qi(s, ai, a−i); para. [0086, 0087]);
a policy for determining how to use the state-action selector to select the action based on the joint state of the agents (i.e. Using centralized-training-decentralized execution, the centralized critic is defined as: Qi(s, ai, a−i) and the decentralized actor is defined as: πi(ai|s). The training may be defined as: Qi(s, ai, a−i)←r+γV(s′); πi(ai|s)←argmaxa i Qi(s, ai, a−i). The centralized critic is thereby utilized to determine how good a particular action is to select an optimum set of actions for the virtual ego agent 102 a and the virtual target agent 104 a; para. [0084-0087]);
a probabilistic behaviour model configured to predict a next state of an environment based on a current state of the environment (i.e. using the simulated model 400, the Markov Game may be specified by (S,{Ai}i=1 n,T,{ri}i=1 n,Reject,s0), where n is the number of agents 102 a, 104 a; S is the state space containing the state for all agents 102 a, 104 a; A″ represent the action space for agent i (where agent i is the virtual ego agent 102 a when determining the reward for the virtual ego agent 102 a, where agent i is the virtual target agent 104 a when determining the reward for the virtual target agent 104 a); T:S×Πi=1 nAi×S→R represents the transition probability conditioned on the current state as well as the actions of all agents 102 a, 104 a; ri: S×Πi=1 nAi×S→R represents the reward for agent i; and s0:S→R represents the initial state distribution of all agents 102 a, 104 a; para. [0084]), wherein the probabilistic behaviour model comprises an artificial neural network (i.e. the multi-agent application 106 may be configured to utilize a neural network 108 and may execute instructions to utilize a recursive reasoning model in a centralized-training-decentralized execution framework to facilitate maneuvers within the multi-agent environment 200 with respect to the agents 102, 104; para. [0030]) configured to predict next actions of items or features in the environment given current joint states of the agents (i.e. Using centralized-training-decentralized execution, the centralized critic is defined as: Qi(s, ai, a−i) and the decentralized actor is defined as: πi(ai|s). The training may be defined as: Qi(s, ai, a−i)←r+γV(s′); πi(ai|s)←argmaxa i Qi(s, ai, a−i). The centralized critic is thereby utilized to determine how good a particular action is to select an optimum set of actions for the virtual ego agent 102 a and the virtual target agent 104 a; para. [0084-0087]); and
an environment model configured to generate a set of next joint states of the agents and rewards (i.e. The ego agent 102 and the target agent 104 are evaluated as the central actors and are treated as nodes to build the recursive reasoning graph to efficiently calculate higher level recursive actions of interacting agents. As central actors, the ego agent 102 and the target agent 104 are analyzed as self-interested agents that are attempting to reach their respective goals 204, 206 in a most efficient manner. The multi-agent central actor critic model includes one or more iterations of Markov Games where one or more critics evaluate one or more actions (output of actor models) taken by a simulated ego agent and a simulated target agent to determine one or more rewards and one or more states related to a goal-specific reward function; para. [0032]), based on the current state of the environment, the current joint states of the agents, and current actions of the agents in the swarm (i.e. using the simulated model 400, the Markov Game may be specified by (S,{Ai}i=1 n,T,{ri}i=1 n,Reject,s0), where n is the number of agents 102 a, 104 a; S is the state space containing the state for all agents 102 a, 104 a; A″ represent the action space for agent i (where agent i is the virtual ego agent 102 a when determining the reward for the virtual ego agent 102 a, where agent i is the virtual target agent 104 a when determining the reward for the virtual target agent 104 a); T:S×Πi=1 nAi×S→R represents the transition probability conditioned on the current state as well as the actions of all agents 102 a, 104 a; ri: S×Πi=1 nAi×S→R represents the reward for agent i; and s0:S→R represents the initial state distribution of all agents 102 a, 104 a; para. [0084]).
Isele does not explicitly teach the policy comprising an e-greedy policy; configured to predict next actions of non-controllable items or features in the environment.
However, Faust teaches the policy comprising an e-greedy policy (i.e. The system can select a candidate action among multiple possible actions using a current state of the reinforcement learning model. In other words, the system can provide the initial environment observation and each of the possible actions to the current reinforcement learning model and receive a respective cumulative action score for each possible action. The system can then select the candidate action having the highest cumulative action score, or select the candidate action using other criteria. For example, the system can select a candidate action that encourages exploration with some measure of randomness, e.g., using an epsilon-greedy selection strategy; col. 7, lines 18-29);
a probabilistic behaviour model (i.e. the vehicle behavior engine 210 is a machine-learning system that generates predictions about how taking a particular action affects the behavior of other actors in the environment, e.g., nearby vehicles, pedestrians, bicyclists, or other moving objects. For example, swerving toward another vehicle tends to cause the other vehicle to swerve away. And braking in front of another vehicle tends to cause the other vehicle to also apply the brakes; col. 5, lines 15-23) configured to predict a next state of an environment based on a current state of the environment (i.e. the vehicle behavior engine 210 receives an initial environment observation 215 and a candidate action 225 and generates a predicted environment observation 235; col. 5, lines 33-45), wherein the probabilistic behaviour model comprises an artificial neural network (i.e. The vehicle behavior engine 210 can implement one or more neural networks for generating predicted environment observations 235; col. 5, lines 53-59) configured to predict next actions of non-controllable items or features in the environment (i.e. the vehicle behavior engine 210 is a machine-learning system that generates predictions about how taking a particular action affects the behavior of other actors in the environment, e.g., nearby vehicles, pedestrians, bicyclists, or other moving objects. For example, swerving toward another vehicle tends to cause the other vehicle to swerve away. And braking in front of another vehicle tends to cause the other vehicle to also apply the brakes; col. 5, lines 15-23) given current joint states of the agents (i.e. The vehicle behavior engine 210 thus models how an initial environment observation changes in response to the agent taking some action in the initial environment. In this context, each environment observation can represent the state of each actor in the environment using, e.g., data representing a type of the actor, a position, an orientation, and a velocity of the actor, and optionally, a stochastic factor that introduces some level of uncertainty in these components; col. 5, lines 24-32).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Isele to include the feature of Faust. One would have been motivated to make this modification because it improves the accuracy and exploration of policy learning process.
Allowable Subject Matter
Claims 1-5 and 7-17 are allowed and if the 35 USC §112(b) is successfully addressed.
Claims 20 and 21 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Cao et al. (Pub. No. US 12561602 B2), A method for controlling an ego agent includes periodically receiving policy information comprising a spatial environment observation and a current state of the ego agent. The method also includes selecting, for each received policy information, a low-level policy from a number of low-level policies. The low-level policy may be selected based on a high-level policy. The method further includes controlling an action of the ego agent based on the selected low-level policy.
Sun et al. (Pub. No. US 12265924 B1), a reinforcement learning policy for controlling the computer-implemented agent. The method also comprises storing a distribution as a latent representation of a belief of the computer-implemented agent about at least one other agent in the environment.
Hofmann et al. (Pub. No. US 12488278 B2), techniques for robust multi-agent reinforcement learning (MARL) are described. An exemplary method includes initializing a plurality of parameters for a plurality of agents including at least policy parameters and action-value (Q) parameters; performing robust multi-agent reinforcement learning to learn polices for the agents, wherein in the learned polices no agent has an incentive to deviate, the agents include an implicit agent that is to select a worst-case at any given time during the learning process; and at least one agent utilizing its learned policy.
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TAN H TRAN/Primary Examiner, Art Unit 2141