DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is made non-final.
This action is made in response to the application and claims filed May 23, 2024. Claims 1-20 are pending in the case and have been examined. Claims 1-20 are rejected.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 3, 4, 5, 13,14, and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 3 recites the limitation "generating the one or more reward functions based on the obtained sequence of state-joint action pairs and observed environment dynamics" in page 1. There is insufficient antecedent basis for this limitation in the claim. Claim 3 depends on claim 1 and claim 1 does not introduce a “sequence of state-joint action pairs”.
Claims 4 and 5 depend from claim 3 and inherit the same deficiencies described above.
Claim 13 recites the limitation "generate the one or more reward functions based on the obtained sequence of state-joint action pairs and observed environment dynamics" in page 3. There is insufficient antecedent basis for this limitation in the claim. Claim 13 depends on claim 11 and claim 11 does not introduce a “sequence of state-joint action pairs”.
Claims 14 and 15 depend from claim 13 and inherit the same deficiencies described above.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
To determine if a claim is directed to patent ineligible subject matter, the Court has guided the Office to apply the Alice/Mayo test, which requires:
Step 1: Determining if the claim falls within a statutory category.
Step 2A: Determining if the claim is directed to a patent ineligible judicial exception consisting of a law of nature, a natural phenomenon, or abstract idea; and Step 2A is a two prong inquiry. MPEP 2106.04(II)(A). Under the first prong, examiners evaluate whether a law of nature, natural phenomenon, or abstract idea is set forth or described in the claim. Abstract ideas include mathematical concepts, certain methods of organizing human activity, and mental processes. MPEP 2104.04(a)(2). The second prong is an inquiry into whether the claim integrates a judicial exception into a practical application. MPEP 2106.04(d).
Step 2B: If the claim is directed to a judicial exception, determining if the claim recites limitations or elements that amount to significantly more than the judicial exception. (See MPEP 2106).
Claims 1-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1-10 are directed to a method (a process), Claims 11-19 are directed to a computing system comprising processing circuitry (a machine), and Claim 20 is directed to a non-transitory computer-readable storage medium (a manufacture). Therefore, Claims 1-20 are directed to a process, machine or manufacture or composition of matter.
Regarding claim 1
Step 2A Prong 1
Claim 1 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., “probability distribution”, “plurality of trajectories” and “reward functions”) [see MPEP 2106.04(a)(2)(I)].
“generating, based on the data indicating the plurality of trajectories, a probability distribution of each agent of the plurality of agents over the plurality of baseline profiles, wherein the probability distribution of each agent describes a behavior of the agent” (determining mathematical probability values associated with the baseline profiles)
“updating, based on one or more observed joint actions performed by the team, the corresponding probability distribution of each agent of the plurality of agents” (mathematically modifying the probability values of the distribution based on the observed joint actions)
Claim 1 further recites the following mental processes, that in each case under the broadest reasonable interpretation, covers performance of the limitation in the mind (including observation, evaluation, judgement, opinion) or with the aid of pencil and paper but for recitation of generic computer components (e.g., “probability distribution”, “plurality of trajectories” and “reward functions”) [see MPEP 2106.04(a)(2)(III)].
“generating, based on the updated probability distributions of the plurality of agents, one or more reward functions that explain the observed one or more joint actions performed by the team, wherein each of the one or more reward functions describes the behavior of a corresponding one of the plurality of agents” (e.g., a human can evaluate observed agent behavior to determine information describing/explaining the agent’s behavior, observe behavior and then determine an explanation of why the agent behaved that way)
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “team modeling”, and “agents” which are recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)). In particular it is merely describing how the data is labeled for use in the claimed process. The Examiner notes that this is used throughout the claim limitations and is rejected thusly for each claim which recites the same language.
Regarding “obtaining data indicating a plurality of trajectories representing a behavior of a team comprising a plurality of agents” which is recited at a high-level of generality such that it amounts to extra-solution activity of receiving data, i.e. pre-solution activity of data gathering for use in the claimed process (see MPEP 2106.05(g)).
Regarding “obtaining a plurality of baseline profiles, wherein each of the plurality of baseline profiles encodes at least one of a preference and/or a goal that is relevant to a task performed by the team” which is recited at a high-level of generality such that it amounts to extra-solution activity of receiving data, i.e. pre-solution activity of data gathering for use in the claimed process (see MPEP 2106.05(g)).
Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, the additional element of “team modeling”, and “agents” which are recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)).
Regarding “obtaining data indicating a plurality of trajectories representing a behavior of a team comprising a plurality of agents” limitation, the additional element is recited at a high-level of generality and amounts to extra-solution activity of pre-solution activity of gathering data to be manipulated. The courts have found limitations directed to obtaining information electronically, recited at a high-level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
Regarding “obtaining a plurality of baseline profiles, wherein each of the plurality of baseline profiles encodes at least one of a preference and/or a goal that is relevant to a task performed by the team” limitation, the additional element is recited at a high-level of generality and amounts to extra-solution activity of pre-solution activity of gathering data to be manipulated. The courts have found limitations directed to obtaining information electronically, recited at a high-level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
Accordingly, at Step 2B, the additional element individually or in combination does not amount to significantly more than the judicial exception.
Regarding claim 2
Step 2A Prong 1
Claim 2 does not recite any additional abstract ideas but is directed to the abstract idea identified in its parents claim(s).
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “wherein each of the plurality of trajectories is represented as a sequence of state-joint action pairs over time” which is recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)). In particular it is merely limits the trajectory information used I the claimed mathematical analysis to a particular representation comprising state joint action pairs over time, and does not meaningfully alter how the recited mathematical concepts are performed or otherwise integrate the judicial exception into a practical application.
Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, the additional element of
“wherein each of the plurality of trajectories is represented as a sequence of state-joint action pairs over time” which is recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)).
Accordingly, at Step 2B, the additional element individually or in combination does not amount to significantly more than the judicial exception.
Regarding claim 3
Step 2A Prong 1
Claim 3 does not recite any additional abstract ideas but is directed to the abstract idea identified in its parents claim(s).
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “generating the one or more reward functions based on the obtained sequence of state-joint action pairs and observed environment dynamics” which is recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). This limitation merely instructs that the obtained information be used to generate the reward functions without reciting a particular manner of generating the reward functions that meaningfully limits the judicial exception.
In accordance with Step 2A, Prong 2, the claim does not include any additional elements, and the judicial exception is not integrated into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional element of “generating the one or more reward functions based on the obtained sequence of state-joint action pairs and observed environment dynamics” which is recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)).
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 4
Step 2A Prong 1
Claim 4 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., “probability distribution”, “plurality of trajectories” and “reward functions”)) [see MPEP 2106.04(a)(2)(I)].
“generating the one or more reward functions using inverse reinforcement learning method” (e.g., mathematical determination or infers a reward function based on observed behavior by optimizing reward function parameters)
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
In accordance with Step 2A, Prong 2, the claim does not include any additional elements, and the judicial exception is not integrated into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 5
Step 2A Prong 1
Claim 5 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., “probability distribution”, “plurality of trajectories” and “reward functions”)) [see MPEP 2106.04(a)(2)(I)].
“generating one or more reward weights that maximize the probability of observing actual trajectories of an agent given the updated probability distributions of the plurality of agents” (e.g., mathematically determining reward weights according to an optimization criterion that maximizes a probability value based on the updated probability distributions)
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
In accordance with Step 2A, Prong 2, the claim does not include any additional elements, and the judicial exception is not integrated into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 6
Step 2A Prong 1
Claim 6 recites the following mathematical concepts, that in each case under the broadest reasonable interpretation, covers performance of mathematical relationships, mathematical formulas or equations, and mathematical calculations but for recitation of generic computer components (e.g., “probability distribution”, “plurality of trajectories” and “reward functions”)) [see MPEP 2106.04(a)(2)(I)].
“wherein the team operates in a Multiagent Partially Observable Markov Decision Processes environment” (e.g., mathematical relationships including states, actions, transition probabilities, reward functions, and associated probability/value relationships for modeling decision making)
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
In accordance with Step 2A, Prong 2, the claim does not include any additional elements, and the judicial exception is not integrated into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 7
Step 2A Prong 1
Claim 7 recites the following mental processes, that in each case under the broadest reasonable interpretation, covers performance of the limitation in the mind (including observation, evaluation, judgement, opinion) or with the aid of pencil and paper but for recitation of generic computer components (e.g., “probability distribution”, “plurality of trajectories” and “reward functions”) [see MPEP 2106.04(a)(2)(III)].
“wherein generating the one or more reward functions comprises generating the one or more reward functions to achieve decentralized equilibrium” (e.g., e.g., a human can evaluate how individual actions affect the group and establish rewards/penalties intended to produce cooperative behavior)
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
In accordance with Step 2A, Prong 2, the claim does not include any additional elements, and the judicial exception is not integrated into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 8
Step 2A Prong 1
Claim 3 does not recite any additional abstract ideas but is directed to the abstract idea identified in its parents claim(s).
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “performing an action that maximizes a team benefit based on the one or more reward functions” which is recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)). This limitation merely requires using the generated reward functions to perform an unspecified action that maximizes a team benefit, without reciting a particular manner of performing the action or otherwise meaningfully limiting application of the judicial exception.
In accordance with Step 2A, Prong 2, the claim does not include any additional elements, and the judicial exception is not integrated into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional element of “performing an action that maximizes a team benefit based on the one or more reward functions” which is recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components (See MPEP 2106.05(f)).
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Regarding claim 9
Step 2A Prong 1
Claim 9 does not recite any additional abstract idea but is directed to the abstract idea identified in its parents claim(s).
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “wherein the team comprises a hybrid team that includes one or more human agents and one or more Artificial Intelligence agents" which is recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)). The limitation does not meaningfully alter how the mathematical concepts are performed or require the mathematical results to be applied to achieve a technological improvement.
Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, the additional element of a “wherein the team comprises a hybrid team that includes one or more human agents and one or more Artificial Intelligence agents” which is recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)).
Accordingly, at Step 2B, the additional element individually or in combination does not amount to significantly more than the judicial exception.
Regarding claim 10
Step 2A Prong 1
Claim 10 does not recite any additional abstract idea but is directed to the abstract idea identified in its parents claim(s).
Accordingly, at Step 2A, prong one, the claim recites an abstract idea.
Step 2A Prong 2
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of “wherein the generated one or more reward functions represent one or more motivations, intentions, goals of the one or more agents" which is recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)). The limitation does not meaningfully alter how the mathematical concepts are performed or require the mathematical results to be applied to achieve a technological improvement.
Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application.
Step 2B
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, the additional element of a “wherein the generated one or more reward functions represent one or more motivations, intentions, goals of the one or more agents” which is recited at a high-level of generality such that they amount to no more than generally linking the use of abstract idea to a particular technological environment or field of use using a generic computer component (See MPEP 2106.05(h)).
Accordingly, at Step 2B, the additional element individually or in combination does not amount to significantly more than the judicial exception.
Regarding claims 11-19
Claims 11-19 recites a computing system. Each of these claims corresponds to the method steps of claims 1-9, respectively, with the addition of generic hardware components such as processing circuitry which are insufficient to render the claims subject matter eligible for the same reasons as described above.
Regarding claim 20
Claim 20 recites a non-transitory computer-readable storage media. Each of these claims corresponds to the method steps of claim 1, respectively, with the addition of generic hardware components such as a processor which are insufficient to render the claims subject matter eligible for the same reasons as described above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-4, 11-14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. provided in IDS filed 09 September 2024, referred to as Fu, in view of Baker et al. (“Goal Inference as Inverse Planning”, referred to as Baker), in view of Ramachandran et al. provided in IDS filed 09 September 2024, referred to as Ramachandran.
Regarding claim 1, Fu teaches A method for team modeling, the method comprising:
obtaining data indicating a plurality of trajectories representing a behavior of a team comprising a plurality of agents (Pages 930-931 Section 3.3 and Algorithm1: Describing demonstration samples representing expert behavior in a multi-agent Markov game having N players, wherein demonstration samples correspond to behavioral rollouts/trajectories of the joint multi-agent behavior.);
generating, based on the data indicating the plurality of trajectories, a probability distribution of each agent of the plurality of agents …, wherein the probability distribution of each agent describes a behavior of the agent (Pages 930-931 Section 3.3 and Algorithm1: Describes that each player’s policy/strategy πi(ai) as a distribution over the player’s action set and teaches estimating the respective expert policies πE1:N form demonstration samples, wherein πE represents the expert behavior.);
…based on one or more observed joint actions performed by the team (Fu teaches, (as discussed above) observed multi-agent behavior resulting from the actions of a plurality of agents participating in a Markov game.;)
Although Fu teaches obtaining data indicating a plurality of trajectories representing a behavior of a team comprising a plurality of agents. It does not teach obtaining a plurality of baseline profiles, wherein each of the plurality of baseline profiles encodes at least one of a preference and/or a goal that is relevant to a task performed by the team.
Baker teaches obtaining a plurality of baseline profiles, wherein each of the plurality of baseline profiles encodes at least one of a preference and/or a goal that is relevant to a task performed by the team (Pages 779-781 Introduction, Model 1, and Model 2: Describes maintaining prior knowledge of a space/set of possible goals for an agent and representing alternative foal structure/hypotheses for characterizing the agent, wherein the candidate representation encode respective goals used in determining the agent’s goal dependent behavior);
… over the plurality of baseline profiles…(Baker Pages 779-781 Introduction, Model 1, and Model 2: describes a Bayesian inference of an agent’s goal based on an observed state sequence, s1:T , expressed as P(g| s1:T,w), wherein the possible goas correspond to the baseline profiles and the resulting distribution characterizes the agent’s observed in terms of the likelihood of the respective possible goals.; As discussed above, Fu teaches the plurality of agents and observed multi-agent demonstration data. );
updating, …, the corresponding probability distribution of each agent of the plurality of agents (Baker Page 781 Model 3: Describes updating corresponding probability distribution of an agent over possible goals based on subsequently observed behavior. It recursively determines a posterior probability distribution over goals based on an observed state sequence. It would have been obvious to apply Baker’s Bayesian updating technique to the respective agents observed in Fu’s multi-agent system such that each agent’s probability distribution is updated as additional joint team behavior is observed.)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the multi-agent behavior of Fu with Bakers goal inference. Doing so would have enabled the system to provide additional information regarding the objectives underlying the demonstrated behaviors for use in determining the agent’s reward functions.
As discussed above, Fu in view of Baker, teaches maintaining corresponding updated probability distributions for the individual agents based on observed team behavior. Fu further teaches generating individual reward/utility functions corresponding to the respective agents that explain or rationalize the observed multi-agent behavior. However, Fu, and Baker, do not teach generating the reward functions based on the updated probability distributions.
Ramachandran teaches generating, based on the updated probability distributions of the plurality of agents, one or more reward functions that explain the observed one or more joint actions performed by the team, wherein each of the one or more reward functions describes the behavior of a corresponding one of the plurality of agents (Pages 2587-2588, Sections 3.1, and 4.1: Describes a Bayesian inverse reinforcement learning in which observations of an agent’s actions are used to update a prior probability distribution to obtain a posterior distribution over possible reward functions, and wherein reward learning is performed based on the posterior distribution to generate an estimate of the reward function, such as by determining the mean, median, or maximum a posteriori estimate of the posterior distribution.)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the system of Fu, in view of Baker, with the reward learning technique of Ramachandran. Doing so would have enabled the system to reduce ambiguity in the inverse learning process and improve the accuracy and reliability of the inferred reward functions.
Regarding claim 2, Fu in view of Baker in view of Ramachandran teaches the method of claim 1.
Fu further teaches wherein each of the plurality of trajectories is represented as a sequence of state-joint action pairs over time (Fu page 930 Section 3.3: Defines an N player Markov game in which, at each time t, the system in in a state st and the N agents perform a joint action, with the transition function depending on the actions of all agents. It describes behavior rollouts or trajectories through the Markov game. The trajectories comprise successive state and corresponding joint action pairs over time.)
Regarding claim 3, Fu in view of Baker in view of Ramachandran teaches the method of claim 1 wherein generating one or more reward functions that explain the observed one or more joint actions performed by the team comprises:
Fu in view of Baker in view of Ramachandran teaches generating the one or more reward functions based on the obtained sequence of state-joint action pairs and observed environment dynamics (Fu pages 930-931 Sections 3.3 and 4: Teaches representing the demonstrated multi-agent behavior using successive states and joint actions of the plurality of agents within a Markov game and generating reward/utility functions for the respective agents based on the demonstrated behavior and the Markov game environment, including transition dynamics thereof.; Ramachandran (Pages 2586-2588, Sections 1, 2 and 3.1) teaches that inverse reinforcement learning determines an unknown reward function based on measurements of an agent’s behavior over time together with the dynamics/model of the environment, wherein the environment is represented by an MDP having a state transition function T.).
Regarding claim 4, Fu in view of Baker in view of Ramachandran teaches the method of claim 3 wherein generating one or more reward functions that explain the observed one or more joint actions performed by the team comprises:
Fu in view of Baker in view of Ramachandran teaches generating the one or more reward functions using inverse reinforcement learning method (Fu Page 929 Section 3.2: Describes inverse reinforcement learning as recovering reward functions that explain demonstrated behavior.; Page 931 Section 4 and Algorithm 1L Descriebs reducing the multi-agent inverse-learning problem into respective single agent IRL problems, wherein an IRL method is used to determine the reward function for each player.)
Regarding claims 11-14, which recite substantially the same limitations as claims 1-4, and further recites a computing system for team modeling, the computing system comprising: processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs.) to implement the method steps of claims 1-4 and is rejected for the same reasons as described above.
Regarding claim 20, which recites substantially the same limitations as claim 1 and further recites a non-transitory computer-readable storage media (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs.) to implement the method steps of claim 1 and is rejected for the same reasons as described above.
Claim(s) 5, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. provided in IDS filed 09 September 2024, referred to as Fu, in view of Baker et al. (“Goal Inference as Inverse Planning”, referred to as Baker), in view of Ramachandran et al. provided in IDS filed 09 September 2024, referred to as Ramachandran, in view of Ziebart et al. provided in IDS filed 09 September 2024, referred to as Ziebart.
Regarding claim 5, Fu in view of Baker in view of Ramachandran teaches the method of claim 4 wherein generating one or more reward functions that explain the observed one or more joint actions performed by the team comprises:
Although Fu in view of Baker in view of Ramachandran teaches the method of claim 4. They do not teach generating one or more reward weights that maximize the probability of observing actual trajectories of an agent given the updated probability distributions of the plurality of agents.
Ziebart teaches generating one or more reward weights that maximize the probability of observing actual trajectories of an agent given the updated probability distributions of the plurality of agents (Pages 1433-1435 Background and Learning from Demonstrated Behavior: Describes a reward function parametrized by reward weights ϑ, wherein demonstrated agent behavior is represented by trajectories ζ and teaches determining the reward weights according to argmax θ X examples logP(˜ζ|θ,T) to maximize the likelihood of the demonstrated trajectories.; As discussed above, Fu in view of Backer in view of Ramachandran provides the updated probability distributions associated with the respective agents and generates reward functions using those probabilistic inferences. Ziebart further provides the known MaxEnt IRL techniques for generating the reward weights of those reward functions by maximizing the likelihood of the observed trajectories.)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the system of Fu, in view of Baker, in view of Ramachandran, with the multiple reward functions of Ziebart. Doing so would have enabled the system to improve the reward learning process by selecting reward weights that maximize the likelihood of the observed trajectories.
Regarding claims 15, which recite substantially the same limitations as claim 5, and further recites a computing system for team modeling, the computing system comprising: processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs.) to implement the method steps of claim 5 and is rejected for the same reasons as described above.
Claim(s) 6, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. provided in IDS filed 09 September 2024, referred to as Fu, in view of Baker et al. (“Goal Inference as Inverse Planning”, referred to as Baker), in view of Ramachandran et al. provided in IDS filed 09 September 2024, referred to as Ramachandran, in view of Nair, Ranjit, and Milind Tambe. ("Hybrid BDI-POMDP framework for multiagent teaming.", referred to as Nair).
Regarding claim 6, Fu in view of Baker in view of Ramachandran teaches the method of claim 1.
Although Fu in view of Baker in view of Ramachandran teaches the method of claim 1. They do not teach wherein the team operates in a Multiagent Partially Observable Markov Decision Processes environment.
Nair teaches wherein the team operates in a Multiagent Partially Observable Markov Decision Processes environment (Pages 367 Introduction, and Page 377 Section 3.1: Describes a distributed POMDP models for multiagent teams operating in uncertain, partially observable environments, and defines a multiagent team decision problem (MTDP) for a team of n agents having a state space, joint action space, transition function, respective observations of the agents, an observation function, and a joint reward, wherein each agent selects actions based on its observation history.)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the system of Fu, in view of Baker, in view of Ramachandran, with the partially observed multiagent framework of Nair. Doing so would have enabled the system to model and analyze team behavior despite incomplete or imperfect observations of the environment.
Regarding claims 16, which recite substantially the same limitations as claim 6, and further recites a computing system for team modeling, the computing system comprising: processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs.) to implement the method steps of claim 6 and is rejected for the same reasons as described above.
Claim(s) 7, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. provided in IDS filed 09 September 2024, referred to as Fu, in view of Baker et al. (“Goal Inference as Inverse Planning”, referred to as Baker), in view of Ramachandran et al. provided in IDS filed 09 September 2024, referred to as Ramachandran, in view of Lin et al. provided in IDS filed 09 September 2024", referred to as Lin.
Regarding claim 7, Fu in view of Baker in view of Ramachandran teaches the method of claim 1.
Although Fu in view of Baker in view of Ramachandran teaches the method of claim 1. They do not teach wherein generating the one or more reward functions comprises generating the one or more reward functions to achieve decentralized equilibrium.
Lin teaches wherein generating the one or more reward functions comprises generating the one or more reward functions to achieve decentralized equilibrium (Page 5 Sectio IV: Describes a decentralized multi-agent inverse reinforcement learning d-MIRL approach in which the agents of a multiagent system are assumed to reach a Markov Perfect Equilibrium, and reward functions are selected subject to conditions requiring the observed policy to satisfy the agents’ equilibrium/best response relationships.)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the system of Fu, in view of Baker, in view of Ramachandran, with the decentralized equilibrium approach of Lin. Doing so would have enabled the generated reward functions to remain consistent with the agents’ respective best response behavior in the environment, to provide reward estimates that account for the effect each agents’ behavior has on the behavior and resulting payoff of the other agents.
Regarding claims 17, which recite substantially the same limitations as claim 7, and further recites a computing system for team modeling, the computing system comprising: processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs.) to implement the method steps of claim 7 and is rejected for the same reasons as described above.
Claim(s) 8, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. provided in IDS filed 09 September 2024, referred to as Fu, in view of Baker et al. (“Goal Inference as Inverse Planning”, referred to as Baker), in view of Ramachandran et al. provided in IDS filed 09 September 2024, referred to as Ramachandran, in view of Barrett, Samuel, et al. ("Making friends on the fly: Cooperating with new teammates.", referred to as Barrett).
Regarding claim 8, Fu in view of Baker in view of Ramachandran teaches the method of claim 1.
Although Fu in view of Baker in view of Ramachandran teaches the method of claim 1. They do not teach performing an action that maximizes a team benefit based on the one or more reward functions.
Barrett teaches performing an action that maximizes a team benefit based on the one or more reward functions (Page 136-137 Section 2.1.3: Describes that an agent cooperating with teammates can select its actions to maximize the team reward.; Page 142 Section 3.1: Defines that a reward function R(S,a), determines the long term reward Q*(s,a) based thereon, and derives an optimal policy by selecting the action that maximizes Q*(S,A) where it maximizes the expected reward in its teamwork setting corresponds to optimally cooperating with thee teammates to accomplish their shared goals.)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the system of Fu, in view of Baker, in view of Ramachandran, with the cooperative action selection of Barrett. Doing so would have enabled the system to use inferred objectives of the agents not only to characterize their observed behavior, and to also select subsequent actions that improve cooperative team performance toward the agents shared goals.
Regarding claims 18, which recite substantially the same limitations as claim 8, and further recites a computing system for team modeling, the computing system comprising: processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs.) to implement the method steps of claim 8 and is rejected for the same reasons as described above.
Claim(s) 9, 10 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. provided in IDS filed 09 September 2024, referred to as Fu, in view of Baker et al. (“Goal Inference as Inverse Planning”, referred to as Baker), in view of Ramachandran et al. provided in IDS filed 09 September 2024, referred to as Ramachandran, in view of Schwartz, Tim, et al. ("Hybrid teams: flexible collaboration between humans, robots and virtual agents.", referred to as Schwartz).
Regarding claim 9, Fu in view of Baker in view of Ramachandran teaches the method of claim 1.
Although Fu in view of Baker in view of Ramachandran teaches the method of claim 1. They do not teach wherein the team comprises a hybrid team that includes one or more human agents and one or more Artificial Intelligence agents.
Schwartz teaches wherein the team comprises a hybrid team that includes one or more human agents and one or more Artificial Intelligence agents (Pages 44-50 Section 3: Describes collaboration between human and artificial intelligence agents and teaches a Hybrid Team comprising humans, robots, virtual characters, and softbots)
It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined the system of Fu, in view of Baker, in view of Ramachandran, with the hybrid team of Schwartz. Doing so would have enabled the system to exploit the complementary capabilities of heterogeneous team members, to increase flexibility and robustness of team operation.
Regarding claim 10, Fu in view of Baker in view of Ramachandran in view of Schwartz teaches the method of claim 9.
Fu in view of Baker in view of Ramachandran in view of Schwartz teaches wherein the generated one or more reward functions represent one or more motivations, intentions, goals of the one or more agents (Baker Pages 779-780 Introduction and Inverse Planning Framework: Describes an inverse planning framework which intentions cause agent behavior and where goals cause behavior through probabilistic planning in Markov Decision Processes having goal dependent reward functions, where the reward/cost function associated with an agent is dependent upon the goal pursued by that agent; Ramachandran
Regarding claims 19, which recite substantially the same limitations as claim 9, and further recites a computing system for team modeling, the computing system comprising: processing circuitry in communication with storage media, the processing circuitry configured to execute a machine learning system (Fu’s paper is directed to computational; methodology using conventional computer hardware such as, processors, GPU’s, CPU’s, and storage devices to execute software programs to implement the method steps of claim 9 and is rejected for the same reasons as described above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See attached PTO-892 for additional art including.
US 20190349251 A1: human and AI agent hybrid teams
US 20220129695 A1: reward functions and agents operating behavior
US 20230042431 A1: observed agent trajectories
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DONALD T RODEN whose telephone number is (571)272-6441. The examiner can normally be reached Mon-Thur 8:00-5:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571) 272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/D.T.R./Examiner, Art Unit 2128
/OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128