Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 7/26/2024 was filed after the mailing date of the first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Status of Claims
The present application is being examined under the claims filed on 1/25/2024.
Claims 1-20 are rejected under 35 U.S.C. 101
Claims 1-6 and 8-20 are rejected under 35 U.S.C. 103
Claims 8-11 are rejected under 35 U.S.C. 112(b)
Drawings are object to
Specification is objected to
Drawings
The drawings are objected to because
In Figure 2, reference character 202, “SELEECTION” should read “SELECTION”
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
In paragraph 87, “observation 206” should read “observation 114”
In paragraph 102, “agents 110-A through 110-N” should read “agents 104-A through 104-N”
In paragraph 107, “target agent returns 510” should read “target agent returns 508”
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 8-11 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “best” in claims 8-11 is a relative term which renders the claim indefinite. The term “best” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The specification merely mentions the term “best” (specification [0014] “the aggregate strategy assignment embedding is implemented by a best response action selection neural network that is conditioned on the aggregate strategy assignment embedding”) without providing a threshold or degree, hence rendering the claim language indefinite.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Claim 1:
Step 1: Claim 1 is a method type claim. Therefore, Claims 1-16 are directed to either a process, machine, manufacture, or composition of matter.
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the "Mathematical Concepts" grouping of abstract ideas.
selecting a target action selection policy from the population of action selection policies (mental process – selecting a target action selection policy from the population of action selection policies may be performed manually by a user with the aid of pen and paper by observing/analyzing a set of policies in a population of policies and accordingly using judgement/evaluation to select a target action selection policy using the received input (specification [0096] “As part of selecting the target action selection policy for the time step, the system can obtain the strategy embedding that corresponds to the target policy” ) )
selecting an action to be performed by the agent at the time step using the action selection output (mental process – selecting an action to be performed from the action selection output may be performed manually by a user with the aid of pen and paper by observing/analyzing actions in the action selection output. For example, if the action selection output is a probability distribution over the set of possible actions (specification [0098] “the action selection output can characterize a probability distribution over a set of actions that the agent can perform”), a user can select the action to be performed by the agent at the time step by selecting the action with the highest probability by observing/analyzing the probability distribution)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
A method performed by one or more computers (recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components)
controlling an agent interacting with an environment […] (adding field of use to the judicial exception – see MPEP 2106.05(h))
using a population of action selection policies that are jointly represented by a population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more)
obtaining an observation characterizing a current state of the environment at the time step (adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy using the population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of using a machine learning model with previously determined data without significantly more)
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
A method performed by one or more computers (mere instructions to apply the exception using generic computer components cannot provide an inventive concept)
controlling an agent interacting with an environment […] (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that an agent is interacting with an environment does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
using a population of action selection policies that are jointly represented by a population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
obtaining an observation characterizing a current state of the environment at the time step (MPEP 2106.05(d)(II) indicates that merely "Receiving or transmitting data over a network" is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as is observing in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer))
processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy using the population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without
significantly more. This rejection applies equally to dependent claims 2-16. The additional limitations of
the dependent claims are addressed below.
Regarding Claim 2:
2A Prong 1:
determining a set of payoff values, wherein each payoff value characterizes a return received as a result of controlling each agent using a respective action selection policy for the agent (mathematical process – determining a set of payoff values characterizing a return received as a result of controlling each agent using a respective action selection policy for the agent may be performed by a mathematical process, for example, an error calculation can be used to determine payoff values of each policy being used to control agents (specification [0162] - [0163] “For example, the expected error can be an expectation of an error, Δ, over a state visitation distribution, pi,j, (e.g., a distribution that determines a probability of each policy being used to control agents of the pair) following:
PNG
media_image1.png
40
282
media_image1.png
Greyscale
”))
processing the set of payoff values to generate a probability distribution over a strategy assignment space […] (mental process – processing the set of payoff values to generate a probability distribution over a strategy assignment space may be performed mentally with the aid of pen and paper, for example, creating a probability distribution using the received input of payoff values (specification [0112] “Each payoff value can characterize a return received as a result of controlling each agent using a respective action selection policy for the agent”))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the population action selection neural network has been trained by operations comprising, at each of a plurality of update iterations (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
wherein the population of action selection policies comprises, for each agent in the collection of agents, a set of action selection policies for the agent that each define a respective policy for selecting actions to be performed by the agent to interact with the environment (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the population of action selection policies comprises a set of action selection policies for each agent for selecting actions to be performed by the agent does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
wherein the agent is one agent in a collection of agents (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the agent is one in a collection of agents does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
and training the population action selection neural network based on the probability distribution over the strategy assignment space (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 3:
2A Prong 1:
selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space (mental process – selecting one or more points from the strategy assignment space using the probability distribution may be performed manually by a user with the aid of pen and paper by observing the probability distribution over the strategy assignment space and selecting the points with the highest probabilities (specification [0010] “selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space includes selecting one or more points in the strategy assignment space having highest probabilities under the probability distribution over the strategy assignment space”))
generating an aggregate strategy assignment embedding of the points selected from the strategy assignment space (mental process – generating an aggregate strategy assignment embedding of the selected points may be performed manually by a user with the aid of pen and paper by receiving input of the selected points from the strategy assignment space and writing out a vector representing the selected points (specification [0006] “an “embedding” can refer to an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values”))
generating a plurality of trajectories representing interaction of the collection of agents with the environment as the target agent is controlled by an action selection policy associated with the aggregate strategy assignment embedding (mathematical process – generating a plurality of trajectories representing interaction of the collection of agents with the environment may be performed by a mathematical process, for example, by selecting pairs of action selection policies from a probability distribution (specification [0167] and [0170] “the pre-defined distribution can be a fictitious play distribution, in which
PNG
media_image2.png
36
106
media_image2.png
Greyscale
PNG
media_image3.png
27
164
media_image3.png
Greyscale
”
“The system can generate trajectories for the update iteration by selecting pairs of action selection policies according to the probability distribution (step 806)… For example, the system can generate a trajectory for the i-th action selection policy, vi, by controlling the first agent using vi and by selecting the j -th action selection policy, vj, to control the second agent with probability Σi,j”)
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
Additional elements:
training the population action selection neural network based on the plurality of trajectories (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 4:
2A Prong 1:
selecting one or more points in the strategy assignment space having highest probabilities under the probability distribution over the strategy assignment space (mental process – selecting one or more points from the strategy assignment space having highest probabilities under the probability distribution over the strategy assignment space may be performed manually by a user with the aid of pen and paper by observing the probability distribution and selection the points with the highest probabilities)
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
There are no additional elements.
Regarding Claim 5:
2A Prong 1:
sampling one or more points from the strategy assignment space in accordance with the probability distribution over the strategy assignment space (mental process – sampling one or more points from the strategy assignment space in accordance with the probability distribution may be performed manually by a user by observing the probability distribution and randomly sampling points from the distribution)
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
There are no additional elements.
Regarding Claim 6:
2A Prong 1:
determining, for each of the points selected from the strategy assignment space, a respective strategy assignment embedding for the point based on the respective strategy embedding of each action selection policy specified by the point in the strategy assignment space other than the action selection policy specified for the target agent (mental process – determining a respective strategy assignment embedding for each selected point based on the respective strategy embedding of each action selection policy specified by the point in the strategy assignment space, other than the policy specified by the target agent may be performed manually by a user by, for example, by receiving an input of the selected points and identifying the corresponding strategy embeddings for each point from their respective strategy embedding of each action selection policy specified by the point in the space)
generating the aggregate strategy assignment embedding based on the respective strategy assignment embedding for each of the points selected from the strategy assignment space (mental process – generating the aggregate strategy assignment embedding based on the strategy assignment embedding for each of the selected points may be performed manually by a user, for example, by receiving input of the selected points from the strategy assignment space and writing out a vector representing the selected points (specification [0006] “an “embedding” can refer to an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values”) based on each point’s corresponding strategy assignment embedding)
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
There are no additional elements.
Regarding Claim 7:
2A Prong 1:
generating the aggregate strategy assignment embedding as a linear combination of the respective strategy assignment embedding for each of the points selected from the strategy assignment space (mathematical process – generating the aggregate strategy assignment embedding as a linear combination of the respective strategy assignment embedding for each of the selected points from the assignment space may be performed by a mathematical process (specification [0134] “the system can generate the aggregate strategy assignment embedding as a linear combination of the respective strategy assignment embedding for each of the selected points that combines the embeddings of the selected points by probabilities of the points under the probability distribution over the strategy assignment space…the system can generate the marginal strategy embedding v¬i, for the i-th agent following:
PNG
media_image4.png
68
176
media_image4.png
Greyscale
”))
wherein for each of the points selected from the strategy assignment space, the strategy assignment embedding for the point is scaled by a probability of the point under the probability distribution over the strategy assignment space (mathematical process – scaling each selected point by a probability of the point under the probability distribution over the strategy assignment space may be performed by a mathematical process utilizing a mathematical equation/algorithm for scaling an embedding (specification [0134] “the system can generate the aggregate strategy assignment embedding as a linear combination of the respective strategy assignment embedding for each of the selected points that combines the embeddings of the selected points by probabilities of the points under the probability distribution over the strategy assignment space, p(a)…the system can generate the marginal strategy embedding v¬i, for the i-th agent following:
PNG
media_image4.png
68
176
media_image4.png
Greyscale
Where aj is the j-th selected point in the strategy assignment space and v(aj) is the strategy embedding for the j-th selected point in the strategy assignment space”) by a probability (i.e. p(a))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
There are no additional elements.
Regarding Claim 8:
2A Prong 1:
See the rejection of Claim 3 above, which Claim 8 depends on.
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
wherein the action selection policy associated with the aggregate strategy assignment embedding is implemented by a best response action selection neural network that is conditioned on the aggregate strategy assignment embedding (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the aggregate strategy assignment embedding is implemented by a best response action selection neural network does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 9:
2A Prong 1:
process the observation and the aggregate strategy assignment embedding, in accordance with values of a set of neural network parameters, to generate an action selection output […] (mental/mathematical process – processing the observation and aggregate strategy assignment embedding, in accordance with parameters, to generate an action selection may be performed by a mathematical process performed by a user where they make an observation of the environment, receive input of the aggregate strategy assignment embedding, and generate a representation of a probability distribution (specification [0092] “As another example, the conditional policy neural network 212 can model a conditional probability distribution of actions given the observation 114 and the strategy embedding 116 and the action selection output 202 can characterize the conditional distribution”))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
receive an observation characterizing a state of the environment (MPEP 2106.05(d)(II) indicates that merely "Receiving or transmitting data over a network" is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as is observing in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 10:
2A Prong 1:
See the rejection of Claim 8 above, which Claim 10 depends on.
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
conditioning the best response action selection neural network on the aggregate strategy assignment embedding (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
training the best response action selection neural network on the plurality of trajectories using a reinforcement learning technique (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
training the population action selection neural network using the best response action selection neural network (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 11:
2A Prong 1:
[…] measures an error between: (i) action selection outputs generated by the population action selection neural network, and (ii) action selection outputs generated by the best response action selection neural network (mathematical process – measuring the error between the action selection outputs generated by the population action selection neural network and the best response action selection neural network may be performed by a mathematical process, for example, the error between the two outputs can be calculated using the Kullback-Leibler divergence (specification [0148] “the distillation loss can be the Kullback-Leibler divergence:
PNG
media_image5.png
34
200
media_image5.png
Greyscale
))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
conditioning the population action selection neural network on a strategy embedding corresponding to an action selection policy of the target agent (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
training the population action selection neural network to optimize a distillation loss […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 12:
2A Prong 1:
See the rejection of Claim 11 above, which Claim 12 depends on.
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
training the strategy embedding corresponding to the action selection policy of the target agent, comprising backpropagating gradients of the distillation loss through the population action selection neural network and into the strategy embedding corresponding to the action selection policy of the target agent (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 13:
2A Prong 1:
See the rejection of Claim 3 above, which Claim 13 depends on.
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
controlling each agent other than the target agent using the population action selection neural network (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that each agent other than the target agent is controlled using the population action selection neural network does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 14:
2A Prong 1:
[…] measures an error between: (i) action selection outputs generated by the population action selection neural network by processing observations from the trajectories, and (ii) action selection outputs generated by a baseline population action selection neural network by processing observations from the trajectories (mathematical process – measuring the error between the action selection outputs generated by the population action selection neural network and the baseline population action selection neural network may be performed by a mathematical process, for example, the error between the two outputs can be calculated using the sum of Kullback-Leibler divergences (specification [0151] “the regularization loss can be a sum of the Kullback-Leibler divergences:
PNG
media_image6.png
42
188
media_image6.png
Greyscale
))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
training the population action selection neural network to optimize a regularization loss […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
wherein the baseline population action selection neural network is a static, lagging copy of the population action selection neural network (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the baseline population action selection neural network is a static, lagging copy of the population action selection neural network does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 15:
2A Prong 1:
[…] a predicted return that is predicted to result from controlling each agent using the corresponding action selection policy (mathematical process – using a payoff prediction model to generate a predicted return for each agent being controlled by the corresponding action selection policy may be performed by a mathematical process, for example, an error calculation can be used to determine payoff values of each policy being used to control agents (specification [0162] - [0163] “For example, the expected error can be an expectation of an error, Δ, over a state visitation distribution, pi,j, (e.g., a distribution that determines a probability of each policy being used to control agents of the pair) following:
PNG
media_image1.png
40
282
media_image1.png
Greyscale
”))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
processing an input that identifies a respective action selection policy for each agent in the collection of agents using a payoff prediction model to generate […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 16:
2A Prong 1:
determining that a termination criterion for the update iteration is not satisfied (mental process – determining that a termination criterion for the update iteration is not satisfied may be performed manually by a user with the aid of pen and paper, for example, by identifying if a pre-determined number of training epochs for the update iteration have not been met (specification [0124] “the termination criterion can be satisfied after a pre-determined number of training epochs for the update iteration”) )
determining, for each of multiple strategy assignments, a delta between: (i) a current payoff value for the strategy assignment, and (ii) a previous payoff value for the strategy assignment […] (mathematical process – determining, for each of multiple strategy assignments, a delta between a current payoff value and a previous payoff value for the strategy assignment may be performed by a mathematical process, for example, calculating the difference between the two payoff values)
determining that the termination criterion for the update iteration is not satisfied based on the deltas (mental process - determining that the termination criterion for the update iteration is not satisfied based on the deltas may be performed manually by a user with the aid of pen and paper by observing the received input of the delta values and termination criterion and identifying if the pre-determined termination criterion has been met)
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
in response, further training the population action selection neural network before starting a next update iteration (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 17:
Step 1: Claim 1 is a system type claim. Therefore, Claim 17 is directed to either a process, machine, manufacture, or composition of matter.
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the "Mathematical Concepts" grouping of abstract ideas.
selecting a target action selection policy from the population of action selection policies (mental process – selecting a target action selection policy from the population of action selection policies may be performed manually by a user with the aid of pen and paper by observing/analyzing a set of policies in a population of policies and accordingly using judgement/evaluation to select a target action selection policy using the received input (specification [0096] “As part of selecting the target action selection policy for the time step, the system can obtain the strategy embedding that corresponds to the target policy” )
selecting an action to be performed by the agent at the time step using the action selection output (mental process – selecting an action to be performed from the action selection output may be performed manually by a user with the aid of pen and paper by observing/analyzing actions in the action selection output. For example, if the action selection output is a probability distribution over the set of possible actions (specification [0098] “the action selection output can characterize a probability distribution over a set of actions that the agent can perform”), a user can select the action to be performed by the agent at the time step by selecting the action with the highest probability by observing/analyzing the probability distribution – See MPEP 2106.04(a)(2)(III)(C))
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
one or more computers (recited at a high level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components)
one or more storage devices communicatively coupled to the one or more computers […] (recited at a high level of generality such that they amount to no more than mere instructions to apply the exception using generic computer components)
[…] wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more)
controlling an agent interacting with an environment […] (adding field of use to the judicial exception – see MPEP 2106.05(h))
using a population of action selection policies that are jointly represented by a population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more)
obtaining an observation characterizing a current state of the environment at the time step (adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy using the population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of using a machine learning model with previously determined data without significantly more)
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
one or more computers (mere instructions to apply the exception using generic computer components cannot provide an inventive concept)
one or more storage devices communicatively coupled to the one or more computers […] (mere instructions to apply the exception using generic computer components cannot provide an inventive concept)
[…] wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
controlling an agent interacting with an environment […] (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that an agent is interacting with an environment does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
using a population of action selection policies that are jointly represented by a population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
obtaining an observation characterizing a current state of the environment at the time step (MPEP 2106.05(d)(II) indicates that merely "Receiving or transmitting data over a network" is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as is observing in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer))
processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy using the population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
For the reasons above, Claim 17 is rejected as being directed to an abstract idea without
significantly more.
Regarding Claim 18:
Step 1: Claim 18 is a system type claim. Therefore, Claims 18-20 are directed to either a process, machine, manufacture, or composition of matter.
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the "Mathematical Concepts" grouping of abstract ideas.
selecting a target action selection policy from the population of action selection policies (mental process – selecting a target action selection policy from the population of action selection policies may be performed manually by a user with the aid of pen and paper by observing/analyzing a set of policies in a population of policies and accordingly using judgement/evaluation to select a target action selection policy using the received input (specification [0096] “As part of selecting the target action selection policy for the time step, the system can obtain the strategy embedding that corresponds to the target policy” )
selecting an action to be performed by the agent at the time step using the action selection output (mental process – selecting an action to be performed from the action selection output may be performed manually by a user with the aid of pen and paper by observing/analyzing actions in the action selection output. For example, if the action selection output is a probability distribution over the set of possible actions (specification [0098] “the action selection output can characterize a probability distribution over a set of actions that the agent can perform”), a user can select the action to be performed by the agent at the time step by selecting the action with the highest probability by observing/analyzing the probability distribution)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations, the operations comprising (recited at a high level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components)
controlling an agent interacting with an environment […] (adding field of use to the judicial exception – see MPEP 2106.05(h))
using a population of action selection policies that are jointly represented by a population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more)
obtaining an observation characterizing a current state of the environment at the time step (adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy using the population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of using a machine learning model with previously determined data without significantly more)
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations, the operations comprising (mere instructions to apply the exception using generic computer components cannot provide an inventive concept)
controlling an agent interacting with an environment […] (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that an agent is interacting with an environment does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
using a population of action selection policies that are jointly represented by a population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
obtaining an observation characterizing a current state of the environment at the time step (MPEP 2106.05(d)(II) indicates that merely "Receiving or transmitting data over a network" is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as is observing in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer))
processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy using the population action selection neural network […] (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f )- Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
For the reasons above, Claim 18 is rejected as being directed to an abstract idea without
significantly more. This rejection applies equally to dependent claims 19-20. The additional limitations of
the dependent claims are addressed below.
Regarding Claim 19:
2A Prong 1:
determining a set of payoff values, wherein each payoff value characterizes a return received as a result of controlling each agent using a respective action selection policy for the agent (mathematical process – determining a set of payoff values characterizing a return received as a result of controlling each agent using a respective action selection policy for the agent may be performed by a mathematical process, for example, an error calculation can be used to determine payoff values of each policy being used to control agents (specification [0162] - [0163] “For example, the expected error can be an expectation of an error, Δ, over a state visitation distribution, pi,j, (e.g., a distribution that determines a probability of each policy being used to control agents of the pair) following:
PNG
media_image1.png
40
282
media_image1.png
Greyscale
”))
processing the set of payoff values to generate a probability distribution over a strategy assignment space […] (mental process – processing the set of payoff values to generate a probability distribution over a strategy assignment space may be performed mentally with the aid of pen and paper, for example, creating a probability distribution using the received input of payoff values (specification [0112] “Each payoff value can characterize a return received as a result of controlling each agent using a respective action selection policy for the agent”))
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the population action selection neural network has been trained by operations comprising, at each of a plurality of update iterations (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
wherein the population of action selection policies comprises, for each agent in the collection of agents, a set of action selection policies for the agent that each define a respective policy for selecting actions to be performed by the agent to interact with the environment (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the population of action selection policies comprises a set of action selection policies for each agent for selecting actions to be performed by the agent does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
wherein the agent is one agent in a collection of agents (field of use - limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the agent is one in a collection of agents does not integrate the exception into a practical application nor amount to significantly more - See MPEP 2106.05(h))
and training the population action selection neural network based on the probability distribution over the strategy assignment space (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 18.
Regarding Claim 20:
2A Prong 1:
selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space (mental process – selecting one or more points from the strategy assignment space using the probability distribution may be performed manually by a user with the aid of pen and paper by observing the probability distribution over the strategy assignment space and selecting the points with the highest probabilities (specification [0010] “selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space includes selecting one or more points in the strategy assignment space having highest probabilities under the probability distribution over the strategy assignment space”))
generating an aggregate strategy assignment embedding of the points selected from the strategy assignment space (mental process – generating an aggregate strategy assignment embedding of the selected points may be performed manually by a user with the aid of pen and paper by receiving input of the selected points from the strategy assignment space and writing out a vector representing the selected points (specification [0006] “an “embedding” can refer to an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values”))
generating a plurality of trajectories representing interaction of the collection of agents with the environment as the target agent is controlled by an action selection policy associated with the aggregate strategy assignment embedding (mathematical process – generating a plurality of trajectories representing interaction of the collection of agents with the environment may be performed by a mathematical process, for example, by selecting pairs of action selection policies from a probability distribution (specification [0167] and [0170] “the pre-defined distribution can be a fictitious play distribution, in which
PNG
media_image2.png
36
106
media_image2.png
Greyscale
PNG
media_image3.png
27
164
media_image3.png
Greyscale
”
“The system can generate trajectories for the update iteration by selecting pairs of action selection policies according to the probability distribution (step 806)… For example, the system can generate a trajectory for the i-th action selection policy, vi, by controlling the first agent using vi and by selecting the j -th action selection policy, vj, to control the second agent with probability Σi,j”)
2A Prong 2 & Step 2B: This judicial exception is not integrated into a practical application.
Additional elements:
training the population action selection neural network based on the plurality of trajectories (adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) - Examiner's note: high level recitation of applying a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 18.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Silver et al. (US 20200244707 A1, herein Silver), in view of Grover et al. (“Learning Policy Representations in Multiagent Systems”, 31 Jul 2018, herein Grover).
Regarding Claim 1:
Silver teaches:
“A method performed by one or more computers, the method comprising:”
(Silver, Paragraph 72, "the process will be described as being performed by a system of one or more computers located in one or more locations”; Examiner’s Note: A method performed by one or more computers (i.e. a system of one or more computers), the method comprising is taught)
“controlling an agent interacting with an environment using a population of action selection policies that are jointly represented by a population action selection neural network, comprising, at each of a plurality of time steps”
(Silver, Claim 1, “A method of training a policy neural network having a plurality of policy parameters and used to select actions to be performed by an agent to control the agent to perform a particular task while interacting with one or more other agents in an environment, the method comprising: maintaining data specifying a pool of candidate action selection policies”; Examiner’s note: controlling an agent interacting with an environment (i.e. control the agent to perform a particular task…in an environment) using a population of action selection policies that are jointly represented by a population action selection neural network (i.e. pool of candidate action selection policies), comprising, at each of a plurality of time steps it taught.)
“obtaining an observation characterizing a current state of the environment at the time step”
(Silver, paragraph 39, “That is, the reinforcement learning system 100 receives observations, with each observation characterizing a respective state of the environment 104, and, in response to each observation, selects an action from a predetermined set of actions to be performed by the reinforcement learning agent 102A in response to the observation”; Examiner’s note: obtaining an observation characterizing a current state (i.e. respective state) of the environment at the time step is taught.)
“selecting a target action selection policy from the population of action selection policies”
(Silver, paragraph 74, “Specifically, for each of one or more of the leaner policies, the system selects one or more policies (302) from the pool of candidate action selection policies”; Examiner’s note: selecting a target action selection policy (i.e. the system selects one or more policies) from the population of action selection policies (i.e. pool of candidate action selection policies) is taught.)
“and selecting an action to be performed by the agent at the time step using the action selection output”
(Silver, paragraph 42, “The network output includes an action selection output and, in some cases, a predicted expected return output. The action selection output defines an action selection policy for selecting an action to be performed by the agent in response to the input observation”; Examiner’s note: and selecting an action to be performed by the agent at the time step using the action selection output (i.e. using the action selection output) is taught.)
Silver teaches using the population action selection neural network to generate an action selection output. Silver fails to teach processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy, using the population action selection neural network to generate an action selection output.
“processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy, using the population action selection neural network to generate an action selection output”
(Silver, Paragraph 41, “Generally, the policy neural network 110 is a neural network that is configured to receive a network input including an observation and to process the network input in accordance with parameters of the policy neural network (“policy parameters”) to generate a network output”; Examiner’s Note: using the population action selection neural network to generate an action selection output (i.e. generate a network output) is taught)
Grover teaches “Learning Policy Representations in Multiagent Systems (title)” comprising:
“processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy, using the population action selection neural network to generate an action selection output”
(Grover, Section 3.1, “We offset this dichotomy by learning a single conditional policy network. To do so, we first specify a representation function, fθ : Ԑ ―› Rd, with parameters θ, where Ԑ represents the space of episodes. We use this embedding to condition the policy network. Formally, the policy network is denoted by πΦ,θ : S x A x Ԑ ―› [0, 1] and Φ are parameters for the function mapping the agent observation and embedding to a distribution over the agent’s actions.… For every agent, the objective function samples two distinct episodes e1 and e2. The observation and action pairs from e2 are used to learn an embedding fθ(e2) that conditions the policy network trained on observation and action pairs from e1”; Examiner’s Note: processing a network input comprising: (i) the observation, and (ii) a strategy embedding representing the target action selection policy (i.e. embedding fθ(e2)) using the population action selection neural network to generate an action selection output (i.e. mapping the agent observation and embedding to a distribution over the agent’s actions) is taught)
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver by applying the strategy embedding representing the target action selection policy feature as taught in Grover as “such embeddings play the role of privileged information and allow us to train a policy network that uses this information to learn faster and generalize better to opponents or cooperators unseen at training time” (Grover, Section 4.2)
Regarding Claim 17:
Claim 17 recites substantially the same limitations as Claim 1, in the form of a system, therefore, it is rejected under the same rationale.
Silver teaches the additional elements of:
“A system comprising: one or more computers”
(Silver, Paragraph 72, "the process will be described as being performed by a system of one or more computers located in one or more locations”; Examiner’s Note: A system comprising: one or more computers (i.e. a system of one or more computers) is taught)
“and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations, the operations comprising:”
(Silver, Paragraph 91 and 93, “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them”; “A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network”; Examiner’s Note: and one or more storage devices (i.e. a machine-readable storage device) communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers (i.e. on one computer or on multiple computers), cause the one or more computers to perform operations (i.e. deployed to be executed), the operations comprising is taught)
Regarding Claim 18:
Claim 18 is a system to perform the method of Claim 1, therefore, it is rejected under the same rationale.
“One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations, the operations comprising:”
(Silver, Paragraph 91, “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus”; Examiner’s Note: One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations, the operations comprising it taught)
Claims 2, 3, 13, 15, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Silver in view of Grover as applied in claim 1, in view of Lanctot et al. (“A Unified Game-Theoretic Approach Multiagent Reinforcement Learning”, 7 Nov 2017, herein Lanctot).
Regarding Claim 2:
The combination of Silver and Grover teaches:
“The method of claim 1, wherein the agent is one agent in a collection of agents”
(Silver, abstract, “a policy neural network having a plurality of policy parameters and used to select actions to be performed by an agent to control the agent to perform a particular task while interacting with one or more other agents in an environment”; Examiner’s note: The method of claim 1, wherein the agent is one agent in a collection of agents (i.e. agent…interacting with one or more other agents) is taught.)
“wherein the population of action selection policies comprises, for each agent in the collection of agents, a set of action selection policies for the agent that each define a respective policy for selecting actions to be performed by the agent to interact with the environment”
(Silver, paragraph 48, “During the training of the policy neural network 110, the system maintains policy data 140 specifying a pool of candidate action selection policies. The pool of candidate action selection policies includes (i) a plurality of learner polices 142A-M for controlling the agent, each learner policy defined by a respective set of values for the policy parameters of the policy neural network 110, and (ii) one or more fixed policies 152 for controlling the agent”; Examiner’s note: wherein the population of action selection policies comprises, for each agent in the collection of agents, a set of action selection policies for the agent (i.e. policies 152 for controlling the agent) that each define a respective policy for selecting actions to be performed by the agent (i.e. pool of candidate action selection policies) to interact with the environment is taught.)
“and wherein the population action selection neural network has been trained by operations comprising, at each of a plurality of update iterations”
(Silver, paragraph 71, “The system trains the policy neural network (206) using an iterative approach. In other words, the system updates one or more of the learner policies at each of a plurality of training iterations”; Examiner’s note: and wherein the population action selection neural network (i.e. policy neural network) has been trained by operations comprising, at each of a plurality of update iterations (i.e. updates one or more of the learner policies at each of a plurality of training iterations) is taught.)
The combination of Silver and Grover fails to teach determining a set of payoff values, wherein each payoff value characterizes a return received as a result of controlling each agent using a respective action selection policy for the agent, processing the set of payoff values to generate a probability distribution over a strategy assignment space, wherein each point in the strategy assignment space represents an assignment of a respective action selection policy to each agent in the collection of agents, and training the population action selection neural network based on the probability distribution over the strategy assignment space.
Lanctot teaches “A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning (title)” comprising:
“determining a set of payoff values, wherein each payoff value characterizes a return received as a result of controlling each agent using a respective action selection policy for the agent”
(Lanctot, Section 3, “Meta-Strategy solve - We refer to ui(σ) as player i’s expected value given all players’ meta-strategies and the current empirical payoff tensor UΠ (computed via multiple tensor dot products.) Similarly, denote ui(πi,k, σ-i) as the expected utility if player i plays their kth ϵ [[K]] U {0} policy and the other players play with their meta-strategy σ-i”; Examiner’s Note: determining a set of payoff values (i.e. ui(σ) and ui(πi,k, σ-i)), wherein each payoff value characterizes a return received as a result of controlling each agent using a respective action selection policy for the agent (i.e. player i plays their kth ϵ [[K]] U {0} policy…other players play with their meta-strategy) is taught)
“processing the set of payoff values to generate a probability distribution over a strategy assignment space, wherein each point in the strategy assignment space represents an assignment of a respective action selection policy to each agent in the collection of agents”
(Lanctot, Section 3.1, “we introduce a new solver we call projected replicator dynamics (PRD). From Appendix A, when using the asymmetric replicator dynamics, e.g. with two players, where UΠ = (A,B), the change in probabilities for the kth component (i.e., the policy πi,k) of meta-strategies (σ1,σ2) = (x,y) are:
PNG
media_image7.png
50
358
media_image7.png
Greyscale
Examiner’s Note: processing the set of payoff values to generate a probability distribution (i.e. probabilities for the kth component) over a strategy assignment space (i.e. meta-strategies), wherein each point in the strategy assignment space represents an assignment of a respective action selection policy to each agent in the collection of agents is taught)
“and training the population action selection neural network based on the probability distribution over the strategy assignment space.”
(Lanctot, Section 3, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i). At the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ, which takes time exponential in |Π|.”; Examiner’s Note: and training the population action selection neural network (i.e. at the end of the epoch, the new oracles are added to their policy sets) based on the probability distribution over the strategy assignment space (i.e. expected utilities for new policy computed) is taught )
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver and Grover by applying the generating a probability distribution over a strategy assignment space using a set of payoff values, wherein each point in the strategy assignment space represents an assignment of a respective action selection policy to each agent in the collection of agents as taught in Lanctot as it is “particularly important in multiagent settings” that “one must react dynamically based on the observed behavior of others” (Lanctot, Section 1).
Regarding Claim 3:
The combination of Silver, Grover, and Lanctot teaches:
“The method of claim 2, wherein training the population action selection neural network based on the probability distribution over the strategy assignment space comprises, for a target agent”
(Lanctot, Section 3 and 3.2.1, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i). At the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ, which takes time exponential in |Π|”; “For decoupled PRD, we maintain running averages for the overall average value an value of each arm (policy). Unlike in PSRO, in the case of DCH, one sample is obtained at a time and the meta-strategy is updated periodically from online estimates”; Examiner’s Note: wherein training the population action selection neural network (i.e. at the end of the epoch, the new oracles are added to their policy sets Πi) based on the probability distribution over the strategy assignment space (i.e. PRD) comprises, for a target agent (i.e. oracle) is taught)
“selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space”
(Lanctot, Section 3.2.1, “For decoupled PRD, we maintain running averages for the overall average value an value of each arm (policy). Unlike in PSRO, in the case of DCH, one sample is obtained at a time and the meta-strategy is updated periodically from online estimates”; Examiner’s Note: selecting one or more points from the strategy assignment space (i.e. one sample is obtained at a time) using the probability distribution over the strategy assignment space (i.e. PRD) is taught)
“generating an aggregate strategy assignment embedding of the points selected from the strategy assignment space”
(Lanctot, Section 3.1, “A meta-strategy solver takes as input the empirical game (Π,UΠ) and produces a meta-strategy σi for each player i. We try three different solvers: regret-matching, Hedge, and projected replicator dynamics. These specific meta-solvers accumulate values for each policy (“arm”) and an aggregate value based on all players’ meta-strategies.”; Examiner’s Note: generating an aggregate strategy assignment embedding of the points selected from the strategy assignment space (i.e. accumulate values for each policy (“arm”) and an aggregate value based on all players’ meta-strategies) is taught)
“generating a plurality of trajectories representing interaction of the collection of agents with the environment as the target agent is controlled by an action selection policy associated with the aggregate strategy assignment embedding”
(Lanctot, Section 3, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i). At the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ, which takes time exponential in |Π|”; Examiner’s Note: generating a plurality of trajectories representing interaction of the collection of agents with the environment (i.e. a fixed policy is sampled from the opponents’ meta-strategies) as the target agent is controlled by an action selection policy associated with the aggregate strategy assignment embedding (i.e. expected utilities for new policy combinations are computed via simulation) is taught)
“and training the population action selection neural network based on the plurality of trajectories.”
(Lanctot, Section 3, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i). At the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ, which takes time exponential in |Π|”; Examiner’s Note: and training the population action selection neural network based on the plurality of trajectories (i.e. at the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ) is taught)
The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein.
Regarding Claim 13:
The combination of Silver, Grover, and Lanctot teaches:
“The method of claim 3, wherein generating the plurality of trajectories representing interaction of the collection of agents with the environment as the target agent is controlled by the action selection policy associated with the aggregate strategy assignment embedding comprises”
(Lanctot, Section 3, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i)”; Examiner’s Note: wherein generating the plurality of trajectories representing interaction of the collection of agents (i.e. fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i)) with the environment as the target agent is controlled by the action selection policy associated with the aggregate strategy assignment embedding (i.e. one player is set to oracle(learning) mode) comprises is taught)
“controlling each agent other than the target agent using the population action selection neural network”
(Lanctot, Section 3, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i). At the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ, which takes time exponential in |Π|”; Examiner’s Note: controlling each agent other than the target agent (i.e. one player is set to oracle(learning) mode… fixed policy is sampled from the opponents’ meta-strategies) using the population action selection neural network is taught)
The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein.
Regarding Claim 15:
The combination of Silver, Grover, and Lanctot teaches:
“The method of claim 2, wherein determining the set of payoff values comprises, for each payoff value: processing an input that identifies a respective action selection policy for each agent in the collection of agents using a payoff prediction model to generate a predicted return that is predicted to result from controlling each agent using the corresponding action selection policy”
(Lanctot, Section 3.1, “We refer to ui(σ) as player i’s expected value given all players’ meta-strategies and the current empirical payoff tensor UΠ”; Examiner’s Note: wherein determining the set of payoff values comprises, for each payoff value: processing an input that identifies a respective action selection policy for each agent in the collection of agents (i.e. given all players’ meta-strategies) using a payoff prediction model (i.e. current empirical payoff tensor UΠ) to generate a predicted return that is predicted to result from controlling each agent using the corresponding action selection policy (i.e. refer to ui(σ) as player i’s expected value) is taught)
The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein.
Regarding Claim 19:
Claim 19 is a system to perform the method of Claim 2, therefore, it is rejected under the same rationale.
Regarding Claim 20:
Claim 20 is a system to perform the method of Claim 3, therefore, it is rejected under the same rationale.
Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Silver in view of Grover in view of Lanctot as applied in claim 3, in view of Badia et al. (CA 3167201 A1, herein Badia).
Regarding Claim 4:
The combination of Silver, Grover, and Lanctot teaches selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space (claim 3).
The combination of Silver, Grover and Lanctot fails to teach selecting one or more points in the strategy assignment space having highest probabilities under the probability distribution over the strategy assignment space.
Badia teaches “REINFORCEMENT LEARNING WITH ADAPTIVE RETURN COMPUTATION SCHEMES (title)” comprising:
“selecting one or more points in the strategy assignment space having highest probabilities under the probability distribution over the strategy assignment space.”
(Badia, Paragraph 50, “Optionally, the system 100 may select the action 108 to be performed by the agent 104 at the time step using an ϵ-greedy exploration policy in which the system 100 selects the action with the highest final return estimate with probability 1 - ϵ and selecting a random action from the set of actions with probability ϵ”; Examiner’s Note: selecting one or more points in the strategy assignment space having highest probabilities under the probability distribution (i.e. selects the action with the highest final return estimate with probability) over the strategy assignment space is taught)
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver, Grover, and Lanctot by selecting one or more points in the assignment space having highest probabilities as taught in Badia as selecting schemes that are more likely to result in higher extrinsic rewards results “higher quality training data being generated for the action selection neural network(s)” (Badia, Paragraph 121).
Regarding Claim 5:
The combination of Silver, Grover, Lanctot, and Badia teaches:
“The method of claim 3, wherein selecting one or more points from the strategy assignment space using the probability distribution over the strategy assignment space comprises”
(Badia, Paragraph 50, “For example, the system 100 may process the action scores 114 to generate a probability distribution over the set of possible actions, and then select the action 108 to be performed by the agent 104 by sampling an action in accordance with the probability distribution”; Examiner’s Note: wherein selecting one or more points from the strategy assignment space (i.e. select the action 108 to be performed) using the probability distribution (i.e. sampling…the probability distribution) over the strategy assignment space comprises is taught)
“sampling one or more points from the strategy assignment space in accordance with the probability distribution over the strategy assignment space.”
(Badia, Paragraph 50, “For example, the system 100 may process the action scores 114 to generate a probability distribution over the set of possible actions, and then select the action 108 to be performed by the agent 104 by sampling an action in accordance with the probability distribution”; Examiner’s Note: sampling one or more points from the strategy assignment space in accordance with the probability distribution (i.e. select the action…by sampling an action in accordance with the probability distribution) over the strategy assignment space is taught)
The reasons of obviousness have been noted in the rejection of Claim 4 above and applicable herein.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Silver in view of Grover in view of Lanctot as applied in claim 3, in view of Jiang et al. (“Metric Policy Representations for Opponent Modeling”, 21 Jun 2022, herein Jiang).
Regarding Claim 6:
The combination of Silver, Grover, and Lanctot teaches the method of claim 3.
The combination of Silver, Grover, and Lanctot fails to teach the method of claim 3, wherein generating the aggregate strategy assignment embedding of the points selected from the strategy assignment space comprises: determining, for each of the points selected from the strategy assignment space, a respective strategy assignment embedding for the point based on the respective strategy embedding of each action selection policy specified by the point in the strategy assignment space other than the action selection policy specified for the target agent and generating the aggregate strategy assignment embedding based on the respective strategy assignment embedding for each of the points selected from the strategy assignment space.
Jiang teaches “Metric Policy Representations for Opponent Modeling (title)” comprising:
“determining, for each of the points selected from the strategy assignment space, a respective strategy assignment embedding for the point based on the respective strategy embedding of each action selection policy specified by the point in the strategy assignment space other than the action selection policy specified for the target agent”
(Jiang, Section 3.3, “In each episode during training, the ego agent interacts with the other N opponents with policy π1 ∈ Πtrain1, π2 ∈ Πtrain 2 ,…, πN ∈ ΠtrainN . These policies form an opponent joint policy πtraino ∈ Πtraino, on which the ego agent should condition its policy to maximize the expected return”; Examiner’s Note: determining, for each of the points selected from the strategy assignment space, a respective strategy assignment embedding (i.e. πtraino) for the point based on the respective strategy embedding of each action selection policy specified by the point in the strategy assignment space (i.e. π1 ∈ Πtrain1, π2 ∈ Πtrain 2 ,…, πN ∈ ΠtrainN ) other than the action selection policy specified for the target agent (i.e. an opponent joint policy) is taught)
“and generating the aggregate strategy assignment embedding based on the respective strategy assignment embedding for each of the points selected from the strategy assignment space.”
(Jiang, Section 3.2, “For generalization to any possible policies, we consider the “policy space” Πi formed by all possible policies of opponent i in. In each episode, opponent i acts according to a policy πi which is sampled from a distribution Pi over Πi. Therefore, the opponent joint policy πo can be viewed as sampled from the joint distribution P over the opponent joint policy space”; Examiner’s Note: and generating the aggregate strategy assignment embedding (i.e. the opponent joint policy) based on the respective strategy assignment embedding for each of the points selected from the strategy assignment space (i.e. the opponent joint policy πo can be viewed as sampled from the joint distribution P over the opponent joint policy space) is taught)
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver, Grover, and Lanctot by determining, for each of the points selected from the strategy assignment space, a respective strategy assignment embedding for the point and generating the aggregate strategy assignment embedding based on the respective strategy assignment embedding for each of the selected points as taught in Jiang as “in multi-agent reinforcement learning, the inherent non-stationarity of the environment caused by other agents’ actions posed significant difficulties for an agent to learn a good policy independently” and by “one ego agent learns while interacting with other agents” by “sampl[ing] from a set of fixed policies at the beginning of each episode… we expect the ego agent to be immediately generalizable, i.e., to adapt quickly and achieve high performance when facing unseen opponents during execution without updating parameters” (Jiang, Section 1).
Claims 8-12 are rejected under 35 U.S.C. 103 as being unpatentable over Silver in view of Grover in view of Lanctot as applied in claim 3, in view of Czarnecki et al. (US 20190354867 A1, herein Czarnecki).
Regarding Claim 8:
The combination of Silver, Grover, and Lanctot teaches the method of claim 3.
The combination of Silver, Grover, and Lanctot fails to teach the method of claim 3, wherein the action selection policy associated with the aggregate strategy assignment embedding is implemented by a best response action selection neural network that is conditioned on the aggregate strategy assignment embedding.
Czarnecki teaches “REINFORCEMENT LEARNING USING AGENT CURRICULA (title)” comprising:
“The method of claim 3, wherein the action selection policy associated with the aggregate strategy assignment embedding is implemented by a best response action selection neural network that is conditioned on the aggregate strategy assignment embedding”
(Czarnecki, Paragraph 73, “The system trains the action policy neural networks in the set in accordance with the mixing data (step 304). During the training, the system updates the values of the parameters of the policy networks to (1) generate combined action selection policies that result in improved performance on the reinforcement learning task and (2) generate action selection policies that are aligned with other action selection policies generated by the other candidate agent policy neural networks by processing the same training network input. Perform an iteration of training the policy neural networks will be described in more detail below with reference to FIG. 4”; Examiner’s Note: wherein the action selection policy associated with the aggregate strategy assignment embedding (i.e. set in accordance with the mixing data) is implemented by a best response action selection neural network (i.e. generate combined action selection policies that result in improved performance on the reinforcement learning task) that is conditioned (i.e. system updates the values of the parameters of the policy networks) on the aggregate strategy assignment embedding is taught)
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver, Grover, and Lanctot wherein the action selection policy associated with the aggregate strategy assignment embedding is implemented by a best response action selection neural network as taught in Czarnecki “because the weights initially favor the least complex networks and the least complex networks can quickly improve their performance on the reinforcement learning task, the more complex agent policy neural network can initially bootstrap (through the matching updates during training) from solutions found by the simpler networks to assist the more complex networks in learning the tasks” (Czarnecki, Paragraph 68).
Regarding Claim 9:
The combination of Silver, Grover, Lanctot, and Czarnecki teaches:
“The method of claim 8, wherein the best response action selection neural network is configured to, when conditioned on the aggregate strategy assignment embedding”
(Czarnecki, Paragraph 43, “Generally, each action policy neural network in the set receives a network input including an observation and generates a network output that defines an action selection policy for selecting an action to be performed by the agent in response to the observation”; Examiner’s Note: wherein the best response action selection neural network is configured to (i.e. receives…an observation), when conditioned on the aggregate strategy assignment embedding (i.e. a network input including) is taught)
“receive an observation characterizing a state of the environment”
(Czarnecki, Claim 12, “training the plurality of candidate agent policy neural networks jointly to perform the reinforcement learning task, comprising, at each of a plurality of training iterations: obtaining a training network input comprising an observation of the environment”; Examiner’s Note: receive an observation characterizing a state of the environment (i.e. input comprising an observation of the environment) is taught)
“and process the observation and the aggregate strategy assignment embedding, in accordance with values of a set of neural network parameters, to generate an action selection output that characterizes an action to be performed by a corresponding agent in response to the observation.”
(Czarnecki, Paragraph 43, “Generally, each action policy neural network in the set receives a network input including an observation and generates a network output that defines an action selection policy for selecting an action to be performed by the agent in response to the observation”; Examiner’s Note: and process the observation and the aggregate strategy assignment embedding, in accordance with values of a set of neural network parameters, (i.e. receives a network input including an observation) to generate an action selection output that characterizes an action to be performed by a corresponding agent (i.e. and generates a network output that defines an action selection policy for selecting an action to be performed by the agent) in response to the observation is taught)
The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein.
Regarding Claim 10:
The combination of Silver, Grover, Lanctot, and Czarnecki teaches:
“The method of claim 8, wherein training the population action selection neural network based on the plurality of trajectories comprises”
(Czarnecki, Paragraph 7 and 92-93, “The final action policy neural network generally defines the most complex policy of any of the networks in the set, i.e., at least one other action policy neural network in the set defines an action selection policy that is less complex than the policy defined by the final action policy neural network”; “In particular, the system generates training data by causing the agent to act in the environment as described above and then stores the training data in a replay memory. The system then samples training data from the replay memory and uses the sampled training data to train the neural networks”; Examiner’s Note: wherein training the population action selection neural network (i.e. final action policy neural network) based on the plurality of trajectories (i.e. the system generates training data by causing the agent to act in the environment… and then stores the training data in a replay memory… uses the sampled training data to train the neural networks) comprises is taught)
“conditioning the best response action selection neural network on the aggregate strategy assignment embedding”
(Czarnecki, Paragraph 90, “To train the neural network, the system computes gradients of a reinforcement learning loss function, 𝓛RL, that is appropriate for the kinds of network outputs that the policy networks are configured to generate and that encourages the combined policies to show improved performance on the reinforcement learning task”; Examiner’s Note: conditioning the best response action selection neural network (i.e. the policy networks are configured to…encourag[e] the combined policies to show improved performance on the reinforcement learning task) on the aggregate strategy assignment embedding (i.e. combined policies) is taught)
“training the best response action selection neural network on the plurality of trajectories using a reinforcement learning technique”
(Czarnecki, Paragraph 90 and 92-93, “To train the neural network, the system computes gradients of a reinforcement learning loss function, 𝓛RL, that is appropriate for the kinds of network outputs that the policy networks are configured to generate and that encourages the combined policies to show improved performance on the reinforcement learning task”; “In particular, the system generates training data by causing the agent to act in the environment as described above and then stores the training data in a replay memory. The system then samples training data from the replay memory and uses the sampled training data to train the neural networks”; Examiner’s Note: training the best response action selection neural network on the plurality of trajectories (i.e. the system generates training data by causing the agent to act in the environment… and then stores the training data in a replay memory… uses the sampled training data to train the neural networks) using a reinforcement learning technique (i.e. reinforcement learning task) is taught)
“and training the population action selection neural network using the best response action selection neural network.”
(Czarnecki, Paragraph 7, “The final action policy neural network generally defines the most complex policy of any of the networks in the set, i.e., at least one other action policy neural network in the set defines an action selection policy that is less complex than the policy defined by the final action policy neural network”; Examiner’s Note: and training the population action selection neural network (i.e. final action policy neural network) using the best response action selection neural network (i.e. at least one other action policy neural network) is taught)
The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein.
Regarding Claim 11:
The combination of Silver, Grover, Lanctot, and Czarnecki teaches:
“The method of claim 10, wherein training the population action selection neural network using the best response action selection neural network comprises”
(Czarnecki, Paragraph 7, “The final action policy neural network generally defines the most complex policy of any of the networks in the set, i.e., at least one other action policy neural network in the set defines an action selection policy that is less complex than the policy defined by the final action policy neural network”; Examiner’s Note: wherein training the population action selection neural network (i.e. final action policy neural network) using the best response action selection neural network (i.e. at least one other action policy neural network) comprises is taught)
“conditioning the population action selection neural network on a strategy embedding corresponding to an action selection policy of the target agent”
(Czarnecki, Paragraphs 7 and 9, “The system trains the final action policy neural network, i.e., the neural network that will be used to control the reinforcement learning agent after training, as part of a set of candidate agent policy neural networks”; “The system then trains the candidate agent policy neural networks jointly to perform the reinforcement learning task. In particular, during the training, the system uses combined action selection policies that are a combination (in accordance with the weights in the mixing data) of individual action selection policies generated by the candidate networks in the set”; Examiner’s Note: conditioning the population action selection neural network (i.e. the final action policy neural network) on a strategy embedding corresponding to an action selection policy of the target agent (i.e. the system uses combined action selection policies that are a combination (in accordance with the weights in the mixing data) of individual action selection policies) is taught)
“training the population action selection neural network to optimize a distillation loss that measures an error between: (i) action selection outputs generated by the population action selection neural network, and (ii) action selection outputs generated by the best response action selection neural network”
(Czarnecki, Paragraph 97-98, “D is a function that measures the differences between the policy outputs generated by policy networks πi and πj… As a particular example, the function D between a policy network π1 and π2 in the set can satisfy:
PNG
media_image8.png
66
256
media_image8.png
Greyscale
where S is the set of observations, s is a trajectory of observations in the set, |s| is the number of observations in the trajectory, |S| is the number observations in the set, DKL is the K−L divergence”; Examiner’s Note: training the population action selection neural network to optimize a distillation loss (i.e. DKL is the K−L divergence) that measures an error between: (i) action selection outputs generated by the population action selection neural network, and (ii) action selection outputs generated by the best response action selection neural network (i.e. D is a function that measures the differences between the policy outputs generated by policy networks πi and πj) is taught)
The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein.
Regarding Claim 12:
The combination of Silver, Grover, Lanctot, and Czarnecki teaches:
“The method of claim 11, wherein training the population action selection neural network to optimize the distillation loss further comprises”
(Czarnecki, Paragraph 97-98, “D is a function that measures the differences between the policy outputs generated by policy networks πi and πj… As a particular example, the function D between a policy network π1 and π2 in the set can satisfy:
PNG
media_image8.png
66
256
media_image8.png
Greyscale
where S is the set of observations, s is a trajectory of observations in the set, |s| is the number of observations in the trajectory, |S| is the number observations in the set, DKL is the K−L divergence”; Examiner’s Note: wherein training the population action selection neural network to optimize the distillation loss (i.e. DKL is the K−L divergence) further comprises is taught)
“training the strategy embedding corresponding to the action selection policy of the target agent, comprising backpropagating gradients of the distillation loss through the population action selection neural network and into the strategy embedding corresponding to the action selection policy of the target agent”
(Czarnecki, Paragraph 90, “In particular, as part of computing gradients, the system backpropagates through the combined policy output into the individual neural networks in the set in order to compute the update to the parameters of the networks”; Examiner’s Note: training the strategy embedding corresponding to the action selection policy of the target agent, comprising backpropagating gradients of the distillation loss (i.e. as part of computing gradients, the system backpropagates through the combined policy output into the individual neural networks) through the population action selection neural network and into the strategy embedding corresponding to the action selection policy of the target agent is taught)
The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Silver in view of Grover in view of Lanctot as applied in claim 3, in view of Mnih et al. (“Human-level control through deep reinforcement learning”, 25 Feb 2015, herein Mnih).
Regarding Claim 14:
The combination of Silver, Grover, and Lanctot teaches the method of claim 3.
The combination of Silver, Grover, and Lanctot fails to teach the method of claim 3, wherein training the population action selection neural network to optimize a regularization loss that measures an error between: (i) action selection outputs generated by the population action selection neural network by processing observations from the trajectories, and (ii) action selection outputs generated by a baseline population action selection neural network by processing observations from the trajectories wherein the baseline population action selection neural network is a static, lagging copy of the population action selection neural network.
Mnih teaches “Human-level control through deep reinforcement learning (title)” comprising:
“The method of claim 3, wherein training the population action selection neural network based on the plurality of trajectories further comprises”
(Lanctot, Section 3, “In each episode, one player is set to oracle(learning) mode to train πi, and a fixed policy is sampled from the opponents’ meta-strategies (π−i ∼ σ−i). At the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ, which takes time exponential in |Π|”; Examiner’s Note: wherein training the population action selection neural network based on the plurality of trajectories (i.e. at the end of the epoch, the new oracles are added to their policy sets Πi, expected utilities for new policy combinations are computed via simulation and added to the empirical tensor UΠ) further comprises is taught)
“training the population action selection neural network to optimize a regularization loss that measures an error between: (i) action selection outputs generated by the population action selection neural network by processing observations from the trajectories, and (ii) action selection outputs generated by a baseline population action selection neural network by processing observations from the trajectories”
(Mnih, Main, “We parameterize an approximate value function using the deep convolutional neural network shown in Fig. 1, in which θi are the parameters (that is, weights) of the Q-network at iteration i. To perform experience replay we store the agent’s experiences et = (st,at,rt,st + 1) at each time-step t in a data set Dt = {e1,…,et}. During learning, we apply Q-learning updates, on samples (or minibatches) of experience (s,a,r,s′) ∼ U(D), drawn uniformly at random from the pool of stored samples. The Q-learning update at iteration uses the following loss function:
PNG
media_image9.png
90
418
media_image9.png
Greyscale
in which γ is the discount factor determining the agent’s horizon, θi are the parameters of the Q-network at iteration i and are the network parameters used to compute the target at iteration i. The target network parameters are only updated with the Q-network parameters (θi) every C steps and are held fixed between individual updates"; Examiner’s Note: training the population action selection neural network to optimize a regularization loss (i.e. loss function) that measures an error between: (i) action selection outputs generated by the population action selection neural network by processing observations from the trajectories (i.e. the agent’s experiences et = (st,at,rt,st + 1) ), and (ii) action selection outputs generated by a baseline population action selection neural network (i.e. samples (or minibatches) of experience (s,a,r,s′)) by processing observations from the trajectories is taught)
“wherein the baseline population action selection neural network is a static, lagging copy of the population action selection neural network.”
(Mnih, Main, “The target network parameters are only updated with the Q-network parameters (θi) every C steps and are held fixed between individual updates (see Methods)”; Examiner’s Note: wherein the baseline population action selection neural network (i.e. the target network) is a static, lagging copy (i.e. parameters are only updated…every C steps and are held fixed between individual updates) of the population action selection neural network is taught)
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver, Grover, and Lanctot by training the population action selection neural network to optimize a regularization loss by comparing outputs with a baseline population action selection neural network that is a static, lagging copy of the population action selection neural network as taught in Mnih as “reinforcement learning is known to be unstable or even to diverge when a nonlinear function approximator such as a neural network is used to represent the action-value (also known as Q) function” and the lagging neural network taught in Mnih “address these instabilities with a novel variant of Q-learning, which uses… an iterative update that adjusts the action-values (Q) towards target values that are only periodically updated, thereby reducing correlations with the target” (Mnih, Main).
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Silver in view of Grover in view of Lanctot as applied in claim 3, in view of Adam et al. (“Double Oracle Algorithm for Computing Equilibria in Continuous Games”, 30 Sep 2020, herein Adam).
Regarding Claim 16:
The combination of Silver, Grover, and Lanctot teaches the method of claim 2.
The combination of Silver, Grover, and Lanctot fails to teach the method of claim 2, further comprising: determining that a termination criterion for the update iteration is not satisfied, comprising: determining, for each of multiple strategy assignments, a delta between: (i) a current payoff value for the strategy assignment, and (ii) a previous payoff value for the strategy assignment, wherein each strategy assignment assigns a respective action selection policy to each agent in the collection of agents, and determining that the termination criterion for the update iteration is not satisfied based on the deltas, in response, further training the population action selection neural network before starting a next update iteration.
Adam teaches “Double Oracle Algorithm for Computing Equilibria in Continuous Games (title)” comprising:
“The method of claim 2, further comprising: determining that a termination criterion for the update iteration is not satisfied, comprising”
(Adam, Algorithm 3.1,
PNG
media_image10.png
374
606
media_image10.png
Greyscale
Examiner’s Note: determining that the termination criterion for the update iteration is not satisfied (i.e. until v̄i - vi ≤ ϵ) is taught)
“determining, for each of multiple strategy assignments, a delta between: (i) a current payoff value for the strategy assignment, and (ii) a previous payoff value for the strategy assignment, wherein each strategy assignment assigns a respective action selection policy to each agent in the collection of agents”
(Adam, Section 1-2 and Algorithm 3.1, “in which the strategies are vectors of real numbers”; “If players implement a mixed strategy profile (p,q) ∈ ∆, the expected utility of Player 1 is:
PNG
media_image11.png
63
380
media_image11.png
Greyscale
”;
PNG
media_image10.png
374
606
media_image10.png
Greyscale
Examiner’s Note: determining, for each of multiple strategy assignments, a delta (i.e. v̄i - vi )between: (i) a current payoff value (i.e. expected utility) for the strategy assignment (i.e. v̄i := U(xi+1,q∗i)), and (ii) a previous payoff value (i.e. vi := U(p*i,yi+1)) for the strategy assignment, wherein each strategy assignment (i.e. strategies are vectors of real numbers) assigns a respective action selection policy to each agent in the collection of agents is taught)
“and determining that the termination criterion for the update iteration is not satisfied based on the deltas”
(Adam, Algorithm 3.1,
PNG
media_image10.png
374
606
media_image10.png
Greyscale
Examiner’s Note: and determining that the termination criterion for the update iteration is not satisfied (i.e. until v̄i - vi ≤ ϵ) based on the deltas (i.e. v̄i - vi)is taught)
“in response, further training the population action selection neural network before starting a next update iteration.”
(Adam, Algorithm 3.1,
PNG
media_image10.png
374
606
media_image10.png
Greyscale
Examiner’s Note: in response, further training the population action selection neural network before starting a next update iteration (i.e. repeat… until v̄i - vi ≤ ϵ) is taught)
It would have been obvious to one having ordinary skill in the art before the effective filing date of the invention was made to modify the invention in Silver, Grover, and Lanctot by determining, for each of multiple strategy assignments, a delta between a current and previous payoff value, and further training the population action selection neural network before starting the next update iteration if the termination criterion for the update is not met as taught by the Double Oracle Algorithm in Adam as “he double oracle algorithm converges faster than fictitious play on several examples” (Adams, Section 1)
Allowable Subject Matter
No prior art rejection is made for claim 7, however, this claim is rejected under 35 U.S.C. 101 - abstract idea.
Examiner has disclosed Lanctot et al. (“A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning”) and Jiang et al. (“Metric Policy Representations for Opponent Modeling”), which are the closest prior art as compared to the instant application. Lanctot discloses training the population action selection neural network based on the probability distribution over the strategy assignment space – more particularly, Lanctot Section 3 teaches generating an aggregate strategy assignment embedding of the points selected from the strategy assignment space. Jiang discloses generating the aggregate strategy assignment embedding based on the respective strategy assignment embedding for each of the points selected from the strategy assignment space – more particularly, Jiang Section 3.2 discloses generating a joint policy by sampling from the joint distribution over the opponent joint policy space. However, the aforementioned prior art references seemingly do not disclose the specific limitations of dependent claim 7 including "generating the aggregate strategy assignment embedding as a linear combination of the respective strategy assignment embedding for each of the points selected from the strategy assignment space” in combination with the remaining limitations of the claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Budden et al. (US 2020/0293883) teaches training an action selection neural network jointly with a distribution Q network that is used to update the parameters of the action selection neural network. Badia et al. (US 2020/0372366) teaches jointly learning exploratory and non-exploratory action selection policies. Anand et al. (US 2023/0342425) teaches optimal sequential decision making with changing action space.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PALLAVI MAMILLAPALLI whose telephone number is (571)270-5953. The examiner can normally be reached Monday-Friday (7:30 - 3:30) ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/P.M./Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147