DETAILED ACTION
This is a non-final, first office action on the merits. Claims 1-28 and 41 are pending. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant is claiming Foreign Priority to Foreign Applications Great Britain Patent Application No. GB2202994.6, filed March 03, 2022.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-28 and 41 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. Specifically, claims 1-28 and 41 are directed to an abstract idea without additional elements amounting to significantly more than the abstract idea.
With respect to Step 2A Prong One of the framework, claims 1 and 41 recite an abstract idea. Claims 1 and 41 include “receive a policy input comprising an observation of a state of an environment and to process the policy input to generate a policy output that defines an action to be performed by an agent in response to the observation, the method comprising: generating a training trajectory the agent to perform a task episode in the environment across a plurality of time steps using the policy, wherein the environment includes an expert agent that is attempting to perform an instance of a task, and wherein generating the training trajectory comprises, at each of the plurality of time steps: determining, in accordance with an expert dropout policy for the task episode, whether to cause the expert agent to become unobservable by the agent at the time step; generating an observation characterizing the environment at the time step, comprising: in response to determining to cause the expert agent to become unobservable, generating an observation that does not include any sensor measurements of the expert agent; processing a policy input comprising the observation using the policy to generate a policy output for the observation; and selecting an action to be performed by the agent at the time step using the policy output; and
training the policy on the training trajectory”.
With respect to Step 2A Prong One of the framework, claim 16 recites an abstract idea. Claim 16 includes receive a policy input comprising an observation of a state of an environment and to process the policy input to generate a policy output that defines an action to be performed by an agent in response to the observation, the method comprising: generating a training trajectory the agent to perform a task episode in the environment across a sequence of plurality of time steps using the policy, wherein the environment includes an expert agent that is attempting to perform an instance of a task, and wherein generating the training trajectory comprises, at each of the plurality of time steps: obtaining an observation of the state of the environment at the time step; processing a policy input comprising the observation using a first subnetwork of the policy to generate a belief representation; processing the belief representation using a policy head of the policy to generate a policy output for the time step; processing the belief representation using an attention head of the policy to generate respective predicted positions of one or more other agents in the environment at a time step that is at a particular position relative to the time step in the sequence of time steps; and selecting an action to be performed by the agent at the time step using the policy output; and training the policy on the training trajectory, comprising, for one or more of the plurality of time steps: receiving a respective ground truth position for each of the one or more other agents at the time step that is at a particular position relative to the time step in the sequence of time steps; and training the policy on the training trajectory to minimize an error between, for each of the one or more time steps, the respective predicted positions and the respective ground truth positions”.
The limitations above recite an abstract idea under Step 2A Prong One. More particularly, the elements above recite mental processes-concepts performed in the human mind (including an observation, evaluation, judgment, opinion) because the elements describe a process for training a policy. As a result, claims 1, 16, and 41 recite an abstract idea under Step 2A Prong One.
Claims 2-15 and 17-28 further describe the process for training a policy. As a result, claims 2-15 and 17-28 recite an abstract idea under Step 2A Prong One for the same reasons as stated above with respect to claims 1, 16, and 41.
With respect to Step 2A Prong Two of the framework, claims 1, 16, and 41 do not include additional elements that integrate the abstract idea into a practical application. Claims 1, 16, and 41 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 1, 16, and 41 include a neural network, one or more computers, one or more storage devices, and one or more computers. When considered in view of the claim as a whole, the additional elements do not integrate the abstract idea into a practical application because the additional computing elements are generic computing elements that are merely used as a tool to perform the recited abstract idea. As a result, claims 1, 16, and 41 do not include additional elements that integrate the abstract idea into a practical application under Step 2A Prong Two.
Claims 3-6, 10-11, 13-15, 17-19, 21-24, and 26-28 do not include any additional elements beyond those recited with respect to claims 1, 16, and 41. As a result, claims 3-6, 10-11, 13-15, 17-19, 21-24, and 26-28 do not include additional elements that integrate the abstract idea into a practical application under Step 2A Prong Two for the same reasons as stated above with respect to claims 1, 16, and 41.
Claims 2, 7-9, 12, 20, and 25 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 2, 7-9, 12, 20, and 25 include a sensor, neural network, and reinforcement learning. When considered in view of the claims as a whole, the additional elements do not integrate the abstract idea into a practical application because the additional computing elements do no more than generally link the use of the recited abstract idea to a particular technological environment. As a result, claims 2, 7-9, 12, 20, and 25 do not include additional elements that integrate the abstract idea into a practical application under Step 2A Prong Two.
With respect to Step 2B of the framework, claims 1, 16, and 41 do not include additional elements amounting to significantly more than the abstract idea. As noted above, claims 1, 16, and 41 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 1, 16, and 41 include a neural network, one or more computers, one or more storage devices, and one or more computers. The additional elements do not amount to significantly more than the abstract idea because the additional computing elements are generic computing elements that are merely used as a tool to perform the recited abstract idea. Further, looking at the additional elements as an ordered combination adds nothing that is not already present when considering the additional elements individually. As a result, independent claims 1, 16, and 41 do not include additional elements that amount to significantly more than the abstract idea under Step 2B.
Claims 3-6, 10-11, 13-15, 17-19, 21-24, and 26-28 do not include any additional elements beyond those recited with respect to claims 1, 16, and 41. As a result, claims 3-6, 10-11, 13-15, 17-19, 21-24, and 26-28 do not include additional elements that amount to significantly more than the abstract idea under Step 2B for the same reasons as stated above with respect to claims 1, 16, and 41.
Claims 2, 7-9, 12, 20, and 25 include additional elements that do not recite an abstract idea under Step 2A Prong One. The additional elements of claims 2, 7-9, 12, 20, and 25 include a proximal sensor, a remote sensor, a sensor, a remote sensing satellite, an airplane, an unmanned aerial vehicle (UAV). The additional elements do not amount to significantly more than the abstract idea because the additional computing elements do no more than generally link the use of the recited abstract idea to a particular technological environment. Further, looking at the additional elements as an ordered combination adds nothing that is not already present when considering the additional elements individually. As a result, claims 2, 7-9, 12, 20, and 25 do not include additional elements that amount to significantly more than the abstract idea under Step 2B.
Therefore, the claims are directed to an abstract idea without additional elements amounting to significantly more than the abstract idea. Accordingly, claims 1-28 and 41 are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-5, 7, 13, and 41 are rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.) in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.).
Regarding claims 1 and 41, Zheng discloses a method performed by one or more computers and for training a policy neural network that is configured to receive a policy input comprising an observation of a state of an environment and to process the policy input to generate a policy output that defines an action to be performed by an agent in response to the observation (see Zheng, paras [0004]-[0005], wherein select the action to be performed by the agent in response to receiving a given observation in accordance with an output of a neural network….. to predict an output for a received input. Some neural networks are deep neural networks that include one or more hidden layers in addition to an output layer), the method comprising:
one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers (see Zheng, paras [0086]-[0091]);
generating a training trajectory by controlling the agent to perform a task episode in the environment across a plurality of time steps using the policy neural network, wherein the environment includes an expert agent that is attempting to perform an instance of a task, and wherein generating the training trajectory comprises, at each of the plurality of time steps (see Zheng, para [0046], wherein the agent may be a mechanical agent that performs or controls the protein folding actions or chemical synthesis steps selected by the system automatically without human interaction; paras [0061]-[0063], wherein the environment 115 may be specific to one or more tasks of the task distribution….the training may comprise N iterations. For each iteration k=l, 2, ... , N, a trajectory may be generated using the current policy• The current policy 110 may be updated based upon the generated intrinsic reward from the trajectory using a policy gradient technique; paras [0004]-[0005], wherein neural networks are deep neural networks that include one or more hidden layers in addition to an output layer; para [0026], wherein encoding knowledge (i.e., an expert agent) relating to "what to do", e.g., explore particular areas of an environment, rather than knowledge relating to "how to do", e.g., the specific action to perform, which may be encoded by the agent's policy; para [0057], wherein an accumulation of multiple reward values over multiple time steps):
determining, in accordance with an expert dropout policy for the task episode, whether to cause the expert agent to become unobservable by the agent at the time step (see Zheng, para [0005], wherein Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks are deep neural networks that include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, (i.e., the next hidden layer which becomes unobservable by the agent);
generating an observation characterizing the environment at the time step (see Zheng, para [0023], wherein receive observations, i.e., data, characterizing a current state of the environment, and in response to select actions to be performed by an agent interacting with an environment, e.g., to perform a task), comprising:
processing a policy input comprising the observation using the policy neural network to generate a policy output for the observation (see Zheng, paras [0005] & [0062], wherein generates an output from a received input in accordance with current values of a respective set of parameters…..the policy distribution may also be received as an input to the reinforcement learning system 100. For example, the agent's policy 110 may be implemented using a neural network and initializing the agent's policy 110 may comprise initializing the parameters of this policy neural network. The policy neural network may comprise one or more convolutional layers followed by one or more fully connected layers); and
selecting an action to be performed by the agent at the time step using the policy output (see Zheng, para [0040], wherein selects actions to be performed by a reinforcement learning agent interacting with an environment. In order for the agent to interact with the environment, the system receives data characterizing the current state of the environment and selects an action to be performed by the agent in response to the received data. Data characterizing a state of the environment is referred to in this specification as an observation); and
training the policy neural network on the training trajectory (see Zheng, para [0015], wherein the policy may be defined by one or more policy parameters and updating the policy may comprise updating the one or more policy parameters. In this regard, the policy may be provided by a policy neural network).
Zheng et al. fails to explicitly disclose generating an observation that does not include any sensor measurements of the expert agent.
Analogous art Vecerik discloses in response to determining to cause the expert agent to become unobservable, generating an observation that does not include any sensor measurements of the expert agent (see Vecerik, para [0004], wherein a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output; and para [0025], wherein enable an agent to perform a task by imitating an "expert", i.e., by using a set of expert demonstrations of the task to train the action selection network to match the behavior of the expert).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Vecerik directed to training an action selection policy neural network. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included in response to determining to cause the expert agent to become unobservable, generating an observation that does not include any sensor measurements of the expert agent because both inventions teach improving learning algorithms. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Regarding claim 2, Zheng discloses the method of claim 1, wherein generating an observation characterizing the environment at the time step comprises:
in response to determining not to cause the expert agent to become unobservable, generating an observation that include-includes sensor measurements of the expert agent (see Zheng, para [0042], wherein sensor data to capture observations as the agent interacts with the environment, for example sensor data from an image, distance, or position sensor or from an actuator. In the case of a robot or other mechanical agent or vehicle the observations may similarly include one or more of the position, linear or angular velocity, force, torque or acceleration, and global or relative pose of one or more parts of the agent).
Regarding claim 3, Zheng discloses the method of any preceding claim 1, wherein the expert dropout policy for the task episode specifies that the expert agent is observable at all of the plurality of time steps (see Zheng, para [0081], wherein Fig. 4 illustrates a plot of episodic return as a function of the number of episodes on various tasks).
Regarding claim 4, Zheng discloses the method of claim 1, wherein the expert dropout policy for the task episode specifies that the expert agent is unobservable at all of the plurality of time steps (see Zheng, para [0026], wherein for many tasks, far more "unlabeled" expert demonstrations (i.e., that do not specify the expert actions) may be available than "labeled" expert demonstrations (i.e., that do specify the expert actions) train the action selection network to select actions to perform a task even if there are no labeled expert demonstrations of the task available).
Regarding claim 5, Zheng discloses the method of claim 1, wherein the expert dropout policy for the task episode specifies that the expert agent is only observable at an initial proper subset of the plurality of time steps (see Zheng, para [0083], wherein the optimal behavior is to explore A and C at the beginning of a lifetime to assess which is better and then to commit to the better one for all subsequent episodes. As shown in FIG. 6, the reinforcement learning system 100 comprising an intrinsic reward system 120 demonstrates such behavior).
Regarding claim 7, Zheng discloses the method of claim 1, wherein generating the training trajectory comprises, at one or more of the plurality of time steps:
receiving a respective reward in response to the agent performing the selected action at the time step (see Zheng, para [0057], wherein calculated based on the rewards received following a given action, in this case, the intrinsic reward. For instance, the return might be an accumulation of multiple reward values over multiple time steps); and
wherein training the policy neural network on the training trajectory comprises training the policy neural network on the training trajectory using the respective rewards through reinforcement learning (see Zheng, para [0057], wherein calculated based on the rewards received following a given action, in this case, the intrinsic reward. For instance, the return might be an accumulation of multiple reward values over multiple time steps; and para [0062], wherein the agent's policy 110 is initialized. The agent's policy 110 may be initialized randomly from a policy distribution. The policy distribution may also be received as an input to the reinforcement learning system 100. For example, the agent's policy 110 may be implemented using a neural network and initializing the agent's policy 110 may comprise initializing the parameters of this policy neural network. The policy neural network may comprise one or more convolutional layers followed by one or more fully connected layers).
Regarding claim 13, Zheng discloses the method of claim 1, wherein the task episode is defined by a respective value for each of a set of task parameters, and wherein generating the training trajectory comprises sampling the respective values for the set of task parameters from a distribution that is parametrized by a set of distribution parameters (see Zheng, para [0057], wherein the plurality of tasks may be randomly sampled from a task distribution. The agent's policy may be randomly initialized by sampling a policy distribution and re-initialized by resampling the policy distribution. An agent having a different policy may be considered to be a different agent; and para [0072], wherein updates and the lifetime return may be approximated using a lifetime value function parameterized by cp for estimating the remaining lifetime return).
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.), in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), and further in view of SA Seif Tabrizi et al. (Control and Characterization of Open Quantum Systems) - 2020 - drum.lib.umd.edu (hereinafter Tabrizi et al.).
Regarding claim 6, Zheng discloses the method of claim 1, wherein the expert dropout policy for the task episode specifies that, for each of the plurality of time steps, the expert agent is observable with a probability p and unobservable (see Zheng, para [0005], wherein Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks are deep neural networks that include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer (i.e., the next hidden layer which becomes unobservable by the agent).
Zheng et al. and Vecerik et al. combined fail to explicitly disclose with probability 1 – p.
Analogous art Tabrizi discloses with probability 1 – p (see Tabrizi,
PNG
media_image1.png
257
752
media_image1.png
Greyscale
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Tabrizi directed to training a neural network prevent from overfitting. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included with probability 1 – p because both inventions teach improving the accuracy of the solution. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claims 8 and are rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.), in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), and further in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.).
Regarding claim 8, Zheng discloses the method of claim 1, wherein processing a policy input comprising the observation using the policy neural network to generate a policy output for the observation comprises:
Zheng et al. and Vecerik et al. combined fail to explicitly disclose processing the policy input using a first subnetwork of the policy neural network to generate a belief representation.
Analogous art Pietquin discloses processing the policy input using a first subnetwork of the policy neural network to generate a belief representation (see Pietquin, para [0058], wherein a sub-network of a neural network refers to a group of one or more neural network layers in the neural network……configured to process the output of the embedding subnetwork and action data defining each action from the set of possible actions (or data derived from the action data or both). The selection sub-network can be configured to process the output of the core sub-network to generate the Q value outputs for the actions; and para [0017], wherein the ground truth action may be an action performed by an expert agent in response to the current observation in the demonstration (i.e., past observations representing the agent's current understanding)); and
Analogous art Pietquin discloses processing the belief representation (see Pietquin, para [0017], wherein the ground truth action may be an action performed by an expert agent in response to the current observation in the demonstration (i.e., past observations representing the agent's current understanding)).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Pietquin directed to optimizing a reinforcement learning objective function that maximizes returns. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included processing the policy input using a first subnetwork of the policy neural network to generate a belief representation; and processing the belief representation because both inventions teach improving the accuracy of the solution. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Zheng et al. and Vecerik et al. combined fail to explicitly disclose using a policy head of the policy neural network to generate the policy output.
Analogous art Cabi discloses using a policy head of the policy neural network to generate the policy output (see Cabi, paras [0030]-[0031], wherein a sub-network of a neural network refers to a group of one or more neural network layers in the neural network……policy neural network 110 can include a convolutional encoder that encodes the high-dimensional data, a fully connected encoder that encodes the lower-dimensional data, and a policy subnetwork that operates on a combination, e.g., a concatenation, of the encoded data to generate the policy output……The policy neural network 110 then processes the final features using a policy head, implemented as a recurrent neural network, to generate a probability distribution or parameters of a probability distribution).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Cabi directed to processing the observation in the experience. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included using a policy head of the policy neural network to generate the policy output because both inventions teach improving the accuracy of the solution. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claims 9 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.), in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), and further in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.).
Regarding claim 9, Zheng discloses the method of claim 8, wherein processing a policy input comprising the observation using the policy neural network to generate a policy output for the observation further comprises:
Zheng et al. and Vecerik et al. combined fail to explicitly disclose processing the belief representation.
Analogous art Pietquin discloses processing the belief representation (see Pietquin, para [0017], wherein the ground truth action may be an action performed by an expert agent in response to the current observation in the demonstration (i.e., past observations representing the agent's current understanding)).
One of ordinary skill in the art would have recognized that applying the known technique of Pietquin would have yielded predictable results and resulted in an improved system for the same reasons as stated above with respect to claim 8.
Zheng et al., Vecerik et al., Pietquin et al., and Cabi et al. combined fail to explicitly disclose using an attention head to generate respective predicted positions of one or more other agents in the environment at a particular time step of the plurality of time steps.
Analogous art Tran discloses processing the belief representation using an attention head to generate respective predicted positions of one or more other agents in the environment at a particular time step of the plurality of time steps (see Tran, para [0253], wherein agent data includes agent grades (which may be determined from grading or ranking agents on desired outcomes), agent demographic data, agent psychographic data; and para [0342], wherein the attention-mechanism can be parallelized into multiple modules and is repeated multiple times with linear projections of Q, K and V. This allows the system to learn from different representations of Q, K and V. These linear representations are done by multiplying Q, K and V by weight matrices W that are learned during the training. Those matrices Q, K and V are different for each position of the attention modules in the structure depending on whether they are in the encoder, decoder or in-between encoder and decoder. The reason is that we want to attend on either the whole encoder input sequence or a part of the decoder input sequence. The multi-head attention module that connects the encoder and decoder will make sure that the encoder input sequence is considered together with the decoder input sequence up to a given position. After the multi-attention heads in both the encoder and decoder, the transformer has a pointwise feed-forward layer. This feed-forward network has identical parameters for each position, which can be described as a separate, identical linear transformation of each element from the given sequence).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Tran directed to generating a description of the document. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included processing the belief representation using an attention head to generate respective predicted positions of one or more other agents in the environment at a particular time step of the plurality of time steps because both inventions teach improving the accuracy of the solution. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Regarding claim 11, Zheng discloses the method of claim 9, wherein the particular time step is the time step (see Zheng, para [0057], wherein an accumulation of multiple reward values over multiple time steps).
Regarding claim 12, Zheng discloses the method of claims 9, wherein generating the training trajectory comprises, at one or more of the plurality of time steps:
wherein training the policy neural network on the training trajectory comprises training the policy neural network on the training trajectory (see Zheng, para [0015], wherein the policy may be defined by one or more policy parameters and updating the policy may comprise updating the one or more policy parameters. In this regard, the policy may be provided by a policy neural network; paras [0061]-[0063], wherein the environment 115 may be specific to one or more tasks of the task distribution….the training may comprise N iterations. For each iteration k=l, 2, ... , N, a trajectory may be generated using the current policy• The current policy 110 may be updated based upon the generated intrinsic reward from the trajectory using a policy gradient technique).
Zheng et al. fails to explicitly disclose minimize an error between the respective predicted positions.
Analogous art Vecerik discloses minimize an error between the respective predicted positions (see Vecerik, para [0025], wherein reducing the accumulation of errors and enabling the agent to recover from mistakes. The system can train the action selection network to reach an acceptable level of performance over fewer training iterations and using fewer expert observations than some conventional training systems).
One of ordinary skill in the art would have recognized that applying the known technique of Vecerik would have yielded predictable results and resulted in an improved system for the same reasons as stated above with respect to claim 1.
Zheng et al. and Vecerik et al. combined fail to explicitly disclose receiving a respective ground truth position for each of the one or more other agents; and wherein the respective ground truth positions.
Analogous art Pietquin discloses receiving a respective ground truth position for each of the one or more other agents; and wherein the respective ground truth positions (see Pietquin, para [0017], wherein received if the agent performed the ground truth action in response to the current observation in the demonstration; and determining, from the gradient of the bonus estimation loss function, an update to the current values of the bonus estimation network parameters); and
Analogous art Pietquin discloses wherein training the policy neural network on the training trajectory comprises training the policy neural network on the training trajectory to minimize an error between the respective predicted positions and the respective ground truth positions (see Pietquin, para [0017], wherein the ground truth action may be an action performed by an expert agent in response to the current observation in the demonstration (i.e., past observations representing the agent's current understanding)).
One of ordinary skill in the art would have recognized that applying the known technique of Pietquin would have yielded predictable results and resulted in an improved system for the same reasons as stated above with respect to claim 8.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.), in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.), and further in view of A Saran et al. (Leveraging multimodal human cues to enhance robot learning from demonstration) - 2021 - repositories.lib.utexas.edu (hereinafter Saran et al.).
Regarding claim 10, Zheng discloses the method of claim 9, wherein the respective prediction positions, as set forth above with claim 9.
Zheng et al., Vecerik et al., Pietquin et al., Cabi et al., and Tran et al. combined fail to explicitly disclose ego-centric relative positions relative to the agent.
Analogous art Saran discloses ego-centric relative positions relative to the agent (see Saran, page 165, wherein user's egocentric view, (2) pixel location of the human's gaze in the egocentric image, (3) gaze time stamps synchronized with keyframe time stamps along the KT demonstration).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Saran directed to optimizing reinforcement learning to learn behavior policies. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included ego-centric relative positions relative to the agent because both inventions teach improving the accuracy of the solution. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.), in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), and further in view of A Noriega Campero et al. (Human and artificial intelligence in decision systems for social development) - 2019 - dspace.mit.edu (hereinafter Campero et al.).
Regarding claim 14, Zheng discloses the method of claim 13, further comprising:
Zheng et al. and Vecerik et al. combined fail to explicitly disclose determining a cultural transmission metric for the agent; and adjusting the set of distribution parameters using the cultural transmission metric for the agent.
Analogous art Campero discloses determining a cultural transmission metric for the agent (see Campero, page 35, wherein Dynamic interaction networks have been shown to promote human cooperation [109], and culture transmission networks over generations enabled human groups to develop technologies above any individual’s capabilities [66]. In the artificial realm, prominent machine learning algorithms rely on similar logics, where dynamically updated networks integrate input signals into useful output [18, 96]. Across the board, networks’ dynamic properties embody key mechanisms that enable systems to adapt to environmental changes; and page 52, wherein agents that interact can exchange information about their degree of confidence on each estimation task); and
Analogous art Campero discloses adjusting the set of distribution parameters using the cultural transmission metric for the agent (see Campero, page 95, wherein adjusting acceptance thresholds is the basic unit of decision support that the platform provides. To adjust the acceptance threshold of the population, or any population segment, the user clicks on the corresponding tree node, which pops-up the thresholds adjustment window (Figure 4-10). In it, the user can see the distribution of the population or segment in terms of the prioritization variable, and how the threshold divides de distribution in accepted/nonaccepted population; and page 35, wherein Dynamic interaction networks have been shown to promote human cooperation [109], and culture transmission networks over generations enabled human groups to develop technologies above any individual’s capabilities [66]. In the artificial realm, prominent machine learning algorithms rely on similar logics, where dynamically updated networks integrate input signals into useful output [18, 96]. Across the board, networks’ dynamic properties embody key mechanisms that enable systems to adapt to environmental changes).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Campero directed to refining individual judgments and harnessing groups’ collective intelligence. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included determining a cultural transmission metric for the agent; and adjusting the set of distribution parameters using the cultural transmission metric for the agent because both inventions teach improving interaction with its human peers. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.) in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), and further in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.).
Regarding claim 16, Zheng discloses a method performed by one or more computers and for training a policy neural network that is configured to receive a policy input comprising an observation of a state of an environment and to process the policy input to generate a policy output that defines an action to be performed by an agent in response to the observation, as set forth above with claim 1, the method comprising:
generating a training trajectory by controlling the agent to perform a task episode in the environment across a sequence of plurality of time steps using the policy neural network, wherein the environment includes an expert agent that is attempting to perform an instance of a task, and wherein generating the training trajectory comprises, at each of the plurality of time steps, as set forth above with claim 1:
obtaining an observation of the state of the environment at the time step; processing a policy input comprising the observation using a first subnetwork of the policy neural network to generate a belief representation, as set forth above with claim 8;
processing the belief representation using a policy head of the policy neural network to generate a policy output for the time step, as set forth above with claim 8;
processing the belief representation using an attention head of the policy neural network to generate respective predicted positions of one or more other agents in the environment at a time step that is at a particular position relative to the time step in the sequence of time steps, as set forth above with claim 9; and
selecting an action to be performed by the agent at the time step using the policy output, as set forth above with claim 1; and
training the policy neural network on the training trajectory, comprising, for one or more of the plurality of time steps, as set forth above with claim 1:
receiving a respective ground truth position for each of the one or more other agents, as set forth above with claim 12
at the time step that is at a particular position relative to the time step in the sequence of time steps, as set forth above with claim 9; and
training the policy neural network on the training trajectory to minimize an error between, for each of the one or more time steps, the respective predicted positions and the respective ground truth positions, as set forth above with claim 12.
Regarding claim 19, Zheng discloses the method of claims 16, wherein obtaining the observation comprises:
determining, in accordance with an expert dropout policy for the task episode, whether to cause the expert agent to become unobservable by the agent at the time step, as set forth above with claim 1;
generating the observation characterizing the environment at the time step, as set forth above with claim 1, comprising:
in response to determining to cause the expert agent to become unobservable, generating an observation that does not include any sensor measurements of the expert agent, as set forth above with claim 1.
Regarding claim 20, Zheng discloses the method of claim 19, wherein generating an observation characterizing the environment at the time step comprises:
in response to determining not to cause the expert agent to become unobservable, generating an observation that include sensor measurements of the expert agent, as set forth above with claim 2.
Regarding claim 21, Zheng discloses the method of claim 19, wherein the expert dropout policy for the task episode specifies that the expert agent is observable at all of the plurality of time steps, as set forth above with claim 3.
Regarding claim 22, Zheng discloses the method of claim 19, wherein the expert dropout policy for the task episode specifies that the expert agent is unobservable at all of the plurality of time steps, as set forth above with claim 4.
Regarding claim 23, Zheng discloses the method of claim 19, wherein the expert dropout policy for the task episode specifies that the expert agent is only observable at an initial proper subset of the plurality of time steps, as set forth above with claim 5.
Regarding claim 25, Zheng discloses the method of claim 16, wherein generating the training trajectory comprises, at one or more of the plurality of time steps:
receiving a respective reward in response to the agent performing the selected action at the time step, as set forth above with claim 7; and
wherein training the policy neural network on the training trajectory comprises training the policy neural network on the training trajectory using the respective rewards through reinforcement learning, as set forth above with claim 7.
Regarding claim 26, Zheng discloses the method of claim 16, wherein the task episode is defined by a respective value for each of a set of task parameters, and wherein generating the training trajectory comprises sampling the respective values for the set of task parameters from a distribution that is parametrized by a set of distribution parameters, as set forth above with claim 13.
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.) in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.), and further in view of A Saran et al. (Leveraging multimodal human cues to enhance robot learning from demonstration) - 2021 - repositories.lib.utexas.edu (hereinafter Saran et al.).
Regarding claim 17, Zheng discloses the method of claim 16, wherein the respective predicted positions are ego-centric relative positions relative to the agent, as set forth above with claim 10.
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.) in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.). and further in view of Haidar et al. (US Pub No. 2022/0335303) (hereinafter Haidar et al.).
Regarding claim 18, Zheng discloses the method of claim 16.
Zheng et al., Vecerik et al., Pietquin et al., Cabi et al., and Tran et al. combined fail to explicitly disclose wherein the particular position is the same position as the position of the time step.
Analogous art Haidar discloses the particular position is the same position as the position of the time step (see Haidar, para [0103], wherein processes each teacher layer vector and its corresponding student layer vector (i.e. the student layer vector at the same position in the order) to compute the intermediate representation loss 316; and para [0057], wherein all of the teacher's intermediate layers are considered over time throughout the whole training period (consisting of multiple training epochs), which
may solve the skip, search and overfitting problems).
Zheng directed to a system for performing actions that are selected by the reinforcement learning system. Haidar directed to optimizing the learnable parameters of the neurons. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Zheng, regarding the System for Reinforcement Learning Using Meta-Learned Intrinsic Rewards, to have included the particular position is the same position as the position of the time step because both inventions teach improving training performance. Further, the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claim 24 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.) in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.). and further in view of SA Seif Tabrizi et al. (Control and Characterization of Open Quantum Systems) - 2020 - drum.lib.umd.edu (hereinafter Tabrizi et al.).
Regarding claim 24, Zheng discloses the method of claim 19, wherein the expert dropout policy for the task episode specifies that, for each of the plurality of time steps, the expert agent is observable with a probability p and unobservable with probability 1 - p, as set forth above with claim 6.
Claim 27 is rejected under 35 U.S.C. 103 as being unpatentable over Zheng et al. (US Pub No. 2021/0089910) (hereinafter Zheng et al.) in view of Vecerik et al. (US Pub No. 2020/0104684) (hereinafter Vecerik et al.), in view of Pietquin et al. (US Pub No. 2021/0397959) (hereinafter Pietquin et al.), in view of Cabi et al. (US Pub No. 2021/0078169) (hereinafter Cabi et al.), in view of Tran et al. (US Pub No. 2023/0252224) (hereinafter Tran et al.). and further in view of A Noriega Campero et al. (Human and artificial intelligence in decision systems for social development) - 2019 - dspace.mit.edu (hereinafter Campero et al.).
Regarding claim 27, Zheng discloses the method of claim 26, further comprising:
determining a cultural transmission metric for the agent, as set forth above with claim 14; and
adjusting the set of distribution parameters using the cultural transmission metric for the agent, as set forth above with claim 14.
Allowable Subject Matter
Regarding claims 15 and 28 objected to as being dependent upon a rejected base claim, but it appears they would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, and rewritten to overcome the 35 USC 101 rejections.
Conclusion
The prior arts made of record and not relied upon is considered pertinent to applicant's disclosure. (US Pub No. 2023/0102544; US Pub No. 2020/0104680; US Pub No. 2020/0117956; US Pub No. 2023/0041501; US Pub No. 2023/0325635; and M Mcloughlin et al. (The development and application of computational multi-agent models for investigating the cultural transmission and cultural evolution of humpback whale song) - 2018 - pearl.plymouth.ac.uk.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAFIZ A KASSIM whose telephone number is (571)272-8534. The examiner can normally be reached 9:00 - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Rutao Wu can be reached at 571-272-6045. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAFIZ A KASSIM/Primary Examiner, Art Unit 3623 09/22/2026