DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1,16, 18-19, and 22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1,
Step 1: Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a method/process.
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The limitations of:
for each of a plurality of iterations, wherein each iteration corresponds to one of the sub- trajectories: determining the latent variable for the corresponding local goal [using a predictor neural network configured to predict the latent variable in accordance with predictor parameter values based on a preceding latent variable]; (mental process, a human can determine the latent variable, or action depending on a given context/goal)
for each iteration, obtaining the corresponding sub-trajectory by, for each of a plurality of time steps: obtaining an observation characterizing a current state of the environment; (mental observation, a human can look at the environment and observe mentally)
[processing the observation using the action selection policy neural network to] generate a policy output, [wherein the action selection policy neural network is conditioned on the latent variable and generates the policy output in accordance with policy parameter values of the action selection policy neural network]; (mental judgement, based on the current observations and actions, an human can determine the policy using their head)
selecting an action to be performed by the agent in response to the observation using the policy output (mental judgement, a human can in their head determine/select what action to take)
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
The limitations of:
a computer-implemented method for controlling an agent while interacting with an environment, wherein the agent is controlled to execute a sequence of actions defining a state- action trajectory for achieving a goal, the goal comprising a sequence of local goals, each local goal being characterized by a corresponding latent variable, wherein each state-action trajectory comprises a set of sub-trajectories and each sub-trajectory is generated according to a corresponding local goal, the method comprising: (applying the abstract idea on generic computing components and applying it to a particular field of use, MPEP 2106.05(f) and 2106.05(h) respectively)
[…] using a predictor neural network configured to predict the latent variable in accordance with predictor parameter values based on a preceding latent variable (invoking generic computer components merely as a tool to perform and existing process, MPEP 2106.05(f))
processing the observation using the action selection policy neural network to [generate a policy output], wherein the action selection policy neural network is conditioned on the latent variable and generates the policy output in accordance with policy parameter values of the action selection policy neural network; (invoking generic computer components merely as a tool to perform and existing process, MPEP 2106.05(f))
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
The limitations of:
a computer-implemented method for controlling an agent while interacting with an environment, wherein the agent is controlled to execute a sequence of actions defining a state- action trajectory for achieving a goal, the goal comprising a sequence of local goals, each local goal being characterized by a corresponding latent variable, wherein each state-action trajectory comprises a set of sub-trajectories and each sub-trajectory is generated according to a corresponding local goal, the method comprising: (applying the abstract idea on generic computing components and applying it to a particular field of use, MPEP 2106.05(f) and 2106.05(h) respectively)
[…] using a predictor neural network configured to predict the latent variable in accordance with predictor parameter values based on a preceding latent variable (invoking generic computer components merely as a tool to perform and existing process, MPEP 2106.05(f))
processing the observation using the action selection policy neural network to [generate a policy output], wherein the action selection policy neural network is conditioned on the latent variable and generates the policy output in accordance with policy parameter values of the action selection policy neural network; (invoking generic computer components merely as a tool to perform and existing process, MPEP 2106.05(f))
Note independent claims 18 and 19 recite the same substantial subject matter as independent claim 1, only differing in embodiment. The differences in embodiment, a system and non-transitory medium do not meaningfully change the above analysis and therefore the claims are subject to the same rejection.
Dependent claim 2 recites determining the latent variable using perturbations, mental evaluation.
Dependent claim 3 recites the perturbations determining from a distribution, applying the abstract idea to a particular field of use MPEP 2106.05(h).
Dependent claim 4 recites determining the latent variable based on a previous one, mental evaluation.
Dependent claim 5 recites determining based on an observation, mental observation.
Dependent claim 6 recites determining based on an observation and predictor network, mental observation and MPEP 2106.05(f) generic computer components.
Dependent claim 7 recites the variables being linearly related, applying the abstract idea to a particular field of use MPEP 2106.05(h).
Dependent claim 8 recites determining based on an observation, mental observation.
Dependent claim 9 recites a Markov chain, applying the abstract idea to a particular field of use MPEP 2106.05(h).
Dependent claim 10 recites updating based on sub-trajectories, mental evaluation.
Dependent claim 11 recites updating parameters, mental judgement.
Dependent claim 12 updating for each iteration, mathematical concepts.
Dependent claim 13 recites minimizing a difference, mental evaluation.
Dependent claim 14 recites a probability distributions, mathematical concepts.
Dependent claim 15 recites the observations relate to a real-world environment, applying the abstract idea to a particular field of use MPEP 2106.05(h).
Dependent claim 16 recites controlling the agent and using the network, applying the abstract idea to a particular field of use MPEP 2106.05(h) and generic computer components to carry out the abstract idea MPEP 2106.05(f).
Dependent claim 22 corresponds to dependent 2.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4-8, 10-11, 13, 15-16, and 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kawano, Hiroshi. "Hierarchical sub-task decomposition for reinforcement learning of multi-robot delivery mission." in view of Reed et al. US 2020/0104680.
Regarding claims 1, 18, and 19, Kawano teaches “a computer-implemented method for controlling an agent while interacting with an environment” (pg. 1 §1 ¶1 “Reinforcement learning (RL) is one of the most promising methods for building sophisticated algorithms for intelligent robots, because the learning agent can automatically acquire the optimum action policy through the interaction with the environment following a Markov decision process (MDP)”), “wherein the agent is controlled to execute a sequence of actions defining a state- action trajectory for achieving a goal” (pg. 2 §2 ¶1 “grid represents the position of the robot and load. The mission is to lead the load to a goal point by having the robots start moving from a start point and push the load.”), “the goal comprising a sequence of local goals, each local goal being characterized by a corresponding latent variable” (pg. 1 right col. last ¶ “In HRL, the task is decomposed into sub-tasks”), “wherein each state-action trajectory comprises a set of sub-trajectories and each sub-trajectory is generated according to a corresponding local goal” (previous citation “Once the learning process of a sub-task is completed, the obtained action policy of the sub-task can be repeatedly used in the upper levels of the hierarchy.”)
The Kawano reference does not explicitly teach the using of neural networks. In the same field of endeavor however, Reed teaches “the method comprising: for each of a plurality of iterations, wherein each iteration corresponds to one of the sub- trajectories: determining the latent variable for the corresponding local goal using a predictor neural network configured to predict the latent variable in accordance with predictor parameter values based on a preceding latent variable” (Reed [0023] “The method includes obtaining an observation characterizing a state of the environment subsequent to the agent performing a selected action. A latent representation of the observation is generated, including: (i) processing the observation using an encoder neural network to generate the latent representation of the observation, where the encoder neural network has been trained to process a given observation to generate a latent representation of the given observation that is predictive of latent representations”); and
“for each iteration, obtaining the corresponding sub-trajectory by, for each of a plurality of time steps: obtaining an observation characterizing a current state of the environment” (Reed [0059] “Each sequence of observations may be, e.g., a sequence of audio data segments from an audio waveform, a sequence of images of an environment that are captured at respective time points, a sequence of regions from a single image, or a sequence of sentences in a natural language.”);
“processing the observation using the action selection policy neural network to generate a policy output, wherein the action selection policy neural network is conditioned on the latent variable and generates the policy output in accordance with policy parameter values of the action selection policy neural network” (Reed abstract “generating a latent representation of the observation; processing the latent representation of the observation using a discriminator neural network to generate an imitation score; determining a reward from the imitation score; and adjusting the current values of the action selection policy neural network parameters based on the reward using a reinforcement learning training technique”); and
“selecting an action to be performed by the agent in response to the observation using the policy output” (Reed [0045] “training the action selection network to imitate expert observations in latent space may enable the action selection network to select actions that enable an agent to perform tasks more effectively, e.g., more quickly.”)
It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Kawano with that of Reed since a combination of known methods would yield predictable results. As shown in Reed, it is known in the art and RL in general to use neural networks to implement the specifics of the model. Thus when combined the references would operate in a known and expected manner.
Note that independent claims 18 and 19 recite the same substantial subject matter as independent claim 1, only differing in embodiment. The differences in embodiments, a system and non-transitory medium compared to a method are obvious variations of another and are taught by Reed abstract.
Regarding claim 4, the Kawano and Reed references have been addressed above. Reed further teaches wherein the previous latent variable is determined by determining the previous latent variable that is most likely to have conditioned the action-selection neural network to generate a previous sub-trajectory for the previous iteration according to the predictor neural network” (Reed [0023] “A latent representation of the observation is generated, including: (i) processing the observation using an encoder neural network to generate the latent representation of the observation, where the encoder neural network has been trained to process a given observation to generate a latent representation of the given observation that is predictive of latent representations generated by the encoder neural network of one or more other observations that are after the given observation in a sequence of observations; or (ii) processing the observation using the action selection policy neural network to generate the latent representation of the observation as an intermediate output of the action selection policy neural network”)
Regarding claim 5, the Kawano and Reed references have been addressed above. Reed further teaches “wherein the previous latent variable is determined based on a last observation from the previous sub-trajectory” (Reed [0004] “This specification describes a system implemented as computer programs on one or more computers in one or more locations that can train an encoder neural network to generate latent representations of observations that robustly represent generic features in the observations without being specialized towards solving a single task.”)
Regarding claim 6, the Kawano and Reed references have been addressed above. Reed further teaches wherein determining the previous latent variable comprises inputting the last observation from the previous sub-trajectory into the predictor neural network to predict the previous latent variable” (Reed [0055] “the training system trains a discriminator neural network to process latent representations of observations to generate imitation scores that discriminate between agent observations and expert observations. In parallel, the training system trains the action selection network to select actions that result in agent observations which the discriminator network misclassifies as being expert observations”)
Regarding claim 7, the Kawano and Reed references have been addressed above. Reed further teaches “wherein each latent variable is linearly related to the preceding latent variable” (Reed [0070] “At each time step, the state of the environment 208 at the time step (as characterized by the observation 210) depends on the state of the environment 208 at the previous time step and the action 204 performed by the agent 206 at the previous time step”)
Regarding claim 8, the Kawano and Reed references have been addressed above. Reed further teaches “wherein the method is initialized by determining a first latent variable based on a first observation characterizing a first state of the environment” (Reed abstract “obtaining an observation characterizing a state of the environment subsequent to the agent performing a selected action; generating a latent representation of the observation”)
Regarding claim 10, the Kawano and Reed references have been addressed above. Kawano further teaches “further comprising updating one or both of the policy parameter values and the predictor parameter values based at least on one or more of the sub-trajectories” (pg. 3 ¶1 “Policy [Symbol font/0x70] is a mapping from each state s to the probability of taking action a when in state s. In the DP calculation based on the MDP, policy [Symbol font/0x70] is decided so as to maximize the value function , which denotes the expected return when starting in s and following [Symbol font/0x70] thereafter for T time steps. T is called the planning horizon”)
Regarding claim 11, the Kawano and Reed references have been addressed above. Kawano further teaches “further comprising one or both of: updating the policy parameter values based on a first objective that aims to maximize an entropy of the policy; and updating the predictor parameter values based on a second objective that aims to minimize a cross-entropy between consecutive latent variables” (pg. 3 ¶1 “Policy [Symbol font/0x70] is a mapping from each state s to the probability of taking action a when in state s. In the DP calculation based on the MDP, policy [Symbol font/0x70] is decided so as to maximize the value function , which denotes the expected return when starting in s and following [Symbol font/0x70] thereafter for T time steps. T is called the planning horizon”)
Regarding claim 13, the Kawano and Reed references have been addressed above. Kawano further teaches “wherein updating the predictor parameter values comprises an update that attempts to minimize a difference between consecutive latent variables” (pg. 1 last ¶ “Once the learning process of a sub-task is completed, the obtained action policy of the sub-task can be repeatedly used in the upper levels of the hierarchy. The reuse of the action policy of the sub-task can reduce the size of the learning space.”)
Regarding claim 15, the Kawano and Reed references have been addressed above. Kawano further teaches “wherein the observations relate to a real-world environment and wherein the selected action relates to an action to be performed by a mechanical agent” (pg. 2 §II ¶1 “Figure 1 shows the environment of the multi-robot delivery mission assumed in this research” )
Regarding claim 16, the Kawano and Reed references have been addressed above. Reed further teaches “wherein the method controls the agent to perform a task while executing the options, the method further comprising using the predictor neural network and the action selection policy neural network to control the mechanical agent to perform the task while interacting with the real-world environment by obtaining the observations from one or more sensors sensing the real-world environment and using the policy output to select actions to control the mechanical agent to perform the task” (Reed [0023] “there is provided a method for training an action selection policy neural network, where the action selection policy neural network has a set of action selection policy neural network parameters, where the action selection policy neural network is configured to process an observation characterizing a state of an environment in accordance with values of the action selection policy neural network parameters to generate an action selection policy output, where the action selection policy output is used to select an action to be performed by an agent interacting with an environment. The method includes obtaining an observation characterizing a state of the environment subsequent to the agent performing a selected action. A latent representation of the observation is generated, including: (i) processing the observation using an encoder neural network to generate the latent representation of the observation, where the encoder neural network has been trained to process a given observation to generate a latent representation of the given observation that is predictive of latent representations generated by the encoder neural network of one or more other observations that are after the given observation in a sequence of observations; or (ii) processing the observation using the action selection policy neural network to generate the latent representation of the observation as an intermediate output of the action selection policy neural network.”)
Claim(s) 2-3 and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kawano in view Reed further in view of Tan, Kai Liang, et al. "Robustifying reinforcement learning agents via action space adversarial training." [herein Tan].
Regarding claims 2 and 22, the Kawano and Reed references have been addressed above. They do not explicitly teach adding perturbations to the latent variable. Tan however teaches “wherein determining the latent variable comprises determining the previous latent variable and adding a perturbation to the previous latent variable to generate the latent variable” (Tan pg. 3 right col. last ¶ “After obtaining the adversarial perturbations and adding it to the nominal actions, we train the DRL agent using standard policy gradient methods,”)
It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Kawano and Reed with that of Tan since “we show that a well performing DRL agent that is initially susceptible to action space perturbations (e.g. actuator attacks) can be robustified against similar perturbations through adversarial training” Tan abstract. This shows that these techniques allow the model and training to be more effective.
Regarding claim 3, the Kawano, Reed, and Tan references have been addressed above. Tan further teaches “wherein the perturbation is determined from a perturbation distribution, the perturbation distribution being one of a uniform distribution or a Gaussian distribution” (Tan pg. 3 left col. last ¶ “The next step is to define the attack model.We introduce a specific perturbation δ ∈ B for each data point x, where B is the set of allowed perturbations (B ∈ Rm).”)
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kawano and Reed further in view of Van Seijen et al. US 2018/0165603 [herein VS].
Regarding claim 9, the Kawano and Reed references have been addressed above. While they generally recite Markov decision processes, VS more specifically teaches “wherein latent variables over the iterations form a Markov chain” (VS [0084] “By assigning a stationary policy to each of the agents, the sequence of random variables Y.sub.0, Y.sub.1, Y.sub.2, . . . , with Y.sub.tϵY, is a Markov chain. This can be formalized by letting μ={π.sup.1 . . . π.sup.n} define a set of stationary policies for all agents, and M=Π.sup.1× . . . ×Π.sup.n be the space of all such sets”)
It would have been obvious to one having ordinary skill in the art at the time that the invention was effectively filed to combine the teachings of Kawano and Reed with that of VS since a combination of known methods would yield predictable results. As shown in VS, it is known in the art that Markov chains are formed in RL and therefore this would apply to the system above in an obvious and predicable manner.
Allowable Subject Matter
No prior art has been cited for claims 12 and 14 however the claims remain rejected under 101.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN W FIGUEROA whose telephone number is (571)272-4623. The examiner can normally be reached Monday-Friday, 10AM-6PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEVIN W FIGUEROA
Primary Examiner
Art Unit 2124
/Kevin W Figueroa/Primary Examiner, Art Unit 2124