Prosecution Insights
Last updated: August 18, 2026
Application No. 17/918,365

LEARNING OPTIONS FOR ACTION SELECTION WITH META-GRADIENTS IN MULTI-TASK REINFORCEMENT LEARNING

Final Rejection §101§103§112
Filed
Oct 12, 2022
Priority
Jun 05, 2020 — provisional 63/035,467 +1 more
Examiner
DAY, ROBERT N
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
GDM Holding LLC
OA Round
2 (Final)
23%
Grant Probability
At Risk
3-4
OA Rounds
3m
Est. Remaining
46%
With Interview

Examiner Intelligence

Grants only 23% of cases
23%
Career Allowance Rate
6 granted / 26 resolved
-31.9% vs TC avg
Strong +23% interview lift
Without
With
+23.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
22 currently pending
Career history
63
Total Applications
across all art units

Statute-Specific Performance

§101
34.1%
-5.9% vs TC avg
§103
39.7%
-0.3% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
11.5%
-28.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 26 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This action is in response to the amendments filed 10 February 2026. Claims 17-22 are cancelled. Claims 1, 3, 23, and 24 are amended. Claims 1-16, 23, and 24 are pending and have been examined. Information Disclosure Statement The information disclosure statement (IDS) submitted on 10 February 2026 is being considered by the examiner. Response to Arguments Applicant' s arguments, see page 10, filed 10 February 2026, with respect to the objection to the specification have been fully considered and are persuasive. The objection to the specification has been withdrawn. APPLICANT'S ARGUMENT: Applicant argues (page 10, paragraph 7) that "Pursuant to MPEP § 1893.03(e), the requirement under 37 CFR 1.72(b) that the abstract commence on a separate sheet does not apply to the copy of the international application furnished by the International Bureau." EXAMINER'S RESPONSE: Examiner agrees. The objection to the specification has been withdrawn in light of arguments and/or amendments. Applicant' s arguments, see page 11, filed 10 February 2026, with respect to the rejections of Claims 18 and 19 under 35 U.S.C. 112(b) have been fully considered and are persuasive. The rejections of Claims 18 and 19 under 35 U.S.C. 112(b) have been withdrawn. APPLICANT'S ARGUMENT: Applicant argues (page 11, paragraph 1) that "Claims 18-19 are being canceled. Withdrawal of the rejection is therefore respectfully requested." EXAMINER'S RESPONSE: Examiner agrees. The rejections of Claims 18 and 19 under 35 U.S.C. 112(b) have been withdrawn in light of arguments and/or amendments. Applicant' s arguments, see page 11, filed 10 February 2026, with respect to the rejections of Claims 23 and 24 under 35 U.S.C. 112(d) have been fully considered and are persuasive. The rejections of Claims 23 and 24 under 35 U.S.C. 112(d) have been withdrawn. APPLICANT'S ARGUMENT: Applicant argues (page X, paragraph X) that "Claims 23-24 have been amended to clarify the claimed subject matter. Withdrawal of the rejection is respectfully requested." EXAMINER'S RESPONSE: Examiner agrees. The rejections of Claims 23 and 24 under 35 U.S.C. 112(d) have been withdrawn in light of arguments and/or amendments. Applicant' s arguments, see pages 11-13, filed 10 February 2026, with respect to the rejections of Claims 1-16, 18-19, and 23-24 under 35 U.S.C. 101 have been fully considered and are persuasive. The rejections of Claims 1-16, 18-19, and 23-24 under 35 U.S.C. 101 have been withdrawn. APPLICANT'S ARGUMENT: Applicant argues (page 11, paragraphs 4-5) that "This rejection should be withdrawn because the claimed invention provides a technical solution to a technical problem. ¶ Applicant submits that, even if the claims recite an abstract idea (which is not conceded), the claims are not directed to an abstract idea because any abstract ideas recited the claims are integrated into a practical application." Applicant argues (page 12, paragraph 1) that "the claimed invention provides an improvement in the functioning of a computer, e.g., by increasing the data efficiency of the reinforcement learning process, allowing complex tasks to be learned faster and with fewer computational resources than existing methods. ¶ The claimed invention achieves these improvements by automating the discovery of 'options' (a set of option policy neural networks). Specifically, by training internal reward networks (option reward neural networks) to optimize the extrinsic return ( defined by task rewards), the system 'meta-learns' sub-goals that are generalizable across tasks, rather than requiring hand-designed sub-goals or computationally expensive relearning for every new task." Applicant argues (page 13, paragraph 5) that ' These claim limitations correspond directly to the technical improvements described in the Specification. ... Thus, the claims include the steps necessary to achieve the described technical improvements." EXAMINER'S RESPONSE: Examiner agrees. The rejections of Claims 1-16, 18-19, and 23-24 under 35 U.S.C. 101 have been withdrawn in light of arguments and/or amendments. Applicant' s arguments, see pages 14-16, filed 10 February 2026, with respect to Claims 1, 3, 4, 6, 16, 18, 23, and 24 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. APPLICANT'S ARGUMENT: Applicant argues (page 14, paragraph 3) that "Florensa and Henderson do not teach or suggest at least this combination of features [of amended Claim 1]." Applicant argues (page 14, paragraph 5) that "in the Florensa framework, the internal reward structures for the skills are static and do not adapt to the specific task at hand. ¶ In contrast, the claimed invention utilizes a coupled training architecture where both the set of option reward neural networks and the manager neural network are jointly trained by a first reinforcement learning training technique using the task rewards. Florensa does not teach or suggest this joint refinement of the set of option reward neural networks and the manager neural network using task rewards." Applicant argues (page 15, paragraph 3) that "the claimed invention utilizes a reinforcement learning framework where the option reward neural networks are themselves trained via a 'first reinforcement learning training technique' specifically to maximize task rewards. Henderson does not teach or suggest training the option reward neural networks using task rewards derived from the environment as recited by the claim." EXAMINER'S RESPONSE: Examiner notes that Applicant's arguments are moot. Amended Claim 1 is now rejected under 35 U.S.C. 103 in view of Florensa in view of Henderson in view of Frans. In the rejection below, Frans is relied on to teach the newly recited limitation of amended Claim 1 where the option reward networks and the manager network are trained jointly by a first reinforcement learning training technique using the task rewards. Specification The objection to the abstract of the disclosure is withdrawn in light of arguments and/or amendments. Claim Rejections - 35 USC § 112(b) The rejections of Claims 18 and 19 under 35 U.S.C. 112(b) are withdrawn in light of amendments. Claim Rejections - 35 USC § 112(d) The rejections of Claims 23 and 24 under 35 U.S.C. 112(d) are withdrawn in light of arguments and/or amendments. Claim Rejections - 35 USC § 101 The rejections of Claims 1-16, 18, 19, 23, and 24 under 35 U.S.C. 101 for abstract idea without significantly more is withdrawn in light of arguments and/or amendments. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 3, 4, 6, 16, 23, and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Florensa, et al., "Stochastic neural networks for hierarchical reinforcement learning" (hereinafter "Florensa") in view of Henderson, et al., "OptionGAN: Learning joint reward-policy options using generative adversarial inverse reinforcement learning" (hereinafter "Henderson") in view of Frans, et al., "Meta Learning Shared Hierarchies" (hereinafter "Frans"). Regarding Claim 1, Florensa teaches: A computer-implemented (Florensa, p. 6, 6 Experiments: "we report our stronger results with the exact same setting in Appendix B. Our hyperparameters for the neural network architectures and algorithms are detailed in the Appendix A and the full code is available" and footnote 4: "Code available at: https://github.com/florensacc/snn4hrl") system for controlling an agent to perform a plurality of tasks (Florensa, p. 1, Abstract: "we propose a general framework that first learns useful skills in a pre-training environment, and then leverages the acquired skills for learning faster in downstream tasks," where Florensa's framework that learns corresponds to the instant agent) while interacting with an environment (Florensa, p. 4, 5.2 Stochastic Neural Networks For Skill Learning: "To learn several skills at the same time, we propose to use Stochastic Neural Networks (SNNs), a general class of neural networks with stochastic units in the computation graph. ... ¶ For our purpose, we use a simple class of SNNs, where latent variables with fixed distributions are integrated with the inputs to the neural network (here, the observations from the environment) to form a joint embedding"), wherein the system is configured to, at each of a plurality of time steps, process an input comprising an observation characterizing a current state of the environment (Florensa, p. 8, 5.4 Learning High-Level Policies: "Given a span of K skills learned during the pre-training task ... . we leverage the provided skills by freezing them and training a high-level policy (Manager Neural Network) that operates by selecting a skill and committing to it for a fixed amount of steps T . ... For any given task M ∈ M we train a new Manager NN on top of the common skills. Given the factored representation of the state space S M as S agent and S rest M , the high-level policy receives the full state as input," where Florensa's state comprises observations, as in p. 4, 5.2 Stochastic Neural Networks For Skill Learning: "latent variables with fixed distributions are integrated with the inputs to the neural network (here, the observations from the environment)") to generate an output for selecting an action to be performed by the agent (Florensa, p. 5, 5.4 Learning High-Level Policies: "Given the factored representation of the state space S M as S agent and S rest M , the high-level policy receives the full state as input, and outputs the ... distribution from which we sample a discrete action"), and receive a task reward in response to the action (Florensa, p. 5, 5.4 Learning High-Level Policies: "Given a span of K skills learned during the pre-training task, we now describe how to use them as basic building blocks for solving tasks where only sparse reward signals are provided"), the system comprising: a manager neural network (Florensa, p. 5, 5.4 Learning High-Level Policies: "Instead of learning from scratch the low-level controls, we leverage the provided skills by freezing them and training a high-level policy (Manager Neural Network) that operates by selecting a skill and committing to it for a fixed amount of steps T "), and a set of option policy neural networks (Florensa, p. 4, Fig. 1(b), depicting an integration of multiple policy networks, each having input of unique latent variable z n , and 5.2 Stochastic Neural Networks For Skill Learning: "we study using a simple bilinear integration, by forming the outer product between the observation and the latent variable (Fig. 1(b)). Note that ... the bilinear integration [effectively corresponds] to changing all the first hidden layer weights. ... Our bilinear integration already yields a large span of skills, hence no other type of SNNs is studied in this work") each for selecting a sequence of actions to be performed by the agent according to a respective option policy (Florensa, p. 3, 3 Preliminaries: "We define a discrete-time finite-horizon discounted Markov decision process (MDP) .... In policy search methods, we typically optimize a stochastic policy ... parametrized by 𝜃. The objective is to maximize its expected discounted return ... where τ = s 0 , a 0 , … denotes the whole trajectory," where Flroensa's trajectory of the optimized policy corresponds to the instant sequence of actions); wherein the manager neural network is configured to, at a time step: process the observation and data identifying one of the tasks currently being performed by the agent, according to parameter values of the manager neural network, to generate an output for selecting a manager action from a set of manager actions (Florensa, p. 6, Fig. 2, depicting a trained manager neural network at a single time step with input S M and output z , and 5.4 Learning High-Level Policies: "For any given task M ∈ M we train a new Manager NN on top of the common skills. Given the factored representation of the state space ..., the high-level policy receives the full state as input, and outputs the parametrization of a categorical distribution from which we sample a discrete action z out of K possible choices, corresponding to the K available skills. ... z dictates the policy to use during the following T time-steps. If the skills are encapsulated in a SNN, z is used in place of the latent variable," where Florensa's sampled action z corresponds to the generated output, and where Florensa's trained manager neural network is trained according to parameters by inherency under BRI), wherein the set of manager actions comprises possible actions that can be performed by the agent and a set of option selection actions, each option selection action selecting one of the option policy neural networks (Florensa, p. 5, 5.4 Learning High-Level Policies: "Given a span of K skills learned during the pre-training task, we now describe how to use them as basic building blocks for solving tasks .... [W]e leverage the provided skills by ... training a high-level policy (Manager Neural Network) that operates by selecting a skill and committing to it for a fixed amount of steps T ," where Florensa's z corresponds to a possible action and distribution corresponds to the set of option selection actions, as in p. 6, 5.4 Learning High-Level Policies: "the high-level policy receives the full state as input, and outputs the parametrization of a categorical distribution from which we sample a discrete action z out of K possible choices, corresponding to the K available skills"); wherein each option policy neural network is configured to (Florensa, p. 4, 5.2 Stochastic Neural Networks For Skill Learning: "we study using a simple bilinear integration, by forming the outer product between the observation and the latent variable (Fig. 1(b)). Note that ... the bilinear integration [effectively corresponds] to changing all the first hidden layer weights," where Florensa's first hidden layer weights corresponds to the instant policy network configuration), at each of a succession of time steps: process the observation for the time step (Florensa, p. 6, 5.4 Learning High-Level Policies: "If those skills are independently trained uni-modal policies, z dictates the policy to use during the following T time-steps"), according to an option policy defined by parameter values of the option policy neural network, to generate an output for selecting an action to be performed by the agent (Florensa, p. 4, 5.2 Stochastic Neural Networks For Skill Learning: "we use a simple class of SNNs, where latent variables with fixed distributions are integrated with the inputs to the neural network (here, the observations from the environment) to form a joint embedding, which is then fed to a standard feed-forward neural network (FNN) with deterministic units, that computes distribution parameters for a uni-modal distribution (e.g. the mean and variance parameters of a multivariate Gaussian)" and p. 10, 8 Discussion And Future Work: "we only used feedforward architectures and hence the decision of what skill to use next only depends on the observation at the moment of switching, not using any sensory information gathered while the previous skill was active"); wherein, when the selected manager action is an option selection action, the option policy neural network selected by the manager action generates the output for selecting an action for successive time steps (Florensa, p. 6, 5.4 Learning High-Level Policies: "For any given task M ∈ M we train a new Manager NN on top of the common skills. Given the factored representation of the state space S M as S agent and S rest M , the high-level policy receives the full state as input, and outputs the parametrization of a categorical distribution from which we sample a discrete action z out of K possible choices, corresponding to the K available skills") until an option termination criterion is met (Florensa, p. 6, 5.4 Learning High-Level Policies: "If those skills are independently trained uni-modal policies, z dictates the policy to use during the following T time-steps," where Florensa's fixed number of time steps corresponds to the instant termination criterion), and when the selected manager action is one of the possible actions that can be performed by the agent the output for selecting the action is the selected manager action (Florensa, p. 6, 5.4 Learning High-Level Policies: "For any given task M ∈ M we train a new Manager NN on top of the common skills. ... [T]he high-level policy receives the full state as input, and outputs the parametrization of a categorical distribution from which we sample a discrete action z out of K possible choices," where Florensa's sampled action z of K possible actions corresponds to the instant manager action); and ... option reward (Florensa, p. 4, 5.1 Constructing the pre-training environment: "we use a generic proxy reward as the only reward signal to guide skill learning. The design of the proxy reward should encourage the existence of locally optimal solutions, which will correspond to different skills the agent should learn. In other words, it encodes the prior knowledge about what high level behaviors might be useful in the downstream tasks, rewarding all of them roughly equally") ... configured to, for a time step: process the observation ... to generate an option reward ... (Florensa, p. 3, 3 Preliminaries: "We define a discrete-time finite-horizon discounted Markov decision process (MDP) ... in which ... r : S × A → - R max , R max a bounded reward function .... The objective is to maximize its expected discounted return, η π θ = E τ ∑ t = 0 T γ t r s t , a t , where τ = s 0 , a 0 , … denotes the whole trajectory") ... according to parameter values (Florensa, p. 5, 5.3 Information-Theoretic Regularization: "To penalize [entropy function] H Z C ... we modify the reward received at every step as specified in Eq. (1), where p ^ ... is an estimate of the posterior probability of the latent code z n sampled on rollout n , given the coordinates c t n at time t of that rollout," where Florensa's Eq. defines reward parameters); wherein the system is configured to ... train ... the manager neural network by a first reinforcement learning training technique using ... task rewards (Florensa, p. 6, 5.5 Policy Optimization: "For both the pre-training phase and the training of the high-level policies, we use Trust Region Policy Optimization (TRPO) as the policy optimization algorithm" and p. 5, 5.4 Learning High-Level Policies: "Given a span of K skills learned during the pre-training task, we now describe how to use them as basic building blocks for solving tasks where only sparse reward signals are provided. Instead of learning from scratch the low-level controls, we leverage the provided skills by freezing them and training a high-level policy (Manager Neural Network) that operates by selecting a skill and committing to it for a fixed amount of steps T "), and to train each of the option policy neural networks by a second reinforcement learning technique using the option reward generated by the respective option reward neural network (Florensa, p. 6, Algorithm 1: Skill training for SNNs with MI bonus, depicting algorithm for training option policy networks (i.e., SNNs), and p. 6, 5.5 Policy Optimization: "For both the pre-training phase and the training of the high-level policies, we use Trust Region Policy Optimization (TRPO) as the policy optimization algorithm (Schulman et al., 2015). ... The training on downstream tasks does not require any modifications, except that the action space is now the set of skills that can be used. For the pre-training phase, due to the presence of categorical latent variables, the marginal distribution of π a s is now a mixture of Gaussians instead of a simple Gaussian," where Florensa's pre-training corresponds to the second learning technique). Florensa teaches a system for controlling an agent, comprising an option reward configured to, for a time step: process an observation to generate an option reward according to parameter values. Florensa may not explicitly teach a set of option reward neural networks, one for each respective option policy neural network, each configured to ... process the observation, according to parameter values of the option reward neural network, to generate an option reward for the respective option policy neural network for a time step. However, Henderson teaches: a set of option reward neural networks, one for each respective option policy neural network, each configured to ... (Henderson, p. 3201, Reward-Policy Options Framework: "we extend the options framework for decomposing rewards as well as policies. ... In this case, an option is formulated by a tuple: I ω , π ω , β ω , r ω . Here, r ω is a reward option from which a corresponding intra-option policy π ω is derived. That is, each policy option is optimized with respect to its own local reward option. The policy-over-options not only chooses the intra-option policy, but the reward option as well: π Ω → r ω , π ω "): process the observation, according to parameter values of the option reward neural network, to generate an option reward for the respective option policy neural network (Henderson, p. 3202, Mixture-of-Experts as Options: "As we can see in Eq. 2, we formulate our discriminator loss in the same way, using each reward option and the policy-over-options as the experts and gating function respectively. This ensures that the policy-over-options specializes over the state space and converges to a deterministic selection of experts, where Henderson's state corresponds to the instant observation) ... for a time step (Henderson, p. 3200, Preliminaries and Notation, The Options framework: "we instead simplify to one-step options, where β ω s = 1 .... we find that our options still converge to temporally extended and interpretable actions," where Henderson's one-step option corresponds to the instant per time step). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Florensa regarding a system for controlling an agent, comprising an option reward configured to, for a time step, process an observation to generate an option reward according to parameter values with those of Henderson regarding a set of option reward neural networks, one for each respective option policy neural network, each configured to process the observation, according to parameter values of the option reward neural network, to generate an option reward for the respective option policy neural network for a time step. The motivation to do so would be to facilitate learning multiple underlying reward functions corresponding to option policies while training policies (Henderson, p. 3201, Reward-Policy Options Framework: "Based on the need to infer a decomposition of underlying reward functions from a wide range of expert demonstrations in one-shot transfer learning, we extend the options framework for decomposing rewards as well as policies. In this way, intra-option policies, decomposed rewards, and the policy-over-options can all be learned in concert in a cohesive framework"). The Florensa/Henderson/Frans combination teaches wherein the system is configured to train the manager neural network by a first reinforcement learning training technique using task rewards. The Florensa/Henderson/Frans combination does not explicitly teach wherein the system is configured to jointly train both the set of option reward neural networks and the manager neural network by a first reinforcement learning training technique. However, Frans teaches: wherein the system is configured to jointly train both the set of option reward neural networks and the manager neural network by a first reinforcement learning training technique using the task rewards (Frans, p. 4, Algorithm 1 Meta Learning Shared Hierarchies, lines 11-12, "Update θ to maximize expected return from 1/N timescale viewpoint" and "Update ϕ to maximize expected return from full timescale viewpoint," and p. 3, 4.1 Policy Update In MLSH: "We enter the joint update period, where both θ and ϕ are updated," where Frans's master policy parameters θ and shared parameters ϕ for sub-policies correspond to the instant manager and option networks, as in p. 3, 3 Problem Statement: "the shared parameter vector ϕ consists of a set of subvectors ϕ 1 , ϕ 2 , … , ϕ K , where each subvector ϕ k defines a sub-policy π k a s . The parameter θ is a separate neural network that switches between the sub-policies," and where the sub-policy parameters also define value networks according to Eq. 1, as in p. 3, 3 Problem Statement: "The meta-learning objective is to optimize the expected return during an agent's entire lifetime, over the sampled tasks. ... [Eq. (1)] ... This objective tries to find a shared parameter vector ϕ that ensures that ... the agent achieves high T time-step returns"). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Florensa/Henderson/Frans combination teaches regarding wherein the system is configured to train the manager neural network by a first reinforcement learning training technique using task rewards regarding wherein the system is configured to jointly train both the set of option reward neural networks and the manager neural network by a first reinforcement learning training technique. The motivation to do so would be to facilitate a training in a reinforcement learning scenario where learning of new tasks may occur quickly by taking advantage of past experiences (Frans, p. 2, 3 Problem Statement: "prior work is mostly focused on the single-task setting and doesn’t account for the multi-task structure as part of the algorithm. On the other hand, our work takes advantage of the multi-task setting as a way to learn temporally extended primitives. ¶ There has also been work in metalearning, where information from past experiences is used to learn quickly on specific tasks. ... While the prior work on metalearning optimizes to learn as much as possible in a small number of gradient updates, MLSH (our method) optimizes to learn quickly over a large number of policy gradient updates in the RL setting --- a regime not yet explored by prior work"). Regarding Claim 23, Florensa teaches: a method performed by one or more computers (Florensa, p. 5, 5.3 Information-Theoretic Regularization: "Given that we use a batch policy optimization method, we use all trajectories of the current batch") to perform precisely those operations recited for the system of Claim 1. Claim 23 is rejected under the same rationale as Claim 1 Regarding Claim 24, Florensa teaches: one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform (Florensa, p. 7, 6 Experiments: "Our hyperparameters for the neural network architectures and algorithms are detailed in the Appendix A and the full code is available [at: https://github.com/florensacc/snn4hrl]" and repository file README.md: "To reproduce the results, you should first have rllab and Mujoco v1.31 configured. Then, run the following commands in the root folder of rllab.... ¶ Then you can ... [t]rain a hierarchical policy on top of that SNN via python," where a non-transitory computer storage media storing instructions is inherent in training by way of the source) precisely those operations recited for the system of Claim 1. Claim 24 is rejected under the same rationale as Claim 1. Regarding Claim 3, the rejection of Claim 1 is incorporated. The Florensa/Henderson/Frans combination has been shown to teach: wherein the system is configured to train both the set of option reward neural networks and the manager neural network by the first reinforcement learning training technique using the task rewards (as recited in the rejection of Claim 1, Frans, p. 4, Algorithm 1 Meta Learning Shared Hierarchies, lines 11-12, "Update θ to maximize expected return from 1/N timescale viewpoint" and "Update ϕ to maximize expected return from full timescale viewpoint," and p. 3, 4.1 Policy Update In MLSH: "We enter the joint update period, where both θ and ϕ are updated," where Frans's master policy parameters θ and shared parameters ϕ for sub-policies correspond to the instant manager and option networks, as in p. 3, 3 Problem Statement: "the shared parameter vector ϕ consists of a set of subvectors ϕ 1 , ϕ 2 , … , ϕ K , where each subvector ϕ k defines a sub-policy π k a s . The parameter θ is a separate neural network that switches between the sub-policies," and where the sub-policy parameters also define value networks according to Eq. 1, as in p. 3, 3 Problem Statement: "The meta-learning objective is to optimize the expected return during an agent's entire lifetime, over the sampled tasks. ... [Eq. (1)] ... This objective tries to find a shared parameter vector ϕ that ensures that ... the agent achieves high T time-step returns"), and to train each of the option policy neural networks by the second reinforcement learning technique using the option reward generated by the respective option reward neural network (Florensa, p. 6, Algorithm 1: Skill training for SNNs with MI bonus, depicting algorithm for training option policy networks (i.e., SNNs), and p. 6, 5.5 Policy Optimization: "For both the pre-training phase and the training of the high-level policies, we use Trust Region Policy Optimization (TRPO) as the policy optimization algorithm (Schulman et al., 2015). ... The training on downstream tasks does not require any modifications, except that the action space is now the set of skills that can be used. For the pre-training phase, due to the presence of categorical latent variables, the marginal distribution of π a s is now a mixture of Gaussians instead of a simple Gaussian," where Florensa's pre-training corresponds to the second learning technique). The Florensa/Henderson/Frans combination teaches: after the option selection action and for a succession of time steps until the termination criterion is met: ... updating the parameter values of the respective option policy neural network selected by the option selection action using the option reward for the respective option policy neural network (Florensa, p. 6, Algorithm 1: Skill training for SNNs with MI bonus, line 5, "Collect rollout with z n fixed," which indicates selection of a single policy/task, and "we add an additional reward bonus .... As entropy is a measure of uncertainty, another interpretation of this bonus is that, given where the robot is, it should be easy to infer which skill the robot is currently performing. To penalize we modify the reward received at every step as specified in Eq. (1), where p ^ ... is an estimate of the posterior probability of the latent code z n sampled on rollout n , given the coordinates cnt at time t of that rollout," where Florensa's reward modified with bonus and rollout correspond to the instant option policy reward and until termination, respectively, and where Florensa indicates that return is optimized according to a reward calculated on a per-step basis, as in p. 3, 3 Preliminaries: "The objective is to maximize its expected discounted return, η π θ = E τ ∑ t = 0 T γ t r s t , a t "); after the termination criterion is met: updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards (Florensa, p. 6, Algorithm 1: Skill training for SNNs with MI bonus, line 8, "Modify R t n ← R t n + … , " where Florensa's z n is a parameter of reward function R n t ). Henderson further teaches: after the option selection action and for a succession of time steps until the termination criterion is met: updating the parameter values of the manager neural network using the task rewards (Henderson, p. 3202, Algorithm 2, OptionGAN, line 4, "Update ... policy-over-options parameters ζ ," where Henderson's policy over options corresponds to the manager, and p. 3201, Learning Joint Reward-Policy Options: "we reformulate our discriminator loss as a weighted mixture of completely specialized experts in Eq. 2. This allows us to update the parameters of the policy-over-options," where Eq. 2 has a term based on Eq. 7, which is calculated based on reward outputs per option policy). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans combination regarding processing the observation and data identifying one of the tasks currently being performed by the agent according to parameter values of the manager neural network with the further teachings of Henderson regarding after the option selection action and for a succession of time steps until the termination criterion is met: updating the parameter values of the manager neural network using the task rewards. The motivation to do so would be to facilitate training the manager network to select an option policy for a task in a more deterministic manner (Henderson, p. 3202, Mixture-of-Experts as Options: "To ensure that our MoE formulation converges to options in the optimal case, we must properly formulate our loss function such that the gating function specializes over experts. ... ¶ This can intuitively be interpreted as encouraging the gating function to increase the likelihood of choosing an expert when its loss is less than the average loss of all the experts. The gating function will thus move toward deterministic selection of experts"). Regarding Claim 4, the rejection of Claim 3 is incorporated. The Florensa/Henderson/Frans combination teaches: wherein updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards comprises: generating a trajectory comprising a sequence of one or more actions ... and corresponding observations and task rewards (Florensa, p. 3, 3 Preliminaries: "We define a discrete-time finite-horizon discounted Markov decision process (MDP) ... in which ... r : S × A → - R max , R max a bounded reward function .... The objective is to maximize its expected discounted return, η π θ = E τ ∑ t = 0 T γ t r s t , a t , where τ = s 0 , a 0 , … denotes the whole trajectory") ... selected by the respective option policy neural network selected by the option selection action (Florensa, p. 6, 5.4 Learning High-Level Policies: "For any given task M ∈ M we train a new Manager NN on top of the common skills. Given the factored representation of the state space S M as S agent and S rest M , the high-level policy receives the full state as input, and outputs the parametrization of a categorical distribution from which we sample a discrete action z out of K possible choices, corresponding to the K available skills"). Henderson further teaches: updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards from the trajectory (Henderson, p. 3202, Algorithm 2: OptionGAN, lines 3 and 4, "Sample trajectories τ N ~ π Θ i " and "Update discriminator parameters θ ^ , ω and policy-over-options parameters ζ " where Henderson's reward parameters are, p. 3201, Learning Joint Reward-Policy Options: "The reward for a given state is composed as: R Ω , Θ ^ s ... where ... θ ^ ∈ Θ ^ are the parameters of the ... reward options"). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans combination regarding generating a trajectory comprising a sequence of one or more actions and corresponding observations and task rewards with the further teachings of Henderson regarding updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards from the trajectory. The motivation to do so would be to facilitate a learning scenario supporting end-to-end training of specialized experts (Henderson, p. 3201, Learning Joint Reward-Policy Options: "The use of one-step options allows us to learn a policy-over-options in an end-to-end fashion as a Mixture-of-Experts formulation. In the one-step case, selecting an option ... using the policy-over-options ... can be viewed as a mixture of completely specialized experts"). Regarding Claim 6, the rejection of Claim 3 is incorporated. The Florensa/Henderson/Frans combination teaches: wherein updating one or more of the parameter values of the manager neural network, the parameter values of the respective option policy neural network, and the parameter values of the option reward neural network, comprises updating based on an n-step return (Florensa, p. 3, 3 Preliminaries: "We define a discrete-time finite-horizon discounted Markov decision process (MDP) ... in which ... r : S × A → - R max , R max a bounded reward function .... The objective is to maximize its expected discounted return, η π θ = E τ ∑ t = 0 T γ t r s t , a t , where τ = s 0 , a 0 , … denotes the whole trajectory," where Florensa's expected value of sum from 0 to T corresponds to the instant n-step return). Regarding Claim 16, the rejection of Claim 1 is incorporated. The Florensa/Henderson/Frans combination teaches: wherein the set of option policy neural networks comprises a set of option policy neural network heads on a shared option policy neural network body (Florensa, p. 4, Fig. 1(b), depicting an integration of multiple policy networks as heads of a single network, each having input of unique latent variable z n , and 5.2 Stochastic Neural Networks For Skill Learning: "we study using a simple bilinear integration, by forming the outer product between the observation and the latent variable (Fig. 1(b)). Note that ... the bilinear integration [effectively corresponds] to changing all the first hidden layer weights. ... Our bilinear integration already yields a large span of skills, hence no other type of SNNs is studied in this work") Henderson has been shown to teach: wherein the set of option reward neural networks comprises a set of option reward neural network heads on a shared option reward neural network body (as recited in the rejection of Claim 1, Henderson, p. 3202, Mixture-of-Experts as Options: "As we can see in Eq. 2, we formulate our discriminator loss in the same way, using each reward option and the policy-over-options as the experts and gating function respectively. This ensures that the policy-over-options specializes over the state space and converges to a deterministic selection of experts," where Henderson's policy-over-options corresponds to the instant shared body). Claims 2, 5, 7, and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Florensa, et al., "Stochastic neural networks for hierarchical reinforcement learning" (hereinafter "Florensa") in view of Henderson, et al., "OptionGAN: Learning joint reward-policy options using generative adversarial inverse reinforcement learning" (hereinafter "Henderson") in view of Frans, et al., "Meta Learning Shared Hierarchies" (hereinafter "Frans") in view of Xu, et al., "Meta-gradient reinforcement learning" (hereinafter "Xu"). Regarding Claim 2, the rejection of Claim 1 is incorporated. Henderson has been shown to teach: wherein the system is configured to train each option reward neural network using the task reward in a ... training technique ... to optimize a return from the environment ...(as recited in the rejection of Claim 1, Henderson, p. 3201, Reward-Policy Options Framework: "an option is formulated by a tuple: I ω , π ω , β ω , r ω . Here, r ω is a reward option from which a corresponding intra-option policy π ω is derived. That is, each policy option is optimized with respect to its own local reward option. The policy-over-options not only chooses the intra-option policy, but the reward option as well: π Ω → r ω , π ω ") in which parameter values of the option reward neural network are adjusted based on the agent's interaction with the environment under control of the respective option policy neural network (as recited in the rejection of Claim 1, Henderson, p. 3202, Mixture-of-Experts as Options: "As we can see in Eq. 2, we formulate our discriminator loss in the same way, using each reward option and the policy-over-options as the experts and gating function respectively. This ensures that the policy-over-options specializes over the state space and converges to a deterministic selection of experts, where Henderson's state corresponds to the instant observation). The Florensa/Henderson/Frans combination teaches using the task reward in a training technique. The Florensa/Henderson/Frans combination does not explicitly teach using the task reward in a meta-gradient training technique. However, Xu teaches: using the task reward in a meta-gradient training technique (Xu, p. 3, 1.1 Applying Meta-Gradients to Returns: "we view the return g as a function parameterised by meta-parameters η , which may be differentiated to understand its dependence on η . This in turn allows us to compute the gradient ∂ f / ∂ η of the update function with respect to the meta-parameters η , and hence the meta-gradient ∂ J ' τ ' , θ ' , η ' / ∂ η ," where Xu's return as a cumulative reward corresponds to the instant reward). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Florensa/Henderson/Frans combination regarding the system being configured to train each option reward neural network using the task reward in a training technique to optimize a return from the environment with those of Xu regarding using the task reward in a meta-gradient training technique. The motivation to do so would be to facilitate a reinforcement learning where learning parameters are adjusted during training based on model performance (Xu, p. 3, 1.1 Applying Meta-Gradients to Returns: "A typical RL algorithm would hand-select the meta-parameters, such as the discount factor γ and bootstrapping parameter λ , and these would be held fixed throughout training. ... In essence, our agent asks itself the question, 'which return results in the best performance?', and adjusts its meta-parameters accordingly"). Regarding Claim 5, the rejection of Claim 4 is incorporated. The Florensa/Henderson/Frans combination teaches updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards from the trajectory. The Florensa/Henderson/Frans combination does not explicitly teach back propagating gradients of an option reward objective function based on the task rewards from the trajectory through the respective option policy neural network and through the option reward neural network for the respective option policy neural network. However, Xu teaches: wherein updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards from the trajectory comprises: back propagating gradients of an option reward objective function (Xu, p. 5, 1.4 Conditioned Value and Policy Functions: "To deal with non-stationarity in the value function and policy, we utilise an idea similar to universal value function approximation .... ¶ The key idea is to provide the metaparameters η as an additional input to condition the value function and policy.... W η is the embedding matrix (or row vector, for scalar η ) that is updated by backpropagation during training," where Xu's value function as cumulative reward corresponds to the instant reward objective) based on the task rewards from the trajectory through the respective option policy neural network and through the option reward neural network for the respective option policy neural network (Xu, p. 2, 1 Meta-Gradient Reinforcement Learning Algorithms: "At the core of the algorithm is an update function ... that adjusts parameters from a sequence of experience τ t = { S t , A t , R t + 1 , … } consisting of states S , actions A and rewards R . The nature of the function is determined by meta-parameters η "). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans combination regarding updating the parameter values of the option reward neural network for the respective option policy neural network using the task rewards from the trajectory with those of Xu regarding back propagating gradients of an option reward objective function based on the task rewards from the trajectory through the respective option policy neural network and through the option reward neural network for the respective option policy neural network. The motivation to do so would be to facilitate adaption in training scenarios where approximation of returns become inaccurate (Xu, p. 5, 1.4 Conditioned Value and Policy Functions: "there is a danger that the value function v θ becomes inaccurate, since it may be approximating old returns. ¶ ... In this way, the agent explicitly learns value functions and policies that are appropriate for various η . The approximation problem becomes a little harder, but the payoff is that the algorithm can freely shift the meta-parameters without needing to wait for the approximator to 'catch up'"). Regarding Claim 7, the rejection of Claim 3 is incorporated. The Florensa/Henderson/Frans combination teaches: wherein the manager objective function (Florensa, p. 3, 3 Preliminaries: "In policy search methods, we typically optimize a stochastic policy ... parametrized by 𝜃. The objective is to maximize its expected discounted return, η π θ = E τ ∑ t = 0 T γ t r s t , a t , where τ = s 0 , a 0 , … denotes the whole trajectory," where Florensa's expected discounted return corresponds to the instant manager objective) and option policy objective function each comprise a respective reinforcement learning objective function (Florensa, p. 5, 5.3 Information-Theoretic Regularization: "It is desirable to have direct control over the diversity of skills that will be learned. To achieve this, we introduce an information-theoretic regularizer, inspired by recent success of similar objectives in encouraging interpretable representation learning in InfoGAN," where Florensa's regularizer corresponds to the instant option policy comprised objective function). The Florensa/Henderson/Frans combination teaches updating the parameter values of the manager neural network using the task rewards and updating the parameter values of the option reward neural network. The Florensa/Henderson/Frans combination does not explicitly teach wherein updating the parameter values of the manager neural network using the task rewards comprises backpropagating gradients of a manager objective function and wherein updating the parameter values of the respective option policy neural network comprises backpropagating gradients of an option policy objective function. However, Xu teaches: wherein updating the parameter values of the manager neural network using the task rewards comprises backpropagating gradients of a manager objective function, wherein updating the parameter values of the respective option policy neural network comprises backpropagating gradients of an option policy objective function (Xu, p. 5, 1.4 Conditioned Value and Policy Functions: "To deal with non-stationarity in the value function and policy, we utilise an idea similar to universal value function approximation .... The key idea is to provide the meta-parameters η as an additional input to condition the value function and policy, as follows: v θ η S = v θ S ; e η π θ η S = π θ S ; e η e η = W η η ... ¶... The key idea is to provide the meta-parameters η as an additional input to condition the value function and policy.... W η is the embedding matrix (or row vector, for scalar η ) that is updated by backpropagation during training," where Xu's value function and policy function correspond to the instant manager network and option policy network, respectively, and η corresponds to the instant parameters, and where Xu reasonably suggests a gradient of the meta-parameter for update, as in p. 3, 1.2 Meta-Gradient Prediction: "The key idea of the meta-gradient prediction algorithm is to adjust meta-parameters in the direction that achieves the best predictive accuracy"). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans combination regarding updating the parameter values of the manager neural network using the task rewards and updating the parameter values of the option reward neural network with those of Xu regarding wherein updating the parameter values of the manager neural network using the task rewards comprises backpropagating gradients of a manager objective function and wherein updating the parameter values of the respective option policy neural network comprises backpropagating gradients of an option policy objective function. The motivation to do so would be to facilitate training under circumstances where the return function is not stationary (Xu, p. 4, 1.4 Conditioned Value and Policy Functions: "One complication of the approach outlined above is that the return function g η τ is non-stationary, adapting along with the meta-parameters throughout the training process. As a result, there is a danger that the value function v θ becomes inaccurate, since it may be approximating old returns. ... ¶ To deal with non-stationarity in the value function and policy, we utilise an idea similar to universal value function approximation"). Regarding Claim 8, the rejection of Claim 7 is incorporated. Xu further teaches: wherein the gradients of the manager objective function (Xu, p. 3, 1.2 Meta-Gradient Prediction: "The objective of the TD( λ ) algorithm ... is to minimise the squared error between the value function approximator ... and the ... return ... [Eq. 8] ... where τ is a sampled trajectory starting with state S, and ∂ J ... is a semi-gradient," where Xu's value function corresponds to the instant manager objective, and comprises the policy gradient ∂ J ) and of the option policy objective function comprise respective policy gradients (Xu, p. 4, 1.3 Meta-Gradient Control: "The semi-gradient of the A2C objective ∂ J ... is defined as ... [Eq. 12]. The first term represents a control objective, encouraging the policy π θ to select actions that maximise the return," where Xu's policy π θ corresponds to the instant option objective). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Xu combination regarding updating the manager objective function and the option policy objective function with the further teachings of Xu regarding the gradients of the manager objective function and of the option policy objective function comprise respective policy gradients. The motivation to do so would be to facilitate learning scenarios where the agent may choose policies according to learned meta-parameters, allowing it to take advantage of an improved return function (Xu, p. 3, 1.1 Applying Meta-Gradients to Returns: "we view the return g as a function parameterised by meta-parameters η , which may be differentiated to understand its dependence on η . This in turn allows us to compute the gradient ... of the update function with respect to the meta-parameters η , and hence the meta-gradient ... In essence, our agent asks itself the question, 'which return results in the best performance?', and adjusts its meta-parameters accordingly"). Claims 9 and 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over Florensa, et al., "Stochastic neural networks for hierarchical reinforcement learning" (hereinafter "Florensa") in view of Henderson, et al., "OptionGAN: Learning joint reward-policy options using generative adversarial inverse reinforcement learning" (hereinafter "Henderson") in view of Frans, et al., "Meta Learning Shared Hierarchies" (hereinafter "Frans") in view of Bacon, et al., "The option-critic architecture" (hereinafter "Bacon"). Regarding Claim 9, the rejection of Claim 1 is incorporated. The Florensa/Henderson/Frans combination teaches that the option policy neural network selected by the manager action generates the output until an option termination criterion is met. The Florensa/Henderson/Frans combination does not explicitly teach a set of option termination neural networks, one for each respective option policy neural network, each configured to, at each of the time steps: process the observation, according to parameter values of the option reward neural network, to generate an option termination value for the respective option policy neural network wherein, for each option reward neural network, the option termination value determines whether the option termination criterion is met. However, Bacon teaches: further comprising a set of option termination neural networks, one for each respective option policy neural network (Bacon, p. 2, Learning Options: "We consider the call-and-return option execution model, in which an agent picks option ω according to its policy over options π Ω , then follows the intra-option policy π ω until termination (as dictated by β ω ), at which point this procedure is repeated. Let π ω , θ denote the intra-option policy of option ω parametrized by θ and β ω , θ , the termination function of ω parameterized by θ " and p. 5, Arcade Learning Environment: "We applied the option-critic architecture in the Arcade Learning Environment ... using a deep neural network to approximate the critic and represent the intra-option policies and termination functions"), each configured to, at each of the time steps: process the observation, according to parameter values of the option reward neural network, to generate an option termination value for the respective option policy neural network (Bacon, p. 4, Algorithm 1: Option-critic with tabular intra-option Q-learning, line 9, terms β ω , θ s ' for observation s ' , and line 15, "if β ω , θ terminates in s ' ", where parameter ω corresponds to the instant option policy), wherein, for each option reward neural network, the option termination value determines whether the option termination criterion is met (Bacon, p. 4, Algorithm 1: Option-critic with tabular intra-option Q-learning, lines 15 and 16, "if β ω , θ terminates in s ' choose new ω ", where ω corresponds to the instant option policy). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans regarding that the option policy neural network selected by the manager action generates the output until an option termination criterion is met with those of Bacon regarding a set of option termination neural networks, one for each respective option policy neural network, each configured to, at each of the time steps: process the observation, according to parameter values of the option reward neural network, to generate an option termination value for the respective option policy neural network wherein, for each option reward neural network, the option termination value determines whether the option termination criterion is met. The motivation to do so would be to facilitate training scenarios where option choices are suboptimal by providing early termination (Bacon p. 3, Learning Options: "when the option choice is suboptimal with respect to the expected value over all options ... it drives the gradient corrections up, which increases the odds of terminating. After termination, the agent has the opportunity to pick a better option using π Ω [i.e., an agent's policy over options]. ... The termination gradient theorem can be interpreted as providing a gradient-based interrupting Bellman operator"). Regarding Claim 11, the rejection of Claim 9 is incorporated. Bacon further teaches: wherein the system is configured to train the set of option termination neural networks by, after the termination criterion is met for a respective option policy neural network: updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards (Bacon, p. 3, Algorithms and Architecture: "we can now design a stochastic gradient descent algorithm for learning options. Using a two-timescale framework ..., we propose to learn ... while updating the ... termination functions at a slower rate" and p. 6, Arcade Learning Environment: "As a consequence of optimizing for the return, the termination gradient tends to shrink options over time. This is expected since in theory primitive actions are sufficient for solving any MDP. We tackled this issue by adding a small ξ = 0.01 term to the advantage function, used by the termination gradient A Ω s , ω + ξ = Q Ω s , ω – V Ω s + ξ . This term has a regularization effect, by imposing an ξ -margin between the value estimate of an option and that of the 'optimal' one reflected in V Ω ," where Bacon's value V is an expected return and based on the result of a reward function). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Bacon combination regarding a set of option termination networks that calculate an option termination value that determines whether the option termination criterion is met with the further teachings of Bacon regarding updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards. The motivation to do so would be to facilitate learning scenarios where the optimizing for the expected return based on reward values may take advantage of early termination of underperforming policies (Bacon, p. 3, Learning Options: "in our case, it follows as a direct consequence of the derivation and gives the theorem an intuitive interpretation: when the option choice is suboptimal with respect to the expected value over all options, the advantage function is negative and it drives the gradient corrections up, which increases the odds of terminating. After termination, the agent has the opportunity to pick a better option using π Ω "). Regarding Claim 12, the rejection of Claim 11 is incorporated. Bacon further teaches: wherein updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards comprises: generating a trajectory comprising a sequence of one or more actions selected by the respective option policy neural network selected by the option selection action, and corresponding observations and task rewards (Bacon, p. 2, Learning Options: "Suppose we aim to optimize directly the discounted return, expected over all the trajectories starting at a designated state s 0 and option ω 0 , the ρ Ω , θ , ϑ , s 0 , ω 0 = E Ω , θ , ω ∑ t = 0 ∞ γ t r t + 1 s 0 , ω 0 . Note that this return depends on the policy over options, as well as the parameters of the option policies and termination functions" and p. 4, Algorithm 1, "Option-critic with tabular intra-option Q-learning," where a non-terminating sequence of s and s ' for an option policy π ω , θ in the repeat loop corresponds to the instant a trajectory); and updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards from the trajectory (Bacon, p. 6, Arcade Learning Environment: "As a consequence of optimizing for the return, the termination gradient tends to shrink options over time. This is expected since in theory primitive actions are sufficient for solving any MDP. We tackled this issue by adding a small ξ = 0.01 term to the advantage function, used by the termination gradient A Ω s , ω + ξ = Q Ω s , ω – V Ω s + ξ . This term has a regularization effect, by imposing an ξ -margin between the value estimate of an option and that of the 'optimal' one reflected in V Ω ," where Bacon's value V is an expected return and based on the result of a reward function). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Bacon combination teaches regarding XXX with the further teachings of Bacon regarding wherein updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards comprises: generating a trajectory comprising a sequence of one or more actions selected by the respective option policy neural network selected by the option selection action, and corresponding observations and task rewards; and updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards from the trajectory. The motivation to do so would be to facilitate learning scenarios where the optimizing for the expected return based on reward values may take advantage of early termination of underperforming policies (Bacon, p. 3, Learning Options: "in our case, it follows as a direct consequence of the derivation and gives the theorem an intuitive interpretation: when the option choice is suboptimal with respect to the expected value over all options, the advantage function is negative and it drives the gradient corrections up, which increases the odds of terminating. After termination, the agent has the opportunity to pick a better option using π Ω "). Regarding Claim 13, the rejection of Claim 12 is incorporated. Bacon further teaches: wherein updating the parameter values of the option termination neural network for the respective option policy neural network using the task rewards from the trajectory comprises: back propagating gradients of an option termination objective function based on the task rewards from the trajectory through the respective option policy neural network and through the option termination neural network for the respective option policy neural network (Bacon, p. 4, Algorithm 1: Option-critic with tabular intra-option Q-learning, lines 13 and 14 in section "2. Options improvement," where the option network parameters θ and termination network parameters ϑ are updated according to the respective gradients for the given option policy π ω , θ , where a non-terminating sequence of s and s ' for the option policy π ω , θ in the repeat loop corresponds to the instant trajectory through the policy and termination networks). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Bacon combination regarding updating the parameter values of the option termination neural network for the respective option policy neural network with the further teachings of Bacon regarding back propagating gradients of an option termination objective function based on the task rewards from the trajectory through the respective option policy neural network and through the option termination neural network for the respective option policy neural network. The motivation to do so would be to facilitate developing learning scenarios without requiring additional rewards or subgoals while remaining flexible and efficient with respect to the learning environment (Bacon, p. 1, Abstract: "We derive policy gradient theorems for options and propose a new option-critic architecture capable of learning both the internal policies and the termination conditions of options, in tandem with the policy over options, and without the need to provide any additional rewards or subgoals. Experimental results in both discrete and continuous environments showcase the flexibility and efficiency of the framework"). Regarding Claim 14, the rejection of Claim 1 is incorporated. The Florensa/Henderson/Frans combination teaches: wherein the system is configured to train the manager neural network dependent on an estimated return comprising the expected task rewards from the environment (Florensa, p. 3, 3 Preliminaries: "We define a discrete-time finite-horizon discounted Markov decision process (MDP) ... in which ... r : S × A → - R max , R max a bounded reward function .... The objective is to maximize its expected discounted return, η π θ = E τ ∑ t = 0 T γ t r s t , a t , where τ = s 0 , a 0 , … denotes the whole trajectory") when selecting manager actions (Florensa, p. 6, Figure 2, "Hierarchical SNN architecture to solve downstream tasks," depicting the manager neural network sampling the discrete action z based on current state) according to current parameter values of the manager neural network (Florensa, p. 5, 5.4 Learning High-Level Policies: "we leverage the provided skills by freezing them and training a high-level policy (Manager Neural Network) that operates by selecting a skill and committing to it for a fixed amount of steps ...." and "The weights of the low level and high level neural networks could also be jointly optimized to adapt the skills to the task at hand. ... Nevertheless, we show in our experiments that frozen low-level policies are already sufficient to achieve good performance in the studied downstream tasks"). The Florensa/Henderson/Frans combination teaches training the manager neural network dependent on an estimated return comprising the expected task rewards from the environment. The Florensa/Henderson/Frans combination does not explicitly teach train the manager neural network dependent ... on a switching cost. However, Bacon teaches: train the manager neural network dependent ... on a switching cost (Bacon, p. 7, Discussion: "if one wanted to use additional pseudo-rewards, the option-critic framework would easily accommodate it. In this case, the internal policies and termination function gradients would simply need to be taken with respect to the pseudo-rewards instead of the task reward. A simple instance of this idea, which we used in some of the experiments, is to use additional rewards to encourage options that are indeed temporally extended by adding a penalty whenever a switching event occurs"). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans combination regarding training the manager neural network dependent on an estimated return comprising the expected task rewards from the environment with those of Bacon regarding training the manager neural network dependent on a switching cost. The motivation to do so would be to facilitate supporting processing scenarios where task options are temporally extended (Bacon, p. 7, Discussion: "We developed a general gradient-based approach ... in order to optimize a performance objective for the task at hand. ... A simple instance of this idea, which we used in some of the experiments, is to use additional rewards to encourage options that are indeed temporally extended by adding a penalty whenever a switching event occurs. Our approach can work seamlessly with any other heuristic for biasing the set of options towards some desirable property (e.g. compositionality or sparsity), as long as it can be expressed as an additive reward structure"). Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Florensa, et al., "Stochastic neural networks for hierarchical reinforcement learning" (hereinafter "Florensa") in view of Henderson, et al., "OptionGAN: Learning joint reward-policy options using generative adversarial inverse reinforcement learning" (hereinafter "Henderson") in view of Frans, et al., "Meta Learning Shared Hierarchies" (hereinafter "Frans") in view of Bacon, et al., "The option-critic architecture" (hereinafter "Bacon") in view of Xu, et al., "Meta-gradient reinforcement learning" (hereinafter "Xu"). Regarding Claim 10, the rejection of Claim 9 is incorporated. Bacon further teaches: wherein the system is configured to train (Bacon, p. 5, Arcade Learning Environment: "We fixed the learning rate for the intra-option policies and termination gradient to 0:00025 and used RMSProp for the critic") the option termination neural networks using the task rewards in a ... gradient training technique in which parameter values of the option termination neural network are adjusted ... to optimize a return from the environment (Bacon, p. 2, Learning Options: "Suppose we aim to optimize directly the discounted return, expected over all the trajectories starting at a designated state s 0 and option ω 0 .... We will take gradients of this objective with respect to θ and ϑ " and "We adopt a continual perspective on the problem of learning options. ... [W]e focus on learning option policies and termination functions, assuming they are represented using differentiable parameterized function approximators. ¶ ... Let ... β ω , ϑ [denote] the termination function of ω parameterized by ϑ ") based on the agents interaction with the environment under control of the respective option policy neural network (Bacon, p. 3, Algorithms and Architecture, Figure 1, "Diagram of the option-critic architecture. The option execution model is depicted by a switch ⊥ over the contacts ⊸ . A new option is selected according to π Ω only when the current option terminates," depicting Environment and observation s t ). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Bacon combination regarding a set of option termination neural networks, one for each respective option policy neural network, with the further teachings of Bacon regarding training the option termination neural networks using the task rewards in a gradient training technique in which parameter values of the option termination neural network are adjusted to optimize a return from the environment based on the agents interaction with the environment under control of the respective option policy neural network. The motivation to do so would be to facilitate scenarios where option policies are trained by distilling all experience available from the environment (Bacon, p. 2, Learning Options: "We adopt a continual perspective on the problem of learning options. At any time, we would like to distill all of the available experience into every component of our system: value function and policy over options, intra-option policies and termination functions. To achieve this goal, we focus on learning option policies and termination functions, assuming they are represented using differentiable parameterized function approximators"). The Florensa/Henderson/Frans/Bacon combination teaches the system being configured to train the option termination neural networks using the task rewards in a gradient training technique. The Florensa/Henderson/Frans/Bacon combination may not explicitly teach training the option ... neural networks using the task rewards in a meta-gradient training technique. However, Xu teaches: train the option ... neural networks using the task rewards in a meta-gradient training technique (Xu, p. 3, 1.1 Applying Meta-Gradients to Returns: "we view the return g as a function parameterised by meta-parameters η , which may be differentiated to understand its dependence on η . This in turn allows us to compute the gradient ∂ f / ∂ η of the update function with respect to the meta-parameters η , and hence the meta-gradient ∂ J ' τ ' , θ ' , η ' / ∂ η ," where Xu's return as a cumulative reward corresponds to the instant reward). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Bacon combination regarding the system being configured to train the option termination neural networks using the task rewards in a gradient training technique with those of Xu regarding train the option neural networks using the task rewards in a meta-gradient training technique. The motivation to do so would be to facilitate a reinforcement learning where learning parameters are adjusted during training based on model performance (Xu, p. 3, 1.1 Applying Meta-Gradients to Returns: "A typical RL algorithm would hand-select the meta-parameters, such as the discount factor γ and bootstrapping parameter λ , and these would be held fixed throughout training. ... In essence, our agent asks itself the question, 'which return results in the best performance?', and adjusts its meta-parameters accordingly"). Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Florensa, et al., "Stochastic neural networks for hierarchical reinforcement learning" (hereinafter "Florensa") in view of Henderson, et al., "OptionGAN: Learning joint reward-policy options using generative adversarial inverse reinforcement learning" (hereinafter "Henderson") in view of Frans, et al., "Meta Learning Shared Hierarchies" (hereinafter "Frans") in view of Bacon, et al., "The option-critic architecture" (hereinafter "Bacon") in view of Han, et al. "Multi-agent hierarchical reinforcement learning with dynamic termination." Regarding Claim 15, the rejection of Claim 14 is incorporated. The Florensa/Henderson/Frans/Bacon combination teaches training the manager neural network dependent on a switching cost. The Florensa/Henderson/Frans/Bacon combination does not explicitly teach wherein the switching cost is configured to reduce the task reward or return used to update the parameter values of the manager neural network. However, Han teaches: wherein the switching cost is configured to reduce the task reward or return used to update the parameter values of the manager neural network (Han, p. 8, 4 Method, Dynamic Option Termination: "agents that utilize the delayed Q-value of Equation 5 will make sub-optimal decisions whenever another agent terminates. To increase the predictability of agents, while allowing them to terminate flexibly when the task demands it, we propose to put a price δ on the decision to terminate the current option. Option termination is therefore no longer hard-coded, but becomes part of the agent's policy, which we call dynamic termination," where Han's price δ corresponds to the instant reward-reducing cost). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the Florensa/Henderson/Frans/Bacon combination regarding training the manager neural network dependent on a switching cost with those of Han regarding the switching cost being configured to reduce the task reward or return used to update the parameter values of the manager neural network. The motivation to do so would be to support processing scenarios where an agent's flexibility and predictability are balanced according to changes in environment (Han, p. 3, 1 Introduction: "We will refer to an agent's flexibility as the ability to switch options in response to changes in others or the environment. Furthermore, we will use predictability to measure how far an agent will commit to its broadcast option. In this paper, we propose an approach called dynamic termination, which allows an agent to choose whether to terminate its current option according to the state and others' options. This approach balances flexibility and predictability, combining the advantages of both" Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Riemer, et al., "Learning Abstract Options," teach a method of reinforcement learning comprising a two-level hierarchy of options and primitive actions that enables learning both simultaneously using a hierarchy of higher and lower-level options over varying extents of time. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT N DAY whose telephone number is (703)756-1519. The examiner can normally be reached M-F 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /R.N.D./Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Oct 12, 2022
Application Filed
Nov 14, 2025
Non-Final Rejection mailed — §101, §103, §112
Feb 03, 2026
Interview Requested
Feb 10, 2026
Examiner Interview Summary
Feb 10, 2026
Applicant Interview (Telephonic)
Feb 10, 2026
Response Filed
May 26, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12668265
Self-Driving Method, Training Method, and Related Apparatus
5y 3m to grant Granted Jun 30, 2026
Patent 12632783
FEDERATED CONTINUAL LEARNING
3y 10m to grant Granted May 19, 2026
Patent 12406181
METHOD, DEVICE, AND COMPUTER PROGRAM PRODUCT FOR UPDATING MODEL
4y 5m to grant Granted Sep 02, 2025
Patent 12229685
MODEL SUITABILITY COEFFICIENTS BASED ON GENERATIVE ADVERSARIAL NETWORKS AND ACTIVATION MAPS
4y 0m to grant Granted Feb 18, 2025
Study what changed to get past this examiner. Based on 4 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
23%
Grant Probability
46%
With Interview (+23.0%)
4y 1m (~3m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 26 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month