Prosecution Insights
Last updated: July 31, 2026
Application No. 17/432,366

CONTROLLING AGENTS USING LATENT PLANS

Non-Final OA §103
Filed
Aug 19, 2021
Priority
Feb 19, 2019 — provisional 62/807,740 +1 more
Examiner
YI, HYUNGJUN B
Art Unit
2146
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
4 (Non-Final)
30%
Grant Probability
At Risk
4-5
OA Rounds
0m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants only 30% of cases
30%
Career Allowance Rate
7 granted / 23 resolved
-24.6% vs TC avg
Strong +46% interview lift
Without
With
+45.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
24 currently pending
Career history
61
Total Applications
across all art units

Statute-Specific Performance

§101
1.4%
-38.6% vs TC avg
§103
94.8%
+54.8% vs TC avg
§102
3.8%
-36.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the claims filed on 02/26/2026. Claims 1-20 are pending for examination. This action is Final. Response to Arguments In response to Applicant’s argument that claim 1 recites “processing a policy input comprising (i) the current observation, (ii) the goal observation, and (iii) the selected latent plan using the policy neural network having a plurality of policy parameters and configured to generate a current action output that defines an action to be performed in response to the current observation,” and that “the proposed combination of Bacon and Zhu does not teach or suggest at least these features of claim 1,” (Remarks, page 9) the Examiner respectfully disagrees. Zhu teaches processing the current observation and goal observation because Zhu discloses that “action a at time t can be drawn by: a ∼ π(st, g|u) where u are the model parameters, st is the image of the current observation, and g is the image of the navigation target” (Zhu, page 4, col. 2, paragraph 1), and further discloses that “the inputs to the network are two images that represent the agent’s current observation and the target” and that “the model generates policy and value outputs” (Zhu, page 4, col. 2, paragraph 2). Bacon teaches the selected latent plan because Bacon discloses that “an agent picks option ω according to its policy over options πΩ, then follows the intra-option policy πω until termination” and that “πω,θ denote[s] the intra-option policy of option ω parametrized by θ” (Bacon, page 1727, col. 1, last paragraph). Thus, Zhu teaches the current-observation and goal-observation portions of the policy input, while Bacon teaches the selected latent plan/option and the corresponding parameterized intra-option policy used to generate action selections; together, the references teach or suggest processing a policy input comprising the current observation, goal observation, and selected latent plan to generate the claimed current action output. In response to Applicant’s argument that “Bacon does not describe processing the selected option ω using the intra-option policy… Instead, Bacon describes following the intra-option policy…” (Remarks, page 9) the Examiner respectfully disagrees. Bacon teaches that “an agent picks option ω according to its policy over options πΩ, then follows the intra-option policy πω until termination” and that “πω,θ denote[s] the intra-option policy of option ω parametrized by θ” (Bacon, page 1727, col. 1, last paragraph). Therefore, once option ω is selected, the selected option determines which intra-option policy πω is followed to generate primitive action selections. Bacon also teaches that “[t]he options framework … formalizes the idea of temporally extended actions” (Bacon, page 1727, col. 1, paragraph 3), and Figure 4 teaches that “[a]ll options (color-coded) are used by the policy over options in successful trajectories” (Bacon, page 1730, Figure 4), supporting the interpretation that each option represents a different path or action-selection constraint. Accordingly, under the broadest reasonable interpretation, the selected option/latent plan is part of the policy input because it conditions, indexes, or constrains the policy used to generate the action output; the claim does not require the selected latent plan to be concatenated as a separate input vector to a single neural-network layer. In response to Applicant’s argument that “Zhu does not describe selecting any latent plans, much less including any selected latent plans in an input to the model to generate policy and value outputs,” (Remarks, page 9) the Examiner respectfully disagrees because Zhu is not relied upon for selecting latent plans. Bacon is relied upon for this feature because Bacon discloses that “an agent picks option ω according to its policy over options πΩ” and then follows “the intra-option policy πω until termination” (Bacon, page 1727, col. 1, last paragraph), which teaches selecting a latent plan/option from a probability distribution over options. Zhu is relied upon for processing the current and goal observations because Zhu teaches that “st is the image of the current observation, and g is the image of the navigation target” (Zhu, page 4, col. 2, paragraph 1), and that “the inputs to the network are two images that represent the agent’s current observation and the target” (Zhu, page 4, col. 2, paragraph 2). Thus, Applicant’s argument addresses Zhu individually for a feature supplied by Bacon and does not rebut the combined Bacon/Zhu teaching. In response to Applicant’s argument that “even if Zhu’s input of ‘two images that represent the agent’s current observation and the target’ are combined with Bacon’s option ω in a model input, Bacon performs the intra-option policy… until the termination function of ω is satisfied, and thus would neither process the image representing the target nor condition the termination of following intra-option policy... on the image representing the target,” (Remarks pages 9-10) the Examiner respectfully disagrees. Claim 1 requires generating a current action output using a policy input comprising the current observation, goal observation, and selected latent plan; it does not require conditioning Bacon’s termination function on the goal image. Zhu supplies the target-image/current-image policy processing because Zhu teaches that “action a at time t can be drawn by: a ∼ π(st, g|u)” where “st is the image of the current observation, and g is the image of the navigation target” (Zhu, page 4, col. 2, paragraph 1), and that “the inputs to the network are two images that represent the agent’s current observation and the target” and the model generates “policy and value outputs” (Zhu, page 4, col. 2, paragraph 2). Bacon supplies the selected option/latent plan because Bacon teaches that the agent selects option ω according to πΩ and follows intra-option policy πω (Bacon, page 1727, col. 1, last paragraph). Thus, the rejection relies on Zhu’s policy network processing of current and goal images, not on Bacon’s termination function processing the goal image. In response to Applicant’s argument that claim 11 recites “for each observation action pair in the sequence, processing an input comprising the observation in the pair, the last observation in the sequence, and the latent plan using the policy neural network and in accordance with current values of the policy parameters to generate an action probability distribution for the pair,” (Remarks page 10) and that the cited combination fails to teach this feature, the Examiner respectfully disagrees. Zhu teaches processing an observation and target/goal observation to generate policy/action outputs because Zhu discloses that “action a at time t can be drawn by: a ∼ π(st, g|u)” where “st is the image of the current observation, and g is the image of the navigation target” (Zhu, page 4, col. 2, paragraph 1), and further teaches that “the inputs to the network are two images that represent the agent’s current observation and the target” and that “the model generates policy and value outputs” (Zhu, page 4, col. 2, paragraph 2). Bacon teaches the latent-plan portion and policy-parameter portion because Bacon discloses selecting option ω according to πΩ and following intra-option policy πω, where “πω,θ denote[s] the intra-option policy of option ω parametrized by θ” (Bacon, page 1727, col. 1, last paragraph). Thus, Zhu teaches the observation/last-observation or target-observation policy processing, and Bacon teaches the latent plan/option that conditions the parameterized policy used to generate the action probability distribution. In response to Applicant’s argument that “Zhu’s model therefore does not generate its output by processing an input that includes a selected latent plan,” (Remarks, page 10) the Examiner respectfully disagrees because Zhu is not relied upon alone for the latent plan. Zhu is relied upon for processing the observation and target/last-observation images because Zhu teaches that “the inputs to the network are two images that represent the agent’s current observation and the target” and that “the model generates policy and value outputs” (Zhu, page 4, col. 2, paragraph 2). Bacon is relied upon for the selected latent plan because Bacon teaches that “an agent picks option ω according to its policy over options πΩ, then follows the intra-option policy πω until termination” (Bacon, page 1727, col. 1, last paragraph). Therefore, the combined references teach or suggest the full limitation even though Zhu alone does not disclose the selected latent plan. In response to Applicant’s argument that “The termination of Bacon’s intra-option policy… is conditioned by using a separate termination function po, not by intaking an image of the target as input,” (Remarks page 10) the Examiner respectfully disagrees. Claim 11 requires using the policy neural network to generate an action probability distribution; it does not require the termination function to receive the target or last observation as an input. Bacon is relied upon for the selected latent plan and option-conditioned policy because Bacon teaches that “an agent picks option ω according to its policy over options πΩ, then follows the intra-option policy πω until termination” and that “πω,θ denote[s] the intra-option policy of option ω parametrized by θ” (Bacon, page 1727, col. 1, last paragraph). Zhu is relied upon for the observation/target-image policy processing because Zhu teaches “action a at time t can be drawn by: a ∼ π(st, g|u)” where “st is the image of the current observation, and g is the image of the navigation target” (Zhu, page 4, col. 2, paragraph 1). Accordingly, Applicant’s argument is directed to an unclaimed termination-function requirement and does not overcome the rejection. In response to Applicant’s argument that “Bacon, alone or in combination with Zhu, does not describe a single policy neural network that processes and conditions the generated output on all three inputs including ‘the observation in the pair, the last observation in the sequence, and the latent plan’ as recited by claim 11,” (Remarks, page 10) the Examiner respectfully disagrees. Zhu teaches a neural-network policy input including the observation and target/goal image because Zhu discloses that “the inputs to the network are two images that represent the agent’s current observation and the target” and that “the model generates policy and value outputs” (Zhu, page 4, col. 2, paragraph 2). Bacon teaches the latent plan/option and parameterized policy because Bacon discloses that “an agent picks option ω according to its policy over options πΩ” and follows “the intra-option policy πω,” where “πω,θ denote[s] the intra-option policy of option ω parametrized by θ” (Bacon, page 1727, col. 1, last paragraph). The claim does not require a particular architecture in which all three inputs are concatenated at the first layer of one neural network; under the broadest reasonable interpretation, the selected latent plan may condition, index, or select the policy component used to generate the action probability distribution, as taught by Bacon, while Zhu supplies the current/goal observation processing. Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/23/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Bacon et al., (Bacon, P. L., Harb, J., & Precup, D. (2017, February). The option-critic architecture. In Proceedings of the AAAI conference on artificial intelligence (Vol. 31, No. 1).), hereafter referred to as Bacon, in view of Zhu et al., (Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., & Farhadi, A. (2017, May). Target-driven visual navigation in indoor scenes using deep reinforcement learning. In 2017 IEEE international conference on robotics and automation (ICRA) (pp. 3357-3364). IEEE.), hereafter referred to as Zhu. Claim 1: Bacon teaches the following limitations: A computer-implemented method of controlling an agent interacting with an environment to perform a task, by selecting actions to be performed by an agent at each of a sequence of time steps (Bacon, page 1726, col. 2, last paragraph, “A Markov Decision Process consists of a set of states S, a set of actions A, a transition function P : S×A → (S → [0, 1]) and a reward function r : S×A→ R. For convenience, we develop our ideas assuming discrete state and action sets. However, our results extend to continuous spaces using usual measure-theoretic assumptions (some of our empirical results are in continuous tasks). A (Markovian stationary) policy is a probability distribution over actions conditioned on states,”, Bacon formalizes time-stepped control by selecting actions from a policy over states—i.e., a computer-implemented agent making discrete action choices at each step.) using a plan proposal neural network having a plurality of plan proposal parameters (Bacon, page 1729, col. 2, paragraph 1, “We chose this approach for our experiment with deep neural networks in the Arcade Learning Environment.” Bacon, page 1727, col. 1, last paragraph, “We consider the call-and-return option execution model, in which an agent picks option ω according to its policy over options πΩ , then follows the intra-option policy πω until termination (as dictated by βω), at which point this procedure is repeated. Let πω,θ denote the intra-option policy of option ω parametrized by θ and βω,ϑ, the termination function of ω parameterized by ϑ.”, In Bacon, the policy over options πΩ is a parametric network head that outputs a distribution over options (latent plans).) and configured to generate data defining a probability distribution over a space comprising a plurality of latent plans; (Bacon, page 1727, col. 1, last paragraph, “We consider the call-and-return option execution model, in which an agent picks option ω according to its policy over options πΩ , then follows the intra-option policy πω until termination (as dictated by βω), at which point this procedure is repeated. Let πω,θ denote the intra-option policy of option ω parametrized by θ and βω,ϑ, the termination function of ω parameterized by ϑ.”, In Bacon, the policy over options is a distribution over latent options. That distribution is the claimed plan-proposal distribution over latent plans.) where each of the plurality of latent plans represents a different path through the environment or a different action selection constraint to be imposed on a policy neural network; (Bacon, page 1727, col. 1, paragraph 3, “The options framework (Sutton, Precup, and Singh 1999; Precup 2000) formalizes the idea of temporally extended actions.”, Bacon, page 1730, figure 4, “All options (color-coded) are used by the policy over options in successful trajectories.”, An option dictates a multi-step course of action (a path) via its intra-option policy until termination; equivalently, conditioning on ω constrains action selection to πω(·|s). Thus options satisfy either prong: (i) different paths, and/or (ii) different action-selection constraints.) selecting, using the probability distribution, a latent plan from the space; (Bacon, page 1727, col. 1, last paragraph, “We consider the call-and-return option execution model, in which an agent picks option ω according to its policy over options πΩ , then follows the intra-option policy πω until termination (as dictated by βω), at which point this procedure is repeated. Let πω,θ denote the intra-option policy of option ω parametrized by θ and βω,ϑ, the termination function of ω parameterized by ϑ.”, Bacon expressly selects a latent plan (option) from a probability distribution (πΩ).) and (iii) the selected latent plan using the policy neural network having a plurality of policy parameters and configured to generate a current action output that defines an action to be performed in response to the current observation; (Bacon, page 1727, col. 1, last paragraph, “We consider the call-and-return option execution model, in which an agent picks option ω according to its policy over options πΩ , then follows the intra-option policy πω until termination (as dictated by βω), at which point this procedure is repeated. Let πω,θ denote the intra-option policy of option ω parametrized by θ and βω,ϑ, the termination function of ω parameterized by ϑ.”, Bacon teaches selecting an option ω according to a policy over options πΩ, and then generating primitive action selections according to the intra-option policy πω associated with the selected option. Under broadest reasonable interpretation, the selected option/latent plan is part of the policy input because it conditions or indexes the action policy used to generate the current action output. The claim does not require the selected latent plan to be concatenated with the observation images as a single numerical input vector to a single neural network layer. Rather, the selected latent plan need only be processed as part of the policy input used by the policy neural network to generate the action output, which Bacon teaches.) Zhu in the same field of neural network implementation teaches the following limitations which Bacon fails to teach: the method comprising, at a current time step in the sequence: receiving a current observation comprising an image of a current state of the environment at the current time step; (Zhu, page 4, col. 2, paragraph 1, “action a at time t can be drawn by: a ∼ π(st ,g|u) where u are the model parameters, st is the image of the current observation, and g is the image of the navigation target.”, Zhu explicitly states the current observation at time t is an RGB image from the agent’s camera.) receiving a goal observation comprising an image of a goal state of the environment that results in the agent successfully performing the task, wherein the task is completed when the environment reaches the goal state that is characterized by the goal observation at a future time step that is after the current time step (Zhu, figure 1, “The goal of our deep reinforcement learning model is to navigate towards a visual target with a minimum number of steps. Our model takes the current observation and the image of the target as input and generates an action in the 3D environment as the output.”, Zhu defines the goal observation as a goal image and states the agent takes additional steps “until reaching the destination”. Since the action at time t uses st, reaching the destination happens at some t’ > 1, i.e., the goal state that is reached at a future time step that is after the current time step.) processing the current observation and the goal observation (Zhu, page 4, col. 2, paragraph 2, “Overall, the inputs to the network are two images that represent the agent’s current observation and the target… the model generates policy and value outputs”, Zhu processes both current image and goal image jointly in a neural network.) processing a policy input comprising (i) the current observation, (ii) the goal observation, (Zhu, page 4, col. 2, paragraph 2, “Overall, the inputs to the network are two images that represent the agent’s current observation and the target… the model generates policy and value outputs”, Zhu supplies the (i) current image and (ii) goal image pathway and produces a fused embedding.) and causing the agent to perform the action defined by current the action output. (Zhu, page 3, col. 2, last paragraph, “For testing, a mobile robot keeps taking actions drawn from the policy distribution until reaching the destination.”) It would have been obvious to one of ordinary skill in the art before the filing date of the invention to modify Bacon’s option-critic architecture to use Zhu’s target-driven visual navigation input structure, including both a current observation image and a goal/target image, because Zhu teaches that providing the visual task objective as an input allows the agent to adapt to different targets without retraining for every new target. A person of ordinary skill would have had reason to apply Zhu’s target-conditioned visual representation to Bacon’s option-conditioned policy architecture so that Bacon’s selected option/intra-option policy would generate actions in view of both the current visual state and the desired visual goal. The modification amounts to using Zhu’s known target-driven visual policy input technique to improve Bacon’s reinforcement-learning control policy in a predictable way. Claims 19 and 20 are directed to a Computer Readable Storage Medium (CRM) or a system that disclose substantially similar limitations to that of claim 1. Therefore, the rejection of claims 1 similarly apply to claims 19 and 20. Claims 2 and 3 are rejected under 35 U.S.C. 103 as being unpatentable over Bacon in view of Zhu as applied for claims 1, 19, and 20 above, and in further view of Choudhury et al., (C. F., Ben-Akiva, M., & Abou-Zeid, M. (2010). Dynamic latent plan models. Journal of Choice Modelling, 3(2), 50-70.), hereafter referred to as Choudhury and Sebbane et al., (Sebbane, Y. B., & Sebbane, Y. B. (2014). Planning and decision making for aerialrobots. Springer.), hereafter referred to as Sebbane. Claim 2: Bacon and Zhu teaches the limitations of claim 1. Choudhury, in the same field of action space analysis, teaches the following limitations which the above fails to teach: The method of claim 1, further comprising, at a subsequent time step: receiving a subsequent observation characterizing a subsequent state of the environment that follows the current state; (Choudhury, page 54, paragraph 1, “Individuals choose among distinct plans (targets/tactics). Their subsequent decisions are based on these choices. The chosen plans and intermediate choices are latent or unobserved and only the final actions (maneuvers) are observed.”, an individual’s current and subsequent observations are recorded.) processing a policy input comprising (i) the subsequent observation, (ii) the goal observation, and (iii) the selected latent plan (Choudhury, page 58, paragraph 3, “According to the HMM assumption, the action observed at a given time period depends on the current plan. The plan and action of previous time periods affect the current action through the current plan.”, formulating the next best action for the agent includes processing the current latent plan as well as previous ones, which inherently involves processing all observations as well as the goal.) to generate a subsequent action output that defines an action to be performed in response to the subsequent observation; (Choudhury, page 58, paragraph 3, “According to the HMM assumption, the action observed at a given time period depends on the current plan. The plan and action of previous time periods affect the current action through the current plan.”, a choice of subsequent action is outputted for the agent to follow, is in response to a subsequent observation.) and causing the agent to perform the action defined by the subsequent action output. (Choudhury, page 53, paragraph 3, “The observed actions of the individuals depend on their latent plans. The utility of actions and the choice set of alternatives may differ depending on the chosen plan.”, an agents/individuals observed action is dependent on their latent plans. This indicates that the action performed by the agent (observable behavior) is determined by the current plan.) It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to have combined the teachings of Bacon and Zhu with that found in Choudhury and environmental observations of action to formulate a plan and action for the agent. A motivation for the combination is to utilize a plan which helps further refine the choice of action. (Choudhury, page 53, section 3, paragraph 1, “The general framework of latent plan models is schematically shown in Figure 2. At any instant, the decision maker makes a plan based on his/her current state. The choice of plan is unobserved and manifested through the choice of actions given the plan. The actions are reflected on the updated states.”, this shows how the plans are formulated from previous actions and observed states to determine a next best action.) Sebbane in the same field of neural network implementation, teaches the following limitation which the above prior art fails to teach: using the policy neural network (Sebbane, page 277, section 4.4.2, paragraph 3, “A third approach is based on the use of supervised artificial neural network models. Supervised Neural Networks require the environment to serve as a teacher by presenting corrective feedback that indicates how the outcome produced deviates from some target or goal state for the system. Back propagation learning algorithms use this feedback to adjust the weights that perform the mappings from inputs to predictions to improve the network’s accuracy.”, supervised neural networks can be used in combination with Choudhury to perform the processing of a policy input as described in the claims language.) It would have been obvious to one of ordinary skill in the art before the filing date of the invention to have combined the teachings of Bacon, Zhu, and Choudhury with that found in Sebbane and use a neural network to process the observations, goals, and plans of an environment to compute a best action to take for an agent. A motivation for the combination is to utilize the adaptive nature of neural networks to refine the prediction and adjustment of latent plans in dynamic uncertain environments, leveraging their ability to learn from real-time feedback. (Sebbane, page 183, paragraph 1, “Artificial Neural Networks are primarily used to generate the control signals directly from sensor data inputs. Artificial neural networks emulate biological neural networks and have been used to learn how to control systems by observing the way a human plans and learn in an on-line fashion how to best control a system by taking control action, rating the quality of the responses achieved when these actions are used, then adjusting the recipe used for generating planning actions so that the response of the system improves [56].”, by integrating neural networks with Choudhury’s latent plan modeling, the system can dynamically adjust plans and actions, enabling more robust decision-making.) Claim 3: Bacon, Zhu, Choudhury, and Sebbane teaches the limitations of claim 2. Choudhury further teaches: The method of claim 2, further comprising, at a subsequent time step: determining that criteria for selecting a new latent plan are not satisfied when the subsequent observation is received; (Choudhury, page 59, paragraph 2, “The plan may evolve dynamically as the immediate execution of the chosen merging plan may not be feasible. A driver may begin with a plan of normal merging and then change to a plan of forced merging as the merging lane comes to an end. The probabilities of transitions from one plan to another are affected by the risk associated with the merge and the characteristics of the driver such as impatience, urgency, and aggressiveness (latent) as well as a strong inertia to continue the previously chosen merging tactic (state dependence).”, plans are dynamically revised when initial criteria for their execution are unmet, such as a lack of feasible merging opportunities.) and processing a policy input comprising (i) the subsequent observation, (ii) the goal observation, and (iii) the selected latent plan (Choudhury, page 58, paragraph 3, “According to the HMM assumption, the action observed at a given time period depends on the current plan. The plan and action of previous time periods affect the current action through the current plan.”, formulating the next best action for the agent includes processing the current latent plan as well as previous ones, which inherently involves processing all observations as well as the goal.) in response to determining that the criteria are not satisfied. (Choudhury, page 59, paragraph 2, “The plan may evolve dynamically as the immediate execution of the chosen merging plan may not be feasible. A driver may begin with a plan of normal merging and then change to a plan of forced merging as the merging lane comes to an end. The probabilities of transitions from one plan to another are affected by the risk associated with the merge and the characteristics of the driver such as impatience, urgency, and aggressiveness (latent) as well as a strong inertia to continue the previously chosen merging tactic (state dependence).”, plans are dynamically revised when initial criteria for their execution are unsatisfied, such as a lack of feasible merging opportunities) Sebbane further teaches: using the policy neural network (Sebbane, page 277, section 4.4.2, paragraph 3, “A third approach is based on the use of supervised artificial neural network models. Supervised Neural Networks require the environment to serve as a teacher by presenting corrective feedback that indicates how the outcome produced deviates from some target or goal state for the system. Back propagation learning algorithms use this feedback to adjust the weights that perform the mappings from inputs to predictions to improve the network’s accuracy.”, supervised neural networks can be used in combination with Choudhury to perform the processing of a policy input as described in the claims language.) Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Bacon in view of Zhu as applied for claims 1, 19, and 20 above, and in further view of Choudhury. Claim 4: Bacon and Zhu teaches the limitations of claim 1. Choudhury further teaches: The method of claim 1, wherein selecting, using the probability distribution, a latent plan from the space of latent plans, comprises sampling a latent plan in accordance with the probability distribution. (Choudhury, page 55, section 4.1, paragraph 1, “At time t for individual n, the probability of observing a particular action j is the sum of probabilities that he/she is observed to execute action j given that the selected plan is l, over all sequences of plans that could have led to plan l.”, the choice of plan is based on sampling the probability distribution over plans. It is interpreted by the examiner that the “probability of observing” is synonymous to the sampling a latent plan in accordance with the probability distribution as disclosed in the claim.) The rationale for the combination of Bacon and Zhu with Choudhury is similar to that applied for claim 2 above. Claim 5: Bacon and Zhu teaches the limitations of claim 1. Choudhury further teaches: The method of claim 1, wherein the current action output defines a probability distribution over a set of actions that can be performed by the agent. (Choudhury, page 54, paragraph 2, “the observable history can often be summarized adequately with a probability distribution over the current state, and policies can be computed as a function of these distributions”, an outline of the system in Choudhury, shows how probability distributions over the current state (actions performed previously that have led up to this state) to compute policies (actions) for the agent to perform.) The rationale for the combination of Bacon and Zhu with Choudhury is similar to that applied for claim 2 above. Claims 6-8, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Bacon in view of Zhu as applied for claims 1, 19, and 20 above, and in further view of Sebbane. Claim 6: Bacon and Zhu teaches the limitations of claim 1. Sebbane further teaches: The method of claim 1, wherein the data defining the probability distribution over the space of latent plans are a mean and a variance (Sebbane, page 272, paragraph 1, “For the Gaussian Belief system, the distribution is a Gaussian with mean mt and covariance Ωt : P(x) = ⊂ (x|mt,Ωt).”, the Gaussian Distribution (a probability distribution) is found using the mean and covariance of a multi-variate distribution.) of a multi-variate distribution. (Sebbane, page 11, section 1.4.1.3, “Approximating the probability distribution of a random variable using samples or particles can lead to tractable algorithms for estimation and control [23], given two multivariate probability distributions p(x) and q(x).”, sampling the multivariate probability distributions p(x) and q(x) help formulate the mean and variance in the Gaussian Distribution above.) The rationale for the combination of Bacon and Zhu with Sebbane is similar to that as applied for claim 2 above. Claim 7: Bacon and Zhu teaches the limitations of claim 1. Sebbane further teaches: The method of claim 1, wherein the plan proposal neural network and the policy neural network have been trained jointly through self-supervised learning. (Sebbane, page 277, section 4.4.2, paragraph 3, “A third approach is based on the use of supervised artificial neural network models. Supervised Neural Networks require the environment to serve as a teacher by presenting corrective feedback that indicates how the outcome produced deviates from some target or goal state for the system. Back propagation learning algorithms use this feedback to adjust the weights that perform the mappings from inputs to predictions to improve the network’s accuracy.”, supervised neural networks are used to train the plan proposal neural network and policy neural network. It is interpreted by the examiner that the plan proposal neural network and policy neural network are synonymous in function in Sebbane when in combination with the plan proposal formulation of Choudhury.) The rationale for the combination of Bacon and Zhu with Sebbane is similar to that as applied for claim 2 above. Claim 8: Bacon and Zhu teaches the limitations of claim 1. Sebbane further teaches: The method of claim 1 wherein the plan proposal neural network is a feed-forward neural network. (Sebbane, page 210, section 3.5.1.3, “The current work uses three layers, feedforward networks with one input layer, one hidden layer and one output layer. The transfer functions used are the linear function and the hyperbolic tangent sigmoid.”, this quote explicitly shows how feed-forward neural networks are used specifically.) The rationale for the combination of Bacon and Zhu with Sebbane is similar to that as applied for claim 2 above. Claim 18: Bacon and Zhu teaches the limitations of claim 1. Sebbane further teaches: The method of claim 1, wherein the environment is a real- world environment and the agent is a mechanical agent interacting with the real-world environment. (Sebbane, page 2, paragraph 2, “In sequential decision making, the aerial robot seeks to choose the best actions based on its observations of the world to optimize an objective function over the course of a series of such decisions, depending on its mission.”) The rationale for the combination of Bacon and Zhu with Sebbane is similar to that as applied for claim 2 above. Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Bacon in view of Zhu, as applied to claims 1 and 2 above, and in further view of of Commons (US 9015093 B1), hereafter referred to as Commons. Claim 9: Bacon and Zhu teaches the limitations of claim 1. Commons, in the same field of neural network implementation, teaches the following which the above fails to teach: The method of claim 8 wherein the plan proposal neural network includes a multi- later perceptron (MLP). (Commons, col. 2, lines 4-7, “networks with the same architecture as the backpropagation network are referred to as Multi-Layer Perceptrons. This name does not impose any limitations on the type of algorithm used for learning.”) It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to have combined the teachings of Bacon and Zhu with that found in Commons and use multiple neural networks to process the observations, goals, and plans of an environment to compute a best action to take for an agent. A motivation for the combination is to utilize different perspectives from different neural network models in order to generate more refined insights necessary for action planning. (Commons, col. 9, lines 23-31, “Biological studies have shown that the human brain functions not as a single massive network, but as a collection of small networks. This realization gave birth to the concept of modular neural networks, in which several small networks cooperate or compete to solve problems. A committee of machines (CoM) is a collection of different neural networks that together "vote" on a given example. This generally gives a much better result compared to other neural network models. Because neural networks suffer from local minima, starting with the same architecture and training but using different initial random weights often gives vastly different networks. A CoM tends to stabilize the result.”, this shows how the hierarchical structure of multiple neural networks are better than having a single end-to-end neural network.) Claim 10: Bacon and Zhu teaches the limitations of claim 1. Commons, in the same field of neural network implementation, teaches the following which the above fails to teach: The method of any preceding claim, wherein the policy neural network is a recurrent neural network. (Commons, col. 8, lines 50-58, “These networks are not arranged in layers. Usually only a subset of the neurons receive external inputs in addition to the inputs from all the other neurons, and another disjunct subset of neurons report their output externally as well as sending it to all the neurons. These distinctive inputs and outputs perform the function of the input and output layers of a feed-forward or simple recurrent network, and also join all the other neurons in the recurrent processing”) The rationale for the combination of Bacon and Zhu with Commons is similar to that applied for claim 9 above. Claims 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over Bisson et al,, (Bisson, F., Larochelle, H., & Kabanza, F. (2015, January). Using a Recursive Neural Network to Learn an Agent's Decision Model for Plan Recognition. In IJCAI (pp. 918-924).), hereafter referred to as Bisson in view of Zhu and in further view of Bacon and Che et al., (Che, T., Li, Y., Zhang, R., Hjelm, R. D., Li, W., Song, Y., & Bengio, Y. (2017). Maximum-likelihood augmented discrete generative adversarial networks. arXiv preprint arXiv:1702.07983.), hereafter referred to as Che. Claim 11: Bisson teaches the following limitations: A method of training a plan proposal neural network having a plurality of plan proposal parameters and a policy neural network having a plurality of policy parameters jointly with a plan recognizer neural network having a plurality of plan recognizer parameters and configured to receive as input a sequence of observation action pairs and to process the sequence of observation action pairs to generate data defining a probability distribution over the space comprising a plurality of latent plans, the method comprising: (Bisson, page 918, col. 2, paragraph 1, “Assuming a given plan library, we cast the problem of learning the probabilistic decision-making model for an agent behaving based upon the library as multinomial logistic regression. We train a recursive neural network to learn the vector representation of plan hypotheses, and compute a score for each one of them. We then use a softmax classifier to train the model to yield high scores for correct hypotheses and a low score for incorrect ones to allow ranking the hypotheses by score.”, Bisson, page 920, col. 1, last paragraph, “The probability of a generated hypothesis is computed using the observee’s HTN decision model, which is in fact a probability distribution over all decision choices conveyed by the HTN plan library:”, Bisson supplies the plan-recognizer NN that, from observed action sequences, outputs probabilities over plan hypotheses (i.e., a distribution over latent plans).) processing at least the observations in the sequence of observation action pairs using the plan recognizer neural network and in accordance with current values of the plurality of plan recognizer parameters to generate first data defining a first probability distribution over the space; (Bisson, page 920, col. 1, last paragraph, “The probability of a generated hypothesis is computed using the observee’s HTN decision model, which is in fact a probability distribution over all decision choices conveyed by the HTN plan library:” Bisson, page 921, section 4.2, “Then, we simply model the probability that η ∈ H is the correct hypothesis as: PNG media_image1.png 39 257 media_image1.png Greyscale ”, Bisson’s recognizer takes observed behavior (observations/actions) and produces a softmax probability distribution over plan hypotheses (latent plans). This is the claimed first distribution over the same latent-plan space) to generate a second probability distribution over the space (Bisson, page 921, section 4.2, “Then, we simply model the probability that η ∈ H is the correct hypothesis as: PNG media_image1.png 39 257 media_image1.png Greyscale ”, Using Zhu’s two-image embedding as inputs to Bisson’s softmax head yields a second softmax probability distribution over the same latent-plan space (plan hypotheses/options). This keeps the “first” and “last” distributions consistent (both are distributions over first and second observations).) Zhu, in the same field of machine learning agent learning, further teaches the following which the above prior art fails to teach: obtaining a sequence of observation action pairs, the sequence of observation action pairs generated as a result of interactions of the agent with the environment and including a respective observation pair at each of a plurality of time steps that includes an observation comprising an image received at a time step and an action performed at a time step (Zhu, page 4, col. 2, paragraph 1, “action a at time t can be drawn by: a ∼ π(st ,g|u) where u are the model parameters, st is the image of the current observation, and g is the image of the navigation target.”, Zhu explicitly states the current observation and action at time t is an RGB image from the agent’s camera. Zhu, page 3, col. 2, last paragraph, “a mobile robot keeps taking actions drawn from the policy distribution until reaching the destination”, Zhu provides the image observations at each timestep and the corresponding actions, i.e., a sequence of (image, action) pairs produced by agent–environment interaction.) processing the first observation in the sequence and the last observation in the sequence using the plan proposal neural network and in accordance with current values of the plan proposal parameters (Zhu, page 4, col. 2, paragraph 2, “Our approach to reasoning about the spatial arrangement between the current location and the target is to project them into the same embedding space, where their geometric relations are preserved. We use two streams of weight-shared siamese layers to transform the current state and the target into the same embedding space. Information from both embeddings is fused to form a joint representation. This joint representation is passed through scene-specific layers (refer to Fig. 4)... Finally, the model generates policy and value outputs”, Interpreting the sequence endpoints as first (start) image and last (goal) image, Zhu’s two-image conditioning matches the claim’s first/last observation inputs to the plan-proposal network.) for each observation action pair in the sequence, processing an input comprising the observation in the pair, the last observation in the sequence, (Zhu, page 4, col. 2, paragraph 2, “Overall, the inputs to the network are two images that represent the agent’s current observation and the target… the model generates policy and value outputs”, Zhu supplies the (i) current image and (ii) goal image pathway and produces a fused embedding.) The rationale for combining Bisson with Zhu is similar to that applied for Bacon with Zhu as shown in claim 1 above. Bacon, in the same field of machine learning agent action spaces, further teaches the following which the above prior art fails to teach: sampling a latent plan from the first probability distribution over the space, (Bacon, page 1729, col. 1, algorithm 1, “Choose ω according to an ε-soft policy over options πΩ(s)”, Bacon provides the explicit choose/sample step. In the combined system, the recognizer’s first distribution (Bisson) can parameterize or calibrate πΩ; a POSITA would sample a latent plan/option ω according to a distribution over plans.) where each of the plurality of latent plans represents a different path through the environment or a different action selection constraint to be imposed on the policy neural network, (Bacon, page 1727, col. 1, paragraph 3, “The options framework (Sutton, Precup, and Singh 1999; Precup 2000) formalizes the idea of temporally extended actions.”, Bacon, page 1730, figure 4, “All options (color-coded) are used by the policy over options in successful trajectories.”, An option dictates a multi-step course of action (a path) via its intra-option policy until termination; equivalently, conditioning on ω constrains action selection to πω(·|s). Thus options satisfy either prong: (i) different paths, and/or (ii) different action-selection constraints.) and the latent plan using the policy neural network and in accordance with current values of the policy parameters to generate an action probability distribution for the pair; (Bacon, page 1727, col. 1, last paragraph, “We consider the call-and-return option execution model, in which an agent picks option ω according to its policy over options πΩ , then follows the intra-option policy πω until termination (as dictated by βω), at which point this procedure is repeated. Let πω,θ denote the intra-option policy of option ω parametrized by θ and βω,ϑ, the termination function of ω parameterized by ϑ.”, Bacon adds (iii) the selected option ω, with the intra-option policy πω functioning as the policy NN (neural network) that outputs the action distribution conditioned on the state embedding (which encodes both images) and on ω.) It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to have combined the teachings of Bisson and Zhu with that found in Bacon and use policy neural network for modeling probability distributions. A motivation for the combination is to train the neural network to allow for linear and non-linear function approximators. (Bacon, page 1726, col. 2, paragraph 1, “Based on the policy gradient theorem (Sutton et al. 2000), we derive new results which enable a gradual learning process of the intra-option policies and termination functions, simultaneously with the policy over them. This approach works naturally with both linear and non-linear function approximators, under discrete or continuous state and action spaces.”) Che, in the same field of machine learning agent optimization, teaches the following limitations which Choudhury, Sebbane, and Commons fail to teach: and determining a gradient with respect to the policy parameters, the plan recognizer parameters, and the plan proposal parameters of a loss function that includes (i) a first term that depends on, for each observation action pair, a probability assigned to the action in the observation action pair in the action probability distribution for the observation action pair and (Che, page 2, col. 2, paragraph 2, PNG media_image2.png 60 243 media_image2.png Greyscale , , the first term shown in the image PNG media_image3.png 25 98 media_image3.png Greyscale is related to an expectation over the real data of the negative log. This term is part of the loss function and is designed to maximize the likelihood that the discriminator assigns a high probability to real data samples, which is being interpreted as synonymous to the maximum likelihood term.) (ii) a second term that measures a difference between the first probability distribution and the second probability distribution. (Che, page 2, col. 2, paragraph 2, “where c(D) is a constant depending only on D. Hence, optimizing the traditional GAN is basically equivalent to optimizing the KL-divergence KL(pθ||q 0 ).”, The Kullback–Leibler (KL) divergence compares the probability distributions between both the first and second distributions. As shown in the loss function above, a KL-divergence term is used as the first term in the image, designated as the function KL().) It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to have combined the teachings of Bisson, Bacon, and Zhu with that found in Che and use a loss function combining KL-divergence and maximum likelihood terms. A motivation for the combination is to train the neural network to account for a moving reward signal, allowing to adjust the reward toward desired behaviors in a dynamic environment. (Che, page 2, paragraph 2, “Our work is related to the viewpoint of casting the GAN training as a reinforcement learning problem with a moving reward signal monotone in D(x).”) Claim 12: Bisson, Zhu, Bacon and Che teaches the limitations claim 11. Che teaches: The method of claim 11, wherein the second term is a KL divergence between the first probability distribution and the second probability distribution. (Che, page 2, col. 2, paragraph 2, PNG media_image2.png 60 243 media_image2.png Greyscale , “where c(D) is a constant depending only on D. Hence, optimizing the traditional GAN is basically equivalent to optimizing the KL-divergence KL(pθ||q 0 ).”, As shown in the loss function above, a KL-divergence term is used as the first term, designated as the function KL(). The Kullback–Leibler (KL) divergence compares the probability distributions of both the first and second distributions found in Commons.) The rationale for the combination is similar to that disclosed in claim 11 above. Claim 13: Bisson, Zhu, Bacon and Che teaches the limitations of claim 11. Che further teaches: The method of claim 11, wherein the first term is a maximum likelihood loss term. (Che, page 2, col. 2, paragraph 2, PNG media_image2.png 60 243 media_image2.png Greyscale , , the first term shown in the image PNG media_image3.png 25 98 media_image3.png Greyscale is related to an expectation over the real data of the negative log. This term is part of the loss function and is designed to maximize the likelihood that the discriminator assigns a high probability to real data samples, which is being interpreted as synonymous to the maximum likelihood term.) The rationale for the combination is similar to that disclosed in claim 11 above. Claim 14: Bisson, Zhu, Bacon and Che teaches the limitations of claim 11. Che teaches: The method of claim 11, wherein the loss function is of the form L1 + BL2, where L1 is the first term, L2 is the second term, and B is a constant weight value. (Che, page 2, col. 2, paragraph 2, PNG media_image2.png 60 243 media_image2.png Greyscale , , the loss function shown in Che explicitly has the form L1 + BL2 where c(D) is the first term, KL() is the second term, and τ is the constant weight value.) The rationale for the combination is similar to that disclosed in claim 11 above. Claims 16 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Bisson in view of Zhu, Bacon, and Che, as applied to claims 11-14 above, and in further view of Commons. Claim 16: Bisson, Zhu, Bacon and Che teaches the limitations of claim 11. Commons teaches: The method of claim 11, wherein the plan recognizer neural network is a recurrent neural network. (Commons, col. 8, lines 50-58, “These networks are not arranged in layers. Usually only a subset of the neurons receive external inputs in addition to the inputs from all the other neurons, and another disjunct subset of neurons report their output externally as well as sending it to all the neurons. These distinctive inputs and outputs perform the function of the input and output layers of a feed-forward or simple recurrent network, and also join all the other neurons in the recurrent processing”, each network previously disclosed 20-24 are recurrent neural networks.) It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to have combined the teachings of Bisson, Zhu, Bacon and Che with that found in Commons and use multiple neural networks to process the observations, goals, and plans of an environment to compute a best action to take for an agent. A motivation for the combination is to utilize different perspectives from different neural network models in order to generate more refined insights necessary for action planning. (Commons, col. 9, lines 23-31, “Biological studies have shown that the human brain functions not as a single massive network, but as a collection of small networks. This realization gave birth to the concept of modular neural networks, in which several small networks cooperate or compete to solve problems. A committee of machines (CoM) is a collection of different neural networks that together "vote" on a given example. This generally gives a much better result compared to other neural network models. Because neural networks suffer from local minima, starting with the same architecture and training but using different initial random weights often gives vastly different networks. A CoM tends to stabilize the result.”, this shows how the hierarchical structure of multiple neural networks are better than having a single end-to-end neural network.) Claim 17: Bisson, Zhu, Bacon and Che teaches the limitations of claim 16. Commons teaches: The method of claim 16 wherein the plan recognizer neural network is a bi- directional recurrent neural network. (Commons, col. 8, lines 50-58, “These networks are not arranged in layers. Usually only a subset of the neurons receive external inputs in addition to the inputs from all the other neurons, and another disjunct subset of neurons report their output externally as well as sending it to all the neurons. These distinctive inputs and outputs perform the function of the input and output layers of a feed-forward or simple recurrent network, and also join all the other neurons in the recurrent processing”, each network previously disclosed 20-24 are recurrent neural networks.) The rationale for the combination of Bisson, Zhu, and Bacon with Commons is similar to that disclosed in claim 16 above. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Bisson in view of Zhu, Bacon, and Che, as applied to claims 11-14 above, and in further view of Asami et al., (US 20200035223 A1), hereafter referred to as Asami. Claim 15: Bisson, Zhu, Bacon, and Che teaches the limitations of claim 14 Asami teaches the following limitation which the above fails to teach: The method of claim 14, wherein B is less than 1. (Asami, paragraph 54, “Note that weight α is a parameter which is set in advance and which is equal to or greater than 0 and equal to or less than 1.”, Asami discloses a similar loss function to that described in the claim (C=(1−α)C 2 +αC 1) where the weight variable α is greater than 0 but less than 1.) It would have been obvious to one of ordinary skill in the art before the filing date of the invention to have combined the teachings of Bisson, Zhu, Bacon, and Che with that found in Asami and use a loss function combining weighted terms with weight between 0-1. A motivation for the combination is to further balance the influence of multiple terms/objectives used in the loss function. (Asami, paragraph 8, “In Related Art 2, an effect is confirmed that a model which has accuracy equivalent to that of the teacher model and which requires a short calculation period can be learned by using a huge model (which has high accuracy but requires a long calculation period) which has been learned in advance as the teacher model and using a small model which is initialized with a random number as the student model, and by setting the temperature T=2, and the weight α=0.5. Note that the huge model means a model which has a large number of intermediate layers of the neural network and a large number of units in the respective intermediate layers.”, having weighted parameters between 0 and 1 help further improve learning accuracy without the need of a large model.) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Andrew, B., & Richard S, S. (2018). Reinforcement learning: an introduction. Gupta, S., Davidson, J., Levine, S., Sukthankar, R., & Malik, J. (2017). Cognitive mapping and planning for visual navigation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2616-2625). Reddy, S., Dragan, A. D., & Levine, S. (2018). Shared autonomy via deep reinforcement learning. arXiv preprint arXiv:1802.01744. Schaul, T., Horgan, D., Gregor, K., & Silver, D. (2015, June). Universal value function approximators. In International conference on machine learning (pp. 1312-1320). PMLR. Tamar, A., Wu, Y., Thomas, G., Levine, S., & Abbeel, P. (2016). Value iteration networks. Advances in neural information processing systems, 29. Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., & Sermanet, P. (2019). Learning Latent Plans from Play. arXiv preprint arXiv:1903.01973. Yu, T., Shevchuk, G., Sadigh, D., & Finn, C. (2019). Unsupervised visuomotor control through distributional planning networks. arXiv preprint arXiv:1902.05542. Haarnoja, T., Hartikainen, K., Abbeel, P., & Levine, S. (2018, July). Latent space policies for hierarchical reinforcement learning. In International Conference on Machine Learning (pp. 1851-1860). PMLR. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /H.B.Y./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Show 8 earlier events
Aug 15, 2025
Request for Continued Examination
Aug 28, 2025
Response after Non-Final Action
Oct 31, 2025
Non-Final Rejection mailed — §103
Feb 26, 2026
Response Filed
May 13, 2026
Final Rejection mailed — §103
Jul 07, 2026
Response after Non-Final Action
Jul 28, 2026
Request for Continued Examination
Jul 30, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12651178
MACHINE LEARNING TECHNIQUES FOR ASSOCIATING NETWORK ADDRESSES WITH INFORMATION OBJECT ACCESS LOCATIONS
5y 4m to grant Granted Jun 09, 2026
Patent 12619888
END-TO-END SYSTEMS AND METHODS FOR CONSTRUCT SCORING
1y 7m to grant Granted May 05, 2026
Patent 12536429
INTELLIGENTLY MODIFYING DIGITAL CALENDARS UTILIZING A GRAPH NEURAL NETWORK AND REINFORCEMENT LEARNING
4y 7m to grant Granted Jan 27, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
30%
Grant Probability
76%
With Interview (+45.8%)
4y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month