Prosecution Insights
Last updated: October 02, 2026
Application No. 18/406,995

CONTROLLING AGENTS USING AMORTIZED Q LEARNING

Non-Final OA §103§112
Filed
Jan 08, 2024
Priority
Nov 16, 2018 — provisional 62/768,788 +2 more
Examiner
CHEN, ALAN S
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
91%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 91% — above average
91%
Career Allowance Rate
1048 granted / 1152 resolved
+31.0% vs TC avg
Moderate +7% lift
Without
With
+6.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
33 currently pending
Career history
1170
Total Applications
across all art units

Statute-Specific Performance

§101
12.8%
-27.2% vs TC avg
§103
22.7%
-17.3% vs TC avg
§102
36.3%
-3.7% vs TC avg
§112
20.5%
-19.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1152 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 504 (specification at page 19, described as the step of sampling a subset of the actions using the proposal neural network). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 505 (Fig. 5, the "Sample subset of actions" step). Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d) If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification The disclosure is objected to because of the following informalities: the state embedding neural network is referred to as "310" (specification at page 13, first sentence introducing the element) and, in the very next sentence, as "210" ("the proposal neural network 112 is made up of the state embedding neural network 210 and the proposal neural network head 320"). Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 2, 19, and 20 each recite that the proposal neural network is configured to process an input "in accordance with current values of the proposal network parameters." There is insufficient antecedent basis for this limitation in the claim. No prior recitation of "a plurality of proposal network parameters" or similar appears anywhere in claim 2, 19, or 20. This is in contrast to the parallel recitation for the Q neural network, which is properly introduced as "a Q neural network having a plurality of Q network parameters" before the claim refers to "the Q network parameters" -- demonstrating that antecedent basis could readily have been established for the proposal network parameters in the same manner. Claims 3-18 depend, directly or indirectly, from claim 2, and are indefinite for at least the same reason. For purposes of examination, "the proposal network parameters" is interpreted under BRI to refer to the parameters of the recited proposal neural network. Claims 2, 19, and 20 each recite "sampling one or more actions from a set of possible actions that can be performed by the agent to interact with the environment" and then, later in the same claim, recite that the proposal neural network generates outputs defining a distribution over "a set of possible actions that can be performed by the agent to interact with the environment" -- reciting the identical descriptive phrase a second time with the indefinite article "a" rather than "the." It is unclear whether this second recitation refers to the previously introduced set of possible actions or introduces a new, second set of possible actions. This ambiguity renders the scope of the claim uncertain and propagates to every dependent claim that subsequently refers to "the set of possible actions" (e.g., claims 3, 4, 5, 10, 11, 12, 13, and 18), because it cannot be determined which antecedent set is being referenced. For purposes of examination, the second recitation of "a set of possible actions that can be performed by the agent to interact with the environment" is interpreted under BRI as referring back to the same, previously-recited set of possible actions, consistent with the specification's consistent treatment of a single set of possible actions available to the agent. Claims 11, 12, and 13 each recite "the proposal output" (e.g., "wherein the proposal output includes a respective probability for each action in the set of possible actions"). There is insufficient antecedent basis for "the proposal output" because claim 2, from which claims 11-13 depend, recites only "one or more outputs" -- it never introduces a singular "proposal output." Claims 14 and 15, by contrast, correctly refer back to "the one or more outputs" as recited in claim 2. For purposes of examination, "the proposal output" in claims 11, 12, and 13 is interpreted under BRI as referring to the "one or more outputs" recited in claim 2, i.e., the output(s) generated by the proposal neural network that define the proposal probability distribution. Claim 18 recites "sampled action values for any sub-actions before the particular sub-action in the ordering." There is insufficient antecedent basis for "the ordering" because no "ordering" of the sub-actions is introduced anywhere in claim 18's dependency chain (claim 2, from which claim 15 depends, from which claim 16 depends, from which claim 18 depends). An "ordering of the sub-actions" is introduced only in claim 17 -- a sibling claim that also depends from claim 16, but from which claim 18 does not depend. Because claim 18 does not incorporate claim 17's limitations, "the ordering" lacks antecedent basis in claim 18. For purposes of examination, "the ordering" in claim 18 is interpreted under BRI as referring to an ordering of the sub-actions that make up an action. Dependent claims not mentioned are rejected as being dependent upon a rejected base claim. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2-11, 14, 19 and 20 are rejected under 35 USC 103 as being unpatentable over Deep Reinforcement Learning in Large Discrete Action Spaces to Dulac-Arnold et al. (hereinafter Dulac-Arnold) in view of Asynchronous Methods for Deep Reinforcement Learning to Mnih et al. (hereinafter Mnih). Per claim 2, Dulac-Arnold discloses A method comprising (Dulac-Arnold: §4, p. 2…Dulac-Arnold proposes the Wolpertinger policy architecture, a computer-implemented reinforcement-learning procedure carried out at each interaction step between an agent and its environment, which constitutes the method recited in the preamble, "We propose a new policy architecture which we call the Wolpertinger architecture"): receiving a current observation characterizing a current state of an environment being interacted with by an agent (Dulac-Arnold: Algorithm 1, p. 2…each iteration of Dulac-Arnold's Wolpertinger policy begins by taking in the state most recently returned by the environment, and a state so received characterizes the environment's present condition as perceived by the agent, which constitutes receiving a current observation characterizing a current state of an environment being interacted with by an agent, "State s previously received from environment"); …using a proposal neural network, wherein the proposal neural network is configured to process an input comprising the current observation in accordance with current values of the proposal network parameters… (Dulac-Arnold: §4.1, p. 3…Dulac-Arnold's actor fθπ is a multi-layer neural network parametrized by θπ that processes the state to emit a proto-action, from which the mapping gk draws a set Ak of k candidate actions out of the full action set A for downstream evaluation, which constitutes a proposal neural network configured to process an input comprising the current observation in accordance with current values of the proposal network parameters, "This function provides a proto-action in Rn for a given state, which will likely not be a valid action, i.e. it is likely that â ∈’ A"; §4.1, p. 3…the mapping then returns the candidate subset that the critic will score, "It returns the k actions in A that are closest to â by L2 distance"); processing the current observation and each sampled action using a Q neural network having a plurality of Q network parameters, wherein, for each sampled action, the Q neural network is configured to process the current observation and the sampled action in accordance with current values of the Q network parameters to generate a Q value for the sampled action that is an estimate of a return that would be received if the agent performed the sampled action in response to the current observation (Dulac-Arnold: Algorithm 1, p. 2…Dulac-Arnold scores every action in the retrieved candidate set Ak with the parametrized critic QθQ, which takes the state and one candidate action as its two inputs and returns that action's estimated return, which constitutes processing the current observation and each sampled action using a Q neural network to generate a Q value for the sampled action, "a = arg maxaj ∈Ak QθQ (s, aj)"; §3, p. 2…Dulac-Arnold confirms the critic is a parametrized function of both state and action and that scoring every action is what the architecture avoids, "In the common case that the value function is a parameterized function which takes both state and action as input, |A| evaluations are necessary to choose an action"); selecting an action using the Q values generated by the Q neural network for the sampled actions (Dulac-Arnold: §4.2, p. 3…the second phase of Dulac-Arnold's policy picks, from among the candidate actions, the one to which the critic assigned the greatest Q value, which constitutes selecting an action using the Q values generated by the Q neural network for the sampled actions, "the second phase of the algorithm, which is described by the top part of Figure 1, refines the choice of action by selecting the highest-scoring action according to QθQ"); and causing the agent to perform the selected action (Dulac-Arnold: Algorithm 1, p. 2…the selected action is applied to the environment, which is then observed to transition and emit a reward, and applying an action to the environment constitutes causing the agent to perform the selected action, "Apply a to environment; receive r, s’"). Dulac-Arnold does not expressly disclose, but with Mnih does teach: sampling one or more actions from a set of possible actions that can be performed by the agent to interact with the environment …to generate one or more outputs that define a proposal probability distribution over a set of possible actions that can be performed by the agent to interact with the environment (Mnih: §4, p. 4…Mnih's policy network consumes the observed state and emits a single softmax output that assigns a respective probability to every action in the action set, and the agent's actions are then drawn from that distribution, so the candidate actions are obtained by a stochastic draw in which each action's likelihood of selection equals the probability the distribution assigns it, which constitutes sampling one or more actions from the set of possible actions using one or more outputs that define a proposal probability distribution over that set, "We typically use a convolutional neural network that has one softmax output for the policy π(at |st ; θv) and one linear output for the value function V (st ; θv ), with all non-output layers shared"). Dulac-Arnold and Mnih are analogous art because they are from both within the same field of endeavor, specifically deep reinforcement learning in which neural network function approximators are trained to select the actions of an agent interacting with an environment. They address the same problem solving area of obtaining a high-value action without paying the computational cost of evaluating a value function over every action in a large action space. Dulac-Arnold cites the actor-critic framework and the policy-gradient training of its action-generating actor (Dulac-Arnold: §4, p. 2…Dulac-Arnold builds its architecture on the actor-critic framework and implements both elements as neural networks, "This policy builds upon the actor-critic (Sutton & Barto, 1998) framework"), which is the core subject of Mnih (Mnih: §3, p. 3…Mnih's asynchronous advantage actor-critic parameterizes the actor as a softmax policy trained by policy gradient alongside a value estimator, "Standard REINFORCE updates the policy parameters θ in the direction ∇θ log π(at |st ; θ)Rt , which is an unbiased estimate of ∇θ E[Rt ]"). Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to modify the action-generating actor of Dulac-Arnold so that it emits a probability distribution over the action set in the manner taught by Mnih, and to obtain the candidate actions that Dulac-Arnold's critic scores by sampling from that distribution, as claim 2 recites. The suggestion/motivation for doing so would have been Dulac-Arnold's own acknowledgement that its deterministic proto-action can land nearest to poor actions and that a candidate set is therefore needed to protect the policy (Dulac-Arnold: §4.2, p. 3…Dulac-Arnold reports that the nearest action to the proto-action may be a low-value outlier, which is the very defect a sampled candidate set mitigates, "Depending on how well the action representation is structured, actions with a low Q-value may occasionally sit closest to â even in a part of the space where most actions have a high Q-value"), together with Mnih's teaching that a stochastic softmax policy explores the action set more effectively than a deterministic one (Mnih: §4, p. 4…Mnih reports that keeping the policy stochastic avoids collapsing onto a deterministic choice prematurely, "We also found that adding the entropy of the policy π to the objective function improved exploration by discouraging premature convergence to suboptimal deterministic policies"). This is the application of a known technique to a known device ready for improvement to yield predictable results, MPEP § 2143(I)(D). Per claim 3, Dulac-Arnold combined with Mnih discloses claim 2. Mnih further teaches sampling (i) one or more actions from the set of possible actions in accordance with the proposal probability distribution and (ii) one or more actions randomly from the set of possible actions (Mnih: §4, p. 4…Mnih draws actions from the softmax policy distribution, "We typically use a convolutional neural network that has one softmax output for the policy π(at |st ; θ) and one linear output for the value function V (st ; θv ), with all non-output layers shared"; Algorithm 1, p. 3…Mnih additionally selects actions uniformly at random under an ε-greedy rule so that the whole action set is explored, and drawing part of the action complement from the policy distribution and part uniformly at random constitutes sampling in accordance with the proposal probability distribution and randomly from the set of possible actions, "Take action a with ε-greedy policy based on Q(s, a; θ)"). Mnih's ε-greedy rule appears in Algorithm 1, which is a different algorithm from the actor-critic of §4, but Mnih teaches applying the two together within its own asynchronous actor-learner framework rather than as alternatives, reporting that ε-greedy exploration is layered onto that framework for robustness (Mnih: §4, p. 4…Mnih adds per-thread ε-greedy exploration on top of its asynchronous framework, "While there are many possible ways of making the exploration policies differ we experiment with using ε-greedy exploration with ε periodically sampled from some distribution by each thread"). A PHOSITA would therefore have drawn part of the candidate complement from the policy distribution and part uniformly at random in a single sampling step, as claim 3 recites. This is the combination of prior art elements according to known methods to yield predictable results, MPEP § 2143(I)(A). Per claim 4, Dulac-Arnold combined with Mnih discloses claim 2. Dulac-Arnold further teaches selecting the sampled action that has the highest Q value (Dulac-Arnold: §4.2, p. 3…Dulac-Arnold's policy takes the arg max of the critic's scores over the retrieved candidate actions, which constitutes selecting the sampled action that has the highest Q value, "the second phase of the algorithm, which is described by the top part of Figure 1, refines the choice of action by selecting the highest-scoring action according to QθQ"). Per claim 5, Dulac-Arnold combined with Mnih discloses claim 2. Mnih further teaches selecting the sampled action that has the highest Q value with probability ε; and selecting a random action from the set of possible actions with probability 1 - ε (Mnih: Algorithm 1, p. 3…Mnih selects the arg max action of the Q network on one branch of an ε-greedy rule and a uniformly random action on the other, the two branches being taken with complementary probabilities, "Take action a with ε-greedy policy based on Q(s, a; θ)"; Dulac-Arnold: §4.2, p. 3…in the combination the arg max is taken over the candidate subset Dulac-Arnold retrieves rather than over the whole action set, so the greedy branch selects the highest-Q action from among the sampled actions as the claim requires, "the second phase of the algorithm, which is described by the top part of Figure 1, refines the choice of action by selecting the highest-scoring action according to QθQ"). Mnih's rule takes the uniformly-random branch with probability ε and the greedy branch with probability 1 - ε. Because ε in claim 5 is a recited variable rather than a fixed value, setting the claim's ε equal to 1 - ε makes the two rules identical, so Mnih's ε-greedy rule selects the highest-Q sampled action with probability ε and a random action with probability 1 - ε exactly as claim 5 recites. The rationale for drawing on Mnih's ε-greedy rule is that set forth for claim 3. Per claim 6, Dulac-Arnold combined with Mnih discloses claim 2. Mnih further teaches determining a gradient of a proposal loss function with respect to the proposal network parameters, wherein the proposal loss function includes a first term that encourages the probability assigned to the sampled action having the highest Q value to be increased in the proposal probability distribution; and determining, from the gradient of the proposal loss function, an update to the current values of the proposal network parameters (Mnih: §3, p. 3…Mnih differentiates the policy objective with respect to the policy parameters and steps them along ∇θ log π(at |st ; θ)Rt, a term that increases the log-probability the policy assigns to each sampled action in proportion to that action's estimated return, so that the sampled action carrying the greatest estimated return necessarily receives the largest positive increment, which constitutes a first term that encourages the probability assigned to the sampled action having the highest Q value to be increased, together with determining an update to the proposal network parameters from that gradient, "Standard REINFORCE updates the policy parameters θ in the direction ∇θ log π(at |st ; θ)Rta, which is an unbiased estimate of ∇θ E[Rt ]"). The rationale for adopting Mnih's policy-gradient training of the action-generating network is that a proposal network emitting a probability distribution, as introduced for claim 2, must be trained by a method that operates on that distribution, and Mnih supplies the policy-gradient method for exactly that object (Mnih: §3, p. 3, the REINFORCE update above). This is the use of a known technique to improve a similar device in the same way, MPEP § 2143(I)(C). Per claim 7, Dulac-Arnold combined with Mnih discloses claim 6. Mnih further teaches the first term is a negative log likelihood for the proposal probability distribution with the sampled action having the highest Q value as a target (Mnih: §4, p. 4…the first term of Mnih's policy objective is the log probability the policy assigns to the selected action, scaled by that action's advantage, "The gradient of the full objective function including the entropy regularization term with respect to the policy parameters takes the form ∇θ’ log π(at |st ; θ’ )(Rt − V (st ; θv )) + β∇θ’ H(π(st ; θ’ )), where H is the entropy"). The rationale to combine is that set forth for claim 6. Per claim 8, Dulac-Arnold combined with Mnih discloses claim 6. Mnih further teaches the proposal loss function includes a second term that encourages uncertainty in the proposal probability distribution (Mnih: §4, p. 4…Mnih adds a second, entropy-regularization term to the same policy objective expressly to keep the policy distribution spread out rather than concentrated on one action, which constitutes a second term that encourages uncertainty in the proposal probability distribution, "We also found that adding the entropy of the policy π to the objective function improved exploration by discouraging premature convergence to suboptimal deterministic policies"). The rationale for adopting Mnih's policy-gradient training is that set forth for claim 6. Per claim 9, Dulac-Arnold combined with Mnih discloses claim 8. Mnih further teaches the second term is a negative entropy of the proposal probability distribution (Mnih: §4, p. 4…Mnih's second term is β∇θ’ H(π(st ; θ’ )) with H the entropy of the policy, ascended so as to raise entropy, which is the same as descending the negative entropy of that distribution, and which constitutes a second term that is a negative entropy of the proposal probability distribution, "The gradient of the full objective function including the entropy regularization term with respect to the policy parameters takes the form ∇θ’ log π(at |st ; θ’ )(Rt − V (st ; θv )) + β∇θ’ H(π(st ; θ’ )), where H is the entropy"). The rationale for adopting Mnih's policy-gradient training is that set forth for claim 6. Per claim 10, Dulac-Arnold combined with Mnih discloses claim 2. Dulac-Arnold combined with Mnih does not expressly disclose the proposal neural network and the Q neural network share some parameters, Mnih's sharing being between its policy head and its state-value head V(st ; θv) rather than with a state-action Q network. Mnih does teach the sharing technique itself and reports it as its standing practice (Mnih: §4, p. 4…Mnih implements the policy network and the value network as one convolutional network whose every non-output layer is common to both, reserving only the output heads to each, "Note that while the parameters θ of the policy and θv of the value function are shown as being separate for generality, we always share some of the parameters in practice"), and Dulac-Arnold's actor and critic both consume the same state input and are both implemented as multi-layer neural networks (Dulac-Arnold: §4, p. 2…both elements of Dulac-Arnold's architecture are neural function approximators over the same state, "We use multi-layer neural networks as function approximators for both our actor and critic functions"). Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to share the state-processing layers between the proposal network and the Q network of Dulac-Arnold as modified by Mnih, reserving only the output heads to each, as claim 10 recites. The suggestion/motivation for doing so would have been Mnih's report that this sharing is what it always does in practice for a policy head and a value head over a common state input (Mnih: §4, p. 4, quoted above), the two networks of the combination presenting the same common-state structure. This is the use of a known technique to improve a similar device in the same way, MPEP § 2143(I)(C). Per claim 11, Dulac-Arnold combined with Mnih discloses claim 2. Mnih further teaches the set of possible actions is discrete, and wherein the proposal output includes a respective probability for each action in the set of possible actions (Mnih: §4, p. 4…Mnih's policy head is a softmax over the discrete action set of the Atari domain, so its output vector carries one probability per available action, which constitutes a discrete set of possible actions with the proposal output including a respective probability for each action in that set, "We typically use a convolutional neural network that has one softmax output for the policy π(at |st ; θ) and one linear output for the value function V (st ; θv ), with all non-output layers shared"). The rationale to combine Mnih with Dulac-Arnold is that set forth for claim 2, this claim reciting a property of the softmax proposal output there introduced. Per claim 14, Dulac-Arnold combined with Mnih discloses claim 2. Mnih further teaches the one or more outputs comprise a single output that defines the proposal probability distribution (Mnih: §4, p. 4…Mnih's policy is defined by exactly one softmax output of the shared network, which constitutes the one or more outputs comprising a single output that defines the proposal probability distribution, "We typically use a convolutional neural network that has one softmax output for the policy π(at |st ; θ) and one linear output for the value function V (st ; θv ), with all non-output layers shared"). The rationale to combine Mnih with Dulac-Arnold is that set forth for claim 2, this claim reciting a property of the softmax proposal output there introduced. Claims 19 and 20 are substantially similar in scope and spirit as claim 2, claim 19 reciting the operations of claim 2 performed by a system of one or more computers and one or more storage devices storing instructions, and claim 20 reciting the same operations stored as instructions on one or more non-transitory computer storage media. Dulac-Arnold implements both of its networks as neural network function approximators (Dulac-Arnold: §4, p. 2…Dulac-Arnold builds the actor and critic as multi-layer neural networks, "We use multi-layer neural networks as function approximators for both our actor and critic"), and Mnih expressly executes its shared policy-and-value network on a conventional multi-core computer (Mnih: §6, p. 7…Mnih runs the asynchronous algorithms on a standard 16-core CPU machine, which constitutes the one or more computers recited in claim 19, "When trained on the Atari domain using 16 CPU cores, the proposed asynchronous algorithms train faster than DQN trained on an Nvidia K40 GPU"). A network so executed is executed under instructions held in the machine's storage, and a PHOSITA would have stored those instructions on non-transitory storage media of the kind claim 20 recites, that being the ordinary means by which the executable instructions of a program run on such a machine are retained. Therefore the rejections of claim 2 are applied accordingly. Claims 12 and 15-18 are rejected under 35 USC 103 as being unpatentable over Dulac-Arnold in view of Mnih and further in view of Discrete Sequential Prediction of Continuous Actions for Deep RL to Metz et al. (hereinafter Metz). Per claim 12, Dulac-Arnold combined with Mnih discloses claim 2. Dulac-Arnold combined with Mnih does not expressly disclose, but with Metz does teach: the set of possible actions is a continuous set of actions selected from a range of actions, and wherein the proposal output includes a respective probability for each choice in a discrete set of choices representing uniformly-spaced values from the range (Metz: §2, p. 2…Metz takes a continuous action range and partitions it into B bins, placing a discrete distribution over those bins, "Here, we use discrete distributions over each dimension (achieved by discretizing each continuous dimension into bins) and apply it using off-policy learning"; Algorithm 1, p. 16…Metz recovers the action value from a bin index by dividing that index by the bin count B, so the B bins tile the range at equal width and therefore represent uniformly-spaced values, which constitutes a discrete set of choices representing uniformly-spaced values from the range, "Convert the integer bin into a continuous value randomly in that bin. Assume B bins"; Appendix G.2, p. 21…Metz's network emits one output unit per bin so that each choice carries its own score, which constitutes the proposal output including a respective probability for each choice, "The output of this network is "quantization bins" wide with no activation"). Dulac-Arnold, Mnih and Metz are analogous art because all three are from within the same field of endeavor, specifically deep reinforcement learning in which neural networks are trained to select an agent's actions, and all three address the same problem solving area of making action selection tractable when the action space is too large or too finely resolved to enumerate. Dulac-Arnold frames that problem as the intractability of evaluating every action (Dulac-Arnold: §3, p. 2…Dulac-Arnold identifies exhaustive evaluation of the action set as the obstacle its architecture removes, "In the common case that the value function is a parameterized function which takes both state and action as input, | A | evaluations are necessary to choose an action"), and Metz addresses the same obstacle by factoring the action across its dimensions (Metz: §1, p.1-2…Metz identifies the exponential blow-up of jointly discretized action spaces as the obstacle its sequential model removes, "For example with M dimensions being discretized into N bins, the problem would balloon to a discrete space with MN possible actions"). Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to form the proposal distribution of Dulac-Arnold as modified by Mnih using the per-dimension autoregressive parameterization of Metz, so that a multi-dimensional action is proposed one sub-action at a time over uniformly-spaced discretized values, as the claim recites. The suggestion/motivation for doing so would have been Metz's express teaching that factoring the action space dimension by dimension is what makes an otherwise exponentially large discretized action space learnable (Metz: §2, p. 2…Metz reports that predicting one dimension at a time is what avoids the exponential blow-up of a jointly discretized action space, "In this paper, we introduce the idea of building continuous control algorithms utilizing sequential, or autoregressive, models that predict over action spaces one dimension at a time"), which is the same scaling obstacle Dulac-Arnold's candidate-set architecture was built to overcome. This is the combination of prior art elements according to known methods to yield predictable results, MPEP § 2143(I)(A). Per claim 15, Dulac-Arnold combined with Mnih discloses claim 2. Dulac-Arnold combined with Mnih does not expressly disclose, but with Metz does teach the one or more outputs define an auto-regressive distribution (Metz: §2, p. 2…Metz's policy over the action space is an autoregressive model that emits its distribution one action dimension at a time, which constitutes outputs that define an auto-regressive distribution, "In this paper, we introduce the idea of building continuous control algorithms utilizing sequential, or autoregressive, models that predict over action spaces one dimension at a time"). Dulac-Arnold, Mnih and Metz are analogous art because all three are from within the same field of endeavor, specifically deep reinforcement learning in which neural networks are trained to select an agent's actions, and all three address the same problem solving area of making action selection tractable when the action space is too large or too finely resolved to enumerate. Dulac-Arnold frames that problem as the intractability of evaluating every action (Dulac-Arnold: §3, p. 2…Dulac-Arnold identifies exhaustive evaluation of the action set as the obstacle its architecture removes, "In the common case that the value function is a parameterized function which takes both state and action as input, | A | evaluations are necessary to choose an action"), and Metz addresses the same obstacle by factoring the action across its dimensions (Metz: §1, p. 2…Metz identifies the exponential blow-up of jointly discretized action spaces as the obstacle its sequential model removes, "For example with M dimensions being discretized into N bins, the problem would balloon to a discrete space with MN possible actions"). Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to form the proposal distribution of Dulac-Arnold as modified by Mnih using the per-dimension autoregressive parameterization of Metz, so that a multi-dimensional action is proposed one sub-action at a time over uniformly-spaced discretized values, as the claim recites. The suggestion/motivation for doing so would have been Metz's express teaching that factoring the action space dimension by dimension is what makes an otherwise exponentially large discretized action space learnable (Metz: §2, p. 2…Metz reports that predicting one dimension at a time is what avoids the exponential blow-up of a jointly discretized action space, "In this paper, we introduce the idea of building continuous control algorithms utilizing sequential, or autoregressive, models that predict over action spaces one dimension at a time"), which is the same scaling obstacle Dulac-Arnold's candidate-set architecture was built to overcome. This is the combination of prior art elements according to known methods to yield predictable results, MPEP § 2143(I)(A). Per claim 16, Dulac-Arnold combined with Mnih and Metz discloses claim 15. Metz further teaches each action in the set of possible actions comprises a respective action value for each of a plurality of sub-actions, and the auto-regressive distribution defines, for each sub-action, a sub-action probability distribution over a set of possible action values for the sub-action (Metz: §2, p. 2…in Metz an action of the underlying environment is a vector of per-dimension action values and the model places a separate discrete distribution over the candidate values of each such dimension, which constitutes each action comprising a respective action value for each of a plurality of sub-actions with the auto-regressive distribution defining a sub-action probability distribution for each sub-action, "Here, we use discrete distributions over each dimension (achieved by discretizing each continuous dimension into bins) and apply it using off-policy learning"). The rationale to combine Metz with Dulac-Arnold and Mnih is that set forth for claim 15, this claim reciting the per-dimension structure of the autoregressive distribution there introduced. Per claim 17, Dulac-Arnold combined with Mnih and Metz discloses claim 16. Metz further teaches sampling action values for the proposal action auto-regressively one after the other according to an ordering of the sub-actions (Metz: Abstract, p. 1…Metz's model produces the action one dimension at a time in a fixed sequence, each dimension issued after the one before it, which constitutes sampling action values auto-regressively one after the other according to an ordering of the sub-actions, "Central to this method is the realization that complex functions over high dimensional spaces can be modeled by neural networks that predict one dimension at a time"). The rationale to combine Metz with Dulac-Arnold and Mnih is that set forth for claim 15, this claim reciting the sampling order of the autoregressive distribution there introduced. Per claim 18, Dulac-Arnold combined with Mnih and Metz discloses claim 16. Metz further teaches processing an input comprising the current observation and sampled action values for any sub-actions before the particular sub-action in the ordering to generate an output defining a probability distribution over possible action values for the particular sub-action; and sampling the action value for the particular sub-action from the probability distribution over possible action values for the particular sub-action (Metz: §2.3, p. 5…each step of Metz's recurrent policy is fed the environment state together with one action dimension and carries the dimensions already chosen forward in its hidden state, so the distribution it emits for the current dimension is conditioned on both the observation and the preceding dimensions' values, which constitutes processing an input comprising the current observation and sampled action values for preceding sub-actions to generate an output defining a probability distribution over possible action values for the particular sub-action, "This model has shared weights and passes information via hidden activations from one action dimension to another. The input at each time step is a function of the current state from the upper MDP, st, and a single action dimension, ai . As it's an LSTM, the hidden state is capable of accumulating the previous actions"; Appendix G.3, p. 23…Metz then draws the value for that dimension from the emitted bin distribution and feeds the drawn value back in for the next dimension, which constitutes sampling the action value for the particular sub-action from that probability distribution, "Each step of the LSTM takes in the previous action value, and outputs some number of "quantization bins." An action is selected, converted back to a one hot representation, and fed into an embedding module"). The rationale to combine Metz with Dulac-Arnold and Mnih is that set forth for claim 15, this claim reciting the conditioning mechanism of the autoregressive distribution there introduced. Claim 13 is rejected under 35 USC 103 as being unpatentable over Dulac-Arnold in view of Mnih and further in view of Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor to Haarnoja et al. (hereinafter Haarnoja). Dulac-Arnold combined with Mnih discloses claim 2. Dulac-Arnold combined with Mnih does not expressly disclose, but with Haarnoja does teach: the set of possible actions is a continuous set of actions, and wherein the proposal output includes distribution parameters that define the proposal probability distribution over the continuous set of actions (Haarnoja: §4.2, p. 4…Haarnoja's policy for a continuous action space is a Gaussian whose mean and covariance are the quantities the policy network emits, so the network's output is a set of distribution parameters that fixes the density over the continuous action set rather than a probability per enumerated action, which constitutes a continuous set of possible actions with the proposal output including distribution parameters that define the proposal probability distribution over that continuous set, "For example, the value functions can be modeled as expressive neural networks, and the policy as a Gaussian with mean and covariance given by neural networks"). Dulac-Arnold, Mnih and Haarnoja are analogous art because all three are from within the same field of endeavor, specifically deep reinforcement learning with neural network function approximators, and all three address the same problem solving area of parameterizing an action-selection policy jointly with an action-value estimator. Mnih parameterizes the policy as a softmax over a discrete action set (Mnih: §4, p. 4…Mnih's policy head is a softmax over enumerated actions, "We typically use a convolutional neural network that has one softmax output for the policy π(at |st ; θ) and one linear output for the value function V (st ; θv ), with all non-output layers shared"), and Haarnoja supplies the corresponding parameterization for the case where the actions are continuous (Haarnoja: §4.2, p. 4…Haarnoja parameterizes the same object as a Gaussian defined by network-emitted parameters, "For example, the value functions can be modeled as expressive neural networks, and the policy as a Gaussian with mean and covariance given by neural networks"). Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to express the proposal distribution of Dulac-Arnold as modified by Mnih through network-emitted distribution parameters in the manner Haarnoja teaches, so that the proposal distribution can be placed over a continuous action set, as claim 13 recites. The suggestion/motivation for doing so would have been that a softmax over enumerated actions cannot be formed at all when the action set is continuous, and Haarnoja teaches that a stochastic policy over such a set is instead obtained by having the network emit the parameters of a density from which actions are then drawn (Haarnoja: Algorithm 1, p. 5…Haarnoja draws each action from the network-parameterized policy density at every environment step, "at ∼ πφ (at |st )"). This is the use of a known technique to improve a similar device in the same way, MPEP § 2143(I)(C). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN CHEN whose telephone number is (571)272-4143. The examiner can normally be reached M-F 10-7. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALAN CHEN/Primary Examiner, Art Unit 2125
Read full office action

Prosecution Timeline

Jan 08, 2024
Application Filed
Oct 07, 2024
Response after Non-Final Action
Sep 01, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737694
Architecture for Classification of a Decision Tree Ensemble and Method
3y 9m to grant Granted Sep 15, 2026
Patent 12737645
GRAPH MACHINE LEARNING MODEL BASED TECHNIQUES FOR EVALUATING KNOWLEDGE GRAPH DATASETS
3y 6m to grant Granted Sep 15, 2026
Patent 12737181
MANAGING OPERATIONAL RESILIENCE OF SYSTEM ASSETS USING AN ARTIFICIAL INTELLIGENCE MODEL
9m to grant Granted Sep 15, 2026
Patent 12731082
SMART COPY OPTIMIZATION IN CUSTOMER ACQUISITION AND CUSTOMER MANAGEMENT PLATFORMS
4y 8m to grant Granted Sep 08, 2026
Patent 12725070
GRADIENT-BASED QUANTUM ASSISTED HAMILTONIAN LEARNING
4y 0m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
91%
Grant Probability
98%
With Interview (+6.7%)
2y 9m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month