Prosecution Insights
Last updated: October 02, 2026
Application No. 18/754,726

DISTRIBUTIONAL REINFORCEMENT LEARNING

Final Rejection §103§DOUBLEPATENT
Filed
Jun 26, 2024
Priority
Apr 14, 2017 — provisional 62/485,720 +3 more
Examiner
HUANG, YAO D
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
DeepMind Technologies Limited
OA Round
2 (Final)
64%
Grant Probability
Moderate
3-4
OA Rounds
1y 9m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 64% of resolved cases
64%
Career Allowance Rate
89 granted / 138 resolved
+9.5% vs TC avg
Strong +32% interview lift
Without
With
+32.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
13 currently pending
Career history
150
Total Applications
across all art units

Statute-Specific Performance

§101
16.0%
-24.0% vs TC avg
§103
48.6%
+8.6% vs TC avg
§102
10.2%
-29.8% vs TC avg
§112
22.9%
-17.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 138 resolved cases

Office Action

§103 §DOUBLEPATENT
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is in response to applicant’s response filed on July 15, 2026, under which claims 21-40 are pending and under consideration. Remarks Applicant’s amendments have overcome the previous § 112(b), § 101 (software per se), and § 103 rejections. The previous grounds of rejection have been withdrawn. However, upon further consideration, new grounds of rejection have been made under § 103. Applicant’s arguments directed to the § 103 rejections have been fully considered but are moot under the new grounds of rejection because the relied upon distinctions are currently addressed by a new reference, Morimura. Please see the rejections below for details. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-2 and 4-5 of U.S. Patent No. 10,860,920 B2 (the “reference patent”). Although the claims at issue are not identical, they are not patentably distinct from each other, as illustrated in the following table for claims 21-27. Claims of the instant application 18/754,726 (pending claims) Claims of US10,860,920B2 (reference patent claims) 21. A method performed by one or more computers for selecting an action to be performed by a reinforcement learning agent interacting with an environment, the method comprising: receiving a current observation characterizing a current state of the environment; processing a network input comprising the current observation using a distributional Q network and in accordance with current values of a set of neural network parameters of the distributional Q network, wherein for each action of a plurality of actions that can be performed by the agent to interact with the environment: the distributional Q network generates a respective network output that defines a probability distribution over a plurality of possible Q returns for the action, wherein the respective network output comprises a respective score for each of the plurality of possible Q returns for the action; and selecting an action from the plurality of actions to be performed by the agent in response to the current observation using the probability distributions over the pluralities of possible Q returns for the actions. 1. A method performed by one or more data processing apparatus for selecting an action to be performed by a reinforcement learning agent interacting with an environment, the method comprising: receiving a current observation characterizing a current state of the environment; for each action of a plurality of actions that can be performed by the agent to interact with the environment: processing the action and the current observation using a distributional Q network having a plurality of network parameters, wherein the distributional Q network is a deep neural network that is configured to process the action and the current observation in accordance with current values of the network parameters to generate a network output comprising a plurality of numerical values that collectively define a probability distribution over possible Q returns for the action current observation pair, wherein the network output comprises: (i) a respective score for each of a plurality of possible Q returns for the action current observation pair, or (ii) a respective value for each of a plurality of parameters of a parametric probability distribution over possible Q returns for the action-current observation pair, and wherein each possible Q return is an estimate of a return that would result from the agent performing the action in response to the current observation, and determining a measure of central tendency of the possible Q returns with respect to the probability distribution for the action-current observation pair; and selecting an action from the plurality of possible actions to be performed by the agent in response to the current observation using the measures of central tendency for the actions. 22. The method of claim 21, wherein selecting an action from the plurality of actions to be performed by the agent comprises: determining, for each action, a measure of central tendency of the probability distribution over the plurality of possible Q returns for the action; and selecting the action from the plurality of actions to be performed by the agent based at least in part on the measures of central tendency for the actions. Claim 1, quoted above, recites: “…determining a measure of central tendency of the possible Q returns with respect to the probability distribution for the action-current observation pair; and selecting an action from the plurality of possible actions to be performed by the agent in response to the current observation using the measures of central tendency for the actions. 23. The method of claim 22, wherein selecting the action from the plurality of actions to be performed by the agent based at least in part on the measures of central tendency for the actions comprises: selecting an action having a highest measure of central tendency. 2. The method of claim 1, wherein selecting an action to be performed by the agent comprises: selecting an action having the highest measure of central tendency. 24. The method of claim 22, wherein for each action, the measure of central tendency is a mean of the probability distribution over the plurality of possible Q returns for the action. 4. The method of claim 1, wherein the measure of central tendency is a mean of the possible Q returns. 25. The method of claim 24, wherein for each action, determining the measure of central tendency of the probability distribution over the plurality of possible Q returns for the action comprises: determining a respective probability for each of the plurality of possible Q returns from the probability distribution over the plurality of possible Q returns for the action; weighting each possible Q return of the plurality of possible Q returns by the probability for the possible Q return; and determining the mean by summing the weighted possible Q returns. 5. The method of claim 4, wherein determining the mean of the possible Q returns with respect to the probability distribution comprises: determining a respective probability for each of the plurality of possible Q returns from the network output; weighting each possible Q return by the probability for the possible Q return; and determining the mean by summing the weighted possible Q returns. Note: The “network output” is recited in claim 1 as defining “a probability distribution over possible Q returns…” so as to correspond to the instant application’s claims. 26. The method of claim 21, wherein for each action, the distributional Q network generates the respective network output that defines the probability distribution over the plurality of possible Q returns by processing both: (i) the current observation, and (ii) data identifying the action. Claim 1, quoted above, recites: “process the action and the current observation in accordance with current values of the network parameters…” Here, “the action” is also regarded as a disclosure of “data identifying the action” since the action is processed. 27. The method of claim 21, wherein the distributional Q network is a deep neural network that comprises a plurality of neural network layers. Claim 1, quoted above, recites “wherein the distributional Q network is a deep neural network.” Furthermore, a “deep neural network” by definition includes a plurality of neural network layers. As shown above, the features of the pending claims are also present in the reference patent claims. Although the instant application’s claims sometimes use wording that is somewhat different from that of the reference patent, the differences in wording do not make the claims patentably distinct over the reference patent claims. For example, while the pending claims use the term “plurality of possible Q returns,” the concept of plurality is already present in the plural term “returns.” Any differences are, at most, within obvious variations of the reference patent claims. In regards to system claims 28-34 (based on the interpretations made in the § 112 rejections below) and computer-readable medium claims 35-40 of the instant application, these claims recite operations that are the same or substantially the same as those of claims 21-27 and claims 21-26 discussed above. Therefore, the above comparisons are also applied to claims 28-34 and 35-40. Furthermore, the addition of a system or computer-readable medium as recited in the instant application’s claims 28-40 does not make these claims patentably distinct from the reference patent claims because these features are merely generic computer components and the reference patent claims already state that its method is performed by a “data processing apparatus” (i.e., a general purpose computer); thus, the system and computer-readable medium claims of the parent application are obvious variations of the method claims of the reference patent. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 1. Claims 21, 26-28, 33-35, and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Van Hasselt et al. (US 2017/0076201 A1) (“Van Hasselt”) (cited by applicant in an IDS) in view of Morimura et al., “Nonparametric Return Distribution Approximation for Reinforcement Learning,” 27th International Conference on Machine Learning, Haifa, Israel, 2010 (“Morimura”) (cited by applicant in an IDS). As to claim 21, Van Hasselt teaches a method performed by one or more computers for selecting an action to be performed by a reinforcement learning agent interacting with an environment, [Abstract: “Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a Q network used to select actions to be performed by an agent interacting with an environment.” [0002]: “This specification relates to selecting actions to be performed by a reinforcement learning agent.”] the method comprising: receiving a current observation characterizing a current state of the environment; [[0005]: “receiving an observation characterizing a current state of the environment and performing an action from a set of actions in response to the observation.” [0035]: “The system receives a current observation characterizing the current state of the environment (step 202).”] processing a network input comprising the current observation using a distributional Q network and in accordance with current values of a set of neural network parameters of the distributional Q network, [[0036]: “For each action in the set of actions, the system processes the current observation and the action using the Q network in accordance with current values of the parameters of the Q network (step 204). As described above, the Q network is a deep neural network that is configured to receive as input an observation and an action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters.” It is noted that the term “distributional” in “distributional Q network” is treated as a label that by itself does not distinguish over the Q network in the prior art reference, since the term “distributional,” standing alone, is not specific as to what distribution is being required and how such a distribution is being used. Here, the network in Van Hasselt outputs values in a certain way and also incorporates probabilistic elements (see [0039]). Therefore, it is considered to be “distributional” at least in the sense that it involves some distribution. The Examiner notes that the element of a probability distribution over possible Q returns is addressed by a different reference.] wherein for each action of a plurality of actions that can be performed by the agent to interact with the environment: [[0036]: “For each action in the set of actions, the system processes the current observation and the action using the Q network in accordance with current values of the parameters of the Q network (step 204).”] the distributional Q network generates a respective network output […] for the action, [[0036]: “As described above, the Q network is a deep neural network that is configured to receive as input an observation and an action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters. Thus, by, for each action, processing the current observation and the action using a Q network in accordance with current values of the parameters of the Q network, the system generates a respective estimated future cumulative reward for each action in the set of actions.”] […] [[0036]: “As described above, the Q network is a deep neural network that is configured to receive as input an observation and an action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters. Thus, by, for each action, processing the current observation and the action using a Q network in accordance with current values of the parameters of the Q network, the system generates a respective estimated future cumulative reward for each action in the set of actions.”] and selecting an action from the plurality of actions to be performed by the agent in response to the current observation […]. Van Hasselt does not explicitly teach: (1) The output being an output that “defines a probability distribution over a plurality of possible Q returns”; (2) “wherein the respective network output comprises a respective score for each of the plurality of possible Q returns for the action” and (3) The selection of the action “using the probability distributions over the pluralities of possible Q returns for the actions.” Morimura teaches an output that “defines a probability distribution over a plurality of possible Q returns” [§ 1, paragraph 3: “In this paper, we describe an approach to handling various risk-sensitive and/or robust criteria in a unified manner, where the distribution of the possible returns is approximated and then the criteria are evaluated based on the approximated distribution.” § 4.2, paragraph 1: “The return distribution PEπ(η|s) is approximated by a distribution of the particles vs, i.e., an estimate of the cumulative probability distribution function of the return defined” as characterized in equation (5) and related parts of the paper.] “wherein the respective network output comprises a respective score for each of the plurality of possible Q returns for the action” [§ 4.2, paragraph 1: “We suppose the usage of look-up table for the states (or the state-action pairs). Each entry of the table has a number of particles… Each particle vs,k ∈ R represents a return η from the state s…“The return distribution PEπ(η|s) is approximated by a distribution of the particles vs.” That is, each particle represents a possible Q return for an action in a particular state, and the value of the particle is a score for the corresponding possible return.] and selection of action “using the probability distributions over the pluralities of possible Q returns for the actions.” [In general, § 5 teaches: “Now we have a means to approximately evaluate the return distributions, so we can formulate RL algorithms with any criterion defined on the basis of a return distribution.” Specifically, this paper teaches action selection using “the greedy, ε-greedy, or soft-max action selection models" (§ 2.1, last sentence), where in the ε-greedy method, “the action with the highest value is selected with probability ε + (1 − ε)/2 and one of the other actions is selected with probability (1− ε)/2.” In the context of this paper, § 5, paragraph 2 teaches computing the QCVaR+ for a given state an action based on maxk{vs,k}, which is the use of the distributions.] It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Van Hasselt with the teachings of Morimura by modifying Van Hasselt to use Morimura’s techniques described above so as to arrive at the claimed invention. The motivation would have been to approximate the distribution of returns, which allows handling various criteria including risk-sensitivities in a unified manner (see Morimura, § 7, paragraph 1: “approximating the distribution of returns, which allows us to handle various criteria including risk-sensitivities in a unified manner.”). As to claim 26, the combination of Van Hasselt and Morimura teaches the method of claim 21, wherein for each action, the distributional Q network generates the respective network output that defines the probability distribution over the plurality of possible Q values for the action by processing both: (i) the current observation, and (ii) data identifying the action. [Van Hasselt, [0036]: “For each action in the set of actions, the system processes the current observation and the action using the Q network in accordance with current values of the parameters of the Q network (step 204).” Note that the “set of actions” is data identifying the actions.] As to claim 27, the combination of Van Hasselt and Morimura teaches the method of claim 21, wherein the distributional Q network is a deep neural network that comprises a plurality of neural network layers. [Van Hasselt, [0036]: “As described above, the Q network is a deep neural network that is configured to receive as input an observation and an action and to generate an estimated future cumulative reward from the input in accordance with a set of parameters.” Van Hasselt, [0004]: “Some neural networks are deep neural networks that include one or more hidden layers in addition to an output layer.”] As to claims 28 and 33-34, these claims are directed to a system for performing operations that are the same or substantially the same as those of claims 21 and 26-27, respectively. Therefore, the rejections made to claims 21 and 26-27 are applied to claims 28 and 33-34, respectively. Furthermore, Van Hasselt teaches “a system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for selecting an action to be performed by a reinforcement learning agent interacting with an environment” [[0017]: “The reinforcement learning system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.” [0055]: “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, data processing apparatus…The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.”]. As to claims 35 and 40, these claims are directed to a storage medium for performing operations that are the same or substantially the same as those of claims 21 and 26. Therefore, the rejections made to claims 21 and 26 are applied to claims 35 and 40, respectively. Furthermore, Van Hasselt teaches “One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for selecting an action to be performed by a reinforcement learning agent interacting with an environment” [[0055]: “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, data processing apparatus…The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.”]. 2. Claims 22-24, 29-31, and 36-38 are rejected under 35 U.S.C. 103 as being unpatentable over Van Hasselt in view of Morimura, and further in view of in view of Dearden et al., “Bayesian Q-learning,” National Conference on Artificial Intelligence, pp. 761–768, 1998 (“Dearden”) (cited by applicant in an IDS). As to claim 22, the combination of Van Hasselt and Morimura teaches the method of claim 21, but does not teach the further limitations of the instant dependent claim. Dearden teaches wherein selecting an action from the plurality of actions to be performed by the agent comprises: determining, for each action, a measure of central tendency of the probability distribution over the plurality of possible Q returns for the action; [§ 3.2, paragraph 1: “Assuming that we have a probability distribution of Q(s,a) = µs,a for all states s and actions a, how do we select an action to perform in the current state.” See also Assumption 3 and text below: “Assumption 3: The prior p(µs,a, τs,a) is a normal-gamma distribution…A normal-gaussian distribution over the mean μ and the precision τ.” That is, the probability distribution’s mean (a measure of central tendency) is determined as a parameter of that distribution.] and selecting the action from the plurality of actions to be performed by the agent based at least in part on the measures of central tendency for the actions. [§ 3.2 (“Action Selection”) teaches that the probability distribution is used for action selection in general. See § 3.2, paragraph 1: “Assuming that we have a probability distribution of Q(s,a) = µs,a for all states s and actions a, how do we select an action to perform in the current state.” In more detail, this section teaches various techniques such as Greedy selection, Q-value sampling, and Myopic-VPI selection, all of which are based on the probability distributions (as represented by µs,a in the formulations) over the possible Q-value returns.] It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Dearden so as to have also arrived at the claimed invention of the instant dependent claim. The motivation would have been to enable the computation of a myopic approximation to the value of information for each action and hence to select the action that best balances exploration and exploitation (Dearden, abstract: “We extend Watkins’ Q-learning by maintaining and propagating probability distributions over the Q-values. These distributions are used to compute a myopic approximation to the value of information for each action and hence to select the action that best balances exploration and exploitation.”). As to claim 23, the combination of Van Hasselt and Dearden teaches the method of claim 22, as set forth above. Dearden further teaches “wherein selecting the action from the plurality of actions to be performed by the agent based at least in part on the measures of central tendency for the actions comprises: selecting an action having a highest measure of central tendency.” [§ 3.2, paragraph 2: “One possible approach is the greedy approach. In this approach, we select the action a that maximizes the expected value E[µs,a]. Unfortunately, it is easy to show that E[µs,a] is simply our estimate of the mean of Rs,a. Thus, the greedy approach would select the action with the greatest mean.] It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far, including the above teachings of Dearden, so as to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the motivation given for Dearden in the rejection of the parent dependent claim 22. As to claim 24, the combination of Van Hasselt and Dearden teaches the method of claim 22, as set forth above. Dearden further teaches “wherein for each action, the measure of central tendency is a mean of the probability distribution over the plurality of possible Q returns for the action.” [Assumption 3 and text below: “Assumption 3: The prior p(µs,a, τs,a) is a normal-gamma distribution…A normal-gaussian distribution over the mean μ and the precision τ.” See also Assumption 1, second paragraph: “to model a distribution over the mean µs,a and precision τs,a.” That is, the probability distribution’s mean (a measure of central tendency) is determined as a parameter of that distribution.] It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far, including the above teachings of Dearden, so as to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the motivation given for Dearden in the rejection of the parent dependent claim 22. As to claims 29-31, the further limitations recited in these claims are the same or substantially the same as those of claims 22-24. Therefore, the rejections made to claims 22-24 are applied to claims 29-31, respectively. As to claims 29-31, the further limitations recited in these claims are the same or substantially the same as those of claims 22-24. Therefore, the rejections made to claims 22-24 are applied to claims 29-31, respectively. 3. Claims 25, 32, and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Van Hasselt in view of Morimura and Dearden, and further in view of “Mean and Variance of Random Variables,” yale.edu, archived on March 7, 2016 (see form PTO-892 of the April 20, 2026, for the full citation) (hereinafter “Yale.edu”). As to claim 25, the combination of Van Hasselt, Morimura, and Dearden teaches the method of claim 24, but does not teach the limitation that for each action, determining the measure of central tendency of the probability distribution over possible Q returns for the action comprises the further steps of the instant dependent claim. Yale.edu teaches “determining a respective probability for each of the plurality of possible Q returns from the probability distribution over possible Q returns for the action weighting each possible Q return by the probability for the possible Q return; and determining the mean by summing the weighted possible Q returns.” [This sequence of operations is taught in the first section titled “Mean” where xi are analogous to the possible Q returns, and pi are analogous to the respective probabilities for the Q returns and serve as their weights. The xi values weighted by pi are then summed to compute the mean, μx. Note that although this reference does not teach the specific application of Q returns, this context is already taught by the existing references, and the technique of this reference is applicable to computing the mean in general, including computing a mean where values of the variable (Q) have associated probabilities.] It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Yale.edu by implementing the computation for computing the mean as disclosed in Yale.edu in order to compute a mean that characterizes the distribution in Dearden. The motivation for doing so would have been to implement a formula that enables a mean to be computed, where such formula is applicable to arbitrary distributions in general, as suggested by Yale.edu (which states “The mean of a discrete random variable X is a weighted average of the possible values that the random variable can take” and provides a well-known formula for computing this mean that is readily understood to be applicable to any arbitrary distribution, including those in Dearden). Additionally, since doing so also would have been obvious as a combination of prior art elements according to known methods to yield predictable results (MPEP § 2143(A)) because: (1) the prior art included the determination of a mean in general and the claimed specific computation for determining the mean, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference; (2) one of ordinary skill in the art could have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately; and (3) one of ordinary skill in the art would have recognized that the results of the combination were predictable (here, the result of computing a mean would have been predictable because the formulation in Yale.edu is applicable to arbitrary functions in which values of a variable has an associated probability). As to claims 35 and 39, these claims recite further limitations that are the same or substantially the same as those of claim 24. Therefore, the rejection made to claim 25 is applied to claims 32 and 39. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The following document depicts the state of the art. Osband et al., US 20170032245 A1 teaches conventional techniques involving reinforcement learning with DQNs. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YAO DAVID HUANG whose telephone number is (571)270-1764. The examiner can normally be reached Monday - Friday 9:00 am - 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Y.D.H./Examiner, Art Unit 2124 /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Jun 26, 2024
Application Filed
Apr 20, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT
May 20, 2026
Interview Requested
Jun 23, 2026
Applicant Interview (Telephonic)
Jun 24, 2026
Examiner Interview Summary
Jul 15, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731043
METHOD FOR TRAINING MULTIVARIATE RELATIONSHIP GENERATION MODEL, ELECTRONIC DEVICE AND MEDIUM
4y 11m to grant Granted Sep 08, 2026
Patent 12725043
SYSTEM, DEVICES AND/OR PROCESSES FOR ADAPTING NEURAL NETWORK PROCESSING DEVICES
5y 2m to grant Granted Sep 01, 2026
Patent 12725025
COMPUTATION IN MEMORY (CIM) ARCHITECTURE AND DATAFLOW SUPPORTING A DEPTH-WISE CONVOLUTIONAL NEURAL NETWORK (CNN)
5y 2m to grant Granted Sep 01, 2026
Patent 12724948
TRAINING GENERATIVE MACHINE LEARNING MODELS FOR 3D MOLECULAR STRUCTURE PREDICTION USING ALIGNMENT OBJECTIVES
1y 6m to grant Granted Sep 01, 2026
Patent 12694267
OPTIMIZATION APPARATUS AND CONTROL METHOD THEREOF
3y 7m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
64%
Grant Probability
97%
With Interview (+32.5%)
4y 0m (~1y 9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 138 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month