Prosecution Insights
Last updated: August 17, 2026
Application No. 18/424,437

METHODS AND SYSTEMS FOR CONSTRAINED REINFORCEMENT LEARNING

Non-Final OA §101§102§103§112
Filed
Jan 26, 2024
Priority
Jan 26, 2023 — provisional 63/441,398
Examiner
LAU, KAITLYN RENEE
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
56%
Grant Probability
Moderate
1-2
OA Rounds
1y 6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
5 granted / 9 resolved
-4.4% vs TC avg
Strong +80% interview lift
Without
With
+80.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
20 currently pending
Career history
37
Total Applications
across all art units

Statute-Specific Performance

§101
32.7%
-7.3% vs TC avg
§103
32.7%
-7.3% vs TC avg
§102
12.8%
-27.2% vs TC avg
§112
21.1%
-18.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 9 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION This action is in response to the application filed 01/26/2026. Claims 1-20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 1-20 are objected to because of the following informalities: Regarding claim 1, the Examiner respectfully notes that claim 1 is a method claim and the limitation of “if the actions of the agent are chosen according to the policy model” is a contingent limitation and therefore under the broadest reasonable interpretation these limitation may not be performed (“The broadest reasonable interpretation of a method (or process) claim having contingent limitations requires only those steps that must be performed and does not include steps that are not required to be performed because the condition(s) precedent are not met.” MPEP 2111.04(II)). Accordingly the Examiner recommends the Applicant positively recite the actions of the agent are chosen to avoid a contingent interpretation of these limitations. Claims 2-3, 9, and 15 recite the same limitation as claim 1. Claims 2-17 depend on claim 1 and are objected to for the same reasons. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 1-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites the limitation "each constraint limiting, to a corresponding threshold…" in line 12. There is insufficient antecedent basis for this limitation in the claim. Claim 1 recites the limitation “an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold” in lines 25-27. It is unclear as to whether this expected value is the same as the expected value claimed in line 23 or if the two expected values are different. Additionally, if they are different, in paragraph 0008 of the specification, it states, “actions are chosen using the policy model generated in the preceding iteration (i.e., the last-but-one iteration)” and thus, this limitation has the same interpretation as the limitation before stating “an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration”. For purposes of examination, Examiner has interpreted this limitation to be repeating the same limitation as recited immediately before. Regarding claims 2-17, claims 2-17 is rejected for at least the same reasons as claim 1 since claims 2-17 depend on claim 1. Claim 2 recites the limitation “for each constraint, a corresponding constraint cost function indicative of expected values of the corresponding constraint reward function”. It is unclear to whether these expected values are the same as the expected values recited in claim 1. For purposes of examination, Examiner has interpreted these expected values to be the same as the expected values in claim 1. Claim 3 recites the limitation “expected values of the corresponding constraint reward function”. It is unclear to whether these expected values are the same as the expected values recited in claim 1. For purposes of examination, Examiner has interpreted these expected values to be the same as the expected values in claim 1. Claim 4 recites the limitation “the expected value under the updated policy model”. It is unclear as to whether this expected value is the same as the expected value recited in claim 1. For the purposes of examination, Examiner has interpreted this expected value to be the same as the expected value in claim 1. The term “higher” in claim 8 is a relative term which renders the claim indefinite. The term “higher” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Thus, the term “higher” renders the limitation of “the weighting higher for later iterations” indefinite. For purposes of examination, Examiner has interpreted this limitation to mean the value of a weight for a current iteration is greater than the value of a weight for a preceding iteration. Claim 9 recites the limitation “the updated value of the multiplier variables” in the first line. It is unclear as to whether there is one updated value or if there is an updated value for each multiplier variable as recited in claim 1. For purposes of Examination, Examiner has interpreted this to be “the updated value of each of the multiplier variables”. Claim 11 recites the limitation “the updated value of the multiplier variables” in the first line. It is unclear as to whether there is one updated value or if there is an updated value for each multiplier variable as recited in claim 1. For purposes of Examination, Examiner has interpreted this to be “the updated value of each of the multiplier variables”. The term “proportional” in claim 14 is a relative term which renders the claim indefinite. The term “proportional” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Thus, the term “proportional” renders the limitation of “the value of the policy model for each combination is generated in the current iteration as a value proportional to the corresponding value of the policy model generated in the preceding iteration, multiplied by the exponent of a term which is proportional to a weight factor” indefinite. For purposes of examination, Examiner has interpreted this limitation to mean the value of the policy model is weighted by a value of the policy model generated in the preceding iteration. The term “proportional” in claim 15 is a relative term which renders the claim indefinite. The term “proportional” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Thus, the term “proportional” renders the limitation of “a term proportional to a first constraint factor” indefinite. For purposes of examination, Examiner has interpreted this limitation to mean the term is a constant and thus is “proportional” to a first constraint factor. Claim 18 recites the limitation "the preceding iteration" in line 9. There is insufficient antecedent basis for this limitation in the claim. For purposes of examination, Examiner has interpreted this preceding iteration to be the first occurrence of a preceding iteration. Claim 18 recites the limitation "the current iteration" in lines 12-13. There is insufficient antecedent basis for this limitation in the claim. For purposes of examination, Examiner has interpreted this current iteration to be the first occurrence of a current iteration. Claim 18 recites the limitation “an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold” in lines 16-18. It is unclear as to whether this expected value is the same as the expected value claimed in line 14 or if the two expected values are different. Additionally, if they are different, in paragraph 0008 of the specification, it states, “actions are chosen using the policy model generated in the preceding iteration (i.e., the last-but-one iteration)” and thus, this limitation has the same interpretation as the limitation before stating “an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration”. For purposes of examination, Examiner has interpreted this limitation to be repeating the same limitation as recited immediately before. Claim 19 recites the limitation "the preceding iteration" in lines 4-5. There is insufficient antecedent basis for this limitation in the claim. For purposes of examination, Examiner has interpreted this preceding iteration to be the first occurrence of a preceding iteration. Claim 18 recites the limitation "the current iteration" in lines 6-7. There is insufficient antecedent basis for this limitation in the claim. For purposes of examination, Examiner has interpreted this current iteration to be the first occurrence of a current iteration. Claim 19 recites the limitation “an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold” in lines 10-12. It is unclear as to whether this expected value is the same as the expected value claimed in line 8 or if the two expected values are different. Additionally, if they are different, in paragraph 0008 of the specification, it states, “actions are chosen using the policy model generated in the preceding iteration (i.e., the last-but-one iteration)” and thus, this limitation has the same interpretation as the limitation before stating “an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration”. For purposes of examination, Examiner has interpreted this limitation to be repeating the same limitation as recited immediately before. Regarding claim 20, claim 20 is rejected for at least the same reasons as claim 19 since claim 20 depends on claim 19. Claim 20 recites the limitation “for each constraint, a corresponding constraint cost function indicative of expected values of the corresponding constraint reward function”. It is unclear to whether these expected values are the same as the expected values recited in claim 19. For purposes of examination, Examiner has interpreted these expected values to be the same as the expected values in claim 19. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the claim is directed to a computer storage medium (e.g. see claim 18, line 1), which includes a signal based on the broadest reasonable interpretation (i.e. The ordinary and customary meaning of a computer readable medium that includes signals per se). While the Specification discloses that the computer storage medium can be a non-transitory computer storage medium (see paragraph [0136]) the Specification is not limiting the computer readable medium to only a non-transitory embodiment. A computer readable medium, or the like, that covers both transitory and non-transitory embodiments may be amended to narrow the claim to cover only statutory embodiments to avoid a rejection under 35 U.S.C 101 by adding the limitation “non-transitory” to the claim and positively reciting that the computer readable medium is a non-transitory computer readable medium. See also In re Nuijten, 500 F.3d 1346, 1356-57 (Fed. Dir. 2007) (transitory embodiments are not directed to statutory subject matter). Examiner notes that if Applicant amends to overcome the signals per se rejection, claim # will still be rejected under 35 U.S.C. 101. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. For purposes of compact prosecution, claim 18 has been examined as if the computer storage media claimed in line 1 of claim 18 is non-transitory. Regarding Claim 1: Subject Matter Eligibility Analysis Step 1: Claim 1 recites a method and is thus a process, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 1 recites select actions of an agent interacting with an environment to perform one or more tasks, each task having at least one respective reward associated with performance of the task, (This limitation is a mental process as it encompasses a human mentally selecting actions and is thus a judgement.) processing the current observation… to select an action to be performed by the agent at the time step; (This limitation is a mental process as it encompasses a human mentally processing an observation.) each constraint limiting, to a corresponding threshold, the value of a corresponding constraint reward function dependent on at least one of the subsequent observation and the selected action, if the actions of the agent are chosen according to the policy model, each constraint being associated with a corresponding multiplier variable; (This limitation is a mental process as it encompasses a human mentally limiting a reward function with constraints and is thus an evaluation.) the method comprising a plurality of iterations, each iteration comprising (This limitation is a mental process as it encompasses a human mentally iterating the method and is thus an evaluation.) generating a mixed reward function based on values for the multiplier variables generated in the preceding iteration, and estimates of the rewards and the values of constraint reward functions if the actions are chosen based on the policy model generated in the preceding iteration; (This limitation is a mental process as it encompasses a human mentally generating a mixed reward function and is thus an evaluation.) generating an updated policy model using the mixed reward function generated in the current iteration and the mixed reward function generated in the preceding iteration; (This limitation is a mental process as it encompasses a human mentally generating a an updated policy model using the functions and is thus an evaluation.) generating an updated value of each multiplier variable based on an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold. (This limitation is a mental process as it encompasses a human mentally generating an updated value of each multiplier variable and is thus an evaluation.) Therefore, claim 1 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 1 further recites additional elements of A computer-implemented method of training a policy model of an action selection system (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) the agent being controlled by a process comprising, at a plurality of time steps: obtaining a current observation characterizing a current state of the environment; (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of data gathering (see MPEP 2106.05(g)).) using the policy model (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) obtaining a subsequent observation characterizing a subsequent state of the environment at a next time step after the agent performs the selected action (This element does not integrate the abstract idea into a practical application because it recites insignificant extra-solution activity of data gathering (see MPEP 2106.05(g)).) Therefore, claim 1 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 1 do not provide significantly more than the abstract idea itself, taken alone and in combination because A computer-implemented method of training a policy model of an action selection system uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). the agent being controlled by a process comprising, at a plurality of time steps: obtaining a current observation characterizing a current state of the environment; is the well understood, routine, and conventional activity of “transmitting or receiving data over a network” (see MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). using the policy model uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). obtaining a subsequent observation characterizing a subsequent state of the environment at a next time step after the agent performs the selected action is the well understood, routine, and conventional activity of “transmitting or receiving data over a network” (see MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). Therefore, claim 1 is subject-matter ineligible. Regarding Claim 2: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 2 recites in which the mixed reward function is based on (i) a return function indicative of expected future rewards if the actions are chosen using policy model generated in the preceding iteration, (This limitation is a mental process as it further describes the mental process of generating a mixed reward function from claim 1 and is thus an evaluation.) (ii) for each constraint, a corresponding constraint cost function indicative of expected values of the corresponding constraint reward function if the actions are chosen using the policy model generated in the preceding iteration, (This limitation is a mental process as it further describes the mental process of limiting a value with a constraint from claim 1 and is thus an evaluation.) (iii) for each constraint, a value of the corresponding multiplier variable generated in the preceding iteration. (This limitation is a mental process as it encompasses a human mentally generating a multiplier variable and is thus an evaluation.) Therefore, claim 2 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 2 does not further recite any additional elements. Therefore, claim 2 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 2 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 2 is subject-matter ineligible. Regarding Claim 3: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 3 recites in which the mixed reward function is indicative of a sum over the constraints, weighted by the respective values of the multiplier variables generated in the preceding iteration, of the corresponding constraint cost function indicative of expected values of the corresponding constraint reward function if the actions are chosen using the policy model generated in the preceding iteration, minus the return function indicative of expected future rewards if the actions are chosen using policy model generated in the preceding iteration. (This limitation is a mathematical concept.) Therefore, claim 3 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 3 does not further recite any additional elements. Therefore, claim 3 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 3 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 3 is subject-matter ineligible. Regarding Claim 4: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 4 recites minimizes an expression comprising the expected value under the updated policy model for the mixed reward function for the preceding iteration, minus a weight factor times the expected value under the updated policy model of the mixed reward function obtained in the current iteration. (This limitation is a mathematical concept.) Therefore, claim 4 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 4 further recites additional elements of in which, in each iteration, the updated policy model is generated as an updated policy model which (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 4 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 4 do not provide significantly more than the abstract idea itself, taken alone and in combination because in which, in each iteration, the updated policy model is generated as an updated policy model which uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 4 is subject-matter ineligible. Regarding Claim 5: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 5 recites in which the weight factor is 2 (This limitation is a mathematical concept since it’s further describing the mathematical concept from claim 4.) Therefore, claim 5 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 5 does not further recite any additional elements. Therefore, claim 5 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 5 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 5 is subject-matter ineligible. Regarding Claim 6: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 6 recites which minimizes an expression comprising a policy stabilization function of the policy model generated in the current observation, the policy stabilization function being indicative of a divergence between the policy model generated in the current iteration and the policy model generated in the preceding iteration. (This limitation is a mathematical concept.) Therefore, claim 6 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 6 further recites additional elements of in which, in each iteration, the policy model is generated as a policy model which (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 6 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 6 do not provide significantly more than the abstract idea itself, taken alone and in combination because in which, in each iteration, the policy model is generated as a policy model which uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 6 is subject-matter ineligible. Regarding Claim 7: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 7 recites which the policy stabilization function is a Kullback-Leibler divergence between the policy model generated in the current iteration and the policy model generated in the preceding iteration. (This limitation is a mathematical concept as it further describes the mathematical concept in claim 6.) Therefore, claim 7 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 7 does not further recite any additional elements. Therefore, claim 7 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 7 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 7 is subject-matter ineligible. Regarding Claim 8: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 8 recites in which the policy stabilization function is weighted by a first step size parameter which is different for different iterations, the weighting being higher for later iterations. (This limitation is a mathematical concept as it further describes the mathematical concept in claim 6.) Therefore, claim 8 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 8 does not further recite any additional elements. Therefore, claim 8 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 8 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 8 is subject-matter ineligible. Regarding Claim 9: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 9 recites in which, in each iteration, the updated value of the multiplier variables are generated as values for the multiplier variables which maximize an expression having a term which is a sum over the constraints of the corresponding multiplier variable multiplied by: a first constraint weight factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, minus a second constraint weight factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, minus the corresponding threshold. (This limitation is a mathematical concept.) Therefore, claim 9 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 9 does not further recite any additional elements. Therefore, claim 9 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 9 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 9 is subject-matter ineligible. Regarding Claim 10: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 10 recites in which the first constraint weight factor is 2 and the second constraint weight factor is 1. (This limitation is a mathematical concept as it further describes the mathematical concept of claim 9.) Therefore, claim 10 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 10 does not further recite any additional elements. Therefore, claim 10 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 10 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 10 is subject-matter ineligible. Regarding Claim 11: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 11 recites in which, in each iteration, the updated value of the multiplier variables are generated as values for the multiplier variables which maximize an expression having a term which is a multiplier stabilization function indicative of a difference between the multiplier variables and the values of the multiplier variables generated in the preceding iteration. (This limitation is a mathematical concept.) Therefore, claim 11 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 11 does not further recite any additional elements. Therefore, claim 11 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 11 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 11 is subject-matter ineligible. Regarding Claim 12: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 12 recites in which the multiplier stabilization function is weighted using a second step size parameter which is different for different iterations, the weighting being higher for later iterations. (This limitation is a mathematical concept as it further describes the mathematical concept of claim 11.) Therefore, claim 12 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 12 does not further recite any additional elements. Therefore, claim 12 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 12 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 12 is subject-matter ineligible. Regarding Claim 13: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 13 recites in which the policy model and mixed reward function are generated as respective tables having a value for each combination of a possible state and possible action. (This limitation is a mental process as it encompasses a human mentally generating tables and is thus an evaluation.) Therefore, claim 13 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 13 does not further recite any additional elements. Therefore, claim 13 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 13 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 13 is subject-matter ineligible. Regarding Claim 14: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 14 recites in which the value of the policy model for each combination is generated in the current iteration as a value proportional to the corresponding value of the policy model generated in the preceding iteration, multiplied by the exponent of a term which is proportional to a weight factor times the mixed reward function obtained in the current iteration, minus the mixed reward function generated in the preceding iteration. (This limitation is a mathematical concept.) Therefore, claim 14 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 14 does not further recite any additional elements. Therefore, claim 14 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 14 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 14 is subject-matter ineligible. Regarding Claim 15: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 15 recites wherein the value of the each multiplier variable is generated in the current iteration as the higher of (i) zero and (ii) the sum of the value of the multiplier variable in the preceding iteration plus a term proportional to a first constraint factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, minus a second constraint factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, minus the corresponding threshold. (This limitation is a mathematical concept.) Therefore, claim 15 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 15 does not further recite any additional elements. Therefore, claim 15 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 15 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 15 is subject-matter ineligible. Regarding Claim 16: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 16 recites in which the policy model and mixed reward function are based on corresponding neural networks, and in each iteration the generating of the mixed reward function and the generating of the updated policy model comprise generating corresponding sets of numerical parameters for the corresponding neural network models. (This limitation is a mental process as it encompasses a human mentally generating sets of parameters and is thus an evaluation.) Therefore, claim 16 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 16 does not further recite any additional elements. Therefore, claim 16 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 16 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 16 is subject-matter ineligible. Regarding Claim 17: Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 17 recites in which the parameters of the neural network generated to generate the mixed reward function in each iteration are employed in the generation of the policy model in the next iteration. (This limitation is a mental process as it encompasses a human mentally employing the generation of parameters and is thus an evaluation.) Therefore, claim 17 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 17 does not further recite any additional elements. Therefore, claim 17 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: Since there are no additional elements, claim 17 does not provide significantly more than the abstract idea itself, taken alone and in combination. Therefore, claim 17 is subject-matter ineligible. Regarding Claim 18: Subject Matter Eligibility Analysis Step 1: Claim 18 has been rejected because claimed invention is directed to non-statutory subject matter that does not fall within at least one of the four categories of patent eligible subject matter. However for the purposes of compact prosecution, claim 18 has been examined below if the claim were to be amended to fall under one of the four categories of patent eligible subject matter. Thus, Claim 18 is examined as if the computer storage media is non-transitory. Claim 18 recites a computer storage media and is thus an article of manufacture, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 18 recites select, based on observations characterizing a current state of an environment, actions to be performed by an agent interacting with the environment to perform one or more tasks subject to one or more constraints, (This limitation is a mental process as it encompasses a human mentally selecting actions to be performed and is thus a judgment.) each task having at least one respective reward associated with performance of the task, and each constraint limiting, to a corresponding threshold, the value of a corresponding constraint reward function dependent on at least one of the observation and the selected action, each constraint being associated with a corresponding multiplier variable; (This limitation is a mental process as it encompasses a human mentally limiting a reward function with constraints and is thus an evaluation.) each iteration comprising: generating a mixed reward function based on values for the multiplier variables generated in the preceding iteration, and estimates of the rewards and the values of constraint reward functions if the actions are chosen based on the policy model generated in the preceding iteration; (This limitation is a mental process as it encompasses a human mentally generating a mixed reward function and is thus an evaluation.) generating an updated policy model using the mixed reward function generated in the current iteration and the mixed reward function generated in the preceding iteration; (This limitation is a mental process as it encompasses a human mentally generating a an updated policy model using the functions and is thus an evaluation.) generating an updated value of each multiplier variable based on an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold. (This limitation is a mental process as it encompasses a human mentally generating an updated value of each multiplier variable and is thus an evaluation.) Therefore, claim 18 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 18 further recites additional elements of One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to train iteratively an action selection neural network system to (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 18 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 18 do not provide significantly more than the abstract idea itself, taken alone and in combination because One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to train iteratively an action selection neural network system to uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 18 is subject-matter ineligible. Regarding Claim 19: Subject Matter Eligibility Analysis Step 1: Claim 19 recites a system and is thus an apparatus, one of the four statutory categories of patentable subject matter. Subject Matter Eligibility Analysis Step 2A Prong 1: Claim 19 recites select, based on observations characterizing a current state of an environment, actions to be performed by an agent interacting with the environment to perform one or more tasks subject to one or more constraints, (This limitation is a mental process as it encompasses a human mentally selecting actions to be performed and is thus a judgment.) each task having at least one respective reward associated with performance of the task, and each constraint limiting, to a corresponding threshold, the value of a corresponding constraint reward function dependent on at least one of the observation and the selected action, each constraint being associated with a corresponding multiplier variable; (This limitation is a mental process as it encompasses a human mentally limiting a reward function with constraints and is thus an evaluation.) each iteration comprising: generating a mixed reward function based on values for the multiplier variables generated in the preceding iteration, and estimates of the rewards and the values of constraint reward functions if the actions are chosen based on the policy model generated in the preceding iteration; (This limitation is a mental process as it encompasses a human mentally generating a mixed reward function and is thus an evaluation.) generating an updated policy model using the mixed reward function generated in the current iteration and the mixed reward function generated in the preceding iteration; (This limitation is a mental process as it encompasses a human mentally generating a an updated policy model using the functions and is thus an evaluation.) generating an updated value of each multiplier variable based on an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold. (This limitation is a mental process as it encompasses a human mentally generating an updated value of each multiplier variable and is thus an evaluation.) Therefore, claim 19 recites an abstract idea. Subject Matter Eligibility Analysis Step 2A Prong 2: Claim 19 further recites additional elements of One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to train iteratively an action selection neural network system to (This element does not integrate the abstract idea into a practical application because it amounts to mere “apply it on a computer” (see MPEP 2106.05(f)).) Therefore, claim 19 is not integrated into a practical application. Subject Matter Eligibility Analysis Step 2B: The additional elements of claim 19 do not provide significantly more than the abstract idea itself, taken alone and in combination because One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to train iteratively an action selection neural network system to uses a computer as a tool to perform the abstract idea and cannot provide significantly more (see MPEP 2106.05(f)). Therefore, claim 19 is subject-matter ineligible. Regarding claim 20, claim 20 recites substantially similar limitations to claim 2, and is therefore rejected under the same analysis. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-2 and 16-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Tessler et al. (“Reward Constrained Policy Optimization”) (hereafter referred to as Tessler). Regarding claim 1, Tessler teaches A computer-implemented method of training a policy model of an action selection system to select actions of an agent interacting with an environment to perform one or more tasks, each task having at least one respective reward associated with performance of the task, the agent being controlled by a process comprising, at a plurality of time steps (Tessler, page 2, “A Markov Decision Processes M is defined by the tuple (S, A, R, P, μ, γ) (Sutton and Barto, 1998). Where S is the set of states, A is the available actions, R: S x A x S [Wingdings font/0xE0] R is the reward function, P: S x A x S [Wingdings font/0xE0] [0,1] is the transition matrix, where P(s’|s,a) is the probability of transitioning from state s to s’ assuming action a was taken, μ:S [Wingdings font/0xE0] [0,1] is the initial state distribution and γ ∈ [0,1) is the discount factor for future rewards. A policy π:S[Wingdings font/0xE0] Δ A is a probability distribution over actions and π(a|s) denotes the probability of taking action a at state s.”) : obtaining a current observation characterizing a current state of the environment (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that line 4 is the current observation characterizing a current state of the environment.); processing the current observation using the policy model to select an action to be performed by the agent at the time step (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that line 6 is the chosen action at the time-step.); and obtaining a subsequent observation characterizing a subsequent state of the environment at a next time step after the agent performs the selected action (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that observing the next state in line 6 is the subsequent state of the environment at a next time step.); each constraint limiting, to a corresponding threshold, the value of a corresponding constraint reward function dependent on at least one of the subsequent observation and the selected action, if the actions of the agent are chosen according to the policy model, each constraint being associated with a corresponding multiplier variable (Tessler, page 3, Section 2.2 Constrained MDPs, “ A Constrained Markov Decision Process (CMDP) extends the MDP framework by introducing a penalty c(s,a), and constraint C(st) = F(c(st, at), …, c (sN, aN)) and a threshold α ∈ [0,1]” where “Given a CMDP (3), the unconstrained problem is PNG media_image2.png 37 593 media_image2.png Greyscale where L is the Lagrangian and λ ≥ 0 is the Lagrange multiplier (a penalty coefficient)” (Tessler, page 3, section 3 Constrained Policy Optimization). Examiner notes that the Lagrangian multiplier is the multiplier variable. Examiner further notes that the value of a corresponding constraint reward function is J.); the method comprising a plurality of iterations, each iteration comprising: generating a mixed reward function based on values for the multiplier variables generated in the preceding iteration, and estimates of the rewards and the values of constraint reward functions if the actions are chosen based on the policy model generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale where “As opposed to (4), for a fixed π and λ, the penalized value (11) can be estimated using TD-learning critic” (Tessler, page 5, 1st paragraph). Examiner notes that the Langrangian multiplier is the multiplier variables and the mixed reward function is the line 7. Examiner further notes that line 8 has the estimates of the rewards and is thus used In iterations.); generating an updated policy model using the mixed reward function generated in the current iteration and the mixed reward function generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the critic and actor update are the steps to update a policy model in which R is the mixed reward. Examiner further notes that since the updates are within the iteration, the current and previous mixed reward update are used.); and generating an updated value of each multiplier variable based on an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the updated value of each multiplier variable is each update to the Lagrange multiplier. Examiner further notes that the expected value for the corresponding constraint reward function is R in line 7 for each iteration. Examiner further notes that the threshold is α in line 10. ). Regarding claim 2, Tessler teaches The method of claim 1, in which the mixed reward function is based on (i) a return function indicative of expected future rewards if the actions are chosen using policy model generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that line 7 is the mixed reward function. Examiner further notes that R is the return function indicative of expected future rewards.), (ii) for each constraint, a corresponding constraint cost function indicative of expected values of the corresponding constraint reward function if the actions are chosen using the policy model generated in the preceding iteration, and (Tessler, page 3, Section 2.2 Constrained MDPs, “ A Constrained Markov Decision Process (CMDP) extends the MDP framework by introducing a penalty c(s,a), and constraint C(st) = F(c(st, at), …, c (sN, aN)) and a threshold α ∈ [0,1]” where “Given a CMDP (3), the unconstrained problem is PNG media_image2.png 37 593 media_image2.png Greyscale where L is the Lagrangian and λ ≥ 0 is the Lagrange multiplier (a penalty coefficient)” (Tessler, page 3, section 3 Constrained Policy Optimization) and Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale . Examiner notes that the constraint cost function is equation 4. Examiner further notes that the expected value is R and is found by updating λ.) (iii) for each constraint, a value of the corresponding multiplier variable generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the multiplier variable is each update to the Lagrange multiplier. Examiner further notes that the expected value for the corresponding constraint reward function is R in line 7 for each iteration.). Regarding claim 16, Tessler teaches The method of claim 1 in which the policy model and mixed reward function are based on corresponding neural networks, and in each iteration the generating of the mixed reward function and the generating of the updated policy model comprise generating corresponding sets of numerical parameters for the corresponding neural network models (Tessler, page 3, 2.3 Parametrized Policies, “In this work we consider parametrized policies, such as neural networks. The parameters of the policy are denoted by θ and a parametrized policy as π θ ” and Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the mixed reward function is line 7 and the parameters are θ. ). Regarding claim 17, Tessler teaches The method of claim 16 in which the parameters of the neural network generated to generate the mixed reward function in each iteration are employed in the generation of the policy model in the next iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the mixed reward function is line 7 and the parameters are θ. Examiner notes that since the algorithm iterates, the parameters and mixed reward function are employed in the next iteration.). Regarding claim 18, Tessler teaches One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to train iteratively an action selection neural network system to select based on observations characterizing a current state of an environment, actions to be performed by an agent interacting with the environment to perform one or more tasks subject to one or more constraints, each task having at least one respective reward associated with performance of the task, the agent being controlled by a process comprising, at a plurality of time steps (Tessler, page 2, “A Markov Decision Processes M is defined by the tuple (S, A, R, P, μ, γ) (Sutton and Barto, 1998). Where S is the set of states, A is the available actions, R: S x A x S [Wingdings font/0xE0] R is the reward function, P: S x A x S [Wingdings font/0xE0] [0,1] is the transition matrix, where P(s’|s,a) is the probability of transitioning from state s to s’ assuming action a was taken, μ:S [Wingdings font/0xE0] [0,1] is the initial state distribution and γ ∈ [0,1) is the discount factor for future rewards. A policy π:S[Wingdings font/0xE0] Δ A is a probability distribution over actions and π(a|s) denotes the probability of taking action a at state s” where “In the following experiments; the aim is to prolong the motor life of the various robots, while still enabling the robot to perform the task at hand. To do so, the robot motors need to be constrained from using high torque values. This is accomplished by defining the constraint C as the average torque the agent has applied to each motor, and the per-state penalty c(s,a) becomes the amount of torque the agent decided to apply at each time step” (Tessler, page 7, 5.2.2) and Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that line 4 is the current observation characterizing a current state of the environment. Examiner notes that line 6 is the chosen action at the time-step. Examiner notes that the robot and corresponding agent is the computer storage media.) : each constraint limiting, to a corresponding threshold, the value of a corresponding constraint reward function dependent on at least one of the observation and the selected action, each constraint being associated with a corresponding multiplier variable (Tessler, page 3, Section 2.2 Constrained MDPs, “ A Constrained Markov Decision Process (CMDP) extends the MDP framework by introducing a penalty c(s,a), and constraint C(st) = F(c(st, at), …, c (sN, aN)) and a threshold α ∈ [0,1]” where “Given a CMDP (3), the unconstrained problem is PNG media_image2.png 37 593 media_image2.png Greyscale where L is the Lagrangian and λ ≥ 0 is the Lagrange multiplier (a penalty coefficient)” (Tessler, page 3, section 3 Constrained Policy Optimization). Examiner notes that the Lagrangian multiplier is the multiplier variable. Examiner further notes that the value of a corresponding constraint reward function is J.); each iteration comprising: generating a mixed reward function based on values for the multiplier variables generated in the preceding iteration, and estimates of the rewards and the values of constraint reward functions if the actions are chosen based on the policy model generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale where “As opposed to (4), for a fixed π and λ, the penalized value (11) can be estimated using TD-learning critic” (Tessler, page 5, 1st paragraph). Examiner notes that the Langrangian multiplier is the multiplier variables and the mixed reward function is the line 7. Examiner further notes that line 8 has the estimates of the rewards and is thus used In iterations.); generating an updated policy model using the mixed reward function generated in the current iteration and the mixed reward function generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the critic and actor update are the steps to update a policy model in which R is the mixed reward. Examiner further notes that since the updates are within the iteration, the current and previous mixed reward update are used.); and generating an updated value of each multiplier variable based on an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the updated value of each multiplier variable is each update to the Lagrange multiplier. Examiner further notes that the expected value for the corresponding constraint reward function is R in line 7 for each iteration. Examiner further notes that the threshold is α in line 10. ). Regarding claim 19, Tessler teaches A system comprising one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to train iteratively an action selection neural network system to select based on an observation characterizing a current state of an environment, an action to be performed by an agent interacting with the environment to perform one or more tasks subject to one or more constraints, each task having at least one respective reward associated with performance of the task, (Tessler, page 2, “A Markov Decision Processes M is defined by the tuple (S, A, R, P, μ, γ) (Sutton and Barto, 1998). Where S is the set of states, A is the available actions, R: S x A x S [Wingdings font/0xE0] R is the reward function, P: S x A x S [Wingdings font/0xE0] [0,1] is the transition matrix, where P(s’|s,a) is the probability of transitioning from state s to s’ assuming action a was taken, μ:S [Wingdings font/0xE0] [0,1] is the initial state distribution and γ ∈ [0,1) is the discount factor for future rewards. A policy π:S[Wingdings font/0xE0] Δ A is a probability distribution over actions and π(a|s) denotes the probability of taking action a at state s” where “In the following experiments; the aim is to prolong the motor life of the various robots, while still enabling the robot to perform the task at hand. To do so, the robot motors need to be constrained from using high torque values. This is accomplished by defining the constraint C as the average torque the agent has applied to each motor, and the per-state penalty c(s,a) becomes the amount of torque the agent decided to apply at each time step” (Tessler, page 7, 5.2.2) and Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that line 4 is the current observation characterizing a current state of the environment. Examiner notes that line 6 is the chosen action at the time-step. Examiner notes that the robot and corresponding agent is the computer storage media.) : each constraint limiting, to a corresponding threshold, the value of a corresponding constraint reward function dependent on at least one of the observation and the selected action, each constraint being associated with a corresponding multiplier variable (Tessler, page 3, Section 2.2 Constrained MDPs, “ A Constrained Markov Decision Process (CMDP) extends the MDP framework by introducing a penalty c(s,a), and constraint C(st) = F(c(st, at), …, c (sN, aN)) and a threshold α ∈ [0,1]” where “Given a CMDP (3), the unconstrained problem is PNG media_image2.png 37 593 media_image2.png Greyscale where L is the Lagrangian and λ ≥ 0 is the Lagrange multiplier (a penalty coefficient)” (Tessler, page 3, section 3 Constrained Policy Optimization). Examiner notes that the Lagrangian multiplier is the multiplier variable. Examiner further notes that the value of a corresponding constraint reward function is J.); each iteration comprising: generating a mixed reward function based on values for the multiplier variables generated in the preceding iteration, and estimates of the rewards and the values of constraint reward functions if the actions are chosen based on the policy model generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale where “As opposed to (4), for a fixed π and λ, the penalized value (11) can be estimated using TD-learning critic” (Tessler, page 5, 1st paragraph). Examiner notes that the Langrangian multiplier is the multiplier variables and the mixed reward function is the line 7. Examiner further notes that line 8 has the estimates of the rewards and is thus used In iterations.); generating an updated policy model using the mixed reward function generated in the current iteration and the mixed reward function generated in the preceding iteration (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the critic and actor update are the steps to update a policy model in which R is the mixed reward. Examiner further notes that since the updates are within the iteration, the current and previous mixed reward update are used.); and generating an updated value of each multiplier variable based on an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, an expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, and the corresponding threshold (Tessler, page 5, Algorithm 1, PNG media_image1.png 347 776 media_image1.png Greyscale Examiner notes that the updated value of each multiplier variable is each update to the Lagrange multiplier. Examiner further notes that the expected value for the corresponding constraint reward function is R in line 7 for each iteration. Examiner further notes that the threshold is α in line 10. ). Regarding claim 20, claim 20 recites substantially similar limitations to claim 2, and is therefore rejected under the same analysis. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 6-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tessler in view of Ghosh et al. (“Divide-and-Conquer Reinforcement Learning”) (hereafter referred to as Ghosh). Regarding claim 6, Tessler teaches the method of claim 1. Tessler does not teach, but Ghosh does teach in which, in each iteration, the policy model is generated as a policy model which minimizes an expression comprising a policy stabilization function of the policy model generated in the current observation, the policy stabilization function being indicative of a divergence between the policy model generated in the current iteration and the policy model generated in the preceding iteration (Ghosh, page 4, 1st paragraph, “Using Jensen’s inequality to bound the KL divergence between π and π--c, we minimize the right hand side of Equation 1 as a bound for minimizing the intractable KL divergence optimization problem. PNG media_image3.png 58 609 media_image3.png Greyscale Equation 1 shows that finding a set of local policies that translates well into a global policy reduces into minimizing a weighted sum of pairwise KL divergence terms between local policies.” Examiner notes that the policy model generated in the current iteration is πi and the policy model generated in the preceding iteration is πj. Examiner further notes that the policy stabilization function is the KL divergence.). Tessler and Ghosh are considered analogous to the claimed invention because they both use reinforcement learning. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Tessler to use the KL-divergence like in Ghosh. Doing so is advantageous because “our experimental results show that divide-and-conquer reinforcement learning substantially outperforms standard RL algorithms that samples initial and goal states from their respective distributions at each trial, as well as previously proposed methods that employ ensembles of policies” (Ghosh, page 9, 3rd paragraph). Regarding claim 7, Tessler in view of Ghosh teaches the method of claim 6. Tessler in view of Ghosh further teach in which the policy stabilization function is a Kullback-Leibler divergence between the policy model generated in the current iteration and the policy model generated in the preceding iteration (Ghosh, page 4, 1st paragraph, “Using Jensen’s inequality to bound the KL divergence between π and π--c, we minimize the right hand side of Equation 1 as a bound for minimizing the intractable KL divergence optimization problem. PNG media_image3.png 58 609 media_image3.png Greyscale Equation 1 shows that finding a set of local policies that translates well into a global policy reduces into minimizing a weighted sum of pairwise KL divergence terms between local policies.” Examiner notes that the policy model generated in the current iteration is πi and the policy model generated in the preceding iteration is πj. Examiner further notes that the policy stabilization function is the KL divergence.). Tessler and Ghosh are considered analogous to the claimed invention because they both use reinforcement learning. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Tessler to use the KL-divergence like in Ghosh. Doing so is advantageous because “our experimental results show that divide-and-conquer reinforcement learning substantially outperforms standard RL algorithms that samples initial and goal states from their respective distributions at each trial, as well as previously proposed methods that employ ensembles of policies” (Ghosh, page 9, 3rd paragraph). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tessler in view of ADL (“An introduction to Q-Learning: reinforcement learning”) (hereafter referred to as ADL). Regarding claim 13, Tessler teaches the method of claim 1. Tessler does not teach, but ADL does teach in which the policy model and mixed reward function are generated as respective tables having a value for each combination of a possible state and possible action (ADL, page 6, Q-Function, “The Q-function uses the Bellman equation and takes two inputs: state (s) and action (a). Using the above function, we get the values of Q for the cells in the table. When we start, all values in the Q-table are zeros. There is an iterative process of updating the values. As we start to explore the environment, the Q-function gives better and better approximations by continuously updating the Q-values in the table” where “In the Q-Table, the columns are the actions and the rows are the states” (ADL, page 4, last paragraph). Examiner notes that the Bellman function is the policy model and mixed reward function.). Tessler and ADL are considered analogous to the claimed invention because they both use reinforcement learning. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Tessler to use the Q-Learning like in ADL. Doing so is advantageous because “the Q table helps us to find the best action for each state. It helps to maximize the expected reward by selecting the best of all possible actions” (ADL, page 12, 3rd and 4th bullet points). Allowable Subject Matter Claims 3-5, 8-12, and 14-15 currently do not have art applied to them under the interpretations given in the 112(b) rejections above. Should the 112(b) rejections be overcome and change the interpretation of the claims, further search and consideration will be required. Thus, Claims 3-5, 8-12, and 14-15 would be allowable over the prior art of record if the 112(b) and 101 rejections are overcome in light of the instant amendments. Specifically, regarding claim 3, “minus the return function indicative of expected future rewards if the actions are chosen using the policy model generated in the preceding iteration” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler. Tessler teaches a mixed reward function indicative of a sum over constraints, weighted by the respective values of the multiplier variables indicative of expected values of the corresponding constraint reward function if the actions are chosen using the policy model generated in the preceding iteration (Tessler, page 5, Algorithm 1). Tessler does not however teach minus the return function indicative of expected future rewards if the actions are chosen using the policy model generated in the preceding iteration. Therefore, the prior art of record, individually, or in combination, does not disclose claim 3 as a whole. Specifically, regarding claim 4, “minus a weight factor times the expected value under the updated policy model of the mixed reward function obtained in the current iteration” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler and Ghosh. Tessler teaches a updated policy model, expected value, mixed reward function, and iterations (Tessler, page 5, Algorithm 1). Tessler does not teach minimizing an expression. Ghosh teaches minimizing an expression (Ghosh page 4, 1st paragraph), but does not teach minus a weight factor times the expected value under the updated policy model of the mixed reward function obtained in the current iteration. Therefore, the prior art of record, individually, or in combination, does not disclose claim 4 as a whole. Claim 5 is allowable at least due to their dependencies on claim 5 if the 101 and 112(b) rejections are overcome. Specifically, regarding claim 8, “the policy stabilization function is weighted by a first step size parameter which is different for different iterations, the weighting being higher for later iterations” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler and Ghosh. Tessler and Ghosh both teach iterations (Tessler, page 5, Algorithm 1). Tessler does not teach the policy stabilization function. Ghosh teaches the policy stabilization function weighted by a first step size parameter which is different for different iterations (Ghosh page 4, 1st paragraph), but does not teach that the weighting is higher for later iterations. Therefore, the prior art of record, individually, or in combination, does not disclose claim 8 as a whole. Specifically, regarding claim 9, “in each iteration, the updated value of the multiplier variables are generated as values for the multiplier variables which maximize an expression having a term which is a sum over the constraints of the corresponding multiplier variable multiplied by: a first constraint weight factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, minus a second constraint weight factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last- but-one iteration, minus the corresponding threshold.” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler. Tessler teaches iterations and updated values of the multiplier variables (Tessler, page 5, Algorithm 1). Tessler does not however teach maximizing and expression having a term which is a sum over the constraints of the corresponding multiplier variable multiplied by a first constraint weight factor times the expected for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, minus a second constraint weight factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last- but-one iteration, minus the corresponding threshold. Therefore, the prior art of record, individually, or in combination, does not disclose claim 9 as a whole. Claim 10 is allowable at least due to their dependencies on claim 9 if the 101 and 112(b) rejections are overcome. Specifically, regarding claim 11, “in each iteration, the updated value of the multiplier variables are generated as values for the multiplier variables which maximize an expression having a term which is a multiplier stabilization function indicative of a difference between the multiplier variables and the values of the multiplier variables generated in the preceding iteration.” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler and Ghosh. Tessler teaches iterations and updated values of the multiplier variables (Tessler, page 5, Algorithm 1). Tessler does not however teach maximizing and expression having a term a multiplier stabilization function indicative of a difference between the multiplier variables and the values of the multiplier variables generated in the preceding iteration. Ghosh teaches a stabilization function (Ghosh page 4, 1st paragraph) but does not teach a difference between the multiplier variables and the values of the multiplier variables generated in the preceding iteration. Therefore, the prior art of record, individually, or in combination, does not disclose claim 11 as a whole. Claim 12 is allowable at least due to their dependencies on claim 11 if the 101 and 112(b) rejections are overcome. Specifically regarding claim 14, “in which the value of the policy model for each combination is generated in the current iteration as a value proportional to the corresponding value of the policy model generated in the preceding iteration, multiplied by the exponent of a term which is proportional to a weight factor times the mixed reward function obtained in the current iteration, minus the mixed reward function generated in the preceding iteration” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler and ADL. Tessler teaches the mixed reward function and iterations (Tessler, page 5, Algorithm 1). Tessler does not disclose each combination nor a value proportional to the corresponding value of the policy model generated in the preceding iteration, multiplied by the exponent of a term which is proportional to a weight factor times the mixed reward function obtained in the current iteration, minus the mixed reward function generated in the preceding iteration. ADL teaches the value of the policy model for each combination (ADL, page 4, last paragraph). ADL does not disclose a value proportional to the corresponding value of the policy model generated in the preceding iteration, multiplied by the exponent of a term which is proportional to a weight factor times the mixed reward function obtained in the current iteration, minus the mixed reward function generated in the preceding iteration. Therefore, the prior art of record, individually, or in combination, does not disclose claim 14 as a whole. Specifically regarding claim 15, “wherein the value of the each multiplier variable is generated in the current iteration as the higher of (i) zero and (ii) the sum of the value of the multiplier variable in the preceding iteration plus a term proportional to a first constraint factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, minus a second constraint factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, minus the corresponding threshold” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Tessler. Tessler teaches each multiplier variable, constraint reward function, policy model, iterations and threshold (Tessler, page 5, Algorithm 1). Tessler does not disclose wherein the value of the each multiplier variable is generated in the current iteration as the higher of (i) zero and (ii) the sum of the value of the multiplier variable in the preceding iteration plus a term proportional to a first constraint factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the preceding iteration, minus a second constraint factor times the expected value for the corresponding constraint reward function if actions are chosen using the policy model generated in the last-but-one iteration, minus the corresponding threshold. Therefore, the prior art of record, individually, or in combination, does not disclose claim 15 as a whole. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Moskovitz et al. (“ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPs”) is by the inventors and within the grace period under 102(b)(1). Moskovitz et al. however discusses last-iterate convergent and constraints in reinforcement learning as well as q-learning. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KAITLYN R LAU whose telephone number is (571)272-1429. The examiner can normally be reached Monday - Thursday: 8:00 am - 6:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.R.L./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Jan 26, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688298
FEATURE SELECTION FOR CYBERSECURITY THREAT DISPOSITION
4y 7m to grant Granted Jul 21, 2026
Patent 12602431
METHODS FOR PERFORMING INPUT-OUTPUT OPERATIONS IN A STORAGE SYSTEM USING ARTIFICIAL INTELLIGENCE AND DEVICES THEREOF
3y 10m to grant Granted Apr 14, 2026
Patent 12572828
METHOD FOR INDUSTRY TEXT INCREMENT AND ELECTRONIC DEVICE
4y 5m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
56%
Grant Probability
99%
With Interview (+80.0%)
4y 0m (~1y 6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 9 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month