Prosecution Insights
Last updated: October 02, 2026
Application No. 18/079,123

PROVIDING TRAINED REINFORCEMENT LEARNING SYSTEMS

Final Rejection §101§103§112
Filed
Dec 12, 2022
Examiner
SIPPEL, MOLLY CLARKE
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Massachusetts Institute of Technology
OA Round
2 (Final)
50%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
14 granted / 28 resolved
-5.0% vs TC avg
Strong +26% interview lift
Without
With
+25.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
18 currently pending
Career history
42
Total Applications
across all art units

Statute-Specific Performance

§101
34.4%
-5.6% vs TC avg
§103
31.6%
-8.4% vs TC avg
§102
10.0%
-30.0% vs TC avg
§112
22.8%
-17.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 28 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This action is responsive to the amendment filed on 06/24/2026. Claims 1-20 are pending in the case. Claims 1-3, 6, 8-10, 13, 15-17, and 20 are currently amended. Claims 1, 8, and 15 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 03/16/2026 is being considered by the examiner. Claim Objections Claim 8 objected to because of the following informalities: “a trained reinforcement learning models” in lines 1-2 should read “a trained reinforcement learning model” as it appears to be a typographical error. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “unstable” in claims 1, 8, and 15 is a relative term which renders the claim indefinite. The term “unstable” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The "system" Claims 2-7, 9-14, and 16-20 are rejected as being dependent upon a rejected base claim without curing any of the deficiencies. Regarding claim 2, the claim recites: “the model” on line 5. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “a reinforcement (RL) model”. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the RL model” referring to the previously recited claim element. Regarding claim 6, the claim recites “the model” on line 8. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “a reinforcement (RL) model”. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the RL model” referring to the previously recited claim element. Further, the claim recites: “the system” on line 7. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “an unstable system” in line 3. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the unstable system” referring to the previously recited claim element. Regarding claim 9, the claim recites: “the system” on line 4. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “an unstable system” in line 2 and “one or more computer systems” in line 5. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the unstable system” referring to the previously recited claim element. Regarding claim 10, the claim recites: “the system” on line 4. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “an unstable system” in line 2 and “one or more computer systems” in line 5. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the unstable system” referring to the previously recited claim element. Regarding claim 13, the claim recites: “the system” on line 5. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “an unstable system” in line 2 and “one or more computer systems” in line 5. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the unstable system” referring to the previously recited claim element. Regarding claim 16, the claim recites: “the system” on line 4. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “an unstable system” in line 2 and “a computer system” in line 1. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the unstable system” referring to the previously recited claim element. Regarding claim 20, the claim recites “the model” on line 6. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “a reinforcement (RL) model”. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the RL model” referring to the previously recited claim element. Further, the claim recites: “the system” on line 5. There is insufficient antecedent basis for this limitation in the claim. The parent claim recites: “an unstable system” in line 2 and “a computer system” in line 1. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes, this limitation is being interpreted as “the unstable system” referring to the previously recited claim element. The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claims 2-3, 6, 9-10, 13, 16-17, and 20 rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 2 recites the “defining”, “training”, and “providing” steps of claim 1 with respect to the “logarithmic loss function” while claim 1 recites the three steps for both the “logarithmic loss function” and “initiation point”. Claim 2 places no further limitations not already present in claim 1. Claim 3 recites the “defining”, “training”, and “providing” steps of claim 1 with respect to the “initiation point” while claim 1 recites the three steps for both the “logarithmic loss function” and “initiation point”. Claim 3 places no further limitations not already present in claim 1. Claim 6 is substantially similar to claim 1 and places no further limitations on claim 1. Claim 9 recites the “defining”, “training”, and “providing” steps of claim 1 with respect to the “logarithmic loss function” while claim 8 recites the three steps for both the “logarithmic loss function” and “initiation point”. Claim 9 places no further limitations not already present in claim 8. Claim 10 recites the “defining”, “training”, and “providing” steps of claim 8 with respect to the “initiation point” while claim 8 recites the three steps for both the “logarithmic loss function” and “initiation point”. Claim 10 places no further limitations not already present in claim 8. Claim 13 is substantially similar to claim 8 and places no further limitations on claim 8. Claim 16 recites the “defining”, “training”, and “providing” steps of claim 15 with respect to the “logarithmic loss function” while claim 15 recites the three steps for both the “logarithmic loss function” and “initiation point”. Claim 16 places no further limitations not already present in claim 15. Claim 17 recites the “defining”, “training”, and “providing” steps of claim 15 with respect to the “initiation point” while claim 15 recites the three steps for both the “logarithmic loss function” and “initiation point”. Claim 17 places no further limitations not already present in claim 15. Claim 20 is substantially similar to claim 15 and places no further limitations on claim 15. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1 Statutory Category: Claim 1 is directed to a method, which falls within one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 1 recites, in part, “formulating, …, a decision process problem for the RL model”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “defining, …, a logarithmic loss function for the RL model and defining an initiation point for the RL model according to an optimized spectral norm of unstable system”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular the claim recites that the method is “computer-implemented” and that each step is performed “by [the] one or more computer processors”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites “a reinforcement learning (RL) model for an unstable system”. This limitation is an additional element that generally link links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “training, …, the RL model using system data to define a policy according to the logarithmic loss function from the initiation point”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Finally, the claim recites: “providing, …, the trained RL model”. This limitation is an additional element that amounts to a post-solution step and as such is considered insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Step 2B Significantly More: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: that the method is “computer-implemented”, that each step is performed “by [the] one or more computer processors”, and “training, …, the RL model using system data to define a policy according to the logarithmic loss function from the initiation point” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional element “a reinforcement learning (RL) model for an unstable system” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Finally, the additional element “providing, …, the trained RL model” amounts to insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 2, the rejection of claim 1 is incorporated, and further, the claim recites: “defining, …, the logarithmic loss function for the RL model”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “training, …, the model according to the logarithmic loss function” and “by the one or more computer processors”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the claim recites: “providing, …, the trained RL model”. This limitation is an additional element that amounts to a post-solution step and as such is considered insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 3, the rejection of claim 1 is incorporated, and further, the claim recites: “defining, …, the initiation point for the RL model according to the optimized spectral norm”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “training, …, the RL model from the initiation point” and “by the one or more computer processors”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 4, the rejection of claim 3 is incorporated, and further, the claim recites: “wherein defining the initiation point comprises regulating a system spectral radius”. This limitation is a continuation of the “defining, …, a logarithmic loss function for the RL model and defining an initiation point for the RL model according to an optimized spectral norm of unstable system” limitation identified as an abstract idea in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 5, the rejection of claim 4 is incorporated, and further, the claim recites: “wherein defining the initiation point comprises defining an initiation point wherein a magnitude of an absolute value of the system spectral radius is less than 1”. This limitation is a continuation of the “defining, …, a logarithmic loss function for the RL model and defining an initiation point for the RL model according to an optimized spectral norm of unstable system” limitation identified as an abstract idea in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 6, the rejection of claim 1 is incorporated, and further, the claim recites: “formulating, …, a decision process problem for the RL model”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “defining, …, the logarithmic loss function for the RL model and the initiation point for the RL model according to an optimized spectral norm of the system”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Further, the claim recites: that each step is performed “by [the] one or more computer processors”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the claim recites: “training, …, the system from the initiation point and according to the logarithmic loss function”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Finally, the claim recites: “providing, …, the trained RL model”. This limitation is an additional element that amounts to a post-solution step and as such is considered insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 7, the rejection of claim 6 is incorporated, and further, the claim recites: “defining, …, the initiation point according to a system spectral radius magnitude of less than 1”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim, and thus the claim recites a judicial exception. Further, the claim recites: “by the one or more computer processors”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 8: Step 1 Statutory Category: Claim 8 is directed to an article of manufacture, which falls within one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 8 recites, in part, “formulate a decision process problem for the RL model”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “define a logarithmic loss function for the RL model and an initiation point for the RL model according to an optimized spectral norm of the unstable system”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular, the claim recites “a computer program product”, “one or more computer readable storage media”, “collectively stored program instructions”, and “one or more computer systems”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites “a trained reinforcement learning models for an unstable system”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “train the RL model, using system data to define a policy according to the logarithmic loss function starting from the initiation point”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Finally, the claim recites: “provide the trained RL model”. This limitation is an additional element that amounts to a post-solution step and as such is considered insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Step 2B Significantly More: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: that the method is “a computer program product”, “one or more computer readable storage media”, “collectively stored program instructions”, “one or more computer systems”, and “train the system according to the logarithmic loss function or from the initiation point” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional element “a trained reinforcement learning models for an unstable system” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Finally, the additional element “provide the trained RL model” amounts to insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 9, the rejection of claim 8 is incorporated, and further, claim 9 is substantially similar to claim 2 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 10, the rejection of claim 8 is incorporated, and further, claim 10 is substantially similar to claim 3 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 11, the rejection of claim 10 is incorporated, and further, claim 11 is substantially similar to claim 4 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 12, the rejection of claim 11 is incorporated, and further, claims 12 is substantially similar to claim 5 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 13, the rejection of claim 8 is incorporated, and further, claim 13 is substantially similar to claim 6 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 14, the rejection of claim 13 is incorporated, and further, claim 14 is substantially similar to claim 7 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 15: Step 1 Statutory Category: Claim 15 is directed to a system, which falls within one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 15 recites, in part, “formulate a decision process problem for the RL model”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Further, the claim recites: “define a logarithmic loss function for the RL model and an initiation point for the RL model according to an optimized spectral norm of the unstable system”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical concept, see MPEP §2106.04(a)(2)(I). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular, the claim recites “a computer system”, “one or more computer processors”, “one or more computer readable storage devices”, and “stored program instructions on the one or more computer readable storage devices”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites “a trained reinforcement learning model for an unstable system”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “train the RL model using system data to define a policy according to the logarithmic loss function starting from the initiation point”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Finally, the claim recites: “provide the trained RL model”. This limitation is an additional element that amounts to a post-solution step and as such is considered insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Step 2B Significantly More: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: that the method is “a computer system”, “one or more computer processors”, “one or more computer readable storage devices”, “stored program instructions on the one or more computer readable storage devices”, and “train the RL model using system data to define a policy according to the logarithmic loss function starting from the initiation point” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional element “a trained reinforcement learning model for an unstable system” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Finally, the additional element “provide the trained RL model” amounts to insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 16, the rejection of claim 15 is incorporated, and further, claim 16 is substantially similar to claim 2 and claim 9 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 17, the rejection of claim 15 is incorporated, and further, claim 17 is substantially similar to claim 3 and claim 10 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 18, the rejection of claim 17 is incorporated, and further, claim 18 is substantially similar to claim 4 and claim 11 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 19, the rejection of claim 18 is incorporated, and further, claim 19 is substantially similar to claim 5 and claim 12 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 20, the rejection of claim 15 is incorporated, and further, claim 20 is substantially similar to claim 6 and claim 13 respectively, and is rejected in the same manner and reasoning applying. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 6, 8-10, 13, 15-17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kim in view of Bjorck et al., Towards Deeper Deep Reinforcement Learning with Spectral Normalization, 01/03/2022, https://arxiv.org/pdf/2106.01151v4, hereinafter referred to as "Bjorck". Regarding claim 1, Kim teaches A … method for training a reinforcement learning (RL) model for an unstable system (Kim, Page 1, Abstract, Lines 4-6, “we propose goal-aware cross-entropy (GACE) loss, that can be utilized in a self-supervised way using auto-labeled goal states alongside reinforcement learning”), the method comprising: formulating, …, a decision process problem for the RL model (Kim, Page 3, Section 3.1, Lines 1-4, “Reinforcement learning (RL) from Sutton and Barto [41] aims to maximize cumulative rewards by trial-and-error in a Markov Decision Process (MDP). An MDP is defined by a tuple (S,A,R,P,γ), where S is the set of states, A is the set of actions, R : S × A → R is the reward function, P : S ×A×S →R is the transition probability distribution, and γ ∈ (0,1] is the discount factor”); defining, …, a logarithmic loss function for the RL model … (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”); training, …, the RL model using system data to define a policy according to the logarithmic loss function… (Kim, Page 4, Section 3.3, Lines 24-26, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9. We focus on improving the policy for performing the main task and assign weight η to L G A C E for performing goal-aware representation learning for the feature extractor”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and providing, …, the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Kim also teaches that the method is computer-implemented and performing the steps of the method by the one or more computer processors (Kim, Page 5, Section 4, Paragraph 1, Lines 3-4, “We develop and conduct experiments on (1) visual navigation tasks based on ViZDoom [23, 18], and (2) robot arm manipulation tasks based on MuJoCo [42]”; Kim, Page 9, Table 3 and Figure 4; A person of ordinary skill in the art would recognize that “ViZDoom” and “MuJoCo” as well as the results displayed in Table 3 and Figure 4 would require the use of a computer, which also provides evidence for a computer processor). Kim does not explicitly teach defining an initiation point for the RL model according to an optimized spectral norm of the unstable system nor training the system starting from the initiation point…. Bjorck teaches defining an initiation point for the RL model according to an optimized spectral norm of the unstable system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initiation point for the RL model”) and training the system starting from the initiation point… (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the reinforcement learning method of Kim to include defining an initiation point according to an optimized spectral norm as taught by Bjorck. The motivation to do so would have been that spectral normalization enables stable training with large modern architectures, which results in performance improvements of the model (Bjorck, Page 1, Abstract, Lines 13-17, “We demonstrate that spectral normalization (SN) can mitigate this issue and enable stable training with large modern architectures. After smoothing with SN, larger models yield significant performance improvements—suggesting that more “easy” gains may be had by focusing on model architectures in addition to algorithmic innovations”). Regarding claim 2, the rejection of claim 1 is incorporated, and further, Kim teaches defining, by the one or more computer processors, the logarithmic loss function for the RL model (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”); training, by the one or more computer processors, the model according to the logarithmic loss function (Kim, Page 4, Section 3.3, Lines 24-25, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and providing, by the one or more computer processors, the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Regarding claim 3, the rejection of claim 1 is incorporated, and further, the proposed combination teaches defining, …, the initiation point for the RL model according to the optimized spectral norm (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”); and training, …, the RL model from the initiation point (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”). Further, Kim teaches performing the steps of the method by the one or more computer processors (Kim, Page 5, Section 4, Paragraph 1, Lines 3-4, “We develop and conduct experiments on (1) visual navigation tasks based on ViZDoom [23, 18], and (2) robot arm manipulation tasks based on MuJoCo [42]”; Kim, Page 9, Table 3 and Figure 4; A person of ordinary skill in the art would recognize that “ViZDoom” and “MuJoCo” as well as the results displayed in Table 3 and Figure 4 would require the use of a computer, which also provides evidence for a computer processor). Regarding claim 6, the rejection of claim 1 is incorporated, and further, the proposed combination teaches formulating, …, a decision process problem for the RL model (Kim, Page 3, Section 3.1, Lines 1-4, “Reinforcement learning (RL) from Sutton and Barto [41] aims to maximize cumulative rewards by trial-and-error in a Markov Decision Process (MDP). An MDP is defined by a tuple (S,A,R,P,γ), where S is the set of states, A is the set of actions, R : S × A → R is the reward function, P : S ×A×S →R is the transition probability distribution, and γ ∈ (0,1] is the discount factor”); defining, …, the logarithmic loss function for the RL model (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”) and the initiation point for the RL model according to an optimized spectral norm of the system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initiation point for the RL model”); training, …, the model from the initiation point (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”) and according to the logarithmic loss function (Kim, Page 4, Section 3.3, Lines 24-25, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and providing, …, the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Kim also teaches performing the steps of the method by the one or more computer processors (Kim, Page 5, Section 4, Paragraph 1, Lines 3-4, “We develop and conduct experiments on (1) visual navigation tasks based on ViZDoom [23, 18], and (2) robot arm manipulation tasks based on MuJoCo [42]”; Kim, Page 9, Table 3 and Figure 4; A person of ordinary skill in the art would recognize that “ViZDoom” and “MuJoCo” as well as the results displayed in Table 3 and Figure 4 would require the use of a computer, which also provides evidence for a computer processor). Regarding claim 8, Kim teaches A computer program product for providing a trained reinforcement learning models for an unstable system, the computer program product comprising one or more computer readable storage media and collectively stored program instructions on the one or more computer readable storage media, the stored program instructions which, when executed (Kim, Page 1, Abstract, Lines 4-6, “we propose goal-aware cross-entropy (GACE) loss, that can be utilized in a self-supervised way using auto-labeled goal states alongside reinforcement learning”; Kim, Page 5, Section 4, Paragraph 1, Lines 3-4, “We develop and conduct experiments on (1) visual navigation tasks based on ViZDoom [23, 18], and (2) robot arm manipulation tasks based on MuJoCo [42]”; Kim, Page 9, Table 3 and Figure 4; A person of ordinary skill in the art would recognize that “ViZDoom” and “MuJoCo” as well as the results displayed in Table 3 and Figure 4 would require the use of a computer, which also provides evidence for a computer program product, computer readable storage media, and instructions), cause one or more computer systems to: formulate a decision process problem for the RL model (Kim, Page 3, Section 3.1, Lines 1-4, “Reinforcement learning (RL) from Sutton and Barto [41] aims to maximize cumulative rewards by trial-and-error in a Markov Decision Process (MDP). An MDP is defined by a tuple (S,A,R,P,γ), where S is the set of states, A is the set of actions, R : S × A → R is the reward function, P : S ×A×S →R is the transition probability distribution, and γ ∈ (0,1] is the discount factor”); define a logarithmic loss function for the RL model … (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”); train the RL model, using system data to define a policy according to the logarithmic loss function… (Kim, Page 4, Section 3.3, Lines 24-26, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9. We focus on improving the policy for performing the main task and assign weight η to L G A C E for performing goal-aware representation learning for the feature extractor”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and provide the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Kim does not explicitly teach defining… an initiation point for the RL model according to an optimized spectral norm of the unstable system nor training the system starting from the initiation point…. Bjorck teaches defining… an initiation point for the RL model according to an optimized spectral norm of the unstable system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”) and training the system starting from the initiation point… (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the reinforcement learning method of Kim to include defining an initiation point according to an optimized spectral norm as taught by Bjorck. The motivation to do so would have been that spectral normalization enables stable training with large modern architectures, which results in performance improvements of the model (Bjorck, Page 1, Abstract, Lines 13-17, “We demonstrate that spectral normalization (SN) can mitigate this issue and enable stable training with large modern architectures. After smoothing with SN, larger models yield significant performance improvements—suggesting that more “easy” gains may be had by focusing on model architectures in addition to algorithmic innovations”). Regarding claim 9, the rejection of claim 8 is incorporated, and further, Kim teaches define the logarithmic loss function for the RL model (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”); train the system according to the logarithmic loss function (Kim, Page 4, Section 3.3, Lines 24-25, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and provide the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Regarding claim 10, the rejection of claim 8 is incorporated, and further, the proposed combination teaches define the initiation point for the RL model according to the optimized spectral norm of the system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”); and train the RL model from the initiation point (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”). Regarding claim 13, the rejection of claim 8 is incorporated, and further, the proposed combination teaches formulate a decision process problem for the RL model (Kim, Page 3, Section 3.1, Lines 1-4, “Reinforcement learning (RL) from Sutton and Barto [41] aims to maximize cumulative rewards by trial-and-error in a Markov Decision Process (MDP). An MDP is defined by a tuple (S,A,R,P,γ), where S is the set of states, A is the set of actions, R : S × A → R is the reward function, P : S ×A×S →R is the transition probability distribution, and γ ∈ (0,1] is the discount factor”); define the logarithmic loss function for the RL model (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”) and the initiation point for the RL model according to an optimized spectral norm of the system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”); train the RL model from the initiation point (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”) and according to the logarithmic loss function (Kim, Page 4, Section 3.3, Lines 24-25, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and provide the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Regarding claim 15, Kim teaches A computer system for providing a trained reinforcement learning model for an unstable system, the computer system comprising: one or more computer processors; one or more computer readable storage devices; and stored program instructions on the one or more computer readable storage devices for execution by the one or more computer processors, the stored program instructions which, when executed (Kim, Page 1, Abstract, Lines 4-6, “we propose goal-aware cross-entropy (GACE) loss, that can be utilized in a self-supervised way using auto-labeled goal states alongside reinforcement learning”; Kim, Page 5, Section 4, Paragraph 1, Lines 3-4, “We develop and conduct experiments on (1) visual navigation tasks based on ViZDoom [23, 18], and (2) robot arm manipulation tasks based on MuJoCo [42]”; Kim, Page 9, Table 3 and Figure 4; A person of ordinary skill in the art would recognize that “ViZDoom” and “MuJoCo” as well as the results displayed in Table 3 and Figure 4 would require the use of a computer, which also provides evidence for a computer processor, computer readable storage devices, and instructions), cause the one or more computer processors to: formulate a decision process problem for the RL model (Kim, Page 3, Section 3.1, Lines 1-4, “Reinforcement learning (RL) from Sutton and Barto [41] aims to maximize cumulative rewards by trial-and-error in a Markov Decision Process (MDP). An MDP is defined by a tuple (S,A,R,P,γ), where S is the set of states, A is the set of actions, R : S × A → R is the reward function, P : S ×A×S →R is the transition probability distribution, and γ ∈ (0,1] is the discount factor”); define a logarithmic loss function for the RL model … (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”); train the RL model using system data to define a policy according to the logarithmic loss function… (Kim, Page 4, Section 3.3, Lines 24-26, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9. We focus on improving the policy for performing the main task and assign weight η to L G A C E for performing goal-aware representation learning for the feature extractor”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and provide the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Kim does not explicitly teach defining… an initiation point for the RL model according to an optimized spectral norm of the unstable system nor training the model starting from the initiation point…. Bjorck teaches defining… an initiation point for the RL model according to an optimized spectral norm of the unstable system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”) and training the system starting from the initiation point… (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the reinforcement learning method of Kim to include defining an initiation point according to an optimized spectral norm as taught by Bjorck. The motivation to do so would have been that spectral normalization enables stable training with large modern architectures, which results in performance improvements of the model (Bjorck, Page 1, Abstract, Lines 13-17, “We demonstrate that spectral normalization (SN) can mitigate this issue and enable stable training with large modern architectures. After smoothing with SN, larger models yield significant performance improvements—suggesting that more “easy” gains may be had by focusing on model architectures in addition to algorithmic innovations”). Regarding claim 16, the rejection of claim 15 is incorporated, and further, the proposed combination teaches define the logarithmic loss function for the RL model (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”); train the system according to the logarithmic loss function (Kim, Page 4, Section 3.3, Lines 24-25, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and provide the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Regarding claim 17, the rejection of claim 15 is incorporated, and further, the proposed combination teaches define the initiation point for the RL model according to the optimized spectral norm (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”); and train the RL model from the initiation point (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”). Regarding claim 20, the rejection of claim 15 is incorporated, and further, Kim teaches formulate a decision process problem for the RL model (Kim, Page 3, Section 3.1, Lines 1-4, “Reinforcement learning (RL) from Sutton and Barto [41] aims to maximize cumulative rewards by trial-and-error in a Markov Decision Process (MDP). An MDP is defined by a tuple (S,A,R,P,γ), where S is the set of states, A is the set of actions, R : S × A → R is the reward function, P : S ×A×S →R is the transition probability distribution, and γ ∈ (0,1] is the discount factor”); define the logarithmic loss function for the RL model (Kim, Page 4, Section 3.3, Lines 5-9, and Equations 3-5, “In our visual navigation experiments, we use asynchronous advantage actor-critic (A3C) [28] as the main algorithm, where the loss L R L   is defined as the following L p = ∇ l o g π a t s t , I R t - V s t , I + β ∇ H π a t s t , I   3   L v = R t - V s t , I 2   4   L R L ≔ L A 3 C = L p + 0.5 ∙ L v   5   where L p and L v respectively denote policy and value loss, R t denotes the sum of decayed rewards from time steps t to T, and H and β denote the entropy term and its coefficient respectively”; Kim, Page 4, Section 3.3, Lines 15-17, “we propose Goal-Aware Cross-Entropy (GACE) loss as our contribution, which trains the goal-discriminator that facilitates semantic understanding of goals alongside the policy in Figure 1a”; Kim, Page 4, Section 3.3, Equation 8, “ L G A C E = - ∑ i = 0 M - 1 o n e h o t z i ∙ log ⁡ g g o a l , i 8 ”) and the initiation point for the RL model according to an optimized spectral norm of the system (Bjorck, Page 6, Lines 1-5, “Equation (5) suggest that the critic could be made smooth if the spectral norms of all layers are bounded. Fortunately, there is a method from the GAN literature which achieves this: spectral normalization [47]. Spectral normalization divides the weight W for each layer by its largest singular value σ m a x which ensures that all layers have operator norm 1”; Bjorck, Page 6, Lines 10-11=2, “By repeating this procedure for all layers, we ensure that the spectral norms of all layers are no larger than one. If that is the case, eq. (5) suggests that the critic should be stable in the forward pass. This would then bound the gradients being propagated into the actor as per eq. (3)”; The state of the model after spectral normalization, with bounded gradients, is considered to be the “initial point for the RL model”); train the model from the initiation point (Bjorck, Page 6, Section 5.1, Lines 3-9, “Specifically, for both the actor and the critic, we apply spectral normalization to each linear layer except the first and last. Otherwise, the setup follows Section 3.1. As before, when learning crashes, we simply use the performance recorded before crashes for future time steps. Learning curves for individual environments are given in Figure 4, again over 10 seeds. We see that after smoothing with spectral normalization, performance is relatively stable across tasks, even when using a deep network with normalization and skip connections. On the other hand, without smoothing, learning is slow and sometimes fails”) according to the logarithmic loss function (Kim, Page 4, Section 3.3, Lines 24-25, “We complete the training procedure by optimizing the overall loss L t o t a l as the weighted sum of the two losses in Eq. 9”; Kim, Page 5, Equation 9, “ L t o t a l = L R L + η L G A C E ”); and provide the trained RL model (Kim, Page 10, Lines 1-6, “To ascertain that an agent trained with GACE and GACE&GDAN indeed becomes goal-aware, we use saliency maps [15] to visualize the operation of three agents within the V2 un seen task, as shown in Figure 5”; Kim, Page 10, Lines 11-14, “The three agents are trained with A3C, GACE, and GACE&GDAN, respectively, for 4M updates”; see also Kim, Page 10, Figure 5; In order for the model to be used in operation and develop saliency maps, it must have been provided). Claims 4-5, 7, 11-12, 14, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Kim in view of Bjorck in further view of HAMBLY et al., "Policy gradient methods for the noisy linear quadratic regulator over a finite horizon". SIAM Journal on Control and Optimization, 59(5): 3359–3391, 2021, June 25, 2021, 49 pps, hereinafter referred to as “Hambly”. Regarding claim 4, the rejection for claim 3 is incorporated. The proposed combination does not explicitly teach wherein defining the initiation point comprises regulating a system spectral radius. Hambly teaches wherein defining the initiation point comprises regulating a system spectral radius (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the method of defining an initiation point as taught by the proposed combination to include regulating a system spectral radius as taught by Hambly. The motivation to do so would have been to guarantee the stability of the system (Hambly, Page 11, Remark 3.11, Lines 4-5). Regarding claim 5, the rejection of claim 4 is incorporated, and further, the proposed combination teaches wherein defining the initiation point comprises defining an initiation point wherein a magnitude of an absolute value of the system spectral radius is less than 1 (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). Regarding claim 7, the rejection of claim 6 is incorporated, and further, the proposed combination teaches performing the step of the method by the one or more computer processors (Kim, Page 5, Section 4, Paragraph 1, Lines 3-4, “We develop and conduct experiments on (1) visual navigation tasks based on ViZDoom [23, 18], and (2) robot arm manipulation tasks based on MuJoCo [42]”; Kim, Page 9, Table 3 and Figure 4; A person of ordinary skill in the art would recognize that “ViZDoom” and “MuJoCo” as well as the results displayed in Table 3 and Figure 4 would require the use of a computer, which also provides evidence for a computer processor). The proposed combination does not explicitly teach defining, …, the initiation point according to a system spectral magnitude of less than 1. Hambly teaches defining, …, the initiation point according to a system spectral magnitude of less than 1 (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the method of defining an initiation point as taught by the proposed combination to include regulating a system spectral radius as taught by Hambly. The motivation to do so would have been to guarantee the stability of the system (Hambly, Page 11, Remark 3.11, Lines 4-5). Regarding claim 11, the rejection of claim 10 is incorporated. The proposed combination does not explicitly teach wherein defining the initiation point comprises regulating a system spectral radius. Hambly teaches wherein defining the initiation point comprises regulating a system spectral radius (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the method of defining an initiation point as taught by the proposed combination to include regulating a system spectral radius as taught by Hambly. The motivation to do so would have been to guarantee the stability of the system (Hambly, Page 11, Remark 3.11, Lines 4-5). Regarding claim 12, the rejection of claim 11 is incorporated, and further, the proposed combination teaches wherein defining the initiation point comprises defining an initiation point wherein a magnitude of an absolute value of the system spectral radius is less than 1 (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). Regarding claim 14, the rejection of claim 13 is incorporated. The proposed combination does not explicitly teach defining the initiation point according to a system spectral magnitude of less than 1. Hambly teaches defining the initiation point according to a system spectral magnitude of less than 1 (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the method of defining an initiation point as taught by the proposed combination to include regulating a system spectral radius as taught by Hambly. The motivation to do so would have been to guarantee the stability of the system (Hambly, Page 11, Remark 3.11, Lines 4-5). Regarding claim 18, the rejection of claim 17 is incorporated. The proposed combination does not explicitly teach wherein defining the initiation point comprises regulating a system spectral radius. Hambly teaches wherein defining the initiation point comprises regulating a system spectral radius (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the method of defining an initiation point as taught by the proposed combination to include regulating a system spectral radius as taught by Hambly. The motivation to do so would have been to guarantee the stability of the system (Hambly, Page 11, Remark 3.11, Lines 4-5). Regarding claim 19, the rejection of claim 18 is incorporated, and further, the proposed combination teaches wherein defining the initiation point comprises defining an initiation point wherein a magnitude of an absolute value of the system spectral radius is less than 1 (Hambly, Page 11, Remark 3.11, Lines 4-5, “Note that for the infinite horizon problem, the spectral radius of A - BK needs to be smaller than 1 to guarantee the stability of the system (see [23])”). Response to Arguments Applicant’s arguments regarding the 35 U.S.C. 101 rejections of the claims have been fully considered but are unpersuasive. Argument 1: Applicant first argues, on page 13, in the final paragraph of the response, that the “formulating” and “defining” steps of claim 1 do not set forth or describe any mathematical relationships, calculations, formulas, or equations using words or mathematical symbols”. Examiner’s Response: Examiner Respectfully disagrees. The “formulating” step of claim 1 recites a “decision process problem”, which in light of applicant’s specification paragraph 0019 includes an “input-to-output stability expression having an output such as y=h(x(t)), where, for example y may be a quadratic function”. Thus, while the decision process problem may not describe mathematical concepts using mathematical symbols, a “decision process problem” sets forth a mathematical concept using words. Further, the “defining” step of claim 1 includes defining an initiation point, which according to applicant’s specification paragraph 0023, recites mathematical concepts; and defining a logarithmic loss function, which according to applicant’s specification paragraph 0021 is a mathematical equation, which is a mathematical concept. Argument 2: Applicant next argues, on page 14, paragraph 1 of the response, that claim 1 provides a technical solution to the problem of providing a trained RL model for unstable systems despite instabilities of such systems by defining an RL model loss function including a logarithmic term and defining an initiation point for the system. Examiner’s Response: Examiner Respectfully disagrees. The “defining” limitation present in claim 1 has been identified as an abstract idea. An inventive concept cannot be furnished by the unpatentable abstract idea itself, see MPEP 2106.05(I). Applicant's arguments regarding the remainder of the claims rely upon the arguments asserted with respect to the independent claims, and are thus unpersuasive. Applicant’s amendments to the claims with respect to the 35 U.S.C. 112(b) rejections to the claims have been fully considered and overcome the 35 U.S.C. 112(b) rejections set forth in the nonfinal office action dated 0312/2026. However, the amendments to the claims raise new 35 U.S.C. 112(b) issues which are reflected in the updated 35 U.S.C. 112(b) rejections above. Applicant’s arguments regarding the prior art rejections of the claims have been fully considered but are unpersuasive. Argument 1: Applicant first argues, on page 15, final paragraph and page 17, final paragraph of the response, that Kim fails to teach “defining, by the one or more computer processors, a logarithmic loss function for the RL model and defining an initiation point for the RL model according to an optimized spectral norm of the system; and training, by the one or more computer processors, the model using system data to define a policy according to the logarithmic loss function starting from the initiation point”. Examiner’s Response: Examiner respectfully disagrees. With regard to Kim failing to teach “defining, by the one or more computer processors, a logarithmic loss function for the RL model and defining an initiation point for the RL model according to an optimized spectral norm of the system; and training, by the one or more computer processors, the model using system data to define a policy according to the logarithmic loss function starting from the initiation point”, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Argument 2: Applicant next argues, on page 16, paragraph 1 and page 18, paragraph 1, of the response, that Bjorck teaches forcing smoothness upon a critic function by bounding the spectral norms of each layer and not the selection of a point according to a spectral norm of the system. Applicant asserts that after such smoothing the system is presumably stable and outside the scope of the claim. Examiner’s Response (Under examiner’s best interpretation in light of the 35 U.S.C. 112(b) rejections on the record): Examiner respectfully disagrees. Applicant asserts “no initiation points of an unstable system would remain for selection”. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., selecting an initiation) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). While claim 1 requires an initiation point is defined “according to an optimized spectral norm of the unstable system”, the claims do not require that the system remain “unstable”; nor do the claims require any “selection” of an initiation point, but rather “defining” an initiation point. Further, under the broadest reasonable interpretation of “initiation point”, Bjorck teaches the limitation as the method performs spectral normalization and the resulting state of the RL model is considered the “initiation point”, the claims do not place any limitations on the interpretation of “initiation point”. Applicant's arguments regarding the remainder of the claims rely upon the arguments asserted with respect to the independent claims, and are thus unpersuasive. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOLLY CLARKE SIPPEL whose telephone number is (571)272-3270. The examiner can normally be reached Monday - Friday, 7:30 a.m. - 4:30 p.m. ET.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571)272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.C.S./Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Show 4 earlier events
Jun 04, 2026
Applicant Interview (Telephonic)
Jun 04, 2026
Examiner Interview Summary
Jun 10, 2026
Response after Non-Final Action
Jun 10, 2026
Response Filed
Jun 24, 2026
Response Filed
Aug 18, 2026
Final Rejection mailed — §101, §103, §112
Sep 23, 2026
Interview Requested
Sep 25, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748955
REINFORCEMENT LEARNING WITH ADAPTIVE RETURN COMPUTATION SCHEMES
4y 1m to grant Granted Sep 29, 2026
Patent 12670387
SYSTEM, METHOD, AND COMPUTER-READABLE MEDIA FOR LEAKAGE CORRECTION IN GRAPH NEURAL NETWORK BASED RECOMMENDER SYSTEMS
4y 1m to grant Granted Jun 30, 2026
Patent 12664398
SYSTEM, METHOD AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 9m to grant Granted Jun 23, 2026
Patent 12657427
Systems, Methods, and Computer Program Products for Determining Uncertainty from a Deep Learning Classification Model
4y 1m to grant Granted Jun 16, 2026
Patent 12632779
HYPERPARAMETER SELECTION USING BUDGET-AWARE BAYESIAN OPTIMIZATION
4y 5m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
50%
Grant Probability
76%
With Interview (+25.7%)
3y 10m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 28 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month