Prosecution Insights
Last updated: August 07, 2026
Application No. 18/475,654

METHOD AND DEVICE WITH REINFORCEMENT LEARNING TRANSFERAL

Non-Final OA §101§103
Filed
Sep 27, 2023
Priority
Feb 02, 2023 — RE 10-2023-0014465
Examiner
MAHARAJ, DEVIKA S
Art Unit
Tech Center
Assignee
Seoul National University R&DB Foundation
OA Round
1 (Non-Final)
56%
Grant Probability
Moderate
1-2
OA Rounds
1y 8m
Est. Remaining
65%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
48 granted / 86 resolved
-4.2% vs TC avg
Moderate +9% lift
Without
With
+9.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
26 currently pending
Career history
111
Total Applications
across all art units

Statute-Specific Performance

§101
30.0%
-10.0% vs TC avg
§103
46.2%
+6.2% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
10.5%
-29.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 86 resolved cases

Office Action

§101 §103
CTNF 18/475,654 CTNF 96779 DETAILED ACTION 1. This communication is in response to the Application No. 18/475,654 filed on September 27, 2023 in which Claims 1-20 are presented for examination. Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia 2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Information Disclosure Statement 3. The information disclosure statement submitted on 09/27/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification 07-29 AIA 4. The disclosure is objected to because of the following informalities: Par. [0026] of Applicant’s specification contains a typographical error, as it recites “The upper bound may be determined based on one of multiple linear combinations of the source task vectorsvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-functionvalue-function ” Appropriate correction is required. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 5. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 6. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a system type claim. Therefore, Claims 1-8 are directed to either a process, machine, manufacture, or composition of matter. 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. approximate an optimal value-function for a task vector […] using a state of an agent and source task vectors (mathematical process/mental process – approximating an optimal value-function for a task vector may be performed by mathematical process utilizing Equation 2 in Par. [0056] of Applicant’s specification, in which when a policy vector space is equal to a task vector space and a given policy vector z is equal to the task vector w, the optimal-value function may be approximated for the task vector w using state of an agent and source task vectors according to the equation 2. Alternatively, approximating an optimal value-function for a task vector may be performed manually by a user observing/analyzing the state of an agent and source task vectors and accordingly using judgement/evaluation to approximate an optimal value-function for a task vector (with the aid of pen and paper) based on said analysis) determine an upper and lower bound of the optimal value-function for the task vector (mathematical process/mental process – determining an upper and lower bound of the optimal value-function for the task vector may be performed by mathematical process utilizing Equations 4, 5, and/or 6 in Par. [0074-0083] of Applicant’s specification, in which the upper and lower bounds are determined. Alternatively, determining an upper and lower bound of the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and accordingly using judgement/evaluation to determine an upper and lower bound based on said analysis) correct the optimal value-function for the task vector based on the upper bound and the lower bound (mathematical process/mental process – correcting the optimal value-function for the task vector based on the upper bound and the lower bound may be performed by mathematical process utilizing the equations/mathematical operations described in Par. [0084-0087] of Applicant’s specification, in which the optimal value-function is corrected based on the bounds. Alternatively, correcting the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and the upper/lower bounds and accordingly using judgement/evaluation to correct the optimal value-function (with the aid of pen and paper) based on the upper/lower bounds) determine an optimal policy for the task vector using the corrected optimal value-function for the task vector (mathematical process/mental process – determining an optimal policy for the task vector using the corrected optimal value-function for the task vector may be performed by mathematical process utilizing the equations/mathematical operations described in Par. [0087-0089] of Applicant’s specification, in which the optimal policy is determined by calculating/solving the corrected optimal value-function. Alternatively, determining an optimal policy may be performed manually by a user observing/analyzing the corrected optimal value-function and accordingly using judgement/evaluation to determine an optimal policy using the corrected optimal value-function (by solving the function with the aid of pen and paper)) 2A Prong 2: This judicial exception is not integrated into a practical application. Additional elements: an electronic device comprising: one or more processors; and a memory electrically connected with the one or more processors and storing instructions configured to cause the one or more processors to: […] (i.e., as a generic device comprising generic processor and memory such that it amounts to no more than mere instructions to apply the exception using generic computer components) using a value approximator trained to output a minimum value-function using a state of an agent and source task vectors (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying an already trained machine learning model using generic data without significantly more) 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional elements: an electronic device comprising: one or more processors; and a memory electrically connected with the one or more processors and storing instructions configured to cause the one or more processors to: […] (mere instructions to apply the exception using generic computer components cannot provide an inventive concept) using a value approximator trained to output a minimum value-function using a state of an agent and source task vectors (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying an already trained machine learning model using generic data without significantly more. This cannot provide an inventive concept) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-8. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. Step 2A Prong 2 & Step 2B: wherein the task vector is represented as a linear combination of the source task vectors (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the task vector is represented as a linear combination of source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 3 depends on. determine a number of combinations of the linear combination of the source task vectors to represent the task vector according to a threshold (mental process – determining a number of combinations of the linear combination of the source task vectors to represent the task vector according to a threshold may be performed manually by a user observing/analyzing the linear combination of the source task vectors and threshold and accordingly using judgement/evaluation to determine a number of combinations of the linear combination (with the aid of pen and paper) to represent the task vector according to the threshold) determine the upper bound according to one of the determined linear combinations of the task vectors (mathematical process/mental process – determining the upper bound may be performed by mathematical process per Applicant’s specification Par. [0083] which specifies how the upper bound may be calculated using the linear combination of source task vectors, as described by the rejection of Claim 1 above. Alternatively, the determining the upper bound may be performed manually by a user observing/analyzing the determined linear combinations of the task vectors and accordingly using judgement/evaluation to determine the upper bound according to said analysis of the linear combinations) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. Step 2A Prong 2 & Step 2B: wherein the lower bound is determined based on an approximation of, and an approximation error of, the optimal value-function (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the lower bound is determined generically based on an approximation of, and an approximation error of, the optimal value-function does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 5 depends on. Step 2A Prong 2 & Step 2B: wherein the upper bound is determined based on an arbitrary linear combination of the source task vectors to represent the task vector (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the upper bound is determined based on an arbitrary linear combination of source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on. correct the optimal value-function in a range of less than or equal to the upper bound and greater than or equal to the lower bound (mathematical process/mental process - correcting the optimal value-function for the task vector based on the upper bound and the lower bound may be performed by mathematical process utilizing the equations/mathematical operations described in Par. [0084-0087] of Applicant’s specification, in which the optimal value-function is corrected based on the bounds. Alternatively, correcting the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and the upper/lower bounds and accordingly using judgement/evaluation to correct the optimal value-function (with the aid of pen and paper) based on the upper/lower bounds) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 7 depends on. approximate a minimum value for the task vector […] (mathematical process/mental process – other than reciting “minimum value approximator”, approximating a minimum value may be performed by mathematical process/calculation. Alternatively, approximating a minimum value for the task vector may be performed manually by a user observing/analyzing the task vector and accordingly using judgement/evaluation to approximate a minimum value for the task vector (with the aid of pen and paper)) determine the upper bound using the minimum value for the task vector (mathematical process/mental process – determining the upper bound using the minimum value for the task vector may be performed by mathematical process utilizing the equations/mathematical operations per Applicant’s specification Par. [0110-0111]. Alternatively, the upper bound may be determined manually by a user observing/analyzing the minimum value for the task vector and accordingly using judgement/evaluation to determine the upper bound based on said analysis) Step 2A Prong 2 & Step 2B: […] using a minimum value approximator trained to output a minimum value for the source task vectors (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner' s note: high level recitation of applying an already trained machine learning model without significantly more. This cannot provide an inventive concept) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 8 depends on. Step 2A Prong 2 & Step 2B: wherein the optimal policy comprises a neural network model (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the optimal policy comprises a neural network model does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. Independent Claim 9 recites substantially the same limitations as Claim 1, in the form of a method . The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale. For the reasons above, Claim 9 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 10-16. The additional limitations of the dependent claims are addressed below. Claim 10 recites substantially the same limitations as Claim 2 , in the form of a method . The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale. Regarding Claim 11: Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 11 depends on. determining a number of combinations of the linear combination of the source task vectors to represent the task vector by a predetermined threshold (mental process – determining a number of combinations of the linear combination of the source task vectors to represent the task vector according to a threshold may be performed manually by a user observing/analyzing the linear combination of the source task vectors and threshold and accordingly using judgement/evaluation to determine a number of combinations of the linear combination (with the aid of pen and paper) to represent the task vector according to the threshold) determining the upper bound according to a combination of the linear combination of the source task vectors (mathematical process/mental process – determining the upper bound may be performed by mathematical process per Applicant’s specification Par. [0083] which specifies how the upper bound may be calculated using the linear combination of source task vectors, as described by the rejection of Claim 1 above. Alternatively, the determining the upper bound may be performed manually by a user observing/analyzing the determined linear combinations of the task vectors and accordingly using judgement/evaluation to determine the upper bound according to said analysis of the linear combinations) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. Regarding Claim 12: Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 12 depends on. Step 2A Prong 2 & Step 2B: wherein the lower bound is determined based on an approximation of, and an approximation error of, the optimal value-function for the task vector of policies based on the source task vectors (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the lower bound is determined generically based on an approximation of, and an approximation error of, the optimal value-function does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. Regarding Claim 13: Step 2A Prong 1: See the rejection of Claim 12 above, which Claim 13 depends on. Step 2A Prong 2 & Step 2B: wherein the policies comprise respective neural networks trained with respect to the source task vectors (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the policies comprise respective neural networks trained with respect to the source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. Regarding Claim 14: Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 14 depends on. Step 2A Prong 2 & Step 2B: wherein the upper bound is determined based on a combination of the linear combination of the source task vectors to represent the task vector (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the upper bound is determined based on a combination of the linear combination of source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. Regarding Claim 15: Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 15 depends on. correcting of the optimal value-function for the task vector comprises correcting the optimal value-function for the task vector in a range of less than or equal to the upper bound and greater than or equal to the lower bound (mathematical process/mental process - correcting the optimal value-function for the task vector based on the upper bound and the lower bound may be performed by mathematical process utilizing the equations/mathematical operations described in Par. [0084-0087] of Applicant’s specification, in which the optimal value-function is corrected based on the bounds. Alternatively, correcting the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and the upper/lower bounds and accordingly using judgement/evaluation to correct the optimal value-function (with the aid of pen and paper) based on the upper/lower bounds) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. Regarding Claim 16: Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 16 depends on. approximating a minimum value for the task vector […] (mathematical process/mental process – other than reciting “minimum value approximator”, approximating a minimum value may be performed by mathematical process/calculation. Alternatively, approximating a minimum value for the task vector may be performed manually by a user observing/analyzing the task vector and accordingly using judgement/evaluation to approximate a minimum value for the task vector (with the aid of pen and paper)) determining the upper bound using the minimum value for the task vector (mathematical process/mental process – determining the upper bound using the minimum value for the task vector may be performed by mathematical process utilizing the equations/mathematical operations per Applicant’s specification Par. [0110-0111]. Alternatively, the upper bound may be determined manually by a user observing/analyzing the minimum value for the task vector and accordingly using judgement/evaluation to determine the upper bound based on said analysis) Step 2A Prong 2 & Step 2B: […] using a minimum value approximator trained to output a minimum value for the source task vectors (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner' s note: high level recitation of applying an already trained machine learning model without significantly more. This cannot provide an inventive concept) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 9. Regarding Claim 17: Step 1: Claim 17 is a method type claim. Therefore, Claims 17-20 are directed to either a process, machine, manufacture, or composition of matter. 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. approximating a task vector […] to output source task vectors using source task information of source tasks (mathematical process/mental process – other than reciting “task vector approximator”, approximating a task vector may be performed by mathematical process/calculation according to Par. [0117] of Applicant’s specification. Alternatively, approximating a task vector may be performed manually by a user observing/analyzing the source task information and accordingly using judgement/evaluation to generate/approximate a task vector (with the aid of pen and paper) based on said analysis) approximating a feature of the task vector […] to output a feature of the source task vector based on a state of an agent (mathematical process/mental process – other than reciting a “feature approximator”, approximating a feature of the task vector may be performed by mathematical process/calculation. Alternatively, approximating a feature of the task vector may be performed manually by a user observing/analyzing the state of an agent and source task vector and accordingly using judgement/evaluation to generate/approximate features of the task vector (with the aid of pen and paper) based on said analysis) approximating an optimal value-function for the task vector using the task vector and the feature of the task vector (mathematical process/mental process – approximating an optimal value-function for a task vector may be performed by mathematical process utilizing Equation 2 in Par. [0056] of Applicant’s specification or the equations/mathematical operations described in Par. [0117) of Applicant’s specification. Alternatively, approximating an optimal value-function for the task vector may be performed manually by a user observing/analyzing the task vector and feature of the task vector and accordingly using judgement/evaluation to approximate an optimal value-function for the task vector (with the aid of pen and paper) based on said analysis) determining an upper bound of the optimal value-function (mathematical process/mental process – determining an upper and lower bound of the optimal value-function for the task vector may be performed by mathematical process utilizing Equations 4, 5, and/or 6 in Par. [0074-0083] of Applicant’s specification, in which the upper and lower bounds are determined. Alternatively, determining an upper and lower bound of the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and accordingly using judgement/evaluation to determine an upper and lower bound based on said analysis) determining a lower bound of the optimal value-function (mathematical process/mental process – determining an upper and lower bound of the optimal value-function for the task vector may be performed by mathematical process utilizing Equations 4, 5, and/or 6 in Par. [0074-0083] of Applicant’s specification, in which the upper and lower bounds are determined. Alternatively, determining an upper and lower bound of the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and accordingly using judgement/evaluation to determine an upper and lower bound based on said analysis) correcting the optimal value-function based on the upper bound and the lower bound (mathematical process/mental process – correcting the optimal value-function for the task vector based on the upper bound and the lower bound may be performed by mathematical process utilizing the equations/mathematical operations described in Par. [0084-0087] of Applicant’s specification, in which the optimal value-function is corrected based on the bounds. Alternatively, correcting the optimal value-function may be performed manually by a user observing/analyzing the optimal value-function and the upper/lower bounds and accordingly using judgement/evaluation to correct the optimal value-function (with the aid of pen and paper) based on the upper/lower bounds) determining an optimal policy for the task vector using the corrected optimal value-function (mathematical process/mental process – determining an optimal policy for the task vector using the corrected optimal value-function for the task vector may be performed by mathematical process utilizing the equations/mathematical operations described in Par. [0087-0089] of Applicant’s specification, in which the optimal policy is determined by calculating/solving the corrected optimal value-function. Alternatively, determining an optimal policy may be performed manually by a user observing/analyzing the corrected optimal value-function and accordingly using judgement/evaluation to determine an optimal policy using the corrected optimal value-function (by solving the function with the aid of pen and paper)) 2A Prong 2: This judicial exception is not integrated into a practical application. Additional elements: method of transferring reinforcement learning […] (recited at a high-level of generality (i.e., as a generic method for transferring reinforcement learning without significantly more) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] using a task vector approximator trained to output source task vectors using source task information of source tasks (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying an already trained machine learning model with generic data without significantly more) […] using a feature approximator trained to output a feature of the source task vector based on a state of an agent (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying an already trained machine learning model with generic data without significantly more) 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional elements: method of transferring reinforcement learning […] (mere instructions to apply the exception using generic computer components cannot provide an inventive concept) […] using a task vector approximator trained to output source task vectors using source task information of source tasks (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying an already trained machine learning model with generic data without significantly more. This cannot provide an inventive concept) […] using a feature approximator trained to output a feature of the source task vector based on a state of an agent (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying an already trained machine learning model with generic data without significantly more. This cannot provide an inventive concept) For the reasons above, Claim 17 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 18-20. The additional limitations of the dependent claims are addressed below. Regarding Claim 18: Step 2A Prong 1: See the rejection of Claim 17 above, which Claim 18 depends on. Step 2A Prong 2 & Step 2B: wherein the task vector is represented as a linear combination of the source task vectors (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the task vector is represented as a linear combination of source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 17. Regarding Claim 19: Step 2A Prong 1: See the rejection of Claim 17 above, which Claim 19 depends on. Step 2A Prong 2 & Step 2B: wherein the lower bound is determined based on an approximation and an approximation error of the optimal value-function, wherein the approximation error is based on the source task vectors (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the lower bound is determined generically based on an approximation of, and an approximation error of, the optimal value-function wherein the approximation error is based on the source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 17. Regarding Claim 20: Step 2A Prong 1: See the rejection of Claim 17 above, which Claim 20 depends on. Step 2A Prong 2 & Step 2B: wherein the upper bound is determined based on one of multiple linear combination of the source task vectors (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the upper bound is determined based on one of multiple linear combinations of source task vectors does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 17. Claim Rejections - 35 USC § 103 07-20-aia AIA 7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA 8. Claim s 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Nemecek et al. (hereinafter Nemecek) ( “Policy Caches with Successor Features” ), in view of Zahavy et al. (hereinafter Zahavy) (US PG-PUB 20240104389) . Regarding Claim 1 , Nemecek teaches: approximate an optimal value-function for a task vector ( Nemecek, Pg. 2, Section 2. Background, “A value function V π(s) is the expected total discounted reward for starting in state s and performing actions according to π. The action-value function Qπ(s,a) gives the value of taking action a from state s and following π afterwards. The goal of a learning agent is to learn the optimal policy π ∗ which maximizes the expected future reward from each state. We can express the optimal value function V ∗ and action-value function Q ∗ via the Bellman equation as: [See equation on Pg. 2] […] The underlying idea of transfer learning for reinforcement learning (Taylor & Stone, 2009) is that through learning to perform well in one or more tasks, the learning process for a new task should be improved in some way, typically by reducing the amount of training in the new task required to reach a given level of performance.”, therefore, an optimal value-function is approximated for a task vector (See Nemecek Pg. 3 which states that the variable T comprises a set of tasks). This is analogous to the equation 2, for approximating an optimal value-function, as presented by Applicant’s specification Par. [0056]) using a value approximator trained ( While Nemecek discloses approximating an optimal value-function for a task vector […] to output a minimum value-function using a state of an agent and source task vectors, Nemecek does not explicitly disclose approximating an optimal value-function for a task vector using a value approximator trained to output a minimum value- function using a state of an agent and source task vectors . See introduction of Zahavy reference below for explicit disclosure of approximating an optimal value-function for a task vector using a value approximator trained to perform these operations of the claim language) to output a minimum value-function using a state of an agent and source task vectors ( Nemecek, Pg. 4, Section 4. Policy Cache Construction, “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound. For this LP, the variables are the coefficients αi, W is the set of reward weight vectors for the tasks in which the policies in the cache were learned, wnew is vector of reward weights for the new task, m is the number of policies in the cache, and s is the start state: [See equation on Pg. 4]”, thus, a minimum value-function may be outputted (based on minimizing the upper bound in accordance to the policy cache) using a state of an agent (start state, also described by the cited action-value function above) and source task vectors (reward weight vectors for tasks)) ; determine an upper and lower bound of the optimal value-function for the task vector ( Nemecek, Pg. 4, Section 6. Experimental Results, “As described in Algorithm 1, the agent starts with a policy cache containing policies for each of the base tasks for the given environment and when presented with each subsequent task, calculates the upper and lower bounds at the start state using the current policy cache and compares them. If the size of the gap exceeds threshold η, then a new policy is added to the cache.”, thus, an upper and lower bound of the optimal value-function for the task vector may be determined. This is similarly supported by Algorithms 1 & 2 on Pg. 4) ; correct the optimal value-function for the task vector based on the upper bound and the lower bound ( Nemecek, Pgs. 3-4, Section 4. Bounds on Policy Cache Performance, “We now consider a task R with reward function rR = i ∈ T αiri such that αi ≥ 0 and i ∈ T αi > 0, i.e., rR is a positive conical combination of the reward functions of the tasks in T. The proof of the theorem below (see appendix) provides an alternate derivation of the lower bound from Barreto et al., but does not improve upon it, while the upper bound is more general than those offered in Hunt et al. (2019) and applies to hard-max Q-learning. [See Theorem 1 on Pg. 3] Theorem 1 turns out to be quite powerful in practice because the optimality of a value function for a particular reward function significantly constrains the value function for other reward functions with similar weights”, therefore, the value function may be corrected/updated to account for the upper and lower bounds) ; and determine an optimal policy for the task vector using the corrected optimal value-function for the task vector ( Nemecek, Pg. 4, Section 5. Policy Cache Construction, “We start with a set of base tasks defined by a set of reward feature weights, and assume all subsequent tasks have rewards within the conical hull of the initial set. Algorithm 1 uses Theorem 1 to decide if the current policy cache is sufficient, given a performance threshold. If the gap between the upper and lower bounds is small enough for a given state, then we know that the value of the best policy in the cache must be within some small factor of the value for the optimal policy.”, thus, based on the corrected optimal value-function (which accounts for the upper/lower bounds as shown by Algorithm 1 & Theorem 1), an optimal policy for the task vector may be determined. This is similarly supported by Pg. 2 which states “Our first contribution is to address these limitations by presenting a new upper bound on the hard-max Q-function of the optimal policy for a new task when that task is a conical combination of previous tasks.”). Although Nemecek teaches transfer in reinforcement learning, algorithms, and neural networks that are seemingly run on a computing device (See Nemecek Pgs. 4 & 7), Nemecek does not explicitly disclose an electronic device comprising: one or more processors; and a memory electrically connected with the one or more processors and storing instructions configured to cause the one or more processors to: […] However, Zahavy teaches an electronic device comprising: one or more processors; and a memory electrically connected with the one or more processors and storing instructions configured to cause the one or more processors to: […] ( Zahavy, Par. [0118], “Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both.”, thus, an electronic device (computer) comprising one or more processors (central processing unit(s)) and a memory electrically connected with the one or more processors and storing instructions to be executed by the one or more processors is disclosed) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the device for determining an optimal policy for a task vector based on a corrected optimal value-function for the task vector, as disclosed by Nemecek to include wherein the device is an electronic device comprising: one or more processors; and a memory electrically connected with the one or more processors and storing instructions configured to cause the one or more processors to perform the operations of the claim language, as disclosed by Zahavy. One of ordinary skill in the art would have been motivated to make this modification to enable the use of an electronic device (computer) which may efficiently perform and execute instructions and facilitate streamlined data transfer for determining an optimal policy for a task vector ( Zahavy, Par. [0118], “The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks.”). While Nemecek discloses approximating an optimal value-function for a task vector […] to output a minimum value-function using a state of an agent and source task vectors (See rejection above), Nemecek does not explicitly disclose approximating an optimal value-function for a task vector using a value approximator trained to output a minimum value-function using a state of an agent and source task vectors However, Zahavy teaches approximating an optimal value-function for a task vector using a value approximator trained to output a minimum value-function using a state of an agent and source task vectors ( Zahavy, Par. [0029], “In implementations the one or more policy neural networks 110 comprise a value function neural network configured to process the observation 106 for the current time step, in accordance with current values of value function neural network parameters, to generate a current value estimate relating to the current state of the environment. The value function neural network may be a state or state-action value function neural network. That is, the current value estimate may be a state value estimate, i.e. an estimate of a value of the current state of the environment, or a state-action value estimate, i.e. an estimate of a value of each of a set of possible actions at the current time step.”, thus, a value approximator (which may comprise a neural network, as supported by Applicant’s specification Par. [0100]) is trained to output a minimum value-function (satisfying a minimum performance criterion as supported by Nemecek Par. [0038-0040]) using a state of an agent and source task vectors to approximate an optimal value-function for a task vector (See Nemecek Par. [0051-0054] which explicitly discloses these vectors)) . It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the electronic device for approximating an optimal value-function for a task vector, as disclosed by Nemecek in view of Zahavy to include wherein the optimal value-function is approximated using a value approximator trained to output a minimum value-function using a state of an agent and source task vectors, as disclosed by Zahavy. One of ordinary skill in the art would have been motivated to make this modification to enable the use of a value approximator (neural network) which may more efficiently and accurately approximate the value function to generate value estimates, allowing for improved generalization across a variety of parameters, datasets, and tasks ( Zahavy, Par. [0029], “In implementations the one or more policy neural networks 110 comprise a value function neural network configured to process the observation 106 for the current time step, in accordance with current values of value function neural network parameters, to generate a current value estimate relating to the current state of the environment. The value function neural network may be a state or state-action value function neural network. That is, the current value estimate may be a state value estimate, i.e. an estimate of a value of the current state of the environment, or a state-action value estimate, i.e. an estimate of a value of each of a set of possible actions at the current time step.”). Regarding Claim 2 , Nemecek in view of Zahavy teaches the electronic device of claim 1, wherein the task vector is represented as a linear combination of the source task vectors ( Nemecek, Pg. 3, Section 3. Representation and Successor Features, “Successor features (SFs) (Barreto et al., 2017) have been proposed as a generalization to the successor representation. SFs assume that the reward function can be decomposed into a linear combination of features, r(s,a,s) = φ(s,a,s)Tw, where φ(s,a,s)T is a row vector of features and w is the weights, changing the representation from one based on state visitation to one based on feature occurrences. Using this assumption, successor features ψπ(s,a) are derived from the action-value function for a given policy π, Qπ(s,a), which we can write as the expected sum of discounted rewards starting from timestep t.”, therefore, the task vector may be represented as a linear combination of source task vectors (a linear combination of features comprising a row vector of features and corresponding weights for a variety of tasks)) . Regarding Claim 3 , Nemecek in view of Zahavy teaches the electronic device of claim 2, wherein the instructions are further configured to cause the one or more processors to: determine a number of combinations of the linear combination of the source task vectors to represent the task vector according to a threshold (Nemecek, Pg. 4, Section 5. Policy Cache Construction, “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.” & Pg. 6, Section 6.2 Terrainworld , “The seven base tasks each assign a weight of -1 to one of the seven terrain types and a weight of 0 to the goal. Subsequent tasks assign some combination of the coefficients 1 and 30 to each base task. Exhausting the possible combinations of this form results in 128 generated tasks. Our results are averaged over 5000 permutations of these generated tasks. Like Gridworld, we used a tabular representation for the SFs computed with MPI.”, thus, a number of combinations of the linear combination of the source task vectors to represent the task vector according to a threshold (valid combination which minimizes the upper bound) may be determined – Pg. 6 Section 6.2 is cited as such an example of determining a number of combinations) ; and determine the upper bound according to one of the determined linear combinations of the task vectors ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, an upper bound may be determined based on one of the determined linear combinations of the task vectors (valid combination which minimizes the upper bound)) . Regarding Claim 4 , Nemecek in view of Zahavy teaches the electronic device of claim 1, wherein the lower bound is determined based on an approximation of, and an approximation error of, the optimal value-function ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Algorithm 1 uses Theorem 1 to decide if the current policy cache is sufficient, given a performance threshold. If the gap between the upper and lower bounds is small enough for a given state, then we know that the value of the best policy in the cache must be within some small factor of the value for the optimal policy.” & Pg. 8, Section 8. Discussion and Future Work , “We have demonstrated empirically that our approach can be combined with function approximation. Although our bounds include approximation error, our experiments assume low approximation error and don’t include approximation error when deciding whether to grow the cache.”, therefore, the lower bound may be determined based on an approximation of, and an approximation error of, the optimal value-function (also shown by Theorem 1 & Algorithm 1)) . Regarding Claim 5 , Nemecek in view of Zahavy teaches the electronic device of claim 1, wherein the upper bound is determined based on an arbitrary linear combination of the source task vectors to represent the task vector ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, an upper bound may be determined based on an arbitrary linear combination of the source task vectors to represent the task vector). Regarding Claim 6 , Nemecek in view of Zahavy teaches the electronic device of claim 1, wherein the instructions are further configured to cause the one or more processors to correct the optimal value-function in a range of less than or equal to the upper bound and greater than or equal to the lower bound ( Nemecek, Pg. 1, Abstract , “We present new bounds for the performance of optimal policies in a new task, as well as an approach to use these bounds to decide, when presented with a new task, whether to use cached policies or learn a new policy.” & Pg. 3 Section 4. Bounds on Policy Cache Performance, “We now consider a task R with reward function rR = i ∈ T αiri such that αi ≥ 0 and i ∈ T αi > 0, i.e., rR is a positive conical combination of the reward functions of the tasks in T . The proof of the theorem below (see ap pendix) provides an alternate derivation of the lower bound from Barreto et al., but does not improve upon it, while the upper bound is more general than those offered in Hunt et al. (2019) and applies to hard-max Q-learning.”, thus, the optimal value-function is corrected/updated to account for the upper bound (such that it is less than or equal to the upper bound) and the lower bound (such that it is greater than or equal to the lower bound). The optimal value-function accounts for and falls between the specified bounds – see also Theorem 1 on Pg. 3) . Regarding Claim 7 , Nemecek in view of Zahavy teaches the electronic device of claim 1, wherein the instructions are further configured to cause the one or more processors to: approximate a minimum value for the task vector ( Nemecek, Pg. 4, Section 4. Policy Cache Construction, “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound. For this LP, the variables are the coefficients αi, W is the set of reward weight vectors for the tasks in which the policies in the cache were learned, wnew is vector of reward weights for the new task, m is the number of policies in the cache, and s is the start state: [See equation on Pg. 4]”, thus, a minimum value-function may be outputted (based on minimizing the upper bound in accordance to the policy cache) using a state of an agent (start state, also described by the cited action-value function above) and source task vectors (reward weight vectors for tasks)) using a minimum value approximator trained to output a minimum value for the source task vectors ( Zahavy, Par. [0029], “In implementations the one or more policy neural networks 110 comprise a value function neural network configured to process the observation 106 for the current time step, in accordance with current values of value function neural network parameters, to generate a current value estimate relating to the current state of the environment. The value function neural network may be a state or state-action value function neural network. That is, the current value estimate may be a state value estimate, i.e. an estimate of a value of the current state of the environment, or a state-action value estimate, i.e. an estimate of a value of each of a set of possible actions at the current time step.”, thus, a value approximator (which may comprise a neural network, as supported by Applicant’s specification Par. [0100]) is trained to output a minimum value-function (satisfying a minimum performance criterion as supported by Nemecek Par. [0038-0040]) using a state of an agent and source task vectors to approximate an optimal value-function for a task vector (See Nemecek Par. [0051-0054] which explicitly discloses these vectors)) ; and determine the upper bound using the minimum value for the task vector ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, the upper bound may be determined using the minimum value for the task vector, as a valid combination of tasks (comprising the task vector) which minimizes the upper bound is utilized). The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 8 , Nemecek in view of Zahavy teaches the electronic device of claim 1, wherein the optimal policy comprises a neural network model ( Zahavy, Par. [0020], “The present application provides the following contributions. An incremental method for discovering a diverse set of near-optimal policies is proposed. Each policy in the set may be trained based on iterative updates that attempt to maximize diversity relative to other policies in the set under a minimum performance constraint.” & Par. [0028], “Each of the policy neural networks 110 is configured to process an input that includes a current observation 106 characterizing the current state of the environment 104, in accordance with the policy parameters 140, to generate a neural network output for selecting the action 112.”, therefore, the optimal policy may comprise a neural network model) . It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the electronic device of claim 1, as disclosed by Nemecek in view of Zahavy to include wherein the optimal policy comprises a neural network model, as disclosed by Zahavy. One of ordinary skill in the art would have been motivated to make this modification to provide improvements for reinforcement learning by providing a diverse set of policies in which the optimal policy comprises a neural network which may provide multiple means of performing a given task, hence improving model robustness ( Zahavy, Par. [0019-0020], “The present disclosure presents an improved reinforcement learning method in which training is based on extrinsic rewards from the environment and intrinsic rewards based on diversity. An objective function is provided that combines both performance and diversity to provide a set of diverse policies for performing a task. By providing a diverse set of policies, the methods described herein provide multiple means of performing a given task, thereby improving robustness. The present application provides the following contributions. An incremental method for discovering a diverse set of near-optimal policies is proposed. Each policy in the set may be trained based on iterative updates that attempt to maximize diversity relative to other policies in the set under a minimum performance constraint.”). Regarding Claim 9 , Nemecek in view of Zahavy teaches a method of transferring reinforcement learning ( Nemecek, Pg. 2, Section 2. Background , “The underlying idea of transfer learning for reinforcement learning (Taylor & Stone, 2009) is that through learning to perform well in one or more tasks, the learning process for a new task should be improved in some way, typically by reducing the amount of training in the new task required to reach a given level of performance. As shown in Section 5, an agent may even be able to avoid learning a new pol icy entirely if it can determine that its cached policies are sufficiently good for the novel task. While there are many notions of transfer, we focus on the case where the difference between tasks lies solely in their reward functions, i.e., the dynamics and other aspects of the environment remain the same.”, thus, methods of transferring reinforcement learning are disclosed) , the method comprising: […] The rest of the claim language in Claim 9 recites substantially the same limitations as Claim 1, in the form of a method, therefore it is rejected under the same rationale. The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Claim 10 recites substantially the same limitations as Claim 2, in the form of a method, therefore it is rejected under the same rationale. Regarding Claim 11 , Nemecek in view of Zahavy teaches the method of claim 10, wherein the determining of the upper bound of the optimal value-function for the task vector comprises: determining a number of combinations of the linear combination of the source task vectors to represent the task vector by a predetermined threshold (Nemecek, Pg. 4, Section 5. Policy Cache Construction, “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.” & Pg. 6, Section 6.2 Terrainworld , “The seven base tasks each assign a weight of -1 to one of the seven terrain types and a weight of 0 to the goal. Subsequent tasks assign some combination of the coefficients 1 and 30 to each base task. Exhausting the possible combinations of this form results in 128 generated tasks. Our results are averaged over 5000 permutations of these generated tasks. Like Gridworld, we used a tabular representation for the SFs computed with MPI.”, thus, a number of combinations of the linear combination of the source task vectors to represent the task vector according to a threshold (valid combination which minimizes the upper bound) may be determined – Pg. 6 Section 6.2 is cited as such an example of determining a number of combinations) ; and determining the upper bound according to a combination of the linear combination of the source task vectors ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, an upper bound may be determined based on one of the determined linear combinations of the task vectors (valid combination which minimizes the upper bound)) . Regarding Claim 12 , Nemecek in view of Zahavy teaches the method of claim 9, wherein the lower bound is determined based on an approximation of, and an approximation error of, the optimal value-function for the task vector of policies based on the source task vectors ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Algorithm 1 uses Theorem 1 to decide if the current policy cache is sufficient, given a performance threshold. If the gap between the upper and lower bounds is small enough for a given state, then we know that the value of the best policy in the cache must be within some small factor of the value for the optimal policy.” & Pg. 8, Section 8. Discussion and Future Work , “We have demonstrated empirically that our approach can be combined with function approximation. Although our bounds include approximation error, our experiments assume low approximation error and don’t include approximation error when deciding whether to grow the cache.”, therefore, the lower bound may be determined based on an approximation of, and an approximation error of, the optimal value-function (also shown by Theorem 1 & Algorithm 1)) . Regarding Claim 13 , Nemecek in view of Zahavy teaches the method of claim 12, wherein the policies comprise respective neural networks trained with respect to the source task vectors ( Zahavy, Par. [0020], “The present application provides the following contributions. An incremental method for discovering a diverse set of near-optimal policies is proposed. Each policy in the set may be trained based on iterative updates that attempt to maximize diversity relative to other policies in the set under a minimum performance constraint.” & Par. [0028], “Each of the policy neural networks 110 is configured to process an input that includes a current observation 106 characterizing the current state of the environment 104, in accordance with the policy parameters 140, to generate a neural network output for selecting the action 112.”, therefore, the policies may comprise respective neural networks trained with respect to source task vectors) . It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of transferring reinforcement learning of claims 9 and 12, as disclosed by Nemecek in view of Zahavy to include wherein the policies comprise respective neural networks trained with respect to the source task vectors, as disclosed by Zahavy. One of ordinary skill in the art would have been motivated to make this modification to provide improvements for reinforcement learning by providing a diverse set of policies in which the policies may comprise respective neural networks which may provide multiple means of performing a given task, hence improving model robustness ( Zahavy, Par. [0019-0020], “The present disclosure presents an improved reinforcement learning method in which training is based on extrinsic rewards from the environment and intrinsic rewards based on diversity. An objective function is provided that combines both performance and diversity to provide a set of diverse policies for performing a task. By providing a diverse set of policies, the methods described herein provide multiple means of performing a given task, thereby improving robustness. The present application provides the following contributions. An incremental method for discovering a diverse set of near-optimal policies is proposed. Each policy in the set may be trained based on iterative updates that attempt to maximize diversity relative to other policies in the set under a minimum performance constraint.”). Regarding Claim 14 , Nemecek in view of Zahavy teaches the method of claim 9, wherein the upper bound is determined based on a combination of the linear combination of the source task vectors to represent the task vector ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache beyond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, an upper bound may be determined based on a combination of the linear combination of the source task vectors to represent the task vector). Regarding Claim 15 , Nemecek in view of Zahavy teaches the method of claim 9, wherein the correcting of the optimal value-function for the task vector comprises correcting the optimal value-function for the task vector in a range of less than or equal to the upper bound and greater than or equal to the lower bound ( Nemecek, Pg. 1, Abstract , “We present new bounds for the performance of optimal policies in a new task, as well as an approach to use these bounds to decide, when presented with a new task, whether to use cached policies or learn a new policy.” & Pg. 3 Section 4. Bounds on Policy Cache Performance, “We now consider a task R with reward function rR = i ∈ T αiri such that αi ≥ 0 and i ∈ T αi > 0, i.e., rR is a positive conical combination of the reward functions of the tasks in T . The proof of the theorem below (see ap pendix) provides an alternate derivation of the lower bound from Barreto et al., but does not improve upon it, while the upper bound is more general than those offered in Hunt et al. (2019) and applies to hard-max Q-learning.”, thus, the optimal value-function is corrected/updated to account for the upper bound (such that it is less than or equal to the upper bound) and the lower bound (such that it is greater than or equal to the lower bound). The optimal value-function accounts for and falls between the specified bounds – see also Theorem 1 on Pg. 3) . Regarding Claim 16 , Nemecek in view of Zahavy teaches the method of claim 9, further comprising: approximating a minimum value for the task vector ( Nemecek, Pg. 4, Section 4. Policy Cache Construction, “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound. For this LP, the variables are the coefficients αi, W is the set of reward weight vectors for the tasks in which the policies in the cache were learned, wnew is vector of reward weights for the new task, m is the number of policies in the cache, and s is the start state: [See equation on Pg. 4]”, thus, a minimum value-function may be outputted (based on minimizing the upper bound in accordance to the policy cache) using a state of an agent (start state, also described by the cited action-value function above) and source task vectors (reward weight vectors for tasks)) using a minimum value approximator trained to output a minimum value for the source task vectors ( Zahavy, Par. [0029], “In implementations the one or more policy neural networks 110 comprise a value function neural network configured to process the observation 106 for the current time step, in accordance with current values of value function neural network parameters, to generate a current value estimate relating to the current state of the environment. The value function neural network may be a state or state-action value function neural network. That is, the current value estimate may be a state value estimate, i.e. an estimate of a value of the current state of the environment, or a state-action value estimate, i.e. an estimate of a value of each of a set of possible actions at the current time step.”, thus, a value approximator (which may comprise a neural network, as supported by Applicant’s specification Par. [0100]) is trained to output a minimum value-function (satisfying a minimum performance criterion as supported by Nemecek Par. [0038-0040]) using a state of an agent and source task vectors to approximate an optimal value-function for a task vector (See Nemecek Par. [0051-0054] which explicitly discloses these vectors)) , wherein the determining of the upper bound comprises determining the upper bound using the minimum value for the task vector ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, the upper bound may be determined using the minimum value for the task vector, as a valid combination of tasks (comprising the task vector) which minimizes the upper bound is utilized). The reasons of obviousness have been noted in the rejection of Claim 9 above and applicable herein. Regarding Claim 17 , Nemecek teaches a method of transferring reinforcement learning ( Nemecek, Pg. 2, Section 2. Background , “The underlying idea of transfer learning for reinforcement learning (Taylor & Stone, 2009) is that through learning to perform well in one or more tasks, the learning process for a new task should be improved in some way, typically by reducing the amount of training in the new task required to reach a given level of performance. As shown in Section 5, an agent may even be able to avoid learning a new pol icy entirely if it can determine that its cached policies are sufficiently good for the novel task. While there are many notions of transfer, we focus on the case where the difference between tasks lies solely in their reward functions, i.e., the dynamics and other aspects of the environment remain the same.”, thus, methods of transferring reinforcement learning are disclosed) , the method comprising: approximating a task vector using a task vector approximator trained to ( While Nemecek discloses approximating a task vector […] to output source task vectors using source task information of source tasks. Nemecek does not explicitly disclose approximating a task vector using a task vector approximator trained to output source task vectors using source task information of source tasks. See introduction of Zahavy reference below for explicit disclosure of approximating a task vector using a task vector approximator trained to perform these operations of the claim language) output source task vectors using source task information of source tasks ( Nemecek, Pg. 3, Section 4. Bounds on Policy Cache Performance, “Let T be a set of tasks which differ only in their reward functions. For any task i ∈ T , let ri be the reward function, πi be an optimal policy for ri, and let V πi j be the value function for policy πi executed in task j. As πi is an optimal policy for task i, it follows that the value under πi in task i is at least as good as under any other policy πj: […] We now consider a task R with reward function rR = i ∈ T αiri such that αi ≥ 0 and i ∈ T αi > 0, i.e., rR is a positive conical combination of the reward functions of the tasks in T”, therefore, T comprises a set of tasks which differ only in reward functions, the tasks of the task vector are approximated using source information of source tasks (reward functions & combination of reward functions/features of the tasks)) ; approximating a feature of the task vector using a feature approximator trained to ( While Nemecek discloses approximating a feature of the task vector […] output a feature of the source task vector based on a state of an agent. Nemecek does not explicitly disclose approximating a feature of the task vector using a feature approximator trained to output a feature of the source task vector based on a state of an agent. See introduction of Zahavy reference below for explicit disclosure of approximating a feature of the task vector using a feature approximator trained to perform these operations of the claim language) output a feature of the source task vector based on a state of an agent ( Nemecek, Pg. 3, Section 4. Bounds on Policy Cache Performance , “Let Λπi be a matrix with each row corresponding to a state s ∈ S, and each column corresponding to a successor feature. Thus, each row corresponds to a row vector of successor features for a state under πi: Λπi(s) = ψπi(s,πi(s)) and is analogous to the relationship between V πi and Qπi.”, thus, features of the task are approximated based on state of an agent) ; approximating an optimal value-function for the task vector using the task vector and the feature of the task vector ( Nemecek, Pg. 2, Section 2. Background, “A value function V π(s) is the expected total discounted reward for starting in state s and performing actions according to π. The action-value function Qπ(s,a) gives the value of taking action a from state s and following π afterwards. The goal of a learning agent is to learn the optimal policy π ∗ which maximizes the expected future reward from each state. We can express the optimal value function V ∗ and action-value function Q ∗ via the Bellman equation as: [See equation on Pg. 2] […] The underlying idea of transfer learning for reinforcement learning (Taylor & Stone, 2009) is that through learning to perform well in one or more tasks, the learning process for a new task should be improved in some way, typically by reducing the amount of training in the new task required to reach a given level of performance.”, therefore, an optimal value-function is approximated for a task vector using the task vector and features of the task vector (This is also depicted by the first equation of Section 4. Bounds on Policy Cache Performance on Nemecek Pg. 3 which shows how the value function considers the task vector and features of the task vector). This is analogous to the equation 2, for approximating an optimal value-function, as presented by Applicant’s specification Par. [0056]) ; determining an upper bound of the optimal value-function ( Nemecek, Pg. 4, Section 6. Experimental Results, “As described in Algorithm 1, the agent starts with a policy cache containing policies for each of the base tasks for the given environment and when presented with each subsequent task, calculates the upper and lower bounds at the start state using the current policy cache and compares them. If the size of the gap exceeds threshold η, then a new policy is added to the cache.”, thus, an upper bound of the optimal value-function for the task vector may be determined. This is similarly supported by Algorithms 1 & 2 on Pg. 4) ; determining a lower bound of the optimal value-function ( Nemecek, Pg. 4, Section 6. Experimental Results, “As described in Algorithm 1, the agent starts with a policy cache containing policies for each of the base tasks for the given environment and when presented with each subsequent task, calculates the upper and lower bounds at the start state using the current policy cache and compares them. If the size of the gap exceeds threshold η, then a new policy is added to the cache.”, thus, a lower bound of the optimal value-function for the task vector may be determined. This is similarly supported by Algorithms 1 & 2 on Pg. 4) ; correcting the optimal value-function based on the upper bound and the lower bound ( Nemecek, Pgs. 3-4, Section 4. Bounds on Policy Cache Performance, “We now consider a task R with reward function rR = i ∈ T αiri such that αi ≥ 0 and i ∈ T αi > 0, i.e., rR is a positive conical combination of the reward functions of the tasks in T. The proof of the theorem below (see appendix) provides an alternate derivation of the lower bound from Barreto et al., but does not improve upon it, while the upper bound is more general than those offered in Hunt et al. (2019) and applies to hard-max Q-learning. [See Theorem 1 on Pg. 3] Theorem 1 turns out to be quite powerful in practice because the optimality of a value function for a particular reward function significantly constrains the value function for other reward functions with similar weights”, therefore, the value function may be corrected/updated to account for the upper and lower bounds) ; and determining an optimal policy for the task vector using the corrected optimal value-function ( Nemecek, Pg. 4, Section 5. Policy Cache Construction, “We start with a set of base tasks defined by a set of reward feature weights, and assume all subsequent tasks have rewards within the conical hull of the initial set. Algorithm 1 uses Theorem 1 to decide if the current policy cache is sufficient, given a performance threshold. If the gap between the upper and lower bounds is small enough for a given state, then we know that the value of the best policy in the cache must be within some small factor of the value for the optimal policy.”, thus, based on the corrected optimal value-function (which accounts for the upper/lower bounds as shown by Algorithm 1 & Theorem 1), an optimal policy for the task vector may be determined. This is similarly supported by Pg. 2 which states “Our first contribution is to address these limitations by presenting a new upper bound on the hard-max Q-function of the optimal policy for a new task when that task is a conical combination of previous tasks.”). While Nemecek discloses approximating a task vector […] to output source task vectors using source task information of source tasks (See rejection above), Nemecek does not explicitly disclose approximating a task vector using a task vector approximator trained to output source task vectors using source task information of source tasks. However, Zahavy teaches approximating a task vector using a task vector approximator trained to output source task vectors using source task information of source tasks ( Nemecek, Par. [0024-0025], “The reinforcement learning neural network system 100 has one or more inputs to receive data from the environment characterizing a state of the environment, e.g. data from one or more sensors of the environment. Data characterizing a state of the environment is referred to herein as an observation 106. The data from the environment can also include extrinsic rewards (or task rewards). Generally an extrinsic reward 108 is represented by a scalar numeric value characterizing progress of the agent towards the task goal and can be based on any event in, or aspect of, the environment. Extrinsic rewards may be received as a task progresses or only at the end of a task, e.g. to indicate successful completion of the task. Alternatively or in addition, the extrinsic rewards 108 may be calculated by the reinforcement learning neural network system 100 based on the observations 106 using an extrinsic reward function”, therefore, a task vector may be approximated using a task vector approximator (neural network) trained to output source task vectors using source information (extrinsic rewards, states, actions, etc.)) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of transferring reinforcement learning, as disclosed by Nemecek to include approximating a task vector using a task vector approximator trained to output source task vectors using source task information of source tasks, as disclosed by Zahavy. One of ordinary skill in the art would have been motivated to make this modification to enable the use of a task vector approximator (neural network) which may more efficiently and accurately approximate a task vector by utilizing and contextualizing source task information to approximate a source task vector ( Nemecek, Par. [0024-0025], “The reinforcement learning neural network system 100 has one or more inputs to receive data from the environment characterizing a state of the environment, e.g. data from one or more sensors of the environment. Data characterizing a state of the environment is referred to herein as an observation 106. The data from the environment can also include extrinsic rewards (or task rewards). Generally an extrinsic reward 108 is represented by a scalar numeric value characterizing progress of the agent towards the task goal and can be based on any event in, or aspect of, the environment. Extrinsic rewards may be received as a task progresses or only at the end of a task, e.g. to indicate successful completion of the task. Alternatively or in addition, the extrinsic rewards 108 may be calculated by the reinforcement learning neural network system 100 based on the observations 106 using an extrinsic reward function”) While Nemecek discloses approximating a feature of the task vector […] output a feature of the source task vector based on a state of an agent (See rejection above), Nemecek does not explicitly disclose approximating a feature of the task vector using a feature approximator trained to output a feature of the source task vector based on a state of an agent. However, Zahavy teaches approximating a feature of the task vector using a feature approximator trained to output a feature of the source task vector based on a state of an agent ( Zahavy, Par. [0051], “The feature vector ϕ(s, a) may be considered an encoding of a given state s and action a. The feature vector ϕ(s, a) may be bounded, e.g. between 0 and 1 (ϕ(s, a) ∈ [0,1] d where d is a dimension of the feature vector ϕ(s, a) and of the weight vector w ∈ PNG media_image1.png 38 29 media_image1.png Greyscale d . The mapping from states and actions to feature vectors can be implemented through a trained approximator (e.g. a neural network). Whilst the above references an encoding of actions and states, a feature vector may alternatively be an encoding of a given state only ϕ(s)”, therefore, a feature of the task vector may be approximated using a feature approximator (neural network) trained to output a feature of the source task vector based on a state of an agent). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of transferring reinforcement learning, as disclosed by Nemecek to include approximating a feature of the task vector using a feature approximator trained to output a feature of the source task vector based on a state of an agent, as disclosed by Zahavy. One of ordinary skill in the art would have been motivated to make this modification to enable the use of a feature approximator (neural network) which may more efficiently and accurately approximate features of the task vector by utilizing and contextualizing given states and actions ( Zahavy, Par. [0051], “The feature vector ϕ(s, a) may be considered an encoding of a given state s and action a. The feature vector ϕ(s, a) may be bounded, e.g. between 0 and 1 (ϕ(s, a) ∈ [0,1] d where d is a dimension of the feature vector ϕ(s, a) and of the weight vector w ∈ PNG media_image1.png 38 29 media_image1.png Greyscale d . The mapping from states and actions to feature vectors can be implemented through a trained approximator (e.g. a neural network). Whilst the above references an encoding of actions and states, a feature vector may alternatively be an encoding of a given state only ϕ(s)”) Regarding Claim 18 , Nemecek in view of Zahavy teaches the method of claim 17, wherein the task vector is represented as a linear combination of the source task vectors ( Nemecek, Pg. 3, Section 3. Representation and Successor Features, “Successor features (SFs) (Barreto et al., 2017) have been proposed as a generalization to the successor representation. SFs assume that the reward function can be decomposed into a linear combination of features, r(s,a,s) = φ(s,a,s)Tw, where φ(s,a,s)T is a row vector of features and w is the weights, changing the representation from one based on state visitation to one based on feature occurrences. Using this assumption, successor features ψπ(s,a) are derived from the action-value function for a given policy π, Qπ(s,a), which we can write as the expected sum of discounted rewards starting from timestep t.”, therefore, the task vector may be represented as a linear combination of source task vectors (a linear combination of features comprising a row vector of features and corresponding weights for a variety of tasks)) . Regarding Claim 19 , Nemecek in view of Zahavy teaches the method of claim 17, wherein the lower bound is determined based on an approximation and an approximation error of the optimal value-function, wherein the approximation error is based on the source task vectors ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Algorithm 1 uses Theorem 1 to decide if the current policy cache is sufficient, given a performance threshold. If the gap between the upper and lower bounds is small enough for a given state, then we know that the value of the best policy in the cache must be within some small factor of the value for the optimal policy.” & Pg. 8, Section 8. Discussion and Future Work , “We have demonstrated empirically that our approach can be combined with function approximation. Although our bounds include approximation error, our experiments assume low approximation error and don’t include approximation error when deciding whether to grow the cache.”, therefore, the lower bound may be determined based on an approximation of, and an approximation error of, the optimal value-function (also shown by Theorem 1 & Algorithm 1)) . Regarding Claim 20 , Nemecek in view of Zahavy teaches the method of claim 17, wherein the upper bound is determined based on one of multiple linear combinations of the source task vectors ( Nemecek, Pg. 4, Section 5. Policy Cache Construction , “Once an additional policy has been added to the cache be yond those for the base tasks, any subsequent task may not be a unique conical combination of previous tasks for which policies are stored. Therefore, we use a linear program in Algorithm 2 to find the valid combination which minimizes the upper bound.”, therefore, an upper bound may be determined based on one of multiple linear combinations of the source task vectors to represent the task vector). Conclusion 9. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Devika S Maharaj whose telephone number is (571)272-0829. The examiner can normally be reached Monday - Thursday 8:30am - 5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVIKA S MAHARAJ/Examiner, Art Unit 2123 Application/Control Number: 18/475,654 Page 2 Art Unit: 2123 Application/Control Number: 18/475,654 Page 3 Art Unit: 2123 Application/Control Number: 18/475,654 Page 4 Art Unit: 2123 Application/Control Number: 18/475,654 Page 5 Art Unit: 2123 Application/Control Number: 18/475,654 Page 6 Art Unit: 2123 Application/Control Number: 18/475,654 Page 7 Art Unit: 2123 Application/Control Number: 18/475,654 Page 8 Art Unit: 2123 Application/Control Number: 18/475,654 Page 9 Art Unit: 2123 Application/Control Number: 18/475,654 Page 10 Art Unit: 2123 Application/Control Number: 18/475,654 Page 11 Art Unit: 2123 Application/Control Number: 18/475,654 Page 12 Art Unit: 2123 Application/Control Number: 18/475,654 Page 13 Art Unit: 2123 Application/Control Number: 18/475,654 Page 14 Art Unit: 2123 Application/Control Number: 18/475,654 Page 15 Art Unit: 2123 Application/Control Number: 18/475,654 Page 16 Art Unit: 2123 Application/Control Number: 18/475,654 Page 17 Art Unit: 2123 Application/Control Number: 18/475,654 Page 18 Art Unit: 2123 Application/Control Number: 18/475,654 Page 19 Art Unit: 2123 Application/Control Number: 18/475,654 Page 20 Art Unit: 2123 Application/Control Number: 18/475,654 Page 21 Art Unit: 2123 Application/Control Number: 18/475,654 Page 22 Art Unit: 2123 Application/Control Number: 18/475,654 Page 23 Art Unit: 2123 Application/Control Number: 18/475,654 Page 24 Art Unit: 2123 Application/Control Number: 18/475,654 Page 25 Art Unit: 2123 Application/Control Number: 18/475,654 Page 26 Art Unit: 2123 Application/Control Number: 18/475,654 Page 27 Art Unit: 2123 Application/Control Number: 18/475,654 Page 28 Art Unit: 2123 Application/Control Number: 18/475,654 Page 29 Art Unit: 2123 Application/Control Number: 18/475,654 Page 30 Art Unit: 2123 Application/Control Number: 18/475,654 Page 31 Art Unit: 2123 Application/Control Number: 18/475,654 Page 32 Art Unit: 2123 Application/Control Number: 18/475,654 Page 33 Art Unit: 2123 Application/Control Number: 18/475,654 Page 34 Art Unit: 2123 Application/Control Number: 18/475,654 Page 35 Art Unit: 2123 Application/Control Number: 18/475,654 Page 36 Art Unit: 2123 Application/Control Number: 18/475,654 Page 37 Art Unit: 2123 Application/Control Number: 18/475,654 Page 38 Art Unit: 2123 Application/Control Number: 18/475,654 Page 39 Art Unit: 2123 Application/Control Number: 18/475,654 Page 40 Art Unit: 2123 Application/Control Number: 18/475,654 Page 41 Art Unit: 2123 Application/Control Number: 18/475,654 Page 42 Art Unit: 2123 Application/Control Number: 18/475,654 Page 43 Art Unit: 2123 Application/Control Number: 18/475,654 Page 44 Art Unit: 2123 Application/Control Number: 18/475,654 Page 45 Art Unit: 2123 Application/Control Number: 18/475,654 Page 46 Art Unit: 2123 Application/Control Number: 18/475,654 Page 47 Art Unit: 2123 Application/Control Number: 18/475,654 Page 48 Art Unit: 2123 Application/Control Number: 18/475,654 Page 49 Art Unit: 2123
Read full office action

Prosecution Timeline

Sep 27, 2023
Application Filed
Apr 28, 2026
Non-Final Rejection mailed — §101, §103
Jul 28, 2026
Applicant Interview (Telephonic)
Jul 28, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694269
SELECTIVE REPORTING OF MACHINE LEARNING PARAMETERS FOR FEDERATED LEARNING
4y 0m to grant Granted Jul 28, 2026
Patent 12682205
DIFFERENTIAL EQUATIONS NETWORK
7y 9m to grant Granted Jul 14, 2026
Patent 12682215
FLEXIBLE MACHINE LEARNING
4y 1m to grant Granted Jul 14, 2026
Patent 12675689
MULTI-DOMAIN FEATURE ENHANCEMENT FOR TRANSFER LEARNING (FTL)
4y 4m to grant Granted Jul 07, 2026
Patent 12657441
SPIKING NEURAL NETWORK DEVICE THAT UPDATES SYNAPIC WEIGHT BASED ON OUTPUT FREQUENCY AND LEARNING METHOD OF SPIKING NEURAL NETWORK DEVICE
5y 9m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
56%
Grant Probability
65%
With Interview (+9.3%)
4y 7m (~1y 8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 86 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month