Prosecution Insights
Last updated: August 16, 2026
Application No. 18/689,823

DYNAMIC REINFORCEMENT LEARNING

Non-Final OA §101§102§103§112§Other
Filed
Mar 06, 2024
Priority
Sep 07, 2021 — CN PCT/CN2021/116892 +1 more
Examiner
KWON, JUN
Art Unit
Tech Center
Assignee
Telefonaktiebolaget LM Ericsson
OA Round
1 (Non-Final)
40%
Grant Probability
Moderate
1-2
OA Rounds
2y 3m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 40% of resolved cases
40%
Career Allowance Rate
31 granted / 77 resolved
-19.7% vs TC avg
Strong +46% interview lift
Without
With
+46.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 8m
Avg Prosecution
27 currently pending
Career history
106
Total Applications
across all art units

Statute-Specific Performance

§101
29.1%
-10.9% vs TC avg
§103
46.7%
+6.7% vs TC avg
§102
9.0%
-31.0% vs TC avg
§112
14.4%
-25.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 77 resolved cases

Office Action

§101 §102 §103 §112 §Other
Detailed Action Claims 1-8, 10-16, 18, 20 and 23-24 are currently pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. PCT/CN2021/116892, filed on 09/07/2021. Information Disclosure Statement The information disclosure statement (IDS) submitted on 03/06/2024 was filed before the mailing date of the first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claim 15 is objected to because of the following informalities: “… the generated reward value is: a sum of the K reward value, a weighted sum of said K reward values, …” missing a comma. Claim 16 is objected to because of the following informalities: “The method . Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 24 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 24 recites “The RL agent of claim 22 been canceled. For purpose of the examination, the claim is interpreted as: The RL agent of claim 23 … Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-8, 10-16, 18, 20 and 23-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1, Step 1: Claim 1 recites a method for dynamic reinforcement learning. Therefore, it is directed to the statutory category of Processes. 2A Prong 1: using R1 and/or a performance indicator, PI, to determine whether an algorithm modification condition is satisfied; (mental process of evaluation – determining whether to modify the algorithm, which can be done in one’s mind) as a result of determining that the algorithm modification condition is satisfied, modifying the RL algorithm to produce a modified RL algorithm. (mathematical concept – See spec para [0025] and [0029]) 2A Prong 2: A method for dynamic reinforcement learning, (RL), the method comprising: (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) using an RL algorithm (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) triggering performance of the selected first action; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) after the first action is performed, obtaining a first reward value, R1, associated with the first action; (an insignificant extra-solution activity MPEP 2106.05(g)(iii) of gathering statistics, based on the spec para [0023]. The reward value is collected from somewhere by the agent) The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are mere insignificant extra solution activity, combination of generic computer functions that are implemented to perform the disclosed abstract idea above. 2B: A method for dynamic reinforcement learning, (RL), the method comprising: (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) using an RL algorithm (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) triggering performance of the selected first action; (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) after the first action is performed, obtaining a first reward value, R1, associated with the first action; (indicated as an insignificant extra-solution activity MPEP 2106.05(g) in Step 2A Prong 2. Therefore, it is re-evaluated in Step 2B as well understood, routine, and conventional activity MPEP 2106.05(d)(II)(iv) of gathering statistics, based on the spec para [0023]. The reward value is collected from somewhere by the agent) The additional elements as disclosed above in combination of the abstract idea are not sufficient to amount to significantly more than the judicial exception as they are well, understood, routine and conventional activity disclosed in combination of generic computer functions that are implemented to perform the disclosed abstract idea above. Regarding claim 2, Step 1: Processes, as above. 2A Prong 1: The method of claim 1, wherein modifying the RL algorithm to produce the modified RL algorithm comprises modifying a parameter of the RL algorithm. (mathematical concept – See spec para [0025], [0028], and [0029]) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 3, Step 1: Processes, as above. 2A Prong 1: The method of claim 2, wherein modifying a parameter of the RL algorithm comprises modifying: an exploration probability of the RL algorithm, a learning rate of the RL algorithm, a discount factor of the RL algorithm, and/or a replay memory capacity of the RL algorithm. (in light of the limitation itself and paragraphs [0025] and [0029], directed to a mathematical concept) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 4, Step 1: Processes, as above. 2A Prong 1: The method of claim 3, wherein using the RL algorithm to select the first action comprises selecting the first action based on the exploration probability, (a mental process of judgment – selecting a random action based on the algorithm, which can be done in one’s mind. See [Algorithm 1, lines 6-7]) and modifying the RL algorithm to produce the modified RL algorithm comprises modifying the exploration probability. (mathematical concept – See spec para [0025] and [0029]) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 5, Step 1: Processes, as above. 2A Prong 1: The method of claim 1, wherein using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises one or more of: (mental process of evaluation) comparing R1 to a first threshold, (mental process of evaluation) comparing AR to a second threshold, wherein AR is a difference between R1 and a reward value associated with a second action selected using the RL algorithm, or (mental process of evaluation) comparing the PI to a third threshold. (mental process of evaluation) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 6, Step 1: Processes, as above. 2A Prong 1: The method of claim 1, further comprising: before using the RL algorithm to select the first action and obtaining R1, using the RL algorithm to select a second action; (mental process of judgment – determining a random action) using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises performing a decision process comprising: (mental process of evaluation) calculating AR = R2 - R1; and (in light of the limitation itself, directed to a mathematical concept) determining whether AR is greater than a drop threshold. (mental process of evaluation) 2A Prong 2: triggering performance of the selected second action; and (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) after the second action is performed, obtaining a second reward value, R2, associated with the second action, wherein (an insignificant extra-solution activity MPEP 2106.05(g)(iii) of gathering statistics, based on the spec para [0023]. The reward value is collected from somewhere by the agent) 2B: triggering performance of the selected second action; and (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) after the second action is performed, obtaining a second reward value, R2, associated with the second action, wherein (indicated as an insignificant extra-solution activity MPEP 2106.05(g) in Step 2A Prong 2. Therefore, it is re-evaluated in Step 2B as well understood, routine, and conventional activity MPEP 2106.05(d)(II)(iv) of gathering statistics, based on the spec para [0023]. The reward value is collected from somewhere by the agent) Regarding claim 7, Step 1: Processes, as above. 2A Prong 1: The method of claim 6, wherein ϵ , the algorithm modification condition is satisfied when AR is greater than the drop threshold, (mental process of evaluation – selecting a random action based on the probability and determining whether the condition is met, can be done in one’s mind) and modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new exploration probability, ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ϵ R e S t a r t , where ϵ R e S t a r t is a predetermined exploration probability. (mathematical concept – See spec para [0025] and [0029]) 2A Prong 2: wherein using the RL algorithm (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) 2B: wherein using the RL algorithm (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) Regarding claim 8, Step 1: Processes, as above. 2A Prong 1: The method of claim 6, wherein the decision process further comprises, as a result of determining that AR is not greater than the drop threshold, then determining whether R1 is less than a lower reward threshold. (mental process of evaluation – determining whether the result is greater or less than the threshold) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 10, Step 1: Processes, as above. 2A Prong 1: The method of claim 6, wherein the decision process further comprises, as a result of determining that AR is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, (mental process of evaluation – determining whether the result is greater or less than the threshold) the algorithm modification condition is satisfied when AR is not greater than the drop threshold and R1 is greater than the upper reward threshold, and (mental process of evaluation – determining whether the result is greater or less than the threshold) modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new exploration probability, ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ϵ e n d , where ϵ e n d is a predetermined ending exploration probability. (mathematical concept – See spec para [0025] and [0029]) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 11, Step 1: Processes, as above. 2A Prong 1: The method of claim 6, wherein the decision process further comprises, as a result of determining that AR is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, (mental process of evaluation – determining whether the result is greater or less than the threshold) the algorithm modification condition is satisfied when AR is not greater than the drop threshold and R1 is not greater than the upper reward threshold, and (mental process of evaluation – determining whether the result is greater or less than the threshold) modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new exploration probability, ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ( ϵ   × c ) , where c is a predetermined constant. (mathematical concept – See spec para [0025] and [0029]) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 12, Step 1: Processes, as above. 2A Prong 1: The method claim 1, further comprising, prior to 2A Prong 2: prior to using the RL algorithm to select … using the RL algorithm (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) triggering the performance of each one of the K-1 actions; and (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) for each one of the K-1 actions, obtaining a reward value associated with the action. (an insignificant extra-solution activity MPEP 2106.05(g)(iii) of gathering statistics, based on the spec para [0023]. The reward value is collected from somewhere by the agent) 2B: prior to using the RL algorithm to select … using the RL algorithm (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) triggering the performance of each one of the K-1 actions; and (mere instructions to apply an exception using a generic computer component MPEP 2106.05(f)) for each one of the K-1 actions, obtaining a reward value associated with the action. (indicated as an insignificant extra-solution activity MPEP 2106.05(g) in Step 2A Prong 2. Therefore, it is re-evaluated in Step 2B as well understood, routine, and conventional activity MPEP 2106.05(d)(II)(iv) of gathering statistics, based on the spec para [0023]. The reward value is collected from somewhere by the agent) Regarding claim 13, Step 1: Processes, as above. 2A Prong 1: The method of claim 12, wherein using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises: (mental process of evaluation) using R1 and said K-1 reward values to generate a reward value that is a function of these K reward values; and (a mathematical concept. See paragraph [0028] which discloses the accumulated reward) comparing the generated reward value to a threshold. (a mathematical concept. See paragraph [0028] which discloses the accumulated reward and the comparison to thresholds) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 14, Step 1: Processes, as above. 2A Prong 1: The method of claim 12, wherein using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises: (mental process of evaluation) using R1 and said K-1 reward values to generate a reward value that is a function of these K reward values; and (a mathematical concept. See paragraph [0028] which discloses the accumulated reward) comparing Δ R to a threshold, wherein Δ R is a difference between the generated reward value and a previously generated reward value. (mental process of evaluation) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 15, Step 1: Processes, as above. 2A Prong 1: The method of claim 13, wherein the generated reward value is: a sum of the K reward value a weighted sum of said K reward values, a weighted sum of a subset of said K reward values, a mean of said K reward values, a mean of a subset of said K reward values, a median of said K reward values, or a median of a subset of said K reward values. (a mathematical concept. See paragraphs [0053]-[0054]) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 16, Step 1: Processes, as above. 2A Prong 1: The method claim 12, wherein the value of K is determined based on a correlation time of the environment and/or application requirements, or the value of K is determined based on a maximum allowed service interruption time. (mental process of evaluation) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Regarding claim 18, Step 1: Processes, as above. 2A Prong 1: The method of claim 1, wherein one or more of the recited thresholds is dynamically changed based on environment changes and/or service requirement changes, and (mental process of evaluation – determining the threshold based on the environment change, which can be done in one’s mind) 2A Prong 2: the method further comprises using the modified RL algorithm to select another action and triggering performance of the another action. (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) 2B: the method further comprises using the modified RL algorithm to select another action and triggering performance of the another action. (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) Regarding claim 20, Step 1: Processes, as above. 2A Prong 1: Incorporates the rejection of claim 1. 2A Prong 2: A non-transitory computer readable storage medium storing a computer program comprising instructions which when executed by processing circuitry of an agent causes the agent to perform the method of claim 1. (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) 2B: A non-transitory computer readable storage medium storing a computer program comprising instructions which when executed by processing circuitry of an agent causes the agent to perform the method of claim 1. (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) Regarding claim 23, Claim 23 is an apparatus claim which recites the similar features as the method claim 1, and is rejected under the same rationale as claim 1. Additional limitations of claim 23 not addressed in claim 1 are addressed below. Step 1: Claim 23 recites a reinforcement learning (RL) agent, the RL agent comprising: processing circuitry; and a memory, the memory containing instructions executable by the processing circuitry, wherein the RL is configured to perform a process. Therefore, it is directed to the statutory category of an Apparatus. 2A Prong 1: Rejected under the same rationale as claim 1. 2A Prong 2: A reinforcement learning (RL) agent, the RL agent comprising: processing circuitry; and a memory, the memory containing instructions executable by the processing circuitry, wherein the RL is configured to perform a process comprising: (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) 2B: A reinforcement learning (RL) agent, the RL agent comprising: processing circuitry; and a memory, the memory containing instructions executable by the processing circuitry, wherein the RL is configured to perform a process comprising: (mere instructions to apply an exception using a generic computer component. See MPEP 2106.05(f)) Regarding claim 24, Step 1: Apparatus, as above. 2A Prong 1: The RL agent of claim 22, modifying the RL algorithm to produce the modified RL algorithm comprises modifying a parameter of the RL algorithm, and (mathematical concept – See spec para [0025], [0028], and [0029]) modifying a parameter of the RL algorithm comprises modifying: an exploration probability of the RL algorithm, a learning rate of the RL algorithm, a discount factor of the RL algorithm, and/or a replay memory capacity of the RL algorithm. (mathematical concept – See spec para [0025], [0028], and [0029]) 2A Prong 2: This judicial exception is not integrated into a practical application. 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 12-13, 15-16, 20 and 23-24 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Rezaee et al. (US 20220035375 A1, hereinafter ‘Rezaee’). Regarding claim 1, Rezaee teaches: A method for dynamic reinforcement learning, (RL), the method comprising: ([Rezaee, 0057] and [0064] The RL training processor 412 runs a RL algorithm to update the parameter of the neural network) using an RL algorithm to select a first action; ([Rezaee, 0100] discloses selecting a trajectory (i.e., a first action) using the parameter based on a probabilistic evaluation value) triggering performance of the selected first action; ([Rezaee, 0102] At 806, the selected trajectory is followed by the vehicle 100 in the (actual or simulated) current state for one time step) after the first action is performed, obtaining a first reward value, R1, associated with the first action; ([Rezaee, 0102] At 806, the selected trajectory is followed by the vehicle 100 in the (actual or simulated) current state for one time step, and a reward is calculated based on the performance of the vehicle 100) using R1 and/or a performance indicator, PI, to determine whether an algorithm modification condition is satisfied; ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. Both paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) as a result of determining that the algorithm modification condition is satisfied, modifying the RL algorithm to produce a modified RL algorithm. ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. Both paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) Regarding claim 2, Rezaee teaches: The method of claim 1, wherein modifying the RL algorithm to produce the modified RL algorithm comprises modifying a parameter of the RL algorithm. ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. Both paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) Regarding claim 3, Rezaee teaches: The method of claim 2, wherein modifying a parameter of the RL algorithm comprises modifying: an exploration probability of the RL algorithm, a learning rate of the RL algorithm, a discount factor of the RL algorithm, and/or a replay memory capacity of the RL algorithm. ([Rezaee, 0079-0080] discloses learning the parameter based on the instantaneous reward, discounted estimated future evaluation value γ V ( S t + 1 , τ t + 1 ) (i.e., discount factor of the RL algorithm), and the current estimated evaluation value, and adjusting the weights based on the learning rate α and gradient. Since γ V ( S t + 1 , τ t + 1 ) and the current estimated evaluation value are ‘estimated’, the values are modified, and the parameter is modified in accordance with the modification) Regarding claim 12, Rezaee teaches: The method claim 1, further comprising, prior to using the RL algorithm to select the first action: ([Rezaee, 0100] The trajectory selector 336 selects the trajectory having the highest calculated evaluation value (i.e., rewards) as the selected trajectory. The value of K may be 2 and 2-1=1 action may be selected. [0105] and [Fig. 8] collectively shows that the action selection and evaluation are performed iteratively. This indicates that the RL algorithm selects actions prior to the current action (i.e., the first action)) using the RL algorithm to select K-1 actions, where K > 1; ([Rezaee, 0100] The trajectory selector 336 selects the trajectory having the highest calculated evaluation value (i.e., rewards) as the selected trajectory. The value of K may be 2 and 2-1=1 action may be selected) triggering the performance of each one of the K-1 actions; and ([Rezaee, 0100] The selected trajectory is followed by the vehicle) for each one of the K-1 actions, obtaining a reward value associated with the action. ([Rezaee, 0100] The trajectory selector 336 selects the trajectory having the highest calculated evaluation value (i.e., rewards) as the selected trajectory. This indicates that the predicted reward is calculated for each trajectory of a set of trajectories generated by the trajectory generator 332) Regarding claim 13, Rezaee teaches: The method of claim 12, wherein using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises: ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. The paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) using R1 and said K-1 reward values to generate a reward value that is a function of these K reward values; and ([Rezaee, 0079] and [0104] The parameters are updated based on the current estimated evaluation value (i.e., R1) and discounted sum of expected future rewards (i.e., a function of the K-1 reward value). [Abstract] The paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) comparing the generated reward value to a threshold. ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. [Abstract] The paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) Regarding claim 15, Rezaee teaches: The method of claim 13, wherein the generated reward value is: a sum of the K reward value a weighted sum of said K reward values, a weighted sum of a subset of said K reward values, a mean of said K reward values, a mean of a subset of said K reward values, a median of said K reward values, or a median of a subset of said K reward values. ([Rezaee, 0079] and [0104] The parameters are updated based on the discounted sum of expected future rewards (i.e., a sum of the K reward value). [Abstract] The paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) Regarding claim 16, Rezaee teaches: The method claim 12, wherein the value of K is determined based on a correlation time of the environment and/or application requirements, or the value of K is determined based on a maximum allowed service interruption time. ([Rezaee, 0071] discloses that the state prediction is performed for a predefined number of time steps. The ‘predefined number of time steps’ is the ‘correlation time of the environment’ as the correlation time merely indicates a changing-scale of time. See the instant spec [0026]) Regarding claim 20, Rezaee teaches: A non-transitory computer readable storage medium storing a computer program comprising instructions which when executed by processing circuitry of an agent causes the agent to perform the method of claim 1. ([Rezaee, 0057] and [0064] The RL training processor 412 runs a RL algorithm to update the parameter of the neural network) Regarding claim 23, Rezaee teaches: A reinforcement learning (RL) agent, the RL agent comprising: ([Rezaee, 0057] and [0064] The RL training processor 412 runs a RL algorithm to update the parameter of the neural network) processing circuitry; and ([Rezaee, 0045] The processors 210 may include a CPU and electronic storage 220) a memory, the memory containing instructions executable by the processing circuitry, wherein the RL is configured to perform a process comprising: ([Rezaee, 0045] The processors 210 may include a CPU and electronic storage 220. [0057] and [0064] The RL training processor 412 runs a RL algorithm to update the parameter of the neural network) Claim 23 is an apparatus claim which recites the same features as the method claim 1, and is rejected for at least the same reasons. Regarding claim 24, Rezaee teaches: The RL agent of claim 22 modifying the RL algorithm to produce the modified RL algorithm comprises modifying a parameter of the RL algorithm, and ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. Both paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) modifying a parameter of the RL algorithm comprises modifying: an exploration probability of the RL algorithm, a learning rate of the RL algorithm, a discount factor of the RL algorithm, and/or a replay memory capacity of the RL algorithm. ([Rezaee, 0079-0080] discloses learning the parameter based on the instantaneous reward, discounted estimated future evaluation value γ V ( S t + 1 , τ t + 1 ) (i.e., discount factor of the RL algorithm), and the current estimated evaluation value, and adjusting the weights based on the learning rate α and gradient. Since γ V ( S t + 1 , τ t + 1 ) and the current estimated evaluation value are ‘estimated’, the values are modified, and the parameter is modified in accordance with the modification) Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 4-7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Rezaee in view of Mignon et al. (“An Adaptive Implementation of ϵ -Greedy in Reinforcement Learning”, 2017, hereinafter ‘Mignon’). Regarding claim 4, Rezaee teaches: The method of claim 3, wherein using the RL algorithm to select the first action comprises selecting the first action based on the exploration probability, and modifying the RL algorithm to produce the modified RL algorithm [Rezaee, 0100] discloses selecting the first action based on the probabilistic evaluation value. [Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’. Both paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) However, Rezaee does not specifically disclose: the modified RL algorithm comprises modifying the exploration probability. Mignon teaches: the modified RL algorithm comprises modifying the exploration probability. ([Mignon, page 1147, lines 1-3] indicates that the ϵ is a probability. [Mignon, page 1147, para under Algorithm 1, lines 1-12] The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is obtained, and compared to 0. If it is greater than zero, a new value is calculated for the ϵ (i.e., modifying the probability)) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Rezaee and Mignon to use the method of modifying the exploration probability based on a threshold to implement the reinforcement learning method of present invention. The suggestion and/or motivation for doing so is to improve the performance of the reinforcement learning method, as the adaptive epsilon greedy method presents better performance as compared to the classical reinforcement learning method [Mignon, Abstract]. Regarding claim 5, Rezaee teaches: The method of claim 1, wherein using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises one or more of: ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’) comparing R1 to a first threshold, ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold (i.e., first threshold) is interpreted as the ‘modification condition’. Both paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) However, Rezaee does not specifically disclose: comparing Δ R to a second threshold, wherein Δ R is a difference between R1 and a reward value associated with a second action selected using the RL algorithm, or comparing the PI to a third threshold. Mignon teaches: comparing Δ R to a second threshold, wherein Δ R is a difference between R1 and a reward value associated with a second action selected using the RL algorithm, or comparing the PI to a third threshold. ([Mignon, page 1147, para under Algorithm 1, lines 1-12] The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is obtained, and compared to 0. If it is greater than zero, a new value is calculated for the ϵ . Since both limitations are connected using ‘or’, the examiner is not required to teach ‘comparing the PI to a third threshold’) Regarding claim 6, Rezaee teaches: The method of claim 1, further comprising: before using the RL algorithm to select the first action and obtaining R1, using the RL algorithm to select a second action; ([Rezaee, 0100] The trajectory selector 336 selects the trajectory having the highest calculated evaluation value (i.e., rewards) as the selected trajectory. The value of K may be 2 and 2-1=1 action may be selected. [0105] and [Fig. 8] collectively shows that the action selection and evaluation are performed iteratively (current state is updated based on the next state data). This indicates that the RL algorithm selects actions (i.e., second action) prior to the current action (i.e., the first action)) triggering performance of the selected second action; and ([Rezaee, 0100-0101] discloses performing the action based on a selected action, and obtaining a reward value based on the performance. [0105] and [Fig. 8] collectively shows that the action selection and evaluation are performed iteratively (current state is updated based on the next state data). This indicates that the RL algorithm selects actions (i.e., second action) prior to the current action (i.e., the first action)) after the second action is performed, obtaining a second reward value, R2, associated with the second action, wherein ([Rezaee, 0100-0101] discloses performing the action based on a selected action, and obtaining a reward value based on the performance. [0105] and [Fig. 8] collectively shows that the action selection and evaluation are performed iteratively (current state is updated based on the next state data). This indicates that the RL algorithm selects actions (i.e., second action) prior to the current action (i.e., the first action)) using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises performing a decision process comprising: ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’) However, Rezaee does not specifically disclose: calculating Δ R = R2 - R1; and determining whether Δ R is greater than a drop threshold. Mignon teaches: calculating Δ R = R2 - R1; and ([Mignon, page 1147, para under Algorithm 1, lines 1-12] The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is obtained) determining whether Δ R is greater than a drop threshold. ([Mignon, page 1147, para under Algorithm 1, lines 1-12] The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is compared to 0. If it is greater than zero, a new value is calculated for the ϵ ) Regarding claim 7, Rezaee in view of Mignon teaches: The method of claim 6, wherein using the RL algorithm to select the first action comprises selecting the first action based on an exploration probability, ϵ , the algorithm modification condition is satisfied when Δ R is greater than the drop threshold, and modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new exploration probability, ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ϵ R e S t a r t , where ϵ R e S t a r t is a predetermined exploration probability. ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) . The sigmoid corresponds to the predetermined ϵ r e s t a r t ) Regarding claim 14, Rezaee teaches: The method of claim 12, wherein using R1 and/or PI to determine whether the algorithm modification condition is satisfied comprises: ([Rezaee, 0064], [0065], [0092] and [0104] collectively disclose update the parameters (e.g., weights) of the neural network until termination criteria is met (e.g., performance has reached a minimum threshold). The performance lower than the minimum threshold is interpreted as the ‘modification condition’) using R1 and said K-1 reward values to generate a reward value that is a function of these K reward values; and ([Rezaee, 0079] and [0104] The parameters are updated based on the current estimated evaluation value (i.e., R1) and discounted sum of expected future rewards (i.e., a function of the K-1 reward value). [Abstract] The paragraphs indicate that the reward (i.e., R1) reflects the performance of the selected trajectory) However, Rezaee does not specifically disclose: comparing Δ R to a threshold, wherein Δ R is a difference between the generated reward value and a previously generated reward value. Mignon teaches: comparing Δ R to a threshold, wherein Δ R is a difference between the generated reward value and a previously generated reward value. ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) ) Claims 8 and 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Rezaee in view of Mignon and further in view of Wu et al. (US 20210150310 A1, hereinafter ‘Wu’). Regarding claim 8, Rezaee in view of Mignon teaches: The method of claim 6, wherein the decision process further comprises, as a result of determining that Δ R is not greater than the drop threshold, then determining whether (a reward)[Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) , and if not, the algorithm determines whether the difference is less than the threshold) Rezaee in view of Mignon does not specifically disclose: as a result of determining that Δ R is not greater than the drop threshold, then determining whether R1 is less than a lower reward threshold. Wu teaches: as a result of determining that Δ R is not greater than the drop threshold, then determining whether R1 is less than a lower reward threshold. ([Wu, 0123-0126] and [0136] collectively disclose determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value. The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process. [0079] indicates that the model may be a RL model) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Rezaee, Mignon and Wu to use the method of determining whether to modify a learning parameter based on two or more threshold values of Wu to implement the reinforcement learning method of present invention. The suggestion and/or motivation for doing so is to improve the efficiency of the reinforcement learning method by reducing the amounts of computational resources required to retrain the model parameters. Regarding claim 10, Rezaee in view of Mignon teaches: The method of claim 6, wherein the decision process further comprises, as a result of determining that ΔR is not greater than the drop threshold, then determining ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) , and if not, the algorithm determines whether the difference is less than the threshold) the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) , and if not, the algorithm determines whether the difference is less than the threshold) the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) , and if not, the algorithm determines whether the difference is less than the threshold) However, Rezaee in view of Mignon does not specifically disclose: as a result of determining that ΔR is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold and R1 is greater than the upper reward threshold, and modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new exploration probability, ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ϵ e n d , where ϵ e n d is a predetermined ending exploration probability. Wu teaches: as a result of determining that ΔR is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, ([Wu, 0123-0126] and [0136] collectively disclose determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value (i.e., upper reward threshold). The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process) the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold and R1 is greater than the upper reward threshold, and ([Wu, 0123-0126] discloses determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value (i.e., upper reward threshold). The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process) modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ϵ e n d , where ϵ e n d is a predetermined ending [Wu, 0123-0126] discloses determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value (i.e., upper reward threshold). The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process. [0079] indicates that the model may be a RL model. Generating new exploration probability is taught by the secondary reference [Mignon, page 1147, Algorithm 1]) Regarding claim 11, Rezaee in view of Mignon teaches: The method of claim 6, wherein the decision process further comprises, as a result of determining that Δ R is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) , and if not, the algorithm determines whether the difference is less than the threshold) the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold ([Mignon, page 1147, Algorithm 1, lines 8-13] and [Mignon, page 1147, para under Algorithm 1, lines 1-12] It first selects the action with highest average reward and modifies the value of ϵ . The difference (Δ) between the highest average rewards ( m a x c u r r ) and the highest previous rewards average ( m a x p r e v ) is first compared to the 0 to evaluate if the difference is greater than the 0 (drop threshold). If so, the ϵ is modified to a new ϵ which is s i g m o i d ( Δ ) , and if not, the algorithm determines whether the difference is less than the threshold) … wherein ϵ n e w equals ( ϵ   × c ) , where c is a predetermined constant ([Mignon, page 1148, lines 3-5] The new ϵ is defined by a function: s i g m o i d x = 1.0 1.0 + e x p ( - 2 * x ) - 0.5 . The x denotes the previous epsilon value and the sigmoid(x) denotes the new epsilon value. -2 is the predetermined constant) However, Rezaee in view of Mignon does not specifically disclose: as a result of determining that Δ R is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold and R1 is not greater than the upper reward threshold, and modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new exploration probability, ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ( ϵ   × c ) , where c is a predetermined constant. Wu teaches: as a result of determining that Δ R is not greater than the drop threshold, then determining whether R1 is greater than an upper reward threshold, ([Wu, 0123-0126] and [0136] collectively disclose determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value (i.e., upper reward threshold). The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process) the algorithm modification condition is satisfied when Δ R is not greater than the drop threshold and R1 is not greater than the upper reward threshold, and ([Wu, 0123-0126] discloses determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value (i.e., upper reward threshold). The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process. Generating new exploration probability is taught by the secondary reference [Mignon, page 1147, Algorithm 1]) modifying the RL algorithm as a result of determining that the algorithm modification condition is satisfied comprises generating a new ϵ n e w , for the RL algorithm, wherein ϵ n e w equals ( ϵ   × c ) , where c is a predetermined constant. ([Wu, 0123-0126] discloses determining whether the value of the first loss function (interpreted as Δ R ) is less or greater than the first threshold value and further determining whether the value of the second loss function (interpreted as R1) is less or greater than the second threshold value (i.e., upper reward threshold). The loss functions are ‘reward’ as the reward indicates the performance of the agent. Both may be performed in response to each other to determine whether to terminate or continue the parameter update process. Generating new exploration probability is taught by the secondary reference [Mignon, page 1147, Algorithm 1]) Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Rezaee in view of Wu. Regarding claim 18, Rezaee teaches: The method of claim 1, wherein the method further comprises using the modified RL algorithm to select another action and triggering performance of the another action. ([Rezaee, 0100] The trajectory selector 336 selects the trajectory having the highest calculated evaluation value (i.e., rewards) as the selected trajectory. The value of K may be 2 and 2-1=1 action may be selected. [0105] and [Fig. 8] collectively shows that the action selection and evaluation are performed iteratively (current state is updated based on the next state data). This indicates that the RL algorithm selects actions (i.e., second action) prior to the current action (i.e., the first action)) However, Rezaee does not specifically disclose: one or more of the recited thresholds is dynamically changed based on environment changes and/or service requirement changes. Wu teaches: one or more of the recited thresholds is dynamically changed based on environment changes and/or service requirement changes. ([Wu, 0109] The thresholds may be determined by the system itself, preset by a user or operator via the terminal. This indicates that the threshold may be changed dynamically based on service requirement changes) Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Rezaee and Wu to use the method of determining whether to modify a learning parameter based on two or more threshold values of Wu to implement the reinforcement learning method of present invention. The suggestion and/or motivation for doing so is to improve the efficiency of the reinforcement learning method by reducing the amounts of computational resources required to retrain the model parameters. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUN KWON whose telephone number is (571)272-2072. The examiner can normally be reached Monday – Friday 8:00AM – 5:00PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JUN KWON/Examiner, Art Unit 2127 /ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Mar 06, 2024
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705504
KNOWLEDGE BASE CONSTRUCTION
8y 5m to grant Granted Aug 11, 2026
Patent 12699879
DEFENSE AGAINST ADVERSARIAL EXAMPLE INPUT TO MACHINE LEARNING MODELS
3y 10m to grant Granted Aug 04, 2026
Patent 12645918
TASK SKEW MANAGEMENT FOR NEURAL PROCESSOR CIRCUIT
5y 10m to grant Granted Jun 02, 2026
Patent 12639581
METHOD AND APPARATUS FOR DATA-FREE NETWORK QUANTIZATION AND COMPRESSION WITH ADVERSARIAL KNOWLEDGE DISTILLATION
5y 8m to grant Granted May 26, 2026
Patent 12632739
TEXT-BASED EVENT DETECTION METHOD AND APPARATUS, COMPUTER DEVICE, AND STORAGE MEDIUM
4y 10m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
40%
Grant Probability
87%
With Interview (+46.4%)
4y 8m (~2y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 77 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month