Prosecution Insights
Last updated: October 04, 2026
Application No. 18/560,859

TEMPORAL EQUILIBRIUM ANALYSIS-BASED MULTI-AGENT MULTI-TASK LAYERED METHOD FOR CONTINUOUS CONTROL

Non-Final OA §101§103
Filed
Nov 14, 2023
Priority
Sep 30, 2022 — CN 202211210483.9 +1 more
Examiner
SACKALOSKY, COREY MATTHEW
Art Unit
Tech Center
Assignee
Changzhou University
OA Round
1 (Non-Final)
63%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
29 granted / 46 resolved
+3.0% vs TC avg
Strong +30% interview lift
Without
With
+30.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
25 currently pending
Career history
72
Total Applications
across all art units

Statute-Specific Performance

§101
41.2%
+1.2% vs TC avg
§103
37.3%
-2.7% vs TC avg
§102
12.8%
-27.2% vs TC avg
§112
7.9%
-32.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 46 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 11/14/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claim 1 objected to because of the following informalities: limitation 3 of Claim 1 recites "constructing connection mechanism". Examiner believes it should recite "constructing . Appropriate correction is required. Claim objected to because of the following informalities: step S211 of Claim 3 recites "finite state automata format for systhesizing". Examiner believes it should recite "finite state automata format for synthesizing". Appropriate correction is required. Allowable Subject Matter Claims 2-6 objected to as being dependent upon a rejected base claim, but would be allowable over the prior art if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-6 rejected under 35 U.S.C. 101 because they are directed to abstract ideas without significantly more. Step 1 analysis: Independent Claim 1 recites, in part, a continuous control method, therefore falling into the statutory category of process. Regarding Claim 1: Step 2A: Prong 1 analysis: Claim 1 recites in part: “constructing a multi-agent multi-task game model based on temporal logic, performing temporal equilibrium analysis and synthesizing multi-agent top-level control policies”. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgement, or opinion) or with the aid of pencil and paper. For example, this limitation encompasses initializing parameters for a model. “constructing connection mechanism between the top-level control policies and bottom-level deep deterministic policy gradient algorithms, and constructing multi-agent continuous task controllers based on the connection mechanism”. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgement, or opinion) or with the aid of pencil and paper. For example, this limitation encompasses defining how agents communicate with each other in a model. Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea. Step 2A: Prong 2 analysis: The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of: “constructing a specification auto-completion mechanism, improving dependent task specification by adding environment assumptions”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. The additional element(s) of “constructing a specification auto-completion mechanism, improving dependent task specification by adding environment assumptions” is/are directed to particular field(s) of use (reinforcement learning) (MPEP 2106.05(h)) and therefore do not provide significantly more than the abstract idea, and thus the claim is subject-matter ineligible. Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception. Regarding Claim 2: Step 2A: Prong 1 analysis: Claim 2 recites in part: “Constructing an infeasible region R i G for each agent i , such that the agent i does not have tendency of deviating from the current policy set in the set in which R i G is located, the infeasible region R i G is expressed as follows: R i G = { s ∣ ∃ σ → ⋅ ∀ σ i ⇒ π s , , σ i ⊭ γ i }   where, there exists a policy set σ → in R i G such that all policies σ i and the combination of other policies       , σ i of agent i cannot satisfy γ i . PNG media_image3.png 66 83 media_image3.png Greyscale represents that the policy set does not include the policy combinations of the ith agent; "∃" represents "existence"; "⊭" represents "incompliance"”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. “then computing ⋀ i ∈ L R i G , determining whether there exists a trajectory π in the intersection that satisfies ψ ∧ ⋀ i ∈ W γ i , and using model-checking method to generate the top-level control policy for each agent”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea. Step 2A: Prong 2 analysis: The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of: “The temporal equilibrium analysis-based multi-agent multi-task continuous control method according to claim 1, characterized in that, in step S1, the constructed multi-agent multi-task game model is G = < N a , S , A , S 0 , T r , λ , γ i i ∈ N , ψ >   where, N a represents the agent set, S and A respectively represent the state set and action set of the game model, S 0 is the initial state, T r ∈ S × A → → S represents the state transition function in which all agents in a single state s ∈   S transit to a next state by taking action set a → ∈ A → , A → represents a vector of the action sets of different agents; λ ∈ S → 2 A P represents a labelling function from state to atomic proposition; γ i i ∈ N represents the specification for each agent i; ψ represents the specification that needs to be completed by the overall system”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. The additional element(s) of “The temporal equilibrium analysis-based multi-agent multi-task continuous control method according to claim 1, characterized in that, in step S1, the constructed multi-agent multi-task game model is G = < N a , S , A , S 0 , T r , λ , γ i i ∈ N , ψ >   where, N a represents the agent set, S and A respectively represent the state set and action set of the game model, S 0 is the initial state, T r ∈ S × A → → S represents the state transition function in which all agents in a single state s ∈   S transit to a next state by taking action set a → ∈ A → , A → represents a vector of the action sets of different agents; λ ∈ S → 2 A P represents a labelling function from state to atomic proposition; γ i i ∈ N represents the specification for each agent i; ψ represents the specification that needs to be completed by the overall system” is/are directed to particular field(s) of use (reinforcement learning) (MPEP 2106.05(h)) and therefore do not provide significantly more than the abstract idea, and thus the claim is subject-matter ineligible. Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception. Regarding Claim 3: Step 2A: Prong 1 analysis: Claim 3 recites in part: “S211, computing policies of the negated form of the original specification which acts as policies of finite state automata format for systhesizing ⋀ e = 1 m G F   Ψ e ∧ ¬ ⋀ f = 1 n G F   φ f ; G represents that the specification is always true from the current moment; F represents that the specification will be eventually true at certain moment in the future”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. “then determining whether all of the specification satisfy E a '   ⇒   E b ; if satisfied, completing the refinement of task specification with dependency; if not satisfied, iteratively constructing E a ' and E b until the following formula are satisfied: ⋀ e = 1 m G F Ψ e k 1 ⇒ ⋀ f = 1 n G F φ f k 1   ,   k 1 ∈ N ⋀ e = 1 m G F Ψ e k 2 ⇒ ⋀ f = 1 n G F φ f k 2 ,   k 2 ∈ M   ⇒ ⋀ e = 1 m G F Ψ e k 1 ⇒ ⋀ f = 1 n G F φ f k 1 ∧ E k 1 '   ,   k 1 ∈ N ⋀ e = 1 m G F Ψ e k 2 ∧   E k 2 ⇒ ⋀ f = 1 n G F φ f k 2 ,   k 2 ∈ M   ∀ a , b ⋅ a ∈ N   ∧   b ∈ M   ⇒ ( E a '   ⇒   E b )   where, W represents the set of agents that can satisfy the specification; ⋀ e = 1 m G F Ψ e k 1 represents the e-th assumed specification of agent k1 in the second agent set N; ⋀ f = 1 n G F φ f k 1 represents the f-th guaranteed specification of agent k1 in the second agent set N; ⋀ e = 1 m G F Ψ e k 2 represents the e-th assumed specification of the agent k2 in the second agent set M; ⋀ f = 1 n G F φ f k 2 represents the f-th guarantee rule of agent k2 in the second agent set M”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea. Step 2A: Prong 2 analysis: The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of: “S21, refining task specification by adding environment assumptions”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). “adding environment constraints Ψ of loser L by selecting E ∈ E , automatically generate a new specification using an anti-policy mode, which is expressed as: PNG media_image4.png 392 3921 media_image4.png Greyscale where, E is the environment constraint set; m represents the number of assumed specification in the specification, n represents the number of guaranteed specification (≥ the number of subsequent GF); the value range of e is [1, m], and the value range of f is [1, n]”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). “S212, designing a pattern on the finite state automata that satisfies the form of F G   Ψ e specification”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). “S213, generating a specification according to the generated pattern and perform negation”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). “S22, for a task of a first agent M ⊆ W which is dependent on a task of a second agent N ⊆ W , under the condition of temporal equilibrium, firstly computing policies for all agents through R i G , synthesizing the finite state automata format; then designing patterns which satisfy the form of F G   Ψ e based on policies and using the pattern to generate E a ' ; searching specification refinement set E b of all agents b ∈ M according to step S21”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. The additional element(s) of “S21, refining task specification by adding environment assumptions”, “adding environment constraints Ψ of loser L by selecting E ∈ E , automatically generate a new specification using an anti-policy mode, which is expressed as: PNG media_image4.png 392 3921 media_image4.png Greyscale where, E is the environment constraint set; m represents the number of assumed specification in the specification, n represents the number of guaranteed specification (≥ the number of subsequent GF); the value range of e is [1, m], and the value range of f is [1, n]”, “S212, designing a pattern on the finite state automata that satisfies the form of F G   Ψ e specification”, “S213, generating a specification according to the generated pattern and perform negation”, and “S22, for a task of a first agent M ⊆ W which is dependent on a task of a second agent N ⊆ W , under the condition of temporal equilibrium, firstly computing policies for all agents through R i G , synthesizing the finite state automata format; then designing patterns which satisfy the form of F G   Ψ e based on policies and using the pattern to generate E a ' ; searching specification refinement set E b of all agents b ∈ M according to step S21” is/are directed to particular field(s) of use (reinforcement learning) (MPEP 2106.05(h)) and therefore do not provide significantly more than the abstract idea, and thus the claim is subject-matter ineligible. Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception. Regarding Claim 4: Step 2A: Prong 1 analysis: Claim 4 recites in part: “if ⋀ e = 1 m G F Ψ e ∧ E is reasonable, but there are situations where the specification cannot be realized by the agent after adding environment assumptions, iteratively constructing E ' , such that ⋀ e = 1 m G F Ψ e ∧ E ∧ E ' can be realized”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea. Step 2A: Prong 2 analysis: The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of: “in the case that new specification is generated, determining whether the specification of all agents are reasonable and realizable after adding environment assumptions: if realizable, completing the refinement of specification”. This limitation merely indicates a field of use or technological environment in which the judicial exception is performed (reinforcement learning) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. The additional element(s) of “in the case that new specification is generated, determining whether the specification of all agents are reasonable and realizable after adding environment assumptions: if realizable, completing the refinement of specification” is/are directed to particular field(s) of use (reinforcement learning) (MPEP 2106.05(h)) and therefore do not provide significantly more than the abstract idea, and thus the claim is subject-matter ineligible. Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception. Regarding Claim 5: Step 2A: Prong 1 analysis: Claim 5 recites in part: “S31, according to temporal equilibrium analysis, acquiring policy σ i = < U i , u i 0 , F i , A C i , δ i u , δ i a > of each agent in the game model, expanding the acquired policy as η i = < U i , u i 0 , F i , A C i , δ i u , δ i r > , where δ i r ∈ U i × 2 A P → R , and using it as a reward function in the expanded Markov decision process in a multi-agent environment; the expression of the expanded Markov decision process in a multi-agent environment is as follows: T = < N a , P , Q , h , ζ , L , < η i > i ∈ N > ,   where, N a represents the agent set, P and Q respectively represent the environment states and action set taken by the multi-agent, h represents probability of state transition; ζ represents attenuation coefficient of T ; L ∈ P × Q × P → 2 A P represents labelling function for state transition to atomic propositions, η i represents benefit that the environment obtains when adopting policy of agent i , transferring to p ' ∈ P after agent i taking action q ∈ Q in p ∈ P , its state on η i will also transfer from u ∈ U i ∪ F i to u ' = δ i u u , L p , q , p ' and obtain the reward δ i r u , L p , q , p ' ; "< >" represents a tuple, "∪" represents a union”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. “S32, expanding η i to Markov decision process format with the attenuation function ζ r determined by the state transition, and initializing all δ i r , so that δ i r is 0 when δ i u u , L p , q , p ' ∉ F ; δ i r is 1 when δ i u u , L p , q , p ' ∈ F ; then determining the value function v u * of each state through the value iteration method, and adding the converged v u * to the reward function as a potential energy function, wherein the reward function r p , q , p ' of T is expressed as follows: r ' p , q , p ' = r p , q , p ' + ζ r ( v δ i u u , L p , q , p ' * ) - v u * ”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. “S33, each agent i has an action network μ p θ i with parameters θ , and shares an evaluation network Q p , q → ω , α , β with parameters ω ; constructing a loss function J ω for the evaluation network parameter ω , and updating the network according to the gradient backpropagation of the network. The expression of the loss function J ω is as follows: J ω = 1 d ∑ t = 1 d r t + ζ Q ' p t + 1 , q t + 1 → + ϵ ω ' , α ' , β ' - Q p t , q t → ω , α , β 2 where, r t is the reward value computed in step S32, Q p , q → ω , α , β = A p , q → ω , α + V p ω , β , A p , q → ω , α and V p ω , β are designed as fully connected layer networks to evaluate the state value and action advantage respectively. α and β are the parameters of the two networks respectively; d is randomly sampled data from experience playback buffer data set D”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea. Step 2A: Prong 2 analysis: The judicial exception is not integrated into practical application. In particular, the claim recites the additional elements of: “finally soft-updating the target evaluation network parameter and action network parameters respectively according to the evaluation network parameters ω and action network parameters θ i ”. This additional element is recited at a high level of generality such that the claim recites only the idea of a solution or outcome (updating a model) i.e., the claim fails to recite details of how a solution to a problem is accomplished. Accordingly at Step 2A: Prong 2, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional element(s) of “finally soft-updating the target evaluation network parameter and action network parameters respectively according to the evaluation network parameters ω and action network parameters θ i ” is/are recited at a high-level of generality such that the claim recites only the idea of a solution or outcome (updating a model) i.e., the claim fails to recite details of how a solution to a problem is accomplished (See MPEP 2106.05(f)). Accordingly, at Step 2B, the additional elements individually or in combination do not amount to significantly more than the judicial exception. Regarding Claim 6: Step 2A: Prong 1 analysis:Claim 6 recites in part: “The temporal equilibrium analysis-based multi-agent multi-task continuous control method according to claim 5, characterized in that, when the hetero-policy algorithm is used for gradient update, estimating the expected value of ∇ q i t Q   ⋅ ∇ θ i μ according to the Monte Carlo method, and substituting the randomly sampled data into the following formula to perform unbiased estimation: ∇ θ i J θ i ≈ 1 d ∑ t = 1 d ∇ q i t Q p t , q t → ω ∇ θ i μ p t θ i   where, ∇ represents the differential operator”. As drafted and under its broadest reasonable interpretation, this limitation covers a mathematical calculation. Accordingly, at Step 2A: Prong 1, the claim is directed to an abstract idea. Step 2A: Prong 2 analysis: The claim does not recite any additional elements that integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1 is/are rejected under 35 U.S.C. 103 as being unpatentable over Leon et al (León, B. G., & Belardinelli, F. (2020). Extended Markov Games to Learn Multiple Tasks in Multi-Agent Reinforcement Learning. arXiv [Cs.AI]. Retrieved from http://arxiv.org/abs/2002.06000, hereinafter Leon), in view of Van Seijen et al (US 20180165603 A1, hereinafter Van Seijen), and in further view of Shalev-Shwartz et al (Shalev-Shwartz, S., Shammah, S., & Shashua, A. (2016). Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving. arXiv [Cs.AI]. Retrieved from http://arxiv.org/abs/1610.03295, hereinafter Shwartz). Regarding Claim 1: Leon teaches A temporal equilibrium analysis-based multi-agent multi-task continuous control method, characterized in comprising the following steps: S1, constructing a multi-agent multi-task game model based on temporal logic, performing temporal equilibrium analysis and synthesizing multi-agent top-level control policies (Leon [Page 1, Section 1, par. 7]: "Our work focus on extending what [8, 34] developed for the single-agent framework to multi-agent systems, where reinforcement learning guided by temporal specifications could prove to be an effective tool to address some of the complex challenges mentioned above, including scalability, stationarity or equilibrium of solutions."); S2, constructing a specification auto-completion mechanism, improving dependent task specification by adding environment assumptions (Leon [Page 5, Section 3.2, Example 2]: "In order to model our problem as an Extended Markov Game, we transform the LTL specification of making shears into a DFA whose states would be given to the agents as an extension of the original observation that they perceive from the environment. The agents would then begin the episodes by perceiving an observation of the map extended by the initial state of the automaton, that represents the whole specification of making shears to be fulfilled. Once the agents have progressed the specification, which in this case means that they got iron or wood, the automaton will transit to a new state that represents the remainder of the specification to be fulfilled."; (EN): the transition to a new state "that represents the remainder of the specification" is analogous to the specification auto completion mechanism as the agents move to a new state representing the remainder of the specification); Leon does not distinctly disclose S3, constructing connection mechanism between the top-level control policies and bottom-level deep deterministic policy gradient algorithms, and constructing multi-agent continuous task controllers based on the connection mechanism. However, the combination of Van Seijen and Shwartz teaches S3, constructing connection mechanism between the top-level control policies and bottom-level deep deterministic policy gradient algorithms, and constructing multi-agent continuous task controllers based on the connection mechanism (Van Seijen [0107]: "In an example, agents can be organized in a way that decomposes a task hierarchically. For instance, there can be three agents where Agent 0 is a top-level agent, and Agent 1 and Agent 2 are each bottom-level agents. The top-level agent only has communication actions, specifying which of the bottom level agents is in control."), (Shwartz [Page 8, Section 5, par. 5]: "An immediate benefit of the options graph is the interpretability of the results. Another immediate benefit is that we rely on the decomposable structure of the set D and therefore the policy at each node should choose between a small number of possibilities. Finally, the structure allows us to reduce the variance of the policy gradient estimator."). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the concept of extended Markov games of Leon with the techniques for multi-agent, hybrid reinforcement learning of Shwartz + Van Seijen in order to provide a set of control policies coupled with a set of deterministic policy gradients in order to construct continuous learning controllers for the multi-agent learning scheme (Van Seijen [0003]: “Specifically, there has been work on hierarchical RL methods, which decompose a task into hierarchical subtasks. Hierarchical learning can help accelerate learning on individual tasks by mitigating the exploration challenge of sparse-reward problems. One popular framework for this is the options framework, which extends the standard RL framework based on Markov decision processes (MDP) to include temporally-extended actions.”; Shwartz [Abstract]: “First, we show how policy gradient iterations can be used, and the variance of the gradient estimation using stochastic gradient ascent can be minimized, without Markovian assumptions. Second, we decompose the problem into a composition of a Policy for Desires (which is to be learned) and trajectory planning with hard constraints (which is not learned).”) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Ivanov, S., & D’yakonov, A. (2019). Modern Deep Reinforcement Learning Algorithms. arXiv [Cs.LG]. Retrieved from http://arxiv.org/abs/1906.10025 – In this work latest DRL algorithms are reviewed with a focus on their theoretical justification, practical limitations and observed empirical properties Any inquiry concerning this communication or earlier communications from the examiner should be directed to COREY M SACKALOSKY whose telephone number is (703)756-1590. The examiner can normally be reached M-F 7:30am-3:30pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571) 272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /COREY SACKALOSKY/Examiner, Art Unit 2128 /OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Nov 14, 2023
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748983
Identifying and Correcting Label Bias in Machine Learning
5y 4m to grant Granted Sep 29, 2026
Patent 12748948
INFERENCE SYSTEM, INFERENCE DEVICE, AND INFERENCE METHOD
4y 5m to grant Granted Sep 29, 2026
Patent 12748959
NEURAL NETWORK SCHEDULING METHOD AND APPARATUS
3y 10m to grant Granted Sep 29, 2026
Patent 12737665
ONLINE MACHINE LEARNING-BASED MODEL FOR DECISION RECOMMENDATION
6y 0m to grant Granted Sep 15, 2026
Patent 12737611
CLASSIFYING ELEMENTS AND PREDICTING PROPERTIES IN AN INFRASTRUCTURE MODEL THROUGH PROTOTYPE NETWORKS AND WEAKLY SUPERVISED LEARNING
5y 4m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
63%
Grant Probability
93%
With Interview (+30.3%)
4y 2m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 46 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month