Prosecution Insights
Last updated: October 02, 2026
Application No. 18/393,464

ALGORITHM SYSTEM OF DEEP REINFORCEMENT LEARNING AND ALGORITHM METHOD THEREOF

Non-Final OA §103
Filed
Dec 21, 2023
Priority
Nov 17, 2023 — TW 112144444
Examiner
KARTHOLY, REJI P
Art Unit
Tech Center
Assignee
Industrial Technology Research Institute
OA Round
1 (Non-Final)
63%
Grant Probability
Moderate
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
104 granted / 164 resolved
+3.4% vs TC avg
Strong +70% interview lift
Without
With
+70.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
12 currently pending
Career history
177
Total Applications
across all art units

Statute-Specific Performance

§101
14.4%
-25.6% vs TC avg
§103
61.5%
+21.5% vs TC avg
§102
10.1%
-29.9% vs TC avg
§112
12.5%
-27.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 164 resolved cases

Office Action

§103
DETAILED ACTION This Office Action is in response to Applicant's Communication received on 12/21/2023 for application number 18/393,464. Claims 1-11 are presented for examination. Claims 1 and 7 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 12/21/2023 and 6/3/2024 have been considered by the Examiner. Claim Objections Claims 1 and 7 are objected to because of the following informalities: In these claims, “the network update processes” and “the experience collection processes” should be “the network update process” and “the experience collection process” to be consistent. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-11 are rejected under 35 U.S.C. 103 as being unpatentable over Soyer et al. (US 2021/0034970 A1 hereinafter Soyer) in view of Mnih et al. (US 2021/0166127 A1 hereinafter Mnih). Regarding Claim 1, Soyer teaches an algorithm system for deep reinforcement learning ([0040] the distributed training system separates acting from learning by using multiple actor computing units to generate experience tuple trajectories which are processed by one or more learner computing units to train an action selection network; reinforcement learning system), comprising: a memory disposed to store a previous state of an environment, a previous policy of a model, an inference program, and a training program ([0017] the actor operations comprise storing the trajectory of experience tuples in a queue, where the queue is accessible to each of the actor computing units, and the queue comprises an ordered sequence of different experience tuple trajectories (i.e., previous state); [0012] the policy used to generate a trajectory can lag behind the policy on the learner by several updates (i.e., previous policy); [0040] the distributed training system separates acting (i.e., inference) from learning (i.e., training) by using multiple actor computing units to generate experience tuple trajectories which are processed by one or more learner computing units to train an action selection network; [0066] the learner computing units each maintain a respective “learner” action selection neural network, and the actor computing units each maintain a respective “actor” action selection neural network; [0068] the actor computing units provide the generated experience tuple trajectories 206 to the learner computing units by storing the experience tuple trajectories 206 in a data store that is accessible to each of the actor and learner computing units; [0106] one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus; [0111] one or more memory devices for storing instructions and data); an input/output interface ([0041] action 104 to be performed by the agent 106 in response to the received data; the state of the environment 108 characterized by observation 112; [0049] the actions may be control inputs to control the robot - thus, observations are received from and actions are output/ applied to environment to control the robot); and a processor coupled to the memory and the input/output interface ([0067] a computing unit may be, e.g., a computer, a core within a computer having multiple cores; the computing units include processor cores, , processors, microprocessors; [0111] the essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data) to: perform initialization of the environment and the model through the input/output interface ([0065] FIG. 2 shows training system 200; [0066] the training system 200 is a distributed computing system which includes one or more learner computing units and multiple actor computing units; the learner computing units each maintain a respective “learner” action selection neural network, and the actor computing units each maintain a respective “actor” action selection neural network; the learner computing units are configured to train a set of shared learner action selection network parameter values using reinforcement learning techniques based on trajectories of experience tuples generated by the actor computing units using the actor action selection networks; [0009] current parameter values of the action selection neural network; [0068] each actor computing unit is configured generate trajectories of experience tuples that characterize the interaction of an agent with an instance of an environment; the actor computing units may provide the generated experience tuple trajectories 206 to the learner computing units - thus, setting up an instance of the environment and network parameters); read the inference program and the training program from the memory, wherein the inference program corresponds to an experience collection process and the training program corresponds to a network update process ([0040] the distributed training system separates acting from learning by using multiple actor computing units to generate experience tuple trajectories which are processed by one or more learner computing units to train an action selection network; [0066] the learner computing units each maintain a respective “learner” action selection neural network, and the actor computing units each maintain a respective “actor” action selection neural network; [0068] each actor computing unit is configured generate trajectories of experience tuples that characterize the interaction of an agent with an instance of an environment by performing actions that are selected using the actor action selection network maintained by the actor computing unit; actor computing units may provide the generated experience tuple trajectories to the learner computing units by storing the experience tuple trajectories (i.e., experience collection process) in a data store that is accessible to each of the actor and learner computing units; [0069] after obtaining a batch of experience tuple trajectories, a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters (i.e., network update process) based on the batch of experience tuple trajectories - thus, the acting/ experience collection code being executed by the actor computing units and the learning/ network update code being executed by the learner computing units, as distinct sets of program instructions; execute the experience collection process and the network update process, and determine whether the network update process meet a termination condition ([0068] actor computing units may provide the generated experience tuple trajectories to the learner computing units by storing the experience tuple trajectories in a data store that is accessible to each of the actor and learner computing units; [0069] after obtaining a batch of experience tuple trajectories, a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters based on the batch of experience tuple trajectories; [0030] each learner computing unit can efficiently process batches of experience tuple trajectories in parallel; [0096] after adjusting the current parameter values of the action selection network and the state value network, the system can determine whether a training termination criterion is met); continue executing the experience collection process and the network update process in parallel in response to neither of the experience collection process and the network update processes has met the termination condition ([0068] actor computing units may provide the generated experience tuple trajectories to the learner computing units by storing the experience tuple trajectories in a data store that is accessible to each of the actor and learner computing units; [0069] after obtaining a batch of experience tuple trajectories, a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters based on the batch of experience tuple trajectories; [0087] the system obtains an experience tuple trajectory (402); [0092] the system adjusts the current parameter values of the action selection network (408); [0096] in response to determining that a training termination criterion is not met, the system returns to step 402 and repeats the preceding steps); and stop executing the experience collection process and the network update process in response to one of the experience collection processes and the network update process having met the termination condition ([0096] in response to determining that a training termination criterion is not met, the system returns to step 402 and repeats the preceding steps; in response to determining that a training termination criterion is met, the system can output the trained values of the action selection network parameters); wherein the experience collection process comprises: obtaining a current state of the environment through the input/output interface; wherein the current state comprises a current reward value and a current observation value ([0017] generating an experience tuple comprise receiving an observation characterizing a current state of an instance of the environment (i.e., current state comprising current observation value); generating an experience tuple may further comprise obtaining transition data including: (i) a subsequent observation characterizing a subsequent state of the environment instance subsequent to the agent performing the selected action and (ii) a reward received subsequent to the agent performing the selected action - thus, the reward paired with that time step's observation and action; [0041] at each time step, the action selection network 102 processes data characterizing the current state of the environment 108 to generate policy scores 110 that are used to select an action 104 to be performed by the agent 106 in response to the received data; [0042] the agent 106 may receive a reward 114 based on the current state of the environment 108); calculating to determine a current action based on the current observation value according to a current policy of the model ([0017] determining, using the actor action selection neural network, in accordance with current parameter values of the actor action selection neural network, and based on the observation, a selected action to be performed by the agent and a policy score for the selected action); and returning the current action to the environment through the input/output interface ([0068] each actor computing unit is configured generate trajectories of experience tuples that characterize the interaction of an agent with an instance of an environment by performing actions that are selected using the actor action selection network maintained by the actor computing unit; [0049] the actions may be control inputs to control the robot - thus, returning current action to environment); wherein the network update process comprises: obtaining the previous state of the environment and the previous policy of the model from the memory, wherein the previous state comprises a previous action, a previous reward value, and a previous observation value ([0069] each of the learner computing units are configured to obtain batches of experience tuple trajectories generated by the actor computing units; [0017] generating an experience tuple comprise receiving an observation characterizing a current state of an instance of the environment; generating an experience tuple may further comprise obtaining transition data including: (i) a subsequent observation characterizing a subsequent state of the environment instance subsequent to the agent performing the selected action and (ii) a reward received subsequent to the agent performing the selected action;[0041] at each time step, the action selection network 102 processes data characterizing the current state of the environment 108 to generate policy scores 110 that are used to select an action 104 to be performed by the agent 106 in response to the received data; [0042] the agent 106 may receive a reward 114 based on the current state of the environment 108; [0060] the observation at any given time step may include data from a previous time step that may be beneficial in characterizing the environment, e.g., the action performed at the previous time step, the reward received at the previous time step, and so on (i.e., previous state comprising previous action, previous reward value, previous observation value); [0012] the policy used to generate a trajectory can lag behind the policy on the learner by several updates (i.e., previous policy)); calculating based on the previous state to determine a current data ([0069] a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters based on the batch of experience tuple trajectories (the determined update being the current data)); and updating the previous policy of the model to the current policy based on the current data ([0066] the learner computing units are configured to train a set of shared learner action selection network parameter values using reinforcement learning techniques; [0069] determine updates to the learner action selection network parameters; [0070] the actor computing units can update the values of the actor action selection network parameters by obtaining the current learner action selection network parameters; [0072] action selection policy as defined by the parameter values of the actor action selection networks; [0092] the system adjusts the current parameter values of the action selection network - thus, updating the policy based on current data). Soyer does not expressly teach wherein execute the experience collection process and the network update process in parallel and determine whether the experience collection process meet a termination condition. However, Soyer discloses that each learner computing unit can efficiently process batches of experience tuple trajectories in parallel; rather than consecutively processing the observations included in each experience tuple, a learner computing unit can use a learner action selection network (or a state value network) to process the observations in parallel (see [0030], [0073]), which implies processing batches of experience tuple trajectories in parallel while the actors generate trajectories. In the same field of endeavor, Mnih teaches wherein execute the experience collection process and the network update process in parallel ([0026] each of the workers executes in a separate thread, process or other hardware or software within the computer capable of independently performing the computation for the worker (i.e., inference program); [0008] parallelizing the training using multiple workers operating independently on a single machine; [0049] the reinforcement learning system uses the deep neural network to select values to be performed by the agent while the workers continue to perform the process 200 (i.e., inference parallel with training); [0097] multitasking and parallel processing advantageous) and determine whether the experience collection process meet a termination condition ([0070] the worker receives observations characterizing the state of the environment replica and selects actions to be performed by the actor in accordance with the current values of the parameters of the policy neural network until the environment replica transitions into a state that satisfies particular criteria (i.e., termination condition)). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated wherein execute the experience collection process and the network update process in parallel and determine whether the experience collection process meet a termination condition, as suggested in Mnih into Soyer. Doing so would be desirable because it would allow for the neural network used by a reinforcement learning system to be trained faster and memory requirements for the training can be reduced (Mnih [0008]). As to dependent Claim 2, Soyer and Mnih teach all the limitations of claim 1. Soyer further teaches wherein an inference processing module disposed to read the inference program from the memory and execute the experience collection process ([0066] the actor computing units each maintain a respective “actor” action selection neural network; [0067] a computing unit may be a computer, a core within a computer having multiple cores, or other hardware or software; [0068] the actor computing units provide the generated experience tuple trajectories 206 to the learner computing units by storing the experience tuple trajectories 206 in a data store that is accessible to each of the actor and learner computing units; [0106] one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus - thus, the actor computing units maintain actor action selection neural network and generate experience (i.e., read the inference program and execute the experience collection process)); and a training processing module disposed to read the training program from the memory and execute the network update process ([0066] the learner computing units are configured to train a set of shared learner action selection network parameter values using reinforcement learning techniques based on trajectories of experience tuples generated by the actor computing units using the actor action selection networks; [0067] a computing unit may be a computer, a core within a computer having multiple cores, or other hardware or software; [0069] after obtaining a batch of experience tuple trajectories, a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters based on the batch of experience tuple trajectories; [0106] one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus - thus, the learner computing units train the learner action selection neural network/ update the learner action selection network parameters (i.e., read the training program and execute the network update process)). As to dependent Claim 3, Soyer and Mnih teach all the limitations of claim 1. Mnih further teaches wherein when the processor executes the experience collection process, the processor is further disposed to: determine whether the number of executions of the experience collection process reaches an execution number threshold ([0070] the worker receives observations characterizing the state of the environment replica and selects actions to be performed by the actor in accordance with the current values of the parameters of the policy neural network until the environment replica transitions into a state that satisfies particular criteria; the particular criteria may be satisfied after a predetermined number of observations have been received (i.e., executions of the experience collection process reaches an execution number threshold)); and determine that the experience collection process reaches the termination condition in response to the number of the executions reaching the execution number threshold ([0070] the worker receives observations characterizing the state of the environment replica and selects actions to be performed by the actor in accordance with the current values of the parameters of the policy neural network until the environment replica transitions into a state that satisfies particular criteria; the particular criteria may be satisfied after a predetermined number of observations have been received). As to dependent Claim 4, Soyer and Mnih teach all the limitations of claim 1. Soyer further teaches wherein when the processor executes the network update process, the processor is further disposed to: determine whether the number of executions of the network update process reaches an execution number threshold ([0069] after obtaining a batch of experience tuple trajectories, a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters based on the batch of experience tuple trajectories; [0096] after adjusting the current parameter values of the action selection network and the state value network, the system can determine whether a training termination criterion is met; determine that a training termination criterion is met if the system has performed a predetermined number of training iterations (i.e., executions of the network update process reaches an execution number threshold)); and determine that the network update process reaches the termination condition in response to the number of the executions reaching the execution number threshold ([0069] after obtaining a batch of experience tuple trajectories, a learner computing unit uses a reinforcement learning technique to determine updates to the learner action selection network parameters based on the batch of experience tuple trajectories; [0096] after adjusting the current parameter values of the action selection network and the state value network, the system can determine whether a training termination criterion is met; determine that a training termination criterion is met if the system has performed a predetermined number of training iterations). As to dependent Claim 5, Soyer and Mnih teach all the limitations of claim 1. Soyer further teaches wherein when the processor executes the experience collection process, the processor is further disposed to: after the environment receives the current action through the input/output interface, determine whether a success rate corresponding to the current state of the environment reaches a success rate threshold ([0068] each actor computing unit is configured generate trajectories of experience tuples that characterize the interaction of an agent with an instance of an environment by performing actions that are selected using the actor action selection network maintained by the actor computing unit; [0049] the actions may be control inputs to control the robot; [0042] the agent 106 may receive a reward 114 based on the current state of the environment 108 and the action 104 of the agent 106 at the time step; the reward 114 may indicate whether the agent 106 has accomplished a task - thus, each action in environment yields a per-state indication of task success; [0096] the system may determine that a training termination criterion is met if the performance of an agent in completing one or more tasks using the current values of the action selection network parameters satisfies a threshold (i.e., success rate threshold)); and determine that the experience collection process reaches the termination condition in response to the success rate reaching the success rate threshold ([0030] each learner computing unit can efficiently process batches of experience tuple trajectories in parallel; [0073] rather than consecutively processing the observations included in each experience tuple, a learner computing unit can use a learner action selection network (or a state value network) to process the observations in parallel; [0096] in response to determining that a training termination criterion is met, the system can output the trained values of the action selection network parameters. The same iterative process obtains an experience tuple trajectory at step 402 each iteration (i.e., collection experience), meeting the performance threshold/ success rate threshold ends the loop that generates the experience - so the experience collection process reaches its termination condition). As to dependent Claim 6, Soyer and Mnih teach all the limitations of claim 1. Soyer further teaches wherein when the processor executes the network update process, the processor is further disposed to: calculate to determine the current action based on the previous observation value according to the current policy of the model ([0009] the adjusting include, for each experience tuple: determining, using the action selection neural network, in accordance with current parameter values of the action selection neural network, and based on the observation included in the experience tuple, a learner policy score; [0017] determining, using the actor action selection neural network, in accordance with current parameter values of the actor action selection neural network (i.e., current policy), and based on the observation, a selected action to be performed by the agent and a policy score for the selected action; [0060] the observation at any given time step may include data from a previous time step that may be beneficial in characterizing the environment, e.g., the action performed at the previous time step, the reward received at the previous time step, and so on - thus, the current action is determined based on data from previous step and current policy); when the environment receives the current action through the input/output interface, determine whether a success rate corresponding to the current state of the environment reaches a success rate threshold ([0068] each actor computing unit is configured generate trajectories of experience tuples that characterize the interaction of an agent with an instance of an environment by performing actions that are selected using the actor action selection network maintained by the actor computing unit; [0049] the actions may be control inputs to control the robot; [0042] the agent 106 may receive a reward 114 based on the current state of the environment 108 and the action 104 of the agent 106 at the time step; the reward 114 may indicate whether the agent 106 has accomplished a task - thus, each action in environment yields a per-state indication of task success; [0096] the system may determine that a training termination criterion is met if the performance of an agent in completing one or more tasks using the current values of the action selection network parameters satisfies a threshold (i.e., success rate threshold)); and determine that the experience collection process reaches the termination condition in response to the success rate reaching the success rate threshold ([0030] each learner computing unit can efficiently process batches of experience tuple trajectories in parallel; [0073] rather than consecutively processing the observations included in each experience tuple, a learner computing unit can use a learner action selection network (or a state value network) to process the observations in parallel; [0096] in response to determining that a training termination criterion is met, the system can output the trained values of the action selection network parameters. The same iterative process obtains an experience tuple trajectory at step 402 each iteration (i.e., collection experience), meeting the performance threshold/ success rate threshold ends the loop that generates the experience - so the experience collection process reaches its termination condition). Claims 7-11 are method claims corresponding to the system claims 1, 3, 5, 4, and 6 respectively and therefore, rejected for the same reasons. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Applicant is required under 37 CFR § 1.111(c) to consider these references fully when responding to this action. Kartal et al. (US 2020/0143206 A1) teaches: A3C is an algorithm that employs a parallelized asynchronous training scheme for efficiency; A3C allows multiple workers to simultaneously interact with the environment and compute gradients locally; all the workers pass their computed local gradients to a global network which performs the optimization and synchronizes the updated actor-critic neural network parameters with the workers asynchronously. A3C method, as an actor critic algorithm, has a policy network (actor) and a value network (critic) where actor is parameterized and critic is parameterized, which are updated (see [0051], [0073]). Any inquiry concerning this communication or earlier communications from the examiner should be directed to REJI KARTHOLY whose telephone number is (571)272-3432. The examiner can normally be reached on Monday - Thursday from 7:30 am to 3:30 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch, can be reached at telephone number 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from Patent Center. Status information for published applications may be obtained from Patent Center. Status information for unpublished applications is available through Patent Center for authorized users only. Should you have questions about access to Patent Center, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form. /REJI KARTHOLY/Primary Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Dec 21, 2023
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744835
CROSS-PLATFORM DATA REFRESH FOR COMMUNICATION PROCESS FLOWS
3y 9m to grant Granted Sep 22, 2026
Patent 12725035
METHOD OF AND APPARATUS FOR MACHINE LEARNING IN A RADIO NETWORK
3y 10m to grant Granted Sep 01, 2026
Patent 12694287
PRIVACY-PRESERVING FEDERATED MACHINE LEARNING
4y 3m to grant Granted Jul 28, 2026
Patent 12694308
NEIGHBORHOOD-BASED LINK PREDICTION FOR RECOMMENDATION SYSTEMS
3y 7m to grant Granted Jul 28, 2026
Patent 12682010
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND STORAGE MEDIUM FOR PERFORMING EMEDDING ON DATA OF A GRAPH STRUCTURE
3y 10m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+70.1%)
3y 2m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 164 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month