Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The Amendment filed 06/29/2026 has been entered. Claims 1-20 remain pending in this application.
Information Disclosure Statement
The information disclosure statements submitted on 03/30/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2 and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Dasgupta et al. (US 11080586 B2 hereinafter Dasgupta) in view of VAN HASSELT et al. (US 20210232970 A1 hereinafter VADORI) and BURHANI et al. (US 20190370649 A1 hereinafter BURHANI)
As to independent claim 1, Dasgupta teaches a computer-implemented method for training a reinforcement learning neural network, the method comprising: [performs reinforcement learning in a neural network Col. 3 ln. 44-58]
obtaining observations of states of an environment; [sensors receive observations from an environment like software or a game Col. 4 ln. 9-19 "The observations may be observed through sensors"]
processing the observations to select actions to be performed by an agent in response to the observations, wherein the agent receives rewards in response to the actions; and [selects and causes actions to be performed and gets rewards Col. 4 ln. 41-54 "select the possible action that yields the largest reward probability from the probability function."]
at each of a plurality of training time steps: [action/observation and update iterations (steps) Col. 9-10 ln. 66-25 "operational flow of FIG. 4 is iteratively performed"]
determining a temporal difference error that is dependent upon at least a difference between one of the rewards and a value estimate generated by the reinforcement learning neural network; [calculates TD error based on SARSA which includes rewards (Rt+1 in Eq. 8) Col. 11 ln. 33-43, Col. 5 ln. 12-16 "(25) Calculating section 109 may calculate parameters. For example, calculating section 109 may be operable to calculate a temporal difference error based on the previous action-value, the current action-value, and the plurality of parameters of neural network 120."]
Dasgupta does not specifically teach determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and updating parameters of the reinforcement learning neural network using the scaled temporal difference error.
However, van Hasselt teaches determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and [scaling rewards via normalization (scaling factor ¶87) ¶95 "computing reward 304 includes computing an un-normalized reward and normalizing the un-normalized reward using the calculated mean and standard deviation of the delta VWAPs"]
updating parameters of the reinforcement learning neural network using the scaled temporal difference error. [train using a normalized by a scaling factor (updates with applied normalizations ¶101-102) in reinforcement learning ¶126 "training engine 118 can train the reinforcement learning network 110 using the normalized order count. The total volume of the order can be normalized by dividing the total volume by a scaling factor"]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta by incorporating the determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and scaling the temporal difference error by the scale factor to determine a scaled temporal difference error, and updating parameters of the reinforcement learning neural network using the scaled temporal difference error disclosed by van Hasselt because both techniques address the same field of reinforcement learning and by incorporating van Hasselt into Dasgupta improves recommendations so they are receive more favorable reactions [van Hasselt ¶20]
Dasgupta and van Hasselt do not specifically teach scaling the temporal difference error by the scale factor to determine a scaled temporal difference error;
However, BURHANI teaches scaling the temporal difference error by the scale factor to determine a scaled temporal difference error; [Normalizes target and neural output then computes difference (scaled temporal difference error) ¶29, Fig. 3 306 ¶45 "determines an error for the training item using the normalized target output and the normalized output (step 308). The manner in which the system calculates the error is dependent on the objective function being optimized. For example, for a mean-squared loss function, the system determines the difference between the normalized output and the target output."]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta and VAN HASSELT by incorporating the scaling the temporal difference error by the scale factor to determine a scaled temporal difference error disclosed by BURHANI because all techniques address the same field of reinforcement learning and by incorporating BURHANI into Dasgupta and VAN HASSELT speeds up reinforcement learning for more optimal and metric based results [BURHANI ¶42]
As to dependent claim 2, the rejection of claim 1 is incorporated Dasgupta, VAN HASSELT and BURHANI further teach wherein the temporal difference error is further dependent upon a time discounted value estimate generated by the reinforcement learning neural network system, and wherein the square of the scale factor includes a second term, the method further comprising determining a value for the second term by: determining an estimate of a variance of a time discount factor of the time discounted value estimate; [Dasgupta time based discount factor Col. 11 ln. 33-45 "taking A.sub.t, γ is the discount factor for future reward, and η is the learning rate"]
determining an estimate of an expectation value of returns-squared, wherein a return comprises a time discounted sum of one or more rewards received after a reinforcement learning time step; and [BURHANI mean and standard deviation of rewards include a sum ¶13-14]
forming a product of the estimate of the variance of a time discount factor and the estimate of the expectation value of the returns-squared. [BURHANI formula with deviation time and estimate performance ¶80-81]
As to dependent claim 16, the rejection of claim 1 is incorporated Dasgupta, VAN HASSELT and BURHANI further teach wherein the agent is a mechanical agent, the environment is a real-world environment, and the actions are actions taken by the mechanical agent in the real-world environment to perform the task. [Dasgupta agent model for robotics or machines (real-work environment and task) Col. 15 ln. 29-47 "A neural network in accordance with the present invention can be used for a myriad of applications"…"robotics"]
As to independent claim 17, Dasgupta teaches One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations [computer, storage media and programs Col. 14 ln. 26-44] for training a reinforcement learning neural network system the operations comprising: [performs reinforcement learning in a neural network Col. 3 ln. 44-58]
obtaining observations of states of the environment; [sensors receive observations from an environment like software or a game Col. 4 ln. 9-19 "The observations may be observed through sensors"]
processing the observations to select actions to be performed by the agent in response to the observations, wherein the agent receives rewards in response to the actions; [selects and causes actions to be performed and gets rewards Col. 4 ln. 41-54 "select the possible action that yields the largest reward probability from the probability function."]
at each of a plurality of training time steps: [action/observation and update iterations (steps) Col. 9-10 ln. 66-25 "operational flow of FIG. 4 is iteratively performed"]
determining a temporal difference error that is dependent upon at least a difference between one of the rewards and a value estimate generated by the reinforcement learning neural network; [calculates TD error based on SARSA which includes rewards (Rt+1 in Eq. 8) Col. 11 ln. 33-43, Col. 5 ln. 12-16 "(25) Calculating section 109 may calculate parameters. For example, calculating section 109 may be operable to calculate a temporal difference error based on the previous action-value, the current action-value, and the plurality of parameters of neural network 120."]
Dasgupta does not specifically teach determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and updating parameters of the reinforcement learning neural network using the scaled temporal difference error.
However, van Hasselt teaches determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and [scaling rewards via normalization (scaling factor ¶87) ¶95 "computing reward 304 includes computing an un-normalized reward and normalizing the un-normalized reward using the calculated mean and standard deviation of the delta VWAPs"]
updating parameters of the reinforcement learning neural network using the scaled temporal difference error. [train using a normalized by a scaling factor (updates with applied normalizations ¶101-102) in reinforcement learning ¶126 "training engine 118 can train the reinforcement learning network 110 using the normalized order count. The total volume of the order can be normalized by dividing the total volume by a scaling factor"]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta by incorporating the determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and scaling the temporal difference error by the scale factor to determine a scaled temporal difference error, and updating parameters of the reinforcement learning neural network using the scaled temporal difference error disclosed by van Hasselt because both techniques address the same field of reinforcement learning and by incorporating van Hasselt into Dasgupta improves recommendations so they are receive more favorable reactions [van Hasselt ¶20]
Dasgupta and van Hasselt do not specifically teach scaling the temporal difference error by the scale factor to determine a scaled temporal difference error;
However, BURHANI teaches scaling the temporal difference error by the scale factor to determine a scaled temporal difference error; [Normalizes target and neural output then computes difference (scaled temporal difference error) ¶29, Fig. 3 306 ¶45 "determines an error for the training item using the normalized target output and the normalized output (step 308). The manner in which the system calculates the error is dependent on the objective function being optimized. For example, for a mean-squared loss function, the system determines the difference between the normalized output and the target output."]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta and VAN HASSELT by incorporating the scaling the temporal difference error by the scale factor to determine a scaled temporal difference error disclosed by BURHANI because all techniques address the same field of reinforcement learning and by incorporating BURHANI into Dasgupta and VAN HASSELT speeds up reinforcement learning for more optimal and metric based results [BURHANI ¶42]
As to independent claim 18, Dasgupta teaches a system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to [computer, storage media and programs Col. 14 ln. 26-44] perform operations for training a reinforcement learning neural network system the operations comprising: [performs reinforcement learning in a neural network Col. 3 ln. 44-58]
obtaining observations of states of the environment; [sensors receive observations from an environment like software or a game Col. 4 ln. 9-19 "The observations may be observed through sensors"]
processing the observations to select actions to be performed by the agent in response to the observations, wherein the agent receives rewards in response to the actions; [selects and causes actions to be performed and gets rewards Col. 4 ln. 41-54 "select the possible action that yields the largest reward probability from the probability function."]
at each of a plurality of training time steps: [action/observation and update iterations (steps) Col. 9-10 ln. 66-25 "operational flow of FIG. 4 is iteratively performed"]
determining a temporal difference error that is dependent upon at least a difference between one of the rewards and a value estimate generated by the reinforcement learning neural network; [calculates TD error based on SARSA which includes rewards (Rt+1 in Eq. 8) Col. 11 ln. 33-43, Col. 5 ln. 12-16 "(25) Calculating section 109 may calculate parameters. For example, calculating section 109 may be operable to calculate a temporal difference error based on the previous action-value, the current action-value, and the plurality of parameters of neural network 120."]
Dasgupta does not specifically teach determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and updating parameters of the reinforcement learning neural network using the scaled temporal difference error.
However, van Hasselt teaches determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and [scaling rewards via normalization (scaling factor ¶87) ¶95 "computing reward 304 includes computing an un-normalized reward and normalizing the un-normalized reward using the calculated mean and standard deviation of the delta VWAPs"]
updating parameters of the reinforcement learning neural network using the scaled temporal difference error. [train using a normalized by a scaling factor (updates with applied normalizations ¶101-102) in reinforcement learning ¶126 "training engine 118 can train the reinforcement learning network 110 using the normalized order count. The total volume of the order can be normalized by dividing the total volume by a scaling factor"]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta by incorporating the determining a value of a scale factor, wherein a square of the scale factor has a first term dependent upon a variance of the rewards; and scaling the temporal difference error by the scale factor to determine a scaled temporal difference error, and updating parameters of the reinforcement learning neural network using the scaled temporal difference error disclosed by van Hasselt because both techniques address the same field of reinforcement learning and by incorporating van Hasselt into Dasgupta improves recommendations so they are receive more favorable reactions [van Hasselt ¶20]
Dasgupta and van Hasselt do not specifically teach scaling the temporal difference error by the scale factor to determine a scaled temporal difference error;
However, BURHANI teaches scaling the temporal difference error by the scale factor to determine a scaled temporal difference error; [Normalizes target and neural output then computes difference (scaled temporal difference error) ¶29, Fig. 3 306 ¶45 "determines an error for the training item using the normalized target output and the normalized output (step 308). The manner in which the system calculates the error is dependent on the objective function being optimized. For example, for a mean-squared loss function, the system determines the difference between the normalized output and the target output."]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta and VAN HASSELT by incorporating the scaling the temporal difference error by the scale factor to determine a scaled temporal difference error disclosed by BURHANI because all techniques address the same field of reinforcement learning and by incorporating BURHANI into Dasgupta and VAN HASSELT speeds up reinforcement learning for more optimal and metric based results [BURHANI ¶42]
As to dependent claim 19, the rejection of claim 18 is incorporated Dasgupta, VAN HASSELT and BURHANI further teach wherein the temporal difference error is further dependent upon a time discounted value estimate generated by the reinforcement learning neural network system, and wherein the square of the scale factor includes a second term, the method further comprising determining a value for the second term by: determining an estimate of a variance of a time discount factor of the time discounted value estimate; [Dasgupta time based discount factor Col. 11 ln. 33-45 "taking A.sub.t, γ is the discount factor for future reward, and η is the learning rate"]
determining an estimate of an expectation value of returns-squared, wherein a return comprises a time discounted sum of one or more rewards received after a reinforcement learning time step; and [BURHANI mean and standard deviation of rewards include a sum ¶13-14]
forming a product of the estimate of the variance of a time discount factor and the estimate of the expectation value of the returns-squared. [BURHANI formula with deviation time and estimate performance ¶80-81]
Claims 3-7, 13-15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Dasgupta in view of VAN HASSELT and BURHANI as applied to the rejection of claim 1-2 and 18 above, and further in view of MACGLASHAN (US 20200302323 A1 MACGLASHAN)
As to dependent claim 3, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 1 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI further teach obtaining the observation for a current time step characterizing a current state of the environment; [Dasgupta sensors receive observations from an environment like software or a game Col. 4 ln. 9-19 "The observations may be observed through sensors"]
processing the observation for the current time step using the value function neural network and in accordance with current values of value function neural network parameters, to generate a current value estimate relating to the current state of the environment; [Dasgupta Fig. 1 107 Col. 4- ln. 66-11 "determine a current action-value from an evaluation of action-value function 112"]
causing the agent to perform the selected action and, in response, receiving a reward for the current time step characterizing progress made in the environment as a result of the agent performing the selected action, the environment transitioning to a next state of the environment; and [Dasgupta selects and causes actions to be performed and gets rewards Col. 4 ln. 41-54 "select the possible action that yields the largest reward probability from the probability function."]
wherein the method further comprises, for each of a plurality of training time steps:
determining a temporal difference error between a first value estimate for a first one of the training time steps generated by processing the observation for the first time step using the value function neural network, and [Dasgupta TD-learning based on rewards and values Col. 11-12 ln. 33-54 " a general class of on-policy TD-learning methods for RL. SARSA stands for State-Action-Reward-State-Action,"… "TD-error modulated reinforcement learning"]
a sum of the reward at the first time step and a time discounted value estimate for a subsequent state of the environment at a subsequent one of the time steps; [BURHANI mean and standard deviation of rewards include a sum ¶13-14]
determining the value of the scale factor; [BURHANI scale factor ¶87, ¶95]
scaling the temporal difference error by the scale factor to determine the scaled temporal difference error; and [VAN HASSELT ¶29, ¶45]
updating the values of the value function neural network parameters using the scaled temporal difference error. [train using a normalized by a scaling factor (updates with applied normalizations ¶101-102) in reinforcement learning ¶126 "training engine 118 can train the reinforcement learning network 110 using the normalized order count. The total volume of the order can be normalized by dividing the total volume by a scaling factor"]
Dasgupta, VAN HASSELT and BURHANI do not specifically teach selecting an action to be performed by the agent in response to the observation, using the current value estimate or using an action selection neural network updated using value estimates generated by the value function neural network.
However, MACGLASHAN teaches selecting an action to be performed by the agent in response to the observation, using the current value estimate or using an action selection neural network updated using value estimates generated by the value function neural network; [critic model called an action-value model for selecting actions ¶46, ¶11 " action-value model estimating, within one or more processors of the agent, an expected future discounted reward that would be received if a hypothetical action was selected under a current observation of the agent and the agent's behavior was followed thereafter"].
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the selecting an action to be performed by the agent in response to the observation, using the current value estimate or using an action selection neural network updated using value estimates generated by the value function neural network disclosed by MACGLASHAN because all techniques address the same field of reinforcement learning and by incorporating MACGLASHAN into Dasgupta, VAN HASSELT and BURHANI reduces overfitting in models for stable results [MACGLASHAN ¶4-5]
As to dependent claim 4, the rejection of claim 3 is incorporated Dasgupta, VAN HASSELT BURHANI and MACGLASHAN further teach wherein determining the value of the scale factor comprises determining an estimate of the variance of the rewards from the rewards received at the time steps, and using the estimate of the variance of the rewards to determine the first term. [BURHANI mean and standard deviation of rewards are an estimate of variance of rewards ¶13-14]
As to dependent claim 5, the rejection of claim 3 is incorporated Dasgupta, VAN HASSELT BURHANI and MACGLASHAN further teach wherein the square of the scale factor includes a second term, the method further comprising determining a value for the second term by: determining an estimate of a variance of a time discount factor, wherein the time discount factor is a multiplier of the time discounted value estimate; [Dasgupta time-based discount factor Col. 11 ln. 33-45 "taking A.sub.t, γ is the discount factor for future reward, and η is the learning rate"]
determining an estimate of an expectation value of returns-squared, wherein a return comprises a time discounted sum of one or more rewards received after a time step; [BURHANI mean and standard deviation of rewards include a sum ¶13-14]
forming a product of the estimate of the variance of a time discount factor and the estimate of the expectation value of the returns-squared. [BURHANI formula with deviation time and estimate performance ¶80-81]
As to dependent claim 6, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 2 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI further teach processing an observation of the next state of the environment using the target value function neural network to determine a value estimate for the subsequent state of the environment; and [Dasgupta Fig. 1 107 Col. 4- ln. 66-11 "determine a current action-value from an evaluation of action-value function 112"]
applying a time discount factor to the value estimate for the subsequent state of the environment to determine the time discounted value estimate for the subsequent state of the environment. [Dasgupta time based discount factor Col. 11 ln. 33-45 "taking A.sub.t, γ is the discount factor for future reward, and η is the learning rate"]
Dasgupta, VAN HASSELT and BURHANI do not specifically teach determining an estimate of an expectation value of a squared difference between the first value estimate and a value estimate for the state of the environment at the first time step determined by processing the observation for the first time step using the target value function neural network.
However, MACGLASHAN teaches determining an estimate of an expectation value of a squared difference between the first value estimate and a value estimate for the state of the environment at the first time step determined by processing the observation for the first time step using the target value function neural network. [MACGLASHAN difference between estimate from critic and target model ¶73 "The loss function may then be represented using any norm of the difference between the critic and target. "]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the determining an estimate of an expectation value of a squared difference between the first value estimate and a value estimate for the state of the environment at the first time step determined by processing the observation for the first time step using the target value function neural network disclosed by MACGLASHAN because all techniques address the same field of reinforcement learning and by incorporating MACGLASHAN into Dasgupta, VAN HASSELT and BURHANI reduces overfitting in models for stable results [MACGLASHAN ¶4-5]
As to dependent claim 7, the rejection of claim 6 is incorporated Dasgupta, VAN HASSELT BURHANI and MACGLASHAN further teach wherein the square of the scale factor includes a third term, the method further comprising determining a value for the third term by: determining an estimate of an expectation value of a squared difference between the first value estimate and a value estimate for the state of the environment at the first time step determined by processing the observation for the first time step using the target value function neural network. [MACGLASHAN difference between estimate from critic and target model ¶73 "The loss function may then be represented using any norm of the difference between the critic and target. "].
As to dependent claim 13, the rejection of claim 3 is incorporated Dasgupta, VAN HASSELT BURHANI and MACGLASHAN further teach wherein the method is performed online, the plurality of training time steps corresponds to the plurality of action selection time steps, and the first value estimate for the first time step is the current value estimate for the current time step. [MACGLASHAN online variant for training ¶10 "an online variant, in which data is collected as the algorithm trains the policy model"; timesteps ¶47]
As to dependent claim 14, the rejection of claim 3 is incorporated Dasgupta, VAN HASSELT BURHANI and MACGLASHAN further teach wherein the value function neural network comprises an action-value function neural network for determining an action value for each of a plurality of possible actions, wherein the current value estimate is used to determine an action value for each of the possible actions, wherein selecting the action to be performed by the agent in response to the observation comprises selecting the action based on the action value for each of the possible actions, and wherein the time discounted value estimate for a subsequent state of the environment at a subsequent one of the time steps comprises a time discounted action value for the subsequent state of the environment. [Dasgupta selects and causes actions from possible actions using possible rewards from each Col. 4 ln. 41-54 "select the possible action that yields the largest reward probability from the probability function."]
As to dependent claim 15, the rejection of claim 3 is incorporated Dasgupta, VAN HASSELT BURHANI and MACGLASHAN further teach processing the observation for the time step using the action selection neural network and in accordance with current values of action selection neural network parameters, to generate an action selection output, and selecting, using the action selection output, an action to be performed by the agent in response to the observation; and further comprising updating the action selection neural network parameters using the first value estimate. [MACGLASHAN action model (neural network) with updates ¶8 "action-value model and the policy model, wherein the stale copy is initialized identically to the fresh copy and is slowly moved to match the fresh copy as learning updates are performed on the fresh copy, wherein the algorithm has both an offline variant, in which the algorithm is trained using previously collected data, and an online variant, in which data is collected as the algorithm trains the policy model."]
As to dependent claim 20, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 18 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI further teach obtaining the observation for a current time step characterizing a current state of the environment; [Dasgupta sensors receive observations from an environment like software or a game Col. 4 ln. 9-19 "The observations may be observed through sensors"]
processing the observation for the current time step using the value function neural network and in accordance with current values of value function neural network parameters, to generate a current value estimate relating to the current state of the environment; [Dasgupta Fig. 1 107 Col. 4- ln. 66-11 "determine a current action-value from an evaluation of action-value function 112"]
causing the agent to perform the selected action and, in response, receiving a reward for the current time step characterizing progress made in the environment as a result of the agent performing the selected action, the environment transitioning to a next state of the environment; and [Dasgupta selects and causes actions to be performed and gets rewards Col. 4 ln. 41-54 "select the possible action that yields the largest reward probability from the probability function."]
wherein the method further comprises, for each of a plurality of training time steps:
determining a temporal difference error between a first value estimate for a first one of the training time steps generated by processing the observation for the first time step using the value function neural network, and [Dasgupta TD-learning based on rewards and values Col. 11-12 ln. 33-54 " a general class of on-policy TD-learning methods for RL. SARSA stands for State-Action-Reward-State-Action,"… "TD-error modulated reinforcement learning"]
a sum of the reward at the first time step and a time discounted value estimate for a subsequent state of the environment at a subsequent one of the time steps; [BURHANI mean and standard deviation of rewards include a sum ¶13-14]
determining the value of the scale factor; [BURHANI normalization ¶87, ¶95]
scaling the temporal difference error by the scale factor to determine the scaled temporal difference error; and [VAN HASSELT ¶29, Fig. 3 306 and ¶45
updating the values of the value function neural network parameters using the scaled temporal difference error. [train using a normalized by a scaling factor (updates with applied normalizations ¶101-102) in reinforcement learning ¶126 "training engine 118 can train the reinforcement learning network 110 using the normalized order count. The total volume of the order can be normalized by dividing the total volume by a scaling factor"]
Dasgupta, VAN HASSELT and BURHANI do not specifically teach selecting an action to be performed by the agent in response to the observation, using the current value estimate or using an action selection neural network updated using value estimates generated by the value function neural network.
However, MACGLASHAN teaches selecting an action to be performed by the agent in response to the observation, using the current value estimate or using an action selection neural network updated using value estimates generated by the value function neural network; [critic model called an action-value model for selecting actions ¶46, ¶11 " action-value model estimating, within one or more processors of the agent, an expected future discounted reward that would be received if a hypothetical action was selected under a current observation of the agent and the agent's behavior was followed thereafter"].
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the selecting an action to be performed by the agent in response to the observation, using the current value estimate or using an action selection neural network updated using value estimates generated by the value function neural network disclosed by MACGLASHAN because all techniques address the same field of reinforcement learning and by incorporating MACGLASHAN into Dasgupta, VAN HASSELT and BURHANI reduces overfitting in models for stable results [MACGLASHAN ¶4-5]
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Dasgupta in view of VAN HASSELT and BURHANI as applied to the rejection of claim 1 above, and further in view of DA SILVA et al. (US 20210073912 A1 herein Da Silva)
As to dependent claim 8, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 1 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI further teach scaling a respective temporal difference error for each head by the scale factor to determine a respective scaled temporal difference error for updating the values of the value function neural network parameters. [VAN HASSELT scales (applies factor to data) for corrected results ¶70-73 "apply the risk-sensitive policy with the correction factor Q.sup.β(s.sub.t, a.sub.t) to the real-time data"]
Dasgupta, VAN HASSELT and BURHANI do not specifically teach wherein the value function neural network has multiple heads each to generate a respective first value estimate, the method comprising: determining a respective value of the scale factor for each head.
However, Da Silva teaches wherein the value function neural network has multiple heads each to generate a respective first value estimate, the method comprising: determining a respective value of the scale factor for each head; [multiple heads with value estimates ¶56 "Each head estimates a value for each action. Due to the aleatoric nature of the exploration and network initialization, each head will output a different estimate of the action values"]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the wherein the value function neural network has multiple heads each to generate a respective first value estimate, the method comprising: determining a respective value of the scale factor for each head disclosed by Da Silva because all techniques address the same field of reinforcement learning and by incorporating Da Silva into Dasgupta, VAN HASSELT and BURHANI addresses concerns for models sample complexity of training data for improved learning [Da Silva ¶3, ¶33]
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Dasgupta in view of VAN HASSELT and BURHANI as applied to the rejection of claim 1 above, and further in view of Wright et al. (US 20180012137 A1 hereinafter Wright)
As to dependent claim 9, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 1 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI do not specifically teach wherein the temporal difference error is an n-step temporal difference error, the method comprising determining the n-step temporal difference error between a sum of i) the reward ii) n−1 subsequent rewards and iii) a time discounted value estimate for the nth subsequent state of the environment, and a or the first value estimate.
However, Wright teaches wherein the temporal difference error is an n-step temporal difference error, the method comprising determining the n-step temporal difference error between a sum of i) the reward ii) n−1 subsequent rewards and iii) a time discounted value estimate for the nth subsequent state of the environment, and a or the first value estimate. [Temporal difference with return estimates and sums in formula ¶35-37 "multi-sample trajectories, the idea behind the 1-step return has been extended by Temporal Difference (TD) [Sutton 1998] to produce the n-step return estimates which may be subject to different variance and bias than the 1-step return depending on the learning contexts."]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the wherein the temporal difference error is an n-step temporal difference error, the method comprising determining the n-step temporal difference error between a sum of i) the reward ii) n−1 subsequent rewards and iii) a time discounted value estimate for the nth subsequent state of the environment, and a or the first value estimate disclosed by Wright because all techniques address the same field of reinforcement learning and by incorporating Wright into Dasgupta, VAN HASSELT and BURHANI maximizes model functions with self-optimizing behaviors [Wright ¶4]
Claims 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Dasgupta in view of VAN HASSELT and BURHANI as applied to the rejection of claim 1 above, and further in view of Lawrence et al. (US 11500337 B2 hereinafter Lawrence)
As to dependent claim 10, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 1 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI do not specifically teach maintaining an experience replay memory that stores experience tuples generated as a result of the agent interacting with the environment, wherein the experience tuples identify, for each of a plurality of the time steps, at least: the observation, the action selected, the reward received, and a next observation; and sampling the experience tuples in the experience replay memory for performing the plurality of training time steps.
However, Lawrence teaches maintaining an experience replay memory that stores experience tuples generated as a result of the agent interacting with the environment, wherein the experience tuples identify, for each of a plurality of the time steps, at least: the observation, the action selected, the reward received, and a next observation; and sampling the experience tuples in the experience replay memory for performing the plurality of training time steps. [Replay memory with tuples with states and rewards Col. 7 ln. 41-50 " Replay Memory can be a fixed-size collection of tuples of the form (s.sub.t, u.sub.t, s.sub.t+1, r.sub.t). "]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the maintaining an experience replay memory that stores experience tuples generated as a result of the agent interacting with the environment, wherein the experience tuples identify, for each of a plurality of the time steps, at least: the observation, the action selected, the reward received, and a next observation; and sampling the experience tuples in the experience replay memory for performing the plurality of training time steps disclosed by Lawrence because all techniques address the same field of reinforcement learning and by incorporating Lawrence into Dasgupta, VAN HASSELT and BURHANI stabilizes gains made by networks for improve models [Lawrence Col. 1 ln. 58-11]
As to dependent claim 11, the rejection of claim 10 is incorporated Dasgupta, VAN HASSELT BURHANI and Lawrence further teach sampling the experience tuples prioritizes the sampling using a priority dependent on a magnitude of the scaled temporal difference error. [Lawrence sampling with tuples Col. 6 ln. 29-57, Col. 9 ln. 1-43 “Uniformly sample M tuples”]
As to dependent claim 12, the combination of Dasgupta, VAN HASSELT and BURHANI teach all the limitations of claim 1 that is incorporated.
Dasgupta, VAN HASSELT and BURHANI do not specifically teach initializing the value of the scale factor to a non-zero value.
However, Lawrence teaches initializing the value of the scale factor to a non-zero value. [non-negative scaling Col. 5-6 ln. 63-5 "nonnegative scaling constant, β"; Initialize the gains (scaling) Col. 4 ln. 39-55 "actor network allows us to use initialize training with pre-existing PID gains as well as incorporate individual constraints on each parameter. The actor can be therefore initialized as an operational, interpretable, and industrially accepted controller that can be then updated in an optimal direction "]
Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filling date of the claimed invention to modify the reinforcement network by Dasgupta, VAN HASSELT and BURHANI by incorporating the initializing the value of the scale factor to a non-zero value disclosed by Lawrence because all techniques address the same field of reinforcement learning and by incorporating Lawrence into Dasgupta, VAN HASSELT and BURHANI stabilizes gains made by networks for improve models [Lawrence Col. 1 ln. 58-11]
Response to Arguments
Applicant's arguments filed 06/29/2026, with respect to 101, these rejections have been withdrawn.
Applicant's arguments filed 06/29/2026. In the remark, applicant argues that:
(1) Dasgupta, VADORI and BURHANI fail to teach " at each of a plurality of training time steps: determining a temporal difference error that is dependent upon at least a difference between one of the rewards and a value estimate generated by the reinforcement learning neural network; " as recited by amended claim 1.
As to point (1), Applicant’s arguments with respect to claim 1 have been considered but are moot in view of a new ground of rejection as set forth above of Dasgupta in view of van Hasselt and BURHANI. Check in particular the new formula cited in Dasgupta (SARSA Eq. 8) which appears to used rewards (R). Also note the applicability of van Hasselt (owned by Google) and Schaul (below) owned by DeepMind which use normalization to scale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Applicant is required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action.
Schaul et al. (US 20200265312 A1) teaches TD error based on values and reward (part of tuple see ¶38, ¶13.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BEAU SPRATT whose telephone number is (571)272-9919. The examiner can normally be reached M-F 8:30-5 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on 5712127212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BEAU D SPRATT/ Primary Examiner, Art Unit 2143