DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This action is responsive to the claims dated 7/15/2026. Claims 1-3, 5, 7-8, 16, and 19-20 are amended. Claims 1-20 are presented for examination.
Specification
The disclosure is objected to because of the following informalities:
[0045] recites "the system trains each auxiliary neural network 124 ,126, and 128". A space precedes the comma following "124". Further, these elements are named "auxiliary neural network" here, but are named "auxiliary prediction neural network" throughout the remainder of the disclosure, including the immediately preceding sentence of [0045], and in every claim. See also [0015], [0043], [0058], [0065], [0075] and [0091], which use the same inconsistent name.
[0045] further recites "on some or all of the same training data as the auxiliary prediction neural network 112". Reference character 112 designates the action selection neural network (FIG. 1; [0029] and [0040]; and the first sentence of [0045]), not an auxiliary prediction neural network.
[0082] recites "The value function estimate 406 can be can be calculated as a dot product". The phrase "can be" is duplicated.
[0093] recites "FIG. 7 show a series of graphs". "show" should read --shows--; compare the following sentence of the same paragraph, "FIG. 7 shows seven graphs".
[0096], as amended 7/15/2026, recites "three trendlines 1002, 1004, and 1006 the represent the performance of other reinforcement learning algorithms". "the" should read --that--; compare [0095], "three trendlines 902, 904, and 906 that represent".
Appropriate correction is required.
The specification is objected to as failing to provide proper antecedent basis for the claimed subject matter. See 37 CFR 1.75(d)(1) and MPEP Section 608.01(o). Correction of the following is required:
Claims 1, 19 and 20, as amended 7/15/2026, each recite "an electromechanical agent". The term "electromechanical" does not appear anywhere in the description. The corresponding element in the description is "a mechanical agent" ([0025], [0028] and [0055]). Amendment of the description to provide antecedent basis for the claimed "electromechanical agent" is required. No new matter may be entered. See 37 CFR 1.121(b).
This objection is directed to the absence of the claim term from the description under 37 CFR 1.75(d)(1); it is not a rejection under 35 U.S.C. 112(a).
Claim Objections
Claims 1, 5, 7, 19 and 20 are objected to because of the following informalities:
Claims 1, 19 and 20 each recite "selecting the action to be performed by the electromechanical agent at the time step". No action is introduced in the body of the claim; the only earlier recitation is the plural "actions" in the preamble. "the action" should read --an action--, the subsequent step then properly referring to "the action selected".
Claims 5 and 7 each recite "the auxiliary reward", claim 5 at "determining, for each auxiliary prediction neural network, the auxiliary reward for the time step", and claim 7 at "an importance of the feature to predicting state values relative to the auxiliary reward". Claim 1, from which both ultimately depend, introduces "a corresponding auxiliary reward". "the auxiliary reward" should read --the corresponding auxiliary reward-- in each claim, consistent with the wording claims 2, 8 and 16 already use.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 16 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Graves et al. (hereinafter Graves), US 2019/0329772 A1, in view of Zheng et al. (hereinafter Zheng) "Learning State Representations from Random Deep Action-conditional Predictions" 2021.
Regarding independent claim 1, Graves teaches a method performed by one or more computers for selecting actions to be performed by an electromechanical agent to interact with a real-world environment to perform a main task, the method comprising, for each time step in a sequence of time steps (Graves: [0163], "The ASP control system 400A of FIG. 8 is modified relative to the ASP control system 400 of FIG. 4 to enable RL to be applied directly to the problem of adaptive spacing when training the AS controller module 412."; this rejection relies on ASP control system 400A, which is ASP control system 400 with the AS controller module 412 implemented as a neural network trained by reinforcement learning and with the exchange of predictions [0164] to [0166] describe, so the passages below that describe ASP control system 400 describe ASP control system 400A except in those respects; [0070], "ASP control system 400 receives inputs from the EM wave based sensors 110 and the vehicle sensors 111, and controls actuators of the drive control system 150"; the adaptive spacing predictive control system runs on the processor system 102 of the vehicle control system 115 (one or more computers) and drives the brake unit 154 and throttle unit 156 of the ego vehicle 105 (an electromechanical agent) as it travels on the road (a real-world environment); [0110], "must implement actions that attempt to avoid both front-end and rear-end collisions and thereby improve safety of the ego vehicle 105 while operating on the road, while attempting to maintain a target speed"; avoiding collisions while holding a target speed is among the target objectives (a main task) the controller is built to pursue, [0110], "The predictions represented in the matrix, at time t, denoted by pt are supplied to the AS controller module 412, which determines a next vehicle action"; it does so by settling on a next action at each successive sample time; in ASP control system 400A that determination is made by the neural network implemented AS controller module 412, [0166], "a neural network implemented function for the AS controller 412 can be trained to perform final mapping" (each time step in a sequence of time steps)): receiving a set of features representing a current observation, wherein the current observation characterizes a current state of the real-world environment at the time step (Graves: [0071], "continuously determines a representation of a current state of the ego vehicle 105 and its environment at a current time"; the state sub-module 410 builds, at each current time, a representation of the current state of the ego vehicle 105 and its environment (a current observation); [0074], "where S represents a set of physical parameters about the current environment of the ego vehicle 105, as measured by other vehicle sensors 120, including for example speed v, engine RPM, current engine gear, current throttle position, current brake position, and distances to any obstacles in any risk zone directions"; the enumerated physical parameters measured by the vehicle sensors 120 are the physical parameters (a set of features representing a current observation) that make up that representation); for each of one or more auxiliary prediction neural networks: determining an auxiliary input to the auxiliary prediction neural network, wherein the auxiliary input comprises a proper subset of the set of features representing the current observation (Graves: [0088], "each of the predictor modules 404, 406 and 408 includes one or more predictors in the form of one or more trained neural networks … In some examples, a separate neural network is used for each predictor"; the speed predictor GVF 604 is a separate trained neural network, and the claim recites one or more such networks, so that predictor is the auxiliary prediction neural network relied on here (each of one or more auxiliary prediction neural networks); [0093], "In example embodiments, the specific vehicle state st information input to speed predictor GVF ƒspeed 604 to predict the speed of the ego vehicle 105 includes: current speed of the vehicle; RPM of the vehicle's engine; gear of the vehicle's transmission; and amount of throttle or braking applied."; the specific vehicle state information (an auxiliary input) routed to the speed predictor GVF 604 is information specified for that predictor; [0097], "the specific information from the state space (st) required to predict safety for the front or back SRZs are, for a given time t"; Graves specifies a different selection for its safety predictors, one that adds, [0099], "Distance from the ego vehicle 105 to the target vehicle in each zone", and each predictor, being a separate neural network, is given the information specified for it; the information specified for the speed predictor names the vehicle's own speed, engine RPM, gear and throttle or braking and omits the distances to obstacles that the state of [0074] also carries, so what that predictor receives is fewer than all of the physical parameters, a selection of specific vehicle state information (a proper subset of the set of features)); processing the auxiliary input using the auxiliary prediction neural network, wherein: the auxiliary prediction neural network is configured to generate a state value estimate for the current state of the real-world environment relative to a corresponding auxiliary reward that measures values of a corresponding target feature from the set of features representing observations for the sequence of time steps (Graves: [0134], "When constructing a GVF to implement predictor GVF ƒspeed 604, a cumulant (pseudo-reward) function, pseudo-termination function, and target policy are each defined."; constructing the speed predictor GVF starts by defining its own cumulant (a corresponding auxiliary reward), that is, its pseudo-reward, together with a pseudo-termination function and a target policy; [0135], "where vt is the current measured velocity of the vehicle, and γ is the discount factor."; the speed predictor GVF 604 takes its cumulant from the current measured velocity of the vehicle (a corresponding target feature), which [0074] lists among the physical parameters sensed at each sample time (from the set of features representing observations for the sequence of time steps); [0135], "with this normalization factor, ƒspeed(st,at) represents a weighted average of all future speeds"; what that predictor returns for the current state is the discounted average of its cumulant over the future, that is, its prediction (a state value estimate). Graves writes that return as ƒspeed(st,at), a return computed from the current state st for a candidate action at; [0145], "a function can learn to predict a probabilistic future speed of the vehicle where the action is assumed to be relatively constant over short periods of time"; Graves defines a target policy for the predictor that does not turn on the action, [0136], "a target policy of π(at|st)=1 for all actions at and states st can be used", collects its training data under an expert policy and fits the predictor to that data, [0146], "state-action-reward-state-action (SARSA) reinforcement temporal difference (TD) learning", and treats the action as held over the short horizon of the prediction ([0145]); a general value function is defined by its cumulant, its discount and its target policy ([0134]), so the return is the value of the current state relative to that cumulant under that policy. Graves labels its predictions action conditioned, [0072], "action conditioned (AC) predictions 416", and evaluates the predictor once for each candidate action ([0108]); the claim recites a state value estimate for the current state of the real-world environment relative to the corresponding auxiliary reward and does not exclude a prediction that is also conditioned on a candidate action, and what the predictor returns remains the accumulation of its cumulant from the current state under the target policy just described. The present specification defines a state value function by the same three elements at [0074] and places no restriction on the policy; its Q-value sentence at [0030] describes the action selection output of the action selection neural network, an estimate of the main task reward, and not what an auxiliary prediction neural network returns. What the predictor returns for the current state relative to its cumulant is its prediction (a state value estimate)); selecting [[the]] an action (interpreted per the Claim Objections set forth above) to be performed by the electromechanical agent at the time step using the action selection output (Graves: [0166], "a neural network implemented function for the AS controller 412 can be trained to perform final mapping"; [0166], "using any number of reinforcement learning methods such as Deep-Q-Network (DQN)"; in ASP control system 400A the AS controller module 412 is a neural network that maps the predictions it receives for the candidate actions to their action values (an action selection output); [0166], "but also to train the control function of the AS controller module 412 to make control decisions"; the next vehicle action is decided on those values at the current sample time (selecting an action to be performed by the electromechanical agent at the time step using the action selection output)); and causing the electromechanical agent to perform the action selected using the action selection output to interact with the real-world environment (Graves: [0133], "the drive control system 150 is instructed to perform the next action"; the action the controller settles on is issued as an instruction to the drive control system 150, which actuates the brake unit 154 and the throttle unit 156; [0060], "The mechanical system 190 effects physical operation of the vehicle 105"; the engine 192, transmission 194 and wheels 196 then carry that instruction out, which changes the speed of the ego vehicle 105 and its spacing from the vehicles around it (causing the electromechanical agent to perform the action selected to interact with the real-world environment)).
Graves does not expressly teach extracting, from each auxiliary prediction neural network, a respective intermediate output generated by one or more hidden layers of the auxiliary prediction neural network during processing of the auxiliary input to the auxiliary prediction neural network; processing an input comprising the respective intermediate outputs generated by the one or more hidden layers of each auxiliary prediction neural network at the time step using an action selection neural network to generate an action selection output.
However, Zheng teaches extracting, from each auxiliary prediction neural network, a respective intermediate output generated by one or more hidden layers of the auxiliary prediction neural network (Zheng: page 4, 3.3 Agent Architecture, "The answer network module, parameterized by θans, maps the state vector St to a set of predictions y(St)"; the pathway made up of the state representation module and answer network (an auxiliary prediction neural network) is the network that answers the prediction questions; page 4, 3.3 Agent Architecture, "In the stop-gradient setting, we stopped the gradients from the RL loss from flowing from LRL to θrepr."; in that setting the state representation module is trained by the prediction loss alone, so it belongs to the prediction network rather than to the policy, page 7, Neural Network Architecture, "the state representation module consists of 3 convolutional layers"; those convolutional layers (one or more hidden layers) stand between the observation history and the predictions, and what they emit is the state vector (a respective intermediate output), an internal quantity of the prediction network rather than one of its predictions, page 4, 3.3 Agent Architecture, "The RL module, parameterized by θRL, maps the state vector St to a policy distribution over the available actions"; the state vector the prediction pathway computes is what the module that selects actions is given, so it is read out of that pathway and handed to that module (extracting, from each auxiliary prediction neural network, a respective intermediate output)) during processing of the auxiliary input to the auxiliary prediction neural network (Zheng: page 4, 3.3 Agent Architecture, "The state representation module, parameterized by θrepr, maps the history of observations and actions (O0, A0, . . . , Ot) to a state vector St"; the history of observations and actions is what the prediction network takes in, and the state vector is produced as that history is processed on its way to the predictions (during processing of the auxiliary input to the auxiliary prediction neural network)); processing an input comprising the respective intermediate outputs generated by the one or more hidden layers of each auxiliary prediction neural network at the time step using an action selection neural network to generate an action selection output (Zheng: page 4, 3.3 Agent Architecture, "The RL module, parameterized by θRL, maps the state vector St to a policy distribution over the available actions π(·|St) and a value function v(St)", page 7, "The RL module has one hidden dense layer and two output heads for the policy and the value function respectively"; the RL module is itself a neural network, with a hidden dense layer and policy and value heads, and it takes the state vector at each time t and maps it to a policy distribution over the available actions from which the agent's action is drawn (processing an input comprising the respective intermediate outputs … at the time step using an action selection neural network to generate an action selection output)).
Because Graves and Zheng are analogous art with both addressing the field of the claimed invention, the use of auxiliary predictions to build the state representation a reinforcement learning controller acts on, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Zheng’s teaching that what is carried from a prediction network to the module that selects actions is that network’s internal state vector (Zheng: page 4) to the method of Graves, with a reasonable expectation of success, by taking from the speed predictor GVF 604, as it processes the vehicle state information routed to it for a candidate action, the output of its hidden layers, and passing that hidden layer output, alongside the prediction, to the neural network implemented AS controller module 412 of ASP control system 400A, which maps what it receives to the action values on which the next vehicle action is decided, so that the module receives both, to teach extracting, from each auxiliary prediction neural network, a respective intermediate output generated by one or more hidden layers of the auxiliary prediction neural network during processing of the auxiliary input to the auxiliary prediction neural network; processing an input comprising the respective intermediate outputs generated by the one or more hidden layers of each auxiliary prediction neural network at the time step using an action selection neural network to generate an action selection output. This modification would have been motivated by the desire to use the representation the prediction tasks learn to accelerate the controller’s learning of its task (Zheng: page 1, "Providing auxiliary tasks to Deep Reinforcement Learning (Deep RL) agents has become an important class of methods for driving the learning of representations that accelerate learning on a main task").
Regarding dependent claim 2, Graves, in view of Zheng, teach the method of claim 1, wherein for each auxiliary prediction neural network, the state value estimate for the current state of the real-world environment relative to the corresponding auxiliary reward defines an estimate of a cumulative measure of the corresponding auxiliary reward to be received over future time steps (Graves: [0135], "The correction factor 1−γ normalizes the sum of all future cumulants such that the total return ƒspeed(st,at) is"; what that predictor returns is the sum of its cumulant's values over the future time steps, each weighted by a successive power of the discount factor; [0135], "with this normalization factor, ƒspeed(st,at) represents a weighted average of all future speeds"; what the speed predictor GVF 604 returns for the current state is a weighted average of all future speeds (a cumulative measure), an average taken over the whole future course of its cumulant rather than its next value).
Regarding dependent claim 3, Graves, in view of Zheng, teach the method of claim 2, wherein for each auxiliary prediction neural network, the cumulative measure of the corresponding auxiliary reward comprises a time-discounted sum of values of the corresponding auxiliary reward (Graves: [0134], "The discount factor controls the time horizon for the predictions"; a discount factor is defined for the speed predictor GVF 604 and it is that discount factor which sets how far ahead the prediction reaches; [0135], "The correction factor 1−γ normalizes the sum of all future cumulants such that the total return ƒspeed(st,at) is"; the return that predictor estimates is accordingly the sum of its cumulant’s values over the future time steps, each weighted by a successive power of the discount factor (a time-discounted sum)).
Regarding dependent claim 4, Graves, in view of Zheng, teach the method of claim 1, further comprising: receiving a respective main task reward for each time step in the sequence of time steps (Graves: [0167], "one example of a reward function for RL training the control function of the AS controller module 412 is"; [0168], "where b1, b2, b3, and b4 are constants defining the relative importance of front safety, back safety, and closeness to target speed"; the reward function for RL training the control function (a main task reward) of the AS controller module 412 is evaluated on the state st at each sample time from the front safety, back safety and closeness to the target speed (receiving a respective main task reward for each time step in the sequence of time steps)); and training the action selection neural network based on the main task rewards using reinforcement learning (Graves: [0164], "ASP control system 400A differs from ASP control system 400 in the manner in which predictions are generated and passed between the predictive perception module 402 and the AS controller module 412, and in the implementation of the AS controller module 412."; ASP control system 400A is the implementation relied on for claim 1; [0166], "a neural network implemented function for the AS controller 412 can be trained to perform final mapping"; [0166], "using any number of reinforcement learning methods such as Deep-Q-Network (DQN)"; in that implementation the AS controller module 412 (an action selection neural network) is a neural network implemented function whose control function is trained by reinforcement learning against that reward function (training the action selection neural network based on the main task rewards using reinforcement learning)).
Regarding dependent claim 5, Graves, in view of Zheng, teach the method of claim 1, further comprising: for each time step in the sequence of time steps: determining, for each auxiliary prediction neural network, the corresponding (interpreted per the Claim Objections set forth above) auxiliary reward for the time step based on the value of the corresponding target feature at the time step (Graves: [0135], "The cumulant for predicting speed is…where vt is the current measured velocity of the vehicle"; the cumulant the speed predictor GVF 604 is fitted against is computed at each sample time from the velocity measured at that time (determining the corresponding auxiliary reward for the time step based on the value of the corresponding target feature at the time step)); and training each auxiliary prediction neural network based on values of the corresponding auxiliary reward using reinforcement learning (Graves: [0089], "RL is used to determine GVFs for each of the NN based predictors (also referred to herein as predictor GVFs)"; every neural-network based predictor, and not one of them only, is fitted to its own cumulant by reinforcement learning, [0134], "the speed predictor GVF ƒspeed 604 is configured using RL based on methodology disclosed in the above identified mentioned paper by R. Sutton et al."; the speed predictor GVF 604 relied on above is so fitted, [0160], "In example embodiments, training occurs offline for greater stability and safety in learning; however, it should be noted that the trained GVFs that implement the predictor functions (once trained) can continuously collect and improve predictions in real-time using off-policy RL learning"; Graves prefers to train offline but provides for the predictors to keep being fitted to their cumulants as the vehicle runs, sample time by sample time, and a reference is relied on for all that it teaches, including its nonpreferred embodiments (MPEP 2123) (training each auxiliary prediction neural network based on values of the corresponding auxiliary reward using reinforcement learning)).
Regarding dependent claim 16, Graves, in view of Zheng, teach the method of claim 1, wherein: each auxiliary prediction neural network generates a respective state value estimate for the current state of the real-world environment relative to the corresponding auxiliary reward (Graves: [0089], "RL is used to determine GVFs for each of the NN based predictors (also referred to herein as predictor GVFs)."; each neural-network based predictor is a general value function of its own cumulant (a corresponding auxiliary reward); [0135], "with this normalization factor, ƒspeed(st,at) represents a weighted average of all future speeds"; what the speed predictor GVF 604 returns for the current state, its prediction (a state value estimate), is the discounted average of its cumulant over the future (each auxiliary prediction neural network generates a respective state value estimate for the current state relative to the corresponding auxiliary reward)); and the input to the action selection neural network further comprises the respective state value estimate generated by each auxiliary prediction neural network (Graves: [0165], "the predictions 416A can be considered as an action conditioned predictive state space that is passed from the predictive perception module 402 to the AS controller module 412"; each predictor GVF's prediction (a state value estimate) is passed to the AS controller module 412, [0166], "a neural network implemented function for the AS controller 412 can be trained to perform final mapping", [0166], "using any number of reinforcement learning methods such as Deep-Q-Network (DQN)"; the neural network implemented AS controller module 412 (an action selection neural network) takes those predictions as its input, alongside the hidden layer outputs the combination set out for claim 1 passes to it, so the input it acts on further comprises the prediction of each predictor GVF).
Regarding dependent claim 18, Graves, in view of Zheng, teach the method of claim 1, further comprising, prior to the first time step in the sequence of time steps and for each auxiliary prediction neural network (Zheng: page 3, 3.2 A Random Question Network Generator, "we designed a generator of random question networks from which we can take samples and evaluate their performance as auxiliary tasks", page 8, 6. Conclusion and Future Work, "In this work, the question network was sampled before learning and was held fixed during learning"; the question network, and with it every prediction target, is drawn from the generator before the agent begins to learn and act (prior to the first time step in the sequence of time steps); page 11, A. Pseudocode for RADAR, "create a new prediction node v in G"; the generator makes the draw below for each prediction node it creates in layers 1 through D, the nodes of layer 0 instead each being tied to a distinct feature (for each auxiliary prediction neural network)): randomly sampling a feature from the set of features (Zheng: page 3, Random Features, "RADAR uses random features, each computed by a scalar function gk with random parameters"; the generator's feature nodes stand for features each computed from an observation and the one before it, page 11, A. Pseudocode for RADAR, "create a new feature node f in G"; one feature node is created for each such feature, page 11, A. Pseudocode for RADAR, "randomly select a node from roots"; for each prediction node of layers 1 through D the generator draws one of those feature nodes at random (randomly sampling a feature from the set of features)); and designating the randomly sampled feature as the target feature corresponding to the auxiliary prediction neural network (Zheng: page 11, A. Pseudocode for RADAR, "add edge < v, f, 1 > to G"; the drawn feature node is wired into the target of that prediction node, which fixes the received signal that node is asked to accumulate, the node's target composing that signal with the prediction of the parent node it is also wired to (designating the randomly sampled feature as the target feature corresponding to the auxiliary prediction neural network)).
Regarding independent claim 19, Graves teaches a system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for selecting actions to be performed by an electromechanical agent to interact with a real-world environment to perform a main task, the operations comprising, for each time step in a sequence of time steps (Graves: [0062], "The memory 126 of the vehicle control system 115 has stored thereon a number of software systems 161 in addition to the GUI, where each software system 161 includes instructions that may be executed by the processor 102"; the memory 126 of the vehicle control system 115 (one or more storage devices) holds instructions that the processor 102 (one or more computers) it is coupled to executes, and those instructions carry out the adaptive spacing control of the ego vehicle 105 (an electromechanical agent) on the road (a real-world environment)). The substantive limitations of claim 19 recite the counterparts of the limitations of claim 1 and are taught for the same reasons set forth above. The motivation to combine Graves and Zheng is the same as that set forth for claim 1.
Regarding independent claim 20, Graves teaches one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for selecting actions to be performed by an electromechanical agent to interact with a real-world environment to perform a main task, the operations comprising, for each time step in a sequence of time steps (Graves: [0185], "the technical solution of the present disclosure may be embodied in a non-volatile or non-transitory machine readable medium (e.g., optical disk, flash memory, etc.) having stored thereon executable instructions tangibly stored thereon that enable a processing device (e.g., a vehicle control system) to execute examples of the methods disclosed herein"; the disclosed control method is offered as executable instructions held on a non-transitory machine readable medium (one or more non-transitory computer storage media storing instructions) that a vehicle control system (one or more computers) runs to control the ego vehicle 105 (an electromechanical agent) on the road (a real-world environment)). The substantive limitations of claim 20 recite the counterparts of the limitations of claim 1 and are taught for the same reasons set forth above. The motivation to combine Graves and Zheng is the same as that set forth for claim 1.
Claims 6-10 are rejected under 35 U.S.C. 103 as being unpatentable over Graves in view of Zheng, as applied in the rejection of claim 1 above, and further in view of Martin et al. (hereinafter Martin) "Adapting the Function Approximation Architecture in Online Reinforcement Learning" 2021.
Martin was disclosed in an IDS dated 8/9/24.
Regarding dependent claim 6, Graves, in view of Zheng, teach the method of claim 1.
Graves and Zheng do not expressly teach further comprising, at each of one or more time steps in the sequence of time steps: updating, for one or more of the auxiliary prediction neural networks, data that defines the proper subset of the set of features that are designated to be included in the auxiliary input to the auxiliary prediction neural network.
However, Martin teaches further comprising, at each of one or more time steps in the sequence of time steps (Martin: page 4, left column, "The selection of the largest absolute weights can be expensive with generic implementations that use sorting as a subroutine, so for efficiency line 12 is optionally restricted to occur periodically"; the recomputation of the selections sits inside Algorithm 1's loop over t = 1, 2, 3, · · · and is carried out repeatedly while the system runs, on a periodic schedule chosen for cost (at each of one or more time steps in the sequence of time steps)): updating, for one or more of the auxiliary prediction neural networks, data that defines the proper subset of the set of features that are designated to be included in the auxiliary input to the auxiliary prediction neural network (Martin: page 3, 4. Prediction Adapted Neighborhoods, "Since the selection matrices Mi are adapted from the auxiliary predictions, we call these prediction adapted neighborhoods."; the selections are not fixed at design time but are recomputed from the running auxiliary predictions (updating), page 3, 4. Prediction Adapted Neighborhoods, "The weights are used to identify the relevant features for each GVF"; one selection is held for each general value function and is derived from that function's own learned weights (for one or more of the auxiliary prediction neural networks), page 2, 3. An Approximation Architecture, "The neighborhood selection matrix Mi is an orthogonal rank-k matrix of zeros and ones that provides an ordered selection of k elements of the observation"; a selection matrix Mi is a stored list naming which components of the observation vector are picked out, page 6, Hyperparameter tuning, "a coarse search led to the neighborhood size k = 10 used in all experiments", page 5, 5.2 The Frog’s Eye Domain, "4000 binary proximity sensors"; each selection names 10 of the 4000 observation components, fewer than all of them (data that defines the proper subset of the set of features); Martin applies each selection to the features it builds for its main value function).
Because Graves, in view of Zheng, and Martin are analogous art, Graves and Zheng being in the field of the claimed invention, the use of auxiliary predictions to build the state representation a reinforcement learning controller acts on, and Martin, which learns its predictions in a Markov reward process (Martin: page 2), being reasonably pertinent to the problem the inventor faced of choosing which received features each auxiliary prediction neural network is given, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Martin’s prediction adapted neighborhoods to the method of Graves, in view of Zheng, with a reasonable expectation of success, by learning alongside each predictor general value function a Martin general value function that is linear in the whole vehicle state and accumulates the same cumulant under the same target policy, and by recomputing, as the vehicle runs, the list of vehicle state parameters routed to that predictor’s input from the largest weights of that linear function, a selection of fewer of the parameters than all of them, Martin page 2, "an ordered selection of k elements of the observation", as Martin’s Algorithm 1 recomputes each selection from the weights of its own general value function, to teach further comprising, at each of one or more time steps in the sequence of time steps: updating, for one or more of the auxiliary prediction neural networks, data that defines the proper subset of the set of features that are designated to be included in the auxiliary input to the auxiliary prediction neural network. This modification would have been motivated by the desire to reduce reliance on expert knowledge in choosing each predictor’s inputs without significant loss of performance (Martin: page 1, right column, "We are motivated by the increasing ability of machine learning algorithms to reduce reliance on expert knowledge", page 7, left column, "can be adapted online and without significant performance loss").
Regarding dependent claim 7, Graves, in view of Zheng and Martin, teach the method of claim 6, wherein updating the data that defines the proper subset of the set of features that are designated to be included in the auxiliary input to the auxiliary prediction neural network comprises: determining, for each feature in the set of features, a respective first importance score characterizing an importance of the feature to predicting state values relative to the corresponding (interpreted per the Claim Objections set forth above) auxiliary reward (Martin: page 4, 5.1 How do GVF weights inform neighborhoods, "Similar to simple linear regression, prediction weights reflect the importance of each feature. High magnitude weights suggest more importance than those with low magnitude"; the prediction weights (a respective first importance score) standing against each observation component in a general value function's own solution measure how much that component matters to predicting that function's cumulant (a corresponding auxiliary reward), which under the combination set out for claim 6 is also the cumulant of the predictor whose input the selection feeds; the linear general value function learned alongside that predictor takes the whole vehicle state, so a weight stands against each feature (for each feature in the set of features)); and updating the data defining the proper subset of the set of features that are designated to be included in the auxiliary input to the auxiliary prediction neural network based on the first importance scores (Martin: page 3, 4. Prediction Adapted Neighborhoods, "The weights are used to identify the relevant features for each GVF"; the selection Mi for that general value function is rebuilt from those weights, the components carrying the largest of them being the ones retained (updating the data defining the proper subset of the set of features based on the first importance scores)).
Regarding dependent claim 8, Graves, in view of Zheng and Martin, teach the method of claim 7, wherein determining the respective first importance score for each feature in the set of features comprises: obtaining a state value function that is configured to process the set of features to generate a state value estimate for the current state of the real-world environment relative to the corresponding auxiliary reward (Martin: page 3, 4. Prediction Adapted Neighborhoods, "Each GVF objective has a separately learned solution, represented as a linear function of the observation inputs"; a separate solution is learned for each general value function, a linear function of the observation inputs (a state value function) that takes the whole observation vector (configured to process the set of features) and returns the cumulant's return (a state value estimate) predicted for that function's cumulant (a corresponding auxiliary reward)); and determining the first importance score for each feature in the set of features using the state value function (Martin: page 3, 4. Prediction Adapted Neighborhoods, "The weights are used to identify the relevant features for each GVF"; the per-component scores are read off that learned solution rather than measured separately (determining the first importance score for each feature using the state value function)).
Regarding dependent claim 9, Graves, in view of Zheng and Martin, teach the method of claim 8, wherein the state value function is a linear function that comprises a respective parameter corresponding to each feature in the set of features, and (Martin: page 3, 4. Prediction Adapted Neighborhoods, "separately learned solution, represented as a linear function of the observation inputs"; the solution learned for each general value function is a linear function of the observation inputs (a state value function) that carries one weight for each observation component, each weight being the parameter that corresponds to that component) wherein determining the first importance score for each feature in the set of features using the state value function comprises: determining the first importance score for each feature based on a value of the corresponding parameter of the state value function (Martin: page 4, 5.1 How do GVF weights inform neighborhoods, "High magnitude weights suggest more importance than those with low magnitude"; the score assigned to a component is the size of that component's own weight (determining the first importance score for each feature based on a value of the corresponding parameter of the state value function)).
Regarding dependent claim 10, Graves, in view of Zheng and Martin, teach the method of claim 8, wherein for each time step in the sequence of time steps, the state value function is trained based on the corresponding (interpreted per the Claim Objections set forth above) auxiliary reward for the time step using reinforcement learning (Martin: page 3, 4. Prediction Adapated Neighborhoods, "A cumulant’s return is learned under the same policy and discount as the main value function"; each general value function’s solution is fitted to the return of that function’s own cumulant, which is the auxiliary reward, under the same policy and discount the main value function uses, page 3, 4. Prediction Adapated Neighborhoods, "The weights are updated with TD(λ), along with the associated eligibility traces"; the weights of that solution, and not those of the main value function, are the ones fitted, and they are fitted by temporal-difference learning with eligibility traces (using reinforcement learning), page 4, Algorithm 1, "# Update GVF weights with TD(λ)."; that update stands inside Algorithm 1’s loop over t = 1, 2, 3, · · · and is therefore applied once per time step as each new observation arrives (for each time step in the sequence of time steps, the state value function is trained based on the corresponding auxiliary reward for the time step)).
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Graves in view of Zheng, as applied in the rejection of claim 1 above, and further in view of Schlegel et al. (hereinafter Schlegel) "General Value Function Networks" 2021.
Schlegel was disclosed in an IDS dated 8/9/24.
Regarding dependent claim 11, Graves, in view of Zheng, teach the method of claim 1.
Graves and Zheng do not expressly teach further comprising, at each of one or more time steps: updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks.
However, Schlegel teaches further comprising, at each of one or more time steps (Schlegel: page 529, "The second experiment, Figure 10 (right), uses the full discovery approach to find a representation useful for learning the evaluation GVFs", page 530, "For ϵ = 0.2, which means about 20% of GVFs are pruned in each pruning phase, the prediction error continues to decrease until it almost reaches 0 and is almost as good as the set of hand-design GVFs used in previous experiments"; in the full discovery experiment the least useful general value functions are pruned and replaced in successive pruning phases while the agent continues to learn (at each of one or more time steps)): updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks (Schlegel: page 528, "The evaluator is responsible for testing GVFs and removing unused GVFs. The generator proposes new GVFs from a set of possible GVFs"; a pruned general value function is removed and a newly generated one takes its place in the bank of predictors, page 528, "For cumulants, we consider stimuli cumulants (the cumulant is one of the observations"; each general value function is specified by its cumulant, the signal whose accumulation it is asked to predict, and a stimuli cumulant is one of the received observations, page 528, "We generate new GVFs randomly from a set of GVF primitives"; each replacement draws its cumulant anew, so the stored specification of which observation each predictor targets is revised (updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks)).
Because Graves, in view of Zheng, and Schlegel are analogous art with all three addressing the field of the claimed invention, the use of auxiliary predictions to build the state representation a reinforcement learning controller acts on, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Schlegel's generate-and-test discovery of general value functions to the method of Graves, in view of Zheng, with a reasonable expectation of success, by adding to the speed and safety predictors Graves designs a bank of further predictor general value functions, each built as the speed predictor GVF 604 is, a separate neural network given a selection of the vehicle state parameters of [0074] made as Graves makes it for the predictors it describes, by routing to a predictor the state information required for what that predictor predicts, Graves: [0093], "the specific vehicle state st information input to speed predictor GVF ƒspeed 604 to predict the speed of the ego vehicle 105", Graves: [0097], "the specific information from the state space (st) required to predict safety for the front or back SRZs are, for a given time t", each such selection being fewer than all of those parameters, and whose hidden layer output is passed, alongside its prediction, to the neural network implemented AS controller module 412 as set out for claim 1, with cumulants generated from the vehicle state parameters, and by re-targeting the least useful of the added predictors as the vehicle runs, the designed predictors being kept, the added predictors being the auxiliary prediction neural networks of claim 1 relied on for this claim, claim 1 reciting one or more such networks and its open transition not excluding the predictors Graves designs (MPEP 2111.03), the bank supplying the several networks claim 11 refers to in the plural, to teach further comprising, at each of one or more time steps: updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks. This modification would have been motivated by the desire to supplement the designed predictions with further predictions found from experience, where which further questions would serve the controller is not apparent in advance (Schlegel: page 526, "An approach to discovery will also enable GVFNs to be applied to problems in which a set of GVF questions is not immediately apparent"; page 527, "the agent should focus on finding a set of questions which is useful for the agents overarching goals-for example, maximizing the return in the control problem").
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Graves in view of Zheng and Schlegel, as applied in the rejection of claim 11 above, and further in view of Schwab et al. (hereinafter Schwab) "Not to Cry Wolf: Distantly Supervised Multitask Learning in Critical Care" 2018.
Regarding dependent claim 12, Graves, in view of Zheng and Schlegel, teach the method of claim 11.
Graves, Zheng and Schlegel do not expressly teach wherein updating the data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks comprises: determining, for each feature in the set of features, a respective second importance score characterizing an importance of the feature to predicting main task rewards; and updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks based on the second importance scores.
However, Schwab teaches determining, for each feature in the set of features, a respective second importance score characterizing an importance of the feature to predicting main task rewards (Schwab: page 4, 3.1. Selection of Auxiliary Tasks, "We statistically test the extracted features for their importance related to the main task in order to rank the features by their estimated predictive potential and determine their relevance"; every extracted feature is scored for how much it bears on the signal its main task is trained on, and the scores rank the features and fix which of them are relevant; each such score is a respective second importance score, page 4, 3.1. Selection of Auxiliary Tasks, "A suitable statistical test is, for example, a hypothesis test for correlation between the labels ytrue and the extracted features using Kendall’s τ"; the score is computed between the feature and the signal the main task is trained on); and updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks based on the second importance scores (Schwab: page 4, 3.1. Selection of Auxiliary Tasks, "we are able to identify a large, ranked list of predictive features suitable for use as target labels for auxiliary tasks"; the features the scores rank are what the auxiliary tasks are then set to predict; page 4, 3.1. Selection of Auxiliary Tasks, "There are two approaches to choosing a subset of those features as auxiliary targets: (i) in order of feature importance or (ii) randomly out of the set of relevant features"; on either approach the targets are taken from the features the scores have made relevant, the first taking them in the order of the scores themselves (updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks based on the second importance scores)).
Because Graves, in view of Zheng and Schlegel, and Schwab are analogous art, Graves, Zheng and Schlegel being in the field of the claimed invention, the use of auxiliary predictions to build the state representation a reinforcement learning controller acts on, and Schwab being reasonably pertinent to the problem the inventor faced of choosing which received features the auxiliary predictions are set on, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Schwab's ranking of the received features by their importance to the main task to the method of Graves, in view of Zheng and Schlegel, with a reasonable expectation of success, by scoring each of the vehicle state parameters of [0074], in the way Schwab scores a feature against the signal its main task is trained on, for its bearing on the main task reward the AS controller module 412 is trained on, Graves: [0168], "where b1, b2, b3, and b4 are constants defining the relative importance of front safety, back safety, and closeness to target speed", and by drawing the cumulant of each newly generated predictor from the parameters those scores make relevant rather than from all of them, the scores being recomputed as the vehicle runs, when Schlegel's generate-and-test re-targets a predictor, to teach wherein updating the data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks comprises: determining, for each feature in the set of features, a respective second importance score characterizing an importance of the feature to predicting main task rewards; and updating data that defines the respective target features that specify the auxiliary rewards for the auxiliary prediction neural networks based on the second importance scores. This modification would have been motivated by the desire to set the added predictors on signals related to the task the controller is trained for (Schwab: page 4, 3.1. Selection of Auxiliary Tasks, "rank the features by their estimated predictive potential and determine their relevance").
Claims 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over Graves in view of Zheng, Schlegel and Schwab, as applied in the rejection of claim 12 above, and further in view of Veeriah et al. (hereinafter Veeriah) "Discovery of Useful Questions as Auxiliary Tasks" 2019 and Martin.
Regarding dependent claim 13, Graves, in view of Zheng, Schlegel and Schwab, teach the method of claim 12, wherein determining the respective second importance score for each feature in the set of features comprises (Schwab: page 4, 3.1. Selection of Auxiliary Tasks, "We statistically test the extracted features for their importance related to the main task in order to rank the features by their estimated predictive potential and determine their relevance"; every extracted feature is scored for how much it bears on the signal its main task is trained on, and the scores rank the features and fix which of them are relevant; each such score is a respective second importance score, page 4, 3.1. Selection of Auxiliary Tasks, "A suitable statistical test is, for example, a hypothesis test for correlation between the labels ytrue and the extracted features using Kendall’s τ"; the score is computed between the feature and the signal the main task is trained on).
Graves, Zheng, Schlegel and Schwab do not expressly teach obtaining a main task reward estimation function that is configured to process the set of features representing an observation for a time step to generate a prediction for a main task reward received at a next time step.
However, Veeriah teaches obtaining a main task reward estimation function that is configured to process the set of features representing an observation for a time step to generate a prediction for a main task reward received at a next time step (Veeriah: page 6, 3.3 Baselines: handcrafted questions as auxiliary tasks, "uses the scalar reward obtained at the next time step as the target for the answer network"; the answer network of that baseline is fitted to the reward received at the following time step (an observation for a time step to generate a prediction), so that answer network (obtaining a main task reward estimation function) returns an estimate of that reward (to generate a prediction for a main task reward received at a next time step), page 4, 2.4 An actor critic agent with discovery of questions for auxiliary tasks, "In this paper, functions π, v and y will be linear functions of state xt"; the estimating function is linear in whatever is given to it as the state (that is configured to process the set of features representing)).
Because Graves, in view of Zheng, Schlegel and Schwab, and Veeriah are analogous art, Graves, Zheng, Schlegel and Veeriah being in the field of the claimed invention, the use of auxiliary predictions to build the state representation a reinforcement learning controller acts on, and Schwab being reasonably pertinent to the problem the inventor faced of choosing which received features the auxiliary predictions are set on, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Veeriah's reward prediction task to the method of Graves, in view of Zheng, Schlegel and Schwab, with a reasonable expectation of success, by learning, as the vehicle runs, a function linear in the vehicle state parameters that is fitted at each sample time to the reward the AS controller module 412 is trained on, the state given to that function being the vehicle state parameters Graves receives, so that it carries one parameter for each of the received features, to teach wherein determining the respective second importance score for each feature in the set of features comprises: obtaining a main task reward estimation function that is configured to process the set of features representing an observation for a time step to generate a prediction for a main task reward received at a next time step. This modification would have been motivated by the desire to obtain a measure of what bears on the reward from a quantity the system already computes rather than from a separate test (Veeriah: page 6, 3.3 Baselines: handcrafted questions as auxiliary tasks, "uses the scalar reward obtained at the next time step as the target for the answer network").
Graves, Zheng, Schlegel, Schwab and Veeriah do not expressly teach determining the second importance score for each feature in the set of features using the main task reward estimation function.
However, Martin teaches determining the second importance score for each feature in the set of features using the main task reward estimation function (Martin: page 4, 5.1 How do GVF weights inform neighborhoods, "Similar to simple linear regression, prediction weights reflect the importance of each feature. High magnitude weights suggest more importance than those with low magnitude"; the score standing against a component is read off the weight that component carries in the fitted function (determining the second importance score for each feature using the main task reward estimation function)).
Because Graves, in view of Zheng, Schlegel, Schwab and Veeriah, and Martin are analogous art, Graves, Zheng, Schlegel and Veeriah being in the field of the claimed invention, the use of auxiliary predictions to build the state representation a reinforcement learning controller acts on, and Schwab and Martin being reasonably pertinent to the problem the inventor faced of choosing which received features the auxiliary predictions are set on, Martin learning its predictions in a Markov reward process (Martin: page 2), accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Martin's reading of importance from fitted weights to the method of Graves, in view of Zheng, Schlegel, Schwab and Veeriah, with a reasonable expectation of success, by taking the magnitude of the weight each vehicle state parameter carries in the estimating function just added as the score by which the cumulant of a newly generated predictor is drawn, to teach determining the second importance score for each feature in the set of features using the main task reward estimation function. This modification would have been motivated by the desire to obtain the scores the selection of Schwab needs without a separate statistical test, by reading them from a function the system is already fitting (Martin: page 4, 5.1 How do GVF weights inform neighborhoods, "prediction weights reflect the importance of each feature").
Regarding dependent claim 14, Graves, in view of Zheng, Schlegel, Schwab, Veeriah and Martin, teach the method of claim 13, wherein the main task reward estimation function is a linear function that comprises a respective parameter corresponding to each feature in the set of features (Veeriah: page 4, 2.4 An actor critic agent with discovery of questions for auxiliary tasks, "In this paper, functions π, v and y will be linear functions of state xt"; the estimating function is linear in what it is given, so it carries one weight for each component of the state it is given), and wherein determining the second importance score for each feature in the set of features using the main task reward estimation function comprises: determining the second importance score for each feature based on a value of the corresponding parameter of the main task reward estimation function (Martin: page 4, 5.1 How do GVF weights inform neighborhoods, "High magnitude weights suggest more importance than those with low magnitude"; the score assigned to a component is the size of that component's own weight, read off the fitted function (determining the second importance score for each feature based on a value of the corresponding parameter)).
Regarding dependent claim 15, Graves, in view of Zheng, Schlegel, Schwab, Veeriah and Martin, teach the method of claim 13, wherein for each time step in the sequence of time steps, the main task reward estimation function is trained based on the main task reward for the time step using supervised learning (Veeriah: page 6, 3.3 Baselines: handcrafted questions as auxiliary tasks, "uses the scalar reward obtained at the next time step as the target for the answer network. The auxiliary task loss function for the reward prediction baseline is"; the estimating function is fitted by minimizing the squared difference between what it returns and the reward supplied as its target, with nothing bootstrapped from its own estimate, which is training on a supplied target (using supervised learning), the target is the reward received at the following step, so a fit is made for each step at which a reward arrives (for each time step in the sequence of time steps, the main task reward estimation function is trained based on the main task reward for the time step)).
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Graves in view of Zheng, as applied in the rejection of claim 1 above, and further in view of Schmid et al. (hereinafter Schmid), US 2011/0087627 A1, and Iqbal et al. (hereinafter Iqbal) "Randomized Entity-wise Factorization for Multi-Agent Reinforcement Learning" 2021.
Regarding dependent claim 17, Graves, in view of Zheng, teach the method of claim 1.
Graves and Zheng do not expressly teach further comprising, prior to the first time step in the sequence of time steps and for each auxiliary prediction neural network: selecting a proper subset of the set of features to be included in the auxiliary input to the auxiliary prediction neural network, comprising: randomly sampling a proper subset of the set of features; and designating the randomly sampled proper subset of the set of features for inclusion in the auxiliary input to the auxiliary prediction neural network.
However, Schmid teaches further comprising, prior to the first time step in the sequence of time steps and for each auxiliary prediction neural network (Schmid: [0027], "One neural network 106a may be trained using a random subset of the available data 102a, for example, the engine temperature, the battery voltage, and the CO2 emissions"; the variables a network is given are drawn for it when it is trained; [0028], "Then the trained neural networks may receive current data (via evaluation point inputs 108a, 108b, . . . 108n) from the vehicle sensors"; the networks receive the vehicle's measurements only once trained, so the selection precedes the first of those measurements (prior to the first time step in the sequence of time steps); [0016], "multiple neural network processes A, B, . . . , N may proceed in parallel, each with corresponding training data (102a, 102b, . . . 102n)"; a selection is made for each of the networks (for each auxiliary prediction neural network)): selecting a proper subset of the set of features to be included in the auxiliary input to the auxiliary prediction neural network, comprising (Schmid: [0027], "The vehicle may be equipped with sensors to continuously measure information that might be related to the fuel consumption, for example: engine temperature, air temperature, engine RPM (revolutions per minute), accelerator position, accessory load, vehicle speed, battery voltage, and carbon dioxide emissions"; the measured information (a set of features representing a current observation) is what each network's inputs are chosen from; [0027], "One neural network 106a may be trained using a random subset of the available data 102a, for example, the engine temperature, the battery voltage, and the CO2 emissions"; the selection made for the first network names three of those eight measured variables, fewer than all of them (selecting a proper subset of the set of features)): randomly sampling a proper subset of the set of features (Schmid: [0027], "One neural network 106a may be trained using a random subset of the available data 102a, for example, the engine temperature, the battery voltage, and the CO2 emissions"; [0027], "Another neural network 106b may be trained using a different random subset of the available data 102b, for example, the vehicle speed, the engine RPM, and the engine temperature"; the variables each network is given are drawn at random from the measured variables, a different draw for each network (randomly sampling a proper subset of the set of features)), and designating the randomly sampled proper subset of the set of features for inclusion in the auxiliary input to the auxiliary prediction neural network (Schmid: [0030], "Neural networks A and B, at this point, are already assumed to be trained using training data 102a, 102b that corresponds to these measured variables"; the inputs a network is given when it runs correspond to the measured variables drawn for its training (designating the randomly sampled proper subset of the set of features for inclusion in the auxiliary input to the auxiliary prediction neural network)).
Because Graves, in view of Zheng, and Schmid are analogous art, all three being in the field of endeavor of the claimed invention, the application of neural networks to the measured variables of a vehicle in operation, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Schmid's drawing of each network's input variables at random to the method of Graves, in view of Zheng, with a reasonable expectation of success, by adding to the speed and safety predictors Graves designs one or more further predictor general value functions of the kind Graves contemplates to teach further comprising, prior to the first time step in the sequence of time steps and for each auxiliary prediction neural network: selecting a proper subset of the set of features to be included in the auxiliary input to the auxiliary prediction neural network, comprising: randomly sampling a proper subset of the set of features; and designating the randomly sampled proper subset of the set of features for inclusion in the auxiliary input to the auxiliary prediction neural network. This modification would have been motivated by the desire to obtain predictors whose inputs differ from one another without an expert selecting them (Schmid: [0015]).
Response to Arguments
Applicant's arguments filed 07/15/2026 have been fully considered. They are addressed below in the order presented. Where an argument is directed to a rejection that is withdrawn in this action, it is answered to the extent it bears on the grounds now of record.
Arguments directed to 35 U.S.C. 101 are persuasive. Applicant argues that amended claims 1, 19 and 20 recite an electromechanical agent acting in a real-world environment and that the claims are directed to a specific architecture rather than to the abstract idea itself. As set out above, the rejection is withdrawn on reconsideration. The claims recite the supply of an intermediate output generated by one or more hidden layers of each auxiliary prediction neural network to the action selection neural network, a limitation present in claim 1 as filed and made express by the amendment, together with the proper subset of received features supplied to each auxiliary prediction neural network, and the specification describes the improvement those limitations produce at [0011] and at FIGS. 7 through 10. The claims therefore integrate the recited judicial exception into a practical application. The rejection of claims 1 to 20 under 35 U.S.C. 101 is withdrawn. The argument answered here is that presented at pages 11 to 14 of the remarks.
Arguments directed to 35 U.S.C. 112(b) are persuasive. Each antecedent basis defect identified in the prior action is cured by the amendment. The rejection of claims 1 to 20 under 35 U.S.C. 112(b) is withdrawn. The argument answered here is that presented at page 15 of the remarks.
Arguments directed to the rejections under 35 U.S.C. 103 as applied are moot in view of new grounds of rejection. Applicant argues, at pages 14 to 15 of the remarks, that Veeriah in view of Jaderberg does not teach extracting an intermediate output generated by one or more hidden layers of an auxiliary prediction neural network and supplying it to the action selection neural network, and that the dependent claims and claims 19 and 20 are allowable for the same reasons. The argument is persuasive as to the rejections as applied, and those rejections are withdrawn. The withdrawal does not place the claims in condition for allowance. New grounds of rejection necessitated by the amendment are entered above over a different combination of art, which includes two references applied in that action against different limitations, and Applicant's arguments are directed to references that those grounds do not rely on for the limitations at issue.
The teaching-away argument is not persuasive. Applicant argues, at page 15 of the remarks, that Veeriah teaches away because it states that the question and answer networks are only used during training and that neither is needed for action selection. A reference teaches away only when it criticizes, discredits or otherwise discourages the claimed solution. See MPEP 2145(X)(D) and In re Fulton, 391 F.3d 1195 (Fed. Cir. 2004). A statement that a component is not required for a particular purpose is not a criticism of using it, and a reference does not teach away merely by disclosing an alternative or preferred embodiment. The sentence Applicant quotes also does not reach the use made of Veeriah in the rejection of claims 13 to 15 entered above, which applies Veeriah's reward prediction baseline, an auxiliary task its agent trains while learning, as the estimator whose weights score the received features, and not as the network that selects the agent's action.
Arguments directed to the objections are persuasive. The amendments to [0011], [0041], [0044], [0096] and [0108] and to claim 7 cure the informalities identified in the prior action, and amended [0096] describes reference character 1000. The objections to the drawings, to the disclosure and to claim 7 set forth in the prior action are withdrawn. The arguments answered here are those presented at page 16 of the remarks. The objections stated above are new and are directed to informalities in the claims and in the specification as they now read.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KUANG FU CHEN whose telephone number is (571)272-1393. The examiner can normally be reached M-F 9:00-5:30pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KC CHEN/Primary Patent Examiner, Art Unit 2143