Prosecution Insights
Last updated: October 02, 2026
Application No. 18/205,781

MULTI-OBJECTIVE MULTI-POLICY REINFORCEMENT LEARNING SYSTEM

Non-Final OA §101§103
Filed
Jun 05, 2023
Examiner
CARDOSO, JUSTIN ALEXANDER
Art Unit
Tech Center
Assignee
Hitachi Ltd.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
11 currently pending
Career history
7
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION This action is in response to the original filing on 06/05/2023. Claims 1-18 are pending and have been considered below. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 4 and 13 are objected to because of the following informalities: “reply buffer” should read as “replay buffer” in keeping with the specification. Appropriate correction is required. Claim 13 is a non-transitory computer readable storage claim corresponding to the method of Claim 4 and is objected to for the same reason. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 4-7, 9, 10, 13-16, and 18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more. Claims 1 and 10 Step 1: Claims 1 and 10 recite a method and non-transitory computer-readable medium. As such, the claims are directed to the statutory categories of a method and product of manufacture. Step 2A Prong 1: The claims recite, inter alia: “learning a value function…wherein the value function is configured to take in an input of a state and an action pair, and” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than calculating a formula through a mathematical process. The claims further recite “determining a sequence of actions iteratively based on the output of the value function, an observation of the current state, and the request” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than a mental process comprising a determination and/or judgement, easily performable in the mind by a human being. Step 2A Prong 2: The judicial exception is not sufficiently integrated into a practical application. The additional elements in Claims 1 and 10 recite: “through reinforcement learning (RL)” which amounts to no more than merely linking a judicial exception to a particular technological environment (See MPEP 2106.05(h)). The claims further recite: “provides a set of vectors as output, each of the set of vectors representing an expected total sum of rewards corresponding to a sequence of future control decisions” This amounts to no more than merely linking a judicial exception to a particular technological environment, particularly limiting an abstract idea of collecting information, analyzing it, and displaying results. (See MPEP 2106.05h FoU vi. Limiting the abstract idea of collecting information, analyzing it, and displaying certain results of the collection and analysis) The claims further recite: “receiving, at an initial stage of a control sequence, a request about a total sum of rewards to be achieved” This amounts to no more than receiving a request, meaning this amounts to no more than mere instruction to apply an exception. (See MPEP 2106.05f invokes computers or other machinery merely as a tool to perform an existing process) Step 2B: The claims do not contain significantly more than the judicial exceptions. The additional elements in Claims 1 and 10 recite: “through reinforcement learning (RL)” which amounts to no more than merely linking a judicial exception to a particular technological environment (See MPEP 2106.05(h)). The claims further recite: “provides a set of vectors as output, each of the set of vectors representing an expected total sum of rewards corresponding to a sequence of future control decisions” This amounts to no more than merely linking a judicial exception to a particular technological environment, particularly limiting an abstract idea of collecting information, analyzing it, and displaying results. (See MPEP 2106.05h FoU vi. Limiting the abstract idea of collecting information, analyzing it, and displaying certain results of the collection and analysis) The claims further recite: “receiving, at an initial stage of a control sequence, a request about a total sum of rewards to be achieved” This amounts to no more than receiving a request, meaning this amounts to no more than mere instruction to apply an exception. (See MPEP 2106.05f invokes computers or other machinery merely as a tool to perform an existing process) Considering the additional elements individually and in combination, the claims are directed to the judicial exceptions without significantly more. Claims 4 and 13 Step 1: Claims 4 and 13 recite a method and non-transitory computer-readable medium. As such, the claims are directed to the statutory categories of a method and product of manufacture. Step 2A Prong 1: The claims recite, inter alia: “determining a temporal difference (TD) error for each data in the mini-batch;”, “determining a loss based on the TD error; and”, and “updating the value function based on a gradient of the loss.”. Under its broadest reasonable interpretation in light of the specification, these limitations amount to no more than a determination and/or judgement easily performable by a human. Determining an error for data in a batch, determining a loss based on that error, and updating a mathematical function is performable by a human with aid of pen and paper. Step 2A Prong 2: The claims recite the additional elements of: “obtaining state transition data from the system and storing the state transition data in a reply buffer” and “drawing a mini-batch of random transitions from the reply buffer”. These limitations amount to no more than data gathering, as all that is being done is obtaining data from a system and ‘drawing’ (obtaining) data from a buffer. (See MPEP 2106.05(g)) Step 2B: The claims do not contain significantly more than the judicial exceptions. The claims recite: “obtaining state transition data from the system and storing the state transition data in a reply buffer” and “drawing a mini-batch of random transitions from the reply buffer”. These limitations amount to no more than data gathering, as all that is being done is obtaining data from a system and ‘drawing’ (obtaining) data from a buffer. (See MPEP 2106.05(g)) Considering the additional elements individually and in combination, the claims are directed to the judicial exceptions without significantly more. Claims 5 and 14 Step 1: Claims 5 and 14 recite a method and non-transitory computer-readable medium. As such, the claims are directed to the statutory categories of a method and product of manufacture. Step 2A Prong 1: The claims recite, inter alia: “wherein the loss is computed from the TD error based on a sample-based distributional distance metric” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than performing the mental process of a mathematical calculation, as computing an amount of loss based on a distance metric is easily performable by a human with aid of a pen and paper. Step 2A Prong 2 and Step B: There are no additional elements recited, so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible. Claims 6 and 15 Step 1: Claims 6 and 15 recite a method and non-transitory computer-readable medium. As such, the claims are directed to the statutory categories of a method and product of manufacture. Step 2A Prong 1: The claims recite, inter alia: “wherein a target of the TD error for each data in the mini-batch is computed by first taking a union set of the output of the value function over all possible actions and then selecting a subset of points in the union set” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than performing the mental process of a mathematical calculation, as computing an error based on a union set of the output of a value function over all possible actions and then selection of a subset of points in the set is a determination easily performable by a human with aid of pen and paper. Step 2A Prong 2 and Step B: There are no additional elements recited, so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible. Claims 7 and 16 Step 1: Claims 7 and 16 recite a method and non-transitory computer-readable medium. As such, the claims are directed to the statutory categories of a method and product of manufacture. Step 2A Prong 1: The claims recite, inter alia: “ranking all points in the union set according to a distance metric from a Pareto front of an entirety of the union set; and” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than performing the mental process of a determination or judgement, as all that is being done is a ranking of points according to distance, easily performable by a human. The claims also recite: “after the ranking of the all points, selecting a portion of top-ranked ones of the all points, the selecting comprising selecting the Pareto front of the entirety of the union set.” Under its broadest reasonable interpretation in light of the specification, this amounts to no more than performing the mental process of a determination or judgement, as all that is being done is a selection of top ranked points, easily performable by a human. Step 2A Prong 2 and Step B: There are no additional elements recited, so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible. Claims 9 and 18 Step 1: Claims 9 and 18 recite a method and non-transitory computer-readable medium. As such, the claims are directed to the statutory categories of a method and product of manufacture. Step 2A Prong 1: The claims recite, inter alia: “after the request is received, repeating a process until a terminating condition is satisfied, the process comprising”, under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than a determination or judgement, as continuation of a performance of steps determinant on a condition is easily performed by a human being. The claims also recite: “observing a current state” under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than a mental process of observation, and is thus easily performed by a human. The claims also recite: “selecting the action such that the output by the value function for that action is within a threshold of the targeted total sum of rewards” under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than a determination or judgement, as selection of an action so that an output of a function is within a threshold is determinant of an observation, and thus easily performed by a human. The claims also recite: “executing the selected action” under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than a determination or judgement, as execution of an action is easily performed by a human. The claims also recite: “updating the targeted sum of rewards by subtracting a latest received reward from the target sum of rewards and then dividing a remainder by a temporal discount factor” under its broadest reasonable interpretation in light of the specification, this limitation amounts to no more than a determination or judgement, as updating a numerical target by subtraction and then division is a determination based on a mathematical process. As such, it is a mental process and an abstract idea. Step 2A Prong 2: The claims recite the additional elements of: “receiving the request for the total sum of rewards through a user interface;”, “collecting output sets of the value function for all actions;”, and “receiving a reward and observing a next state; and” amount to no more than mere data gathering, as all three limitations state the collection of reward values or outputs of a value function for actions (See MPEP 2106.05(g)). Step 2B: The claims do not recite significantly more than the judicial exception. The additional elements of: “receiving the request for the total sum of rewards through a user interface;”, “collecting output sets of the value function for all actions;”, and “receiving a reward and observing a next state; and” amount to no more than mere data gathering, as all three limitations state the collection of reward values or outputs of a value function for actions (See MPEP 2106.05(g)). Considering the additional elements individually and in combination, the claims are directed to the judicial exceptions without significantly more. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 9, 10, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Abdolmaleki et al. (US 20230082326 A1, hereinafter Abdolmaleki 326) in view of Abdolmaleki et al. (US 20240185084 A1, hereinafter Abdolmaleki 084). Regarding Claim 1, Abdolmaleki 326 teaches a method for obtaining Pareto optimal solutions through making sequential decisions in a system that has multi-dimensional rewards and a continuous state space, and is controllable through a finite discrete set of actions, (Paragraph [0006] Each trajectory comprises a state of an environment, an action applied by the agent to the environment according to a previous policy in response to the state, and a set of rewards for the action, each reward relating to a corresponding objective of the plurality of objectives. Paragraph [0076] Ultimately, the present methodology provides a distributional view on multi-objective reinforcement learning (MORL), which enables scale-invariant encoding of preferences. This is a theoretically-grounded approach, that arises from taking an RL-as-inference perspective of MORL. Empirically, the mechanics of MO-MPO have been analysed and shown that it finds all Pareto-optimal policies in a popular MORL benchmark task. MO-MPO and MO-V-MPO outperform scalarized approaches on multi-objective tasks across several challenging high-dimensional continuous control domains. (Obtaining Pareto optimal policies through reinforcement learning)) the method comprising: learning a value function through reinforcement learning (RL), wherein the value function is configured to take in an input of a state and an action pair, and (Paragraph [0005] This specification generally describes methods for training a reinforcement learning system that selects actions to be performed by a reinforcement learning agent interacting with an environment. The methods can be used to train a reinforcement learning system which has multiple, potentially conflicting objectives Paragraph [0036] Alternatively, each probability distribution may be an objective-specific policy (an action distribution) defining a distribution of probabilities of actions over states and the value function may be an action-value function representing an expected return according to the corresponding objective that would result from the agent performing a given action in response to a given state according to the previous policy. Paragraph [0093] where {circumflex over (Q)}.sub.k (S, A, R) is a target action-value function (a target Q-function) based on state S, action A and reward R vectors. (Learning a value function with inputs of a state S and action A through reinforcement learning)) provides a set of vectors as output, each of the set of vectors representing an expected total sum of rewards corresponding to a sequence of future control decisions; (Paragraph [0010] Each action-value function may output an action-value representing the expected cumulative discounted reward for the corresponding objective when choosing a given action in response to a given state. This cumulative discounted reward may be calculated over a number of subsequent actions, implemented in accordance with the previous policy. The action-value function for each objective may be considered an objective-specific action-value function. Paragraph [0071] Multi-objective reinforcement learning (MORL) methods aim to tackle such problems. One approach is scalarization: based on preferences across objectives, transform the multi-objective reward vector into a single scalar reward (e.g., by taking a convex combination), and then use standard RL to optimize this scalar reward. (Outputting an action-value representing a cumulative (sum) of rewards corresponding to subsequent (future) decisions, wherein the rewards are held in a vector before then being combined to form a singular scalar reward. See also paragraph [0115] for specificity in denoting the reward vector used for N objectives)) determining a sequence of actions iteratively based on the output of the value function, an observation of the current state, and the request. (Paragraph [0078] The neural network system 100 includes an action selection policy neural network 110 that determines actions 102 that are output to an agent 104 for application to an environment. The neural network system 100 operates over a number of time steps t. Each action a.sub.t 102 is determined based on an observation characterizing a current state s.sub.t 106 of the environment. Following the input of an initial observation characterizing an initial state s.sub.0 106 of the environment, the neural network system 100 determines an action a.sub.0 102 and outputs this action 102 to the agent 104. After the agent 104 has applied the action 102 to the environment 104, an observation of an updated state s.sub.1 106 is input into the neural network 100. The neural network 100 therefore operates over multiple time steps t to select actions a.sub.t 102 in response to input observations s.sub.t 106. For each time step, a set of rewards r.sub.t 108 for the previous action a.sub.t−1 is also received. The set of rewards r.sub.t 108 includes a reward for each objective. Paragraph [0082] In addition to objective-specific rewards, an action-value function (Q-function) is provided for each objective. Paragraph [0141] In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser. (Selection of actions based on rewards (output of value function) and observation of the current state as a result from a user device request)) Abdolmaleki 326 does not teach receiving, at an initial stage of a control sequence, a request about a total sum of rewards to be achieved; and In the same field of endeavor, Abdolmaleki 084 teaches receiving, at an initial stage of a control sequence, a request about a total sum of rewards to be achieved; and (Paragraph [0033] In some implementations the system is configured to train the action selection policy neural network 120 online whilst optimizing multiple different Q-values for the task, each representing a different target objective, such as a different reward, or a cost (i.e. a negative reward), as the agent attempts to perform the task, e.g. to maximize the reward(s) or to minimize the cost(s). Example costs in a real-world environment can include a penalty, e.g. for power or energy use, or for mechanical wear-and-tear. (A target reward objective is specified before steps are taken to achieve that objective, i.e. at the initial stage before the subsequent steps (control sequence) are undertaken)) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have combined receiving a requested total sum of rewards to be achieved as taught by Abdolmaleki 084 into Abdolmaleki 326 as both inventions are in the same field of reinforcement learning when dealing with multi-objective optimization problems, and doing so would provide obvious benefits, one being allowing for neural network to solve multi-objective problems in a more efficient way through the use of cumulative, time-discounted rewards (Abdolmaleki 084 Paragraphs [0002] – [0004]). Regarding Claim 9, the combination of Abdolmaleki 326 and Abdolmaleki 084 teaches the invention as claimed in Claim 1, including: receiving the request for the total sum of rewards through a user interface; and (Abdolmaleki 084 Paragraph [0033] In some implementations the system is configured to train the action selection policy neural network 120 online whilst optimizing multiple different Q-values for the task, each representing a different target objective, such as a different reward, or a cost (i.e. a negative reward), as the agent attempts to perform the task, e.g. to maximize the reward(s) or to minimize the cost(s). Example costs in a real-world environment can include a penalty, e.g. for power or energy use, or for mechanical wear-and-tear. (A target reward objective is specified before steps are taken to achieve that objective, i.e. at the initial stage before the subsequent steps (control sequence) are undertaken through a user interface)) after the request is received, repeating a process until a terminating condition is satisfied, the process comprising: (Abdolmaleki 326 Paragraph [0085] FIG. 2 shows a method for training via multi-objective reinforcement learning according to an arrangement. The method splits the reinforcement learning problem into two sub-problems and iterates until convergence: Paragraph [0089] The method determines if an end criterion is reached 240 (e.g. a fixed number of iterations have been performed, or the policy meets a given level of performance). ()) observing a current state; (Abdolmaleki 326 Paragraph [0006] In one aspect there is provided a method for training a neural network system by reinforcement learning, the neural network system being configured to receive an input observation characterizing a state of an environment interacted with (observing a state)) collecting output sets of the value function for all actions; (Abdolmaleki 326 Paragraph [0099] Algorithm 2 is a method of obtaining trajectories based on a given policy π.sub.θ. The system fetches the current policy parameters θ defining the current policy π.sub.θ. The system then collects a set of trajectories over a number of time steps T. Each trajectory τ includes a state s.sub.t, an action a.sub.t and a set of rewards r for each of the timesteps. (Collecting trajectories (outputs) of all actions)) selecting the action such that the output by the value function for that action is within a threshold of the targeted total sum of rewards; (Abdolmaleki 326 Paragraph [0006] In one aspect there is provided a method for training a neural network system by reinforcement learning, the neural network system being configured to receive an input observation characterizing a state of an environment interacted with by an agent and to select and output an action in accordance with a policy that aims to satisfy a plurality of objectives. Each trajectory comprises a state of an environment, an action applied by the agent to the environment according to a previous policy in response to the state, and a set of rewards for the action, each reward relating to a corresponding objective of the plurality of objectives. Paragraph [0014] The minimization of the difference between the updated policy and the combination of the objective-specific policies may be constrained such that a difference between the updated policy and the previous policy does not exceed a trust region threshold. In other words, the set of policy parameters for the updated policy may be constrained such that a difference between the updated policy and the previous policy cannot exceed the trust region threshold. The trust region threshold may be considered a hyperparameter that limits the overall change in the policy to improve stability of learning. Paragraph [0016] In some implementations determining an objective-specific policy for each objective comprises determining objective-specific policy parameters for the objective-specific policy that increase the expected return according to the action-value function for the corresponding objective relative to the previous policy. (Selection and output of an action, such that the output does not violate a hyperparameter threshold that would disturb the learning in unintended ways)) executing the selected action; (Abdolmaleki 326 Paragraph [0067] The agent performs the selected action which results in a change in the state of the environment.) receiving a reward and observing a next state; and (Abdolmaleki 326 Paragraph [0010] Each action-value function may output an action-value representing the expected cumulative discounted reward for the corresponding objective when choosing a given action in response to a given state. This cumulative discounted reward may be calculated over a number of subsequent actions, implemented in accordance with the previous policy. (Receiving a reward and observing a next state (subsequent actions))) updating the targeted sum of rewards by subtracting a latest received reward from the target sum of rewards and then dividing a remainder by a temporal discount factor. (Abdolmaleki 326 Paragraph [0006] The method further comprises determining an updated policy based on a combination of the action-value functions for the plurality of objectives. [0012] The relative contribution of each objective to the updated policy can be scaled through use of constraints on the determination of the objective-specific policies. Paragraph [0081] Accordingly, a reward function {r.sub.k(s,a)∈custom-character}.sub.k=1.sup.N is assigned for each objective k. A discount factor γ∈[0,1) is provided for application to the rewards. Paragraph [0082] In addition to objective-specific rewards, an action-value function (Q-function) is provided for each objective. The action-value function maps states and actions to a value. The action-value function for objective k is defined as the expected return (i.e., the cumulative discounted reward) from choosing action a in state s for objective k and then following policy π: Q.sub.k.sup.π(s, a)=custom-character.sub.π[Σ.sub.t=0.sup.∞γ.sup.tr.sub.k(s.sub.t, a.sub.t)|s.sub.0=s, a.sub.0=a]. (Iterative selection of actions that result in cumulative rewards, which are updated after every iteration and a discount factor is applied to each reward)) Regarding Claims 10 and 18, Claims 10 and 18 are non-statutory computer readable storage medium claim corresponding to the method claim of Claims 1 and 9. As such, they are rejected for the same reasons. Claims 2-3 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Abdolmaleki 326 in view of Abdolmaleki 084 as applied in Claims 1 and 10 above, and further in view of Budan et al. (US 20240371264 A1, hereinafter Budan). Regarding Claim 2, the combination of Abdolmaleki 326 and Abdolmaleki 084 teaches the invention as claimed in Claim 1, including: the value function is parameterized by a neural network model that samples multiple random variables from a prefixed probability distribution, (Abdolmaleki 084 Paragraph [0030] The policy output 122 may define the action directly, e.g., it may comprise a value used to define a continuous value for an action such as a torque or velocity, or it may parameterize a continuous or categorical distribution from which a value defining the action may be selected, or it may define a set of scores, one for each action of a set of possible actions, for use in selecting the action. Merely as one example the policy output 122 may define a multivariate Gaussian distribution with a diagonal covariance matrix. Paragraph [0069] For example weights may be randomly sampled from a uniform or other distribution over the range [0,1] or systematically sampled, e.g. step-by-step over the range. (The function is parameterized from a prefixed probability distribution (Gaussian curve), and wherein the second action policies output is defined by random sampling of weights)) generates the set of vectors as output representing the expectation values of a total sum of rewards for the state and the action pair input. (Abdolmaleki 326 Paragraph [0010] Each action-value function may output an action-value representing the expected cumulative discounted reward for the corresponding objective when choosing a given action in response to a given state. (The output is a reward in the form of multi-objective vectors that represent the expected reward or cost for a given action-state input)). The combination of Abdolmaleki 326 and Abdolmaleki 084 does not teach: merges the multiple random variables through an embedding layer with the state and the action pair input, and In the same field of endeavor, Budan teaches: merges the multiple random variables through an embedding layer with the state and the action pair input, and (Paragraph [0091] The input portion 510 is configured to receive an observation tensor, e.g. as determined above based on kinematics data for a set of lead connected vehicles in each in-use lane and/or aggregate kinematics data for vehicles in addition to the lead connected vehicles, and to pre-process this prior to the spatio-temporal feature extraction portion 520. This is shown schematically by the illustrated state of the intersection in FIG. 5A. The observation tensor may comprise a joined or concatenated set of observation vectors for all the lanes of the intersection (e.g., an input vector) and the input portion 510 may include a normalization operation as described above. In FIG. 5A the input portion 510 comprises an input layer 512. This may comprise a linear transformation layer as explained above and may function similar to an embedding layer in other neural network architectures. The output of the input layer 512 is then passed to the spatio-temporal feature extraction portion 520. (The input layer (embedding layer) pre-processes (concatenates, aka merges) the observation tensor (defined as an action-state pair) and aggregated vehicle data (variables))) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have combined receiving a requested total sum of rewards to be achieved as taught by Budan into the combination of Abdolmaleki 326 and Abdolmaleki 084 as all three inventions are in the same field of reinforcement learning when dealing with multi-objective optimization problems, and doing so would provide obvious benefits, one such being for allowing the increased capacity of machine learning models handling optimization of traffic flow, in particular easing the integration of connected and autonomous vehicles into traffic flows when dealing with largely unpredictable human drivers (Budan Paragraphs [0002]-[0003]). Regarding Claim 3, the combination of Abdolmaleki 326, Abdolmaleki 084, and Budan teaches the invention as claimed in Claim 2, including: during a training phase of the neural network, synchronizing parameters of a target neural network with parameters of the neural network with the neural network model either periodically or continuously through a Polyak update scheme. (Budan Paragraph [0145] In certain examples, the target network parameters may be updated with a so-called “soft” update, where the weights of the actor and critic neural network architectures 830, 840 are aggregated with the existing weights of the actor and the critic target neural network architectures 832, 842, e.g. via polyak averaging into the target networks. Similar to the policy learning, the target network updates may be performed every other step to improve training performance stability. (Performing updates to a target network via polyak averaging (Polyak update))) Regarding Claims 11-12, Claims 11-12 are non-statutory computer readable storage medium claims corresponding to the method claims of Claims 2 - 3. As such, they are rejected for the same reasons. Claims 4-8 and 13-17 are rejected under 35 U.S.C. 103 as being unpatentable over Abdolmaleki 326 in view of Abdolmaleki 084 in additional view of Budan as applied in Claims 3 and 12 above, and in further view of Yeh et al. (US 20220014963 A1, hereinafter Yeh). Regarding Claim 4, the combination of Abdolmaleki 326, Abdolmaleki 084, and Budan teaches the invention as claimed in Claim 1, including: updating the value function based on a gradient of the loss. (Budan Paragraph [0036] Typically, the update is propagated according to the derivative of the weights of the neural network layers. For example, a gradient of the loss function with respect to the weights of the neural network layers may be determined and used to determine an update to the weights that minimizes the loss function. Paragraph [0110] At first, the neural network architecture may be initialised with a random parameter set and then may be updated every epoch in the direction that minimizes the loss value until a training stop criterion is reached, e.g. the loss decrement in one epoch reaches a plateau. Paragraph [0124] In the context of an actor-critic architecture, the actor version of the neural network architecture updates a policy parameter set, θ, for a policy π.sub.θ(a|s) as guided by a critic neural network architecture, where the parameter set 0 comprise the weights and/or biases of the neural network architecture that implements the policy. (Using a gradient loss to update the value function (Q-value function). See also Equation [00015], which exemplifies the mean-squared error loss function used to determine the loss which is used by the critic control to update the parameters.)) The combination of Abdolmaleki 326, Abdolmaleki 084, and Budan does not teach obtaining state transition data from the system and storing the state transition data in a reply buffer; drawing a mini-batch of random transitions from the reply buffer; determining a temporal difference (TD) error for each data in the mini-batch; determining a loss based on the TD error In the same field of endeavor, Yeh teaches: obtaining state transition data from the system and storing the state transition data in a reply buffer; (Paragraph [0143] Sixth, the agent then deploys the action for the target UE 601, and the reward is collected. The action, state, and reward transition are stored (e.g., in an experience data structure) in the replay buffer 930. (Storing transitions in a replay buffer)) drawing a mini-batch of random transitions from the reply buffer; (Paragraph [0144] Seventh, a minibatch of size N is randomly sampled from the replay buffer 930, with the i index referring to the i-th sample. (Random sampling of items (such as transitions) from the buffer)) determining a temporal difference (TD) error for each data in the mini-batch; (Paragraph [0144] Seventh, a minibatch of size N is randomly sampled from the replay buffer 930, with the i index referring to the i-th sample. The target for the temporal difference (TD) error computation y.sub.i is computed from the immediate reward for the contextual bandit problem. (Determining TD error based on data of a minibatch)) determining a loss based on the TD error; and (Paragraph [0144] Then, the critic network (θ.sup.Q) 905 is trained to minimize the following loss function (LF): (Equation [00003]) [0145] In the loss function (LF), y.sub.i is the TD target where y.sub.i=r.sub.i, r.sub.i is the reward, and c.sub.i is the encoded context (e.g., encoded version of contexts 921). Eighth, the actor network (θ.sup.μ) 903 is trained based on policy gradient to maximize the following reward function (RF2). (Equation 00003 calculates a loss using the TD error)) It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to have included Yeh into the combination of Abdolmaleki 326, Abdolmaleki 084, and Budan as all four references are in the related field of finding optimal solutions for multi-objective problems, and doing so would provide various benefits, such as increasing scalability and offering improvements concerning traffic management techniques, a form of multi-objective problem solving (Yeh Paragraph [0032]-[0034]). Regarding Claim 5, the combination of Abdolmaleki 326, Abdolmaleki 084, Budan, and Yeh teaches the invention as claimed in Claim 4, including: wherein the loss is computed from the TD error based on a sample-based distributional distance metric. (Yeh Paragraph [0056] Both the actor network 303 and the critic network 305 compute action predictions for a current state and generate a temporal difference (TD) error signal each time step. The input to the actor network 303 is a current state (which is a representation produced by the representation network 301), and the actor network 303 outputs a value representing an action chosen from an action space (e.g., a set of potential actions to take). The input to the critic network 305 is the current state (which is the representation produced by the representation network 301), and the critic network 305 outputs feedback or an approximation for the predicted action. The feedback may be a grade, rank, degree, rating, quality, or some other value indicating the adequacy or superiority of the predicted action in relation to other potential actions. Additionally or alternatively, the critic network 305 outputs one or more values for each state-action pair. In either case, the actor network 303 in turn attempts to improve its policy based on the approximation provided by the critic network 305. In other words, the actor network 303 determines the action a, and the critic network 305 evaluates the action a and determines adjustments to be made in the action determination/learning process. (The loss is computed from the temporal difference (TD) based on discrepancies (distance) between predicted and current states)) Regarding Claim 6, the combination of Abdolmaleki 326, Abdolmaleki 084, Budan, and Yeh teaches the invention as claimed in Claim 4, including: wherein a target of the TD error for each data in the mini-batch is computed by first taking a union set of the output of the value function over all possible actions and then selecting a subset of points in the union set. (Yeh Paragraph [0056] The input to the actor network 303 is a current state (which is a representation produced by the representation network 301), and the actor network 303 outputs a value representing an action chosen from an action space (e.g., a set of potential actions to take). Paragraph [0101] Additionally, a current state s.sub.i, current action a.sub.t, current reward r.sub.i, and a next state s.sub.i+1 may be stored in an experience replay buffer 530. The replay buffer 530 is a buffer of past experiences, which is used to stabilize training by decorrelating the training examples in each batch used to update the NN 503. The replay buffer 530 records past states s.sub.i, the actions a.sub.i taken at those states s.sub.i, the received reward r.sub.i from those actions a.sub.i, and the next state s.sub.i+1 that was observed. The replay buffer 530 (or the data stored therein) can be used to prevent action values from oscillating or diverging catastrophically by buffering past experience and sampling (training) data 531 from the replay buffer 530, instead of using the latest experience. The act of sampling a small batch of tuples 531 from the replay buffer 530 in order to learn is known as experience replay. (The replay buffer stores past observations (from a collated set, acting as a union of actions) (such as the outputs of value functions). Then a subset of points in the buffer are selected to form the minibatch)) Regarding Claim 7, the combination of Abdolmaleki 326, Abdolmaleki 084, Budan, and Yeh teaches the invention as claimed in Claim 6, including: ranking all points in the union set according to a distance metric from a Pareto front of an entirety of the union set; and (Yeh Paragraph [0056] The input to the actor network 303 is a current state (which is a representation produced by the representation network 301), and the actor network 303 outputs a value representing an action chosen from an action space (e.g., a set of potential actions to take). The input to the critic network 305 is the current state (which is the representation produced by the representation network 301), and the critic network 305 outputs feedback or an approximation for the predicted action. The feedback may be a grade, rank, degree, rating, quality, or some other value indicating the adequacy or superiority of the predicted action in relation to other potential actions. (The actions in a set of potential actions receive feedback such as rank, grade, degree etc. to show how adequate that particular item is compared to others. This is comparable to a Pareto front, in that optimal solutions are placed in a set.)) after the ranking of the all points, selecting a portion of top-ranked ones of the all points, the selecting comprising selecting the Pareto front of the entirety of the union set. (Yeh Paragraph [0056] The input to the actor network 303 is a current state (which is a representation produced by the representation network 301), and the actor network 303 outputs a value representing an action chosen from an action space (e.g., a set of potential actions to take). The input to the critic network 305 is the current state (which is the representation produced by the representation network 301), and the critic network 305 outputs feedback or an approximation for the predicted action. The feedback may be a grade, rank, degree, rating, quality, or some other value indicating the adequacy or superiority of the predicted action in relation to other potential actions. The feedback may be a state-value from a state-value function (e.g., V(s)) or a quality value (a “Q value” or “action-value”) from a Q value function (e.g., Q(s, a)). [0104] By interacting with the environment and observing the outcome from past actions, the RL agent 501 adapts its policy π so that the expected long-term reward is maximized. (Based on the expected return (highest ranked actions in the set), the system selects these, which comprise a Pareto front of most optimal solutions)) Regarding Claim 8, the combination of Abdolmaleki 326, Abdolmaleki 084, Budan, and Yeh teaches the invention as claimed in Claim 4, including: wherein the obtaining the state transition data is conducted by executing actions according to a multi-objective e-greedy policy in which a uniformly random action is selected with some probability 0<E<1 and a greedy action is selected with probability 1-e where a greedy action is an action whose expected future total reward belongs to the Pareto front of the union set of the value function output over all actions in a given state. (Yeh Paragraph [0097] The real-time agent 501 determines an action a by applying the state s to a policy π (e.g., a=π(s) where “π(s)” is an action taken in state s under deterministic policy π). The state s is provided to the GE 502, which outputs an exploration action (“a_exploration”) based on the state s and additional information, and feeds the a_exploration to the EvE 504. The GE 502 may employ one or more suitable exploration algorithms such as, for example, random exploration (e.g., epsilon-greedy, softmax function, etc.), curiosity based exploration, upper confidence bounds (UCB), Boltzman exploration, Thompson sampling, entropy loss term, noise-based exploration, intrinsic reward (bonus) exploration (e.g., count-based exploration, prediction-based exploration, etc.), optimism in face of uncertainty exploration, Hoeffding's inequality, state-action exploration, parameter exploration, memory-based exploration (e.g., episodic memory exploration, direct exploration, etc.), Q-value exploration, and/or the like. (Using a multi-objective epsilon greedy function to find state transition data (applying state s to policy π))) Regarding Claims 13-17, Claims 13-17 are non-statutory computer readable storage medium claims corresponding to the method claims of Claims 4-8. As such, they are rejected for the same reasons. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Hu et al (US 20190266489 A1) discusses learning value functions with state-action inputs in order to find Pareto optima concerning the results of the value function. Kompella et al. (US 20230368041 A1) discusses the evaluation of temporal difference errors concerning value functions with state-action input pairs. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN A CARDOSO whose telephone number is (571)272-8512. The examiner can normally be reached M-F 7:30 - 5:00, alternate Friday's off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JUSTIN CARDOSO/ Patent Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jun 05, 2023
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month