DETAILED ACTION
This action is in response to the filing on 02/02/2026. Claims 1-3, 5-16, and 18-21, are pending and have been considered below.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 5-11, 13-16, and 18-21 are rejected under 35 U.S.C. 103 as being unpatentable over Ostrovski et al. (US 2020/0364557 A1, first cited in office action mailed 02/26/2025), hereinafter Ostrovski, in view of Bodnar et al. (US 11,571,809 B1, first cited in office action mailed 08/15/2025), hereinafter Bodnar.
Regarding claim 1, Ostrovski teaches A method of determining an action of a device for a given situation and controlling the device based on the action, implemented by a computer system, the method comprising (Ostrovski discloses a method of selecting an action for an agent interacting with an environment [see Ostrovski, para. 5] and performing the action [see Ostrovski, para. 8], implemented by a computer system [see Ostrovski, para. 12]):
for a learning model that has learned a distribution of rewards according to the action of the device for the given situation using a first risk-measure parameter associated with control of the device (Ostrovski discloses a reinforcement learning system that learns a probability distribution over possible returns if an agent performs a particular action in response to an observation [see Ostrovski, para. 15], using an observation-action-probability tuple as input to the quantile network [see Ostrovski, para. 49], that implements a risk-sensitive action selection policy using a risk measure function [see Ostrovski, para. 58], that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Thus, the system can be used when it has finished learning);
determining the action of the device for the given situation based on the first risk-measure parameter (Ostrovski discloses a method for selecting an action to be performed by a reinforcement learning agent interacting with an environment [see Ostrovski, para. 5] using an observation-action-probability tuple as input to the quantile network [see Ostrovski, para. 49], that implements a risk-sensitive action selection policy using a risk measure function [see Ostrovski, para. 58], that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]);
controlling the device in the environment based on the action, (Ostrovski discloses a method of selecting an action for an agent interacting with an environment [see Ostrovski, para. 5] and performing the action [see Ostrovski, para. 8]);
wherein the learning model has learned, using a quantile regression method, the distribution of rewards obtainable according to the action of the device for the given situation by learning values of the rewards (Ostrovski discloses that quantile regression is used for the quantile network [see Ostrovski, para. 11] that learns the distribution of rewards obtainable according to an action for a given situation [see Ostrovski, para. 5]) corresponding to first parameter values belonging to a first range and sampling a second risk-measure parameter that belongs to a second range corresponding to the first range (Ostrovski discloses using a first set of probability values and a second set of probability values [see Ostrovski, para. 79] such that the first and second probability values are randomly sampled within the range [0,1] [see Ostrovski, para. 78], then distorted by the risk measure function [see Ostrovski, para. 59], and equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]) and learning a value of a reward corresponding to the second risk-measure parameter within the distribution of rewards (Ostrovski discloses that the quantile network learns the distribution of rewards obtainable according to an action for a given situation [see Ostrovski, para. 5]. Thus, when using the observation-action-probability tuple of the second probability value tuple [see Ostrovski, para. 78], wherein the probability value is a transformed risk-measure probability [see Ostrovski, para. 70] it would learn the corresponding value of a reward).
However, Ostrovski fails to teach selectively setting the first risk-measure parameter based on a value input through a user terminal of a user device or a user interface of the device, in consideration of an environment in which the device is controlled; and such that the first risk-measure parameter for the learning model is able to be re-set to a different value via the user interface based on a characteristic of the environment, without requiring the learning model to be retrained.
In the same field of endeavor, Bodnar teaches:
selectively setting the first risk-measure parameter based on a value input through a user terminal of a user device or a user interface of the device, in consideration of an environment in which the device is controlled (Bodnar discloses a plurality of "out-of-band" signals, including user preferences and environmental attributes, that can be used to determine the value of the desired risk-measure [see Bodnar, Col. 10, line 61-Col. 11, line 5], that the user preference may be explicitly input/selected by a user [see Bodnar, Col. 13, lines 16-20], and user interface input devices that allow user interaction with the computing device [see Bodnar, Col. 17, lines 39-42 and FIG. 7]);
such that the first risk-measure parameter for the learning model is able to be re-set to a different value via the user interface based on a characteristic of the environment, without requiring the learning model to be retrained (Bodnar discloses a plurality of "out-of-band" signals, including user preferences and environmental attributes, that can be used to determine the value of the desired risk-measure [see Bodnar, Col. 10, line 61-Col. 11, line 5], that the user preference may be explicitly input/selected by a user [see Bodnar, Col. 13, lines 16-20], and user interface input devices that allow user interaction with the computing device [see Bodnar, Col. 17, lines 39-42 and FIG. 7]. Thus, if the user inputs different preferences or environmental attributes change, the value of the risk-measure can be changed accordingly within the spectrum. Further, because the risk-measure value is on constrained to a spectrum, and the model was trained on the spectrum, any value derived from the out-of-band signals within the spectrum can be used without retraining the model).
It would have been obvious to one of ordinary skill, in the art at the time before the effective filing date of the invention to incorporate selectively setting the first risk-measure parameter based on a value input through a user terminal of a user device or a user interface of the device, in consideration of an environment in which the device is controlled; and such that the first risk-measure parameter for the learning model is able to be re-set to a different value via the user interface based on a characteristic of the environment, without requiring the learning model to be retrained as suggested in Bodnar into Ostrovski because both methods perform machine learning to control a device in a given environment (see Ostrovski, Abstract; see Bodnar, Abstract). Incorporating the teaching of Bodnar into Ostrovski would allow out-of-band signals aside from input to the network to determine how conservatively or risk-seeking the device should act [see Bodnar, Col. 2, lines 34-48].
Regarding claim 2, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the determining of the action of the device comprises determining the action of the device to be more risk-averse or risk-seeking for the given situation based on the first risk-measure parameter or a range indicated by the first risk-measure parameter, the range being less than or equal to the first risk-measure parameter or less than the first risk-measure parameter (Ostrovski discloses selecting an action of the device according to a risk-sensitive action policy that can be made more risk-averse or more risk-seeking depending on the risk-measure [see Ostrovski, para. 18]. Bodnar discloses the value of the risk-measure on a spectrum ranging from risk-averse to risk-seeking behavior [see Bodnar, Col. 10, lines 61–63] and gives an example of a spectrum from 0.3 to 0.8 with values below 0.55 being risk averse and values of 0.55 and above being risk-seeking being distorted by a risk-measure such that the value of 0.55 either decreases or increases based on how the risk-measure is weighted towards being risk-averse or risk-seeking [see Bodnar, Col. 11, lines 12-29]. Thus, depending on the set value of the risk-measure, or a range indicated by the set value of the risk-measure (i.e. less than or equal to the value, or greater than or equal to the value), the device may select an action that is more risk-averse or risk-seeking).
Regarding claim 3, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 2 and further teaches:
wherein the device is an autonomous driving robot, and the determining of the action of the device further comprises determining run-forward or acceleration of the autonomous driving robot as a more risk-seeking action of the robot if the first risk-measure parameter is greater than or equal to a desired value or if the first risk-measure parameter is greater than or within a desired range (Ostrovski discloses the actions including navigation controls such as steering, movement, braking and/or acceleration [see Ostrovski, para. 38], and that the system can implement a risk-averse or risk-seeking action policy [see Ostrovski, para. 58] based on the risk-measure [see Ostrovski, para. 18]. Bodnar discloses the value of the risk-measure on a spectrum ranging from risk-averse to risk-seeking behavior [see Bodnar, Col. 10, lines 61–63] and gives an example of a spectrum from 0.3 to 0.8 with values below 0.55 being risk averse and values of 0.55 and above being risk-seeking being distorted by a risk-measure such that the value of 0.55 either decreases or increases based on how the risk-measure is weighted towards being risk-averse or risk-seeking [see Bodnar, Col. 11, lines 12-29]. Thus, the disclosed actions may be determined as more risk-averse or more risk-seeking based on the set value of the risk-measure being greater than or equal to a desired value, or if the set value is greater than or within a desired range).
Regarding claim 5, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein a minimum value among the first parameter values corresponds to a minimum value among values of the rewards and a maximum value among the first parameter values corresponds to a maximum value among the values of the rewards (In an implementation each of the probability values is transformed by a distortion risk measure function prior to being processed by the quantile function network. In an implementation the distortion risk measure function is a non-decreasing function mapping a domain [0,1] to a range [0,1]; the distortion risk measure function maps the point 0 in the domain to the point 0 in the range; and the distortion risk measure function maps the point 1 in the domain to the point 1 in the range. [see Ostrovski, para. 7]).
Regarding claim 6, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the first range is 0-1 and the second range is 0-1, and (In an implementation each of the probability values is transformed by a distortion risk measure function prior to being processed by the quantile function network. In an implementation the distortion risk measure function is a non-decreasing function mapping a domain [0,1] to a range [0,1]; the distortion risk measure function maps the point 0 in the domain to the point 0 in the range; and the distortion risk measure function maps the point 1 in the domain to the point 1 in the range. [see Ostrovski, para. 7]);
wherein the second risk-measure parameter belonging to the second range is randomly sampled at a time of learning of the learning model (Ostrovski discloses sampling second probability values at a time of learning [see Ostrovski, para. 10], and using an observation-action-probability tuple as input to the quantile network [see Ostrovski, para. 49], that implements a risk-sensitive action selection policy using a risk measure function [see Ostrovski, para. 58], that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]).
Regarding claim 7, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 6 and further teaches:
wherein each of the first parameter values represents a percentile, and wherein each of the first parameter values corresponds to a value of a corresponding reward at a corresponding percentile. (The quantile value for a probability value with respect to a return distribution refers to a threshold return value below which random draws from the return distribution would fall with probability given by the probability value. Put another way, the quantile value for a probability value with respect to a return distribution can be obtained by evaluating the inverse of the cumulative distribution function (CDF) for the return distribution at the probability value. That is, integrating a probability density function for a return distribution up to the quantile value for a probability value would yield the probability value itself. [see Ostrovski, para. 50]).
Regarding claim 8, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the learning model comprises: (The reinforcement learning system described in this specification includes a quantile function neural network that implicitly models the quantile function of the probability distribution over possible returns that would be received if an agent performs a particular action in response to an observation. [see Ostrovski, para. 15; FIG. 1]);
a first model configured to predict the action of the device for the given situation, and a second model configured to predict a reward according to the action (By training the quantile function network 112, the system 100 may cause the quantile function network 112 to generate outputs that result in the selection of actions 102 to be performed by the agent 104 which increase a cumulative measure of reward received by the system 100. [see Ostrovski, para. 60]. Further, Bodnar discloses that reinforcement learning uses an actor network and a critic network, with the actor network being trained to predict the probabilities of actions based on a current state, and the critic network being trained to predict the return for the state-action pairs [see Bodnar, Col. 1, lines 6-32]);
wherein each of the first model and the second model is trained using the first risk-measure parameter (Ostrovski discloses training the quantile network [see Ostrovski, para. 60] using an observation-action-probability tuple as input [see Ostrovski, para. 49], and implementing a risk-sensitive policy using a risk measure function [see Ostrovski, para. 58] that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Further, Bodnar discloses that reinforcement learning uses an actor network and a critic network, with the actor network being trained to predict the probabilities of actions based on a current state, and the critic network being trained to predict the return for the state-action pairs [see Bodnar, Col. 1, lines 6-32]. Thus, the combination of Ostrovski and Bodnar could use a combination of actor and critic networks in reinforcement learning of the quantile network);
wherein the first model is trained to predict an action that maximizes the reward predicted from the second model as a next action of the device (Therefore, the system 100 selects the actions 102 to be performed by the agent based on respective measures of central tendency (e.g., means) of the return distribution corresponding to each possible action. [see Ostrovski, para. 58]. Bodnar discloses that reinforcement learning uses an actor network and a critic network, with the actor network being trained to predict the probabilities of actions based on a current state, and the critic network being trained to predict the return for the state-action pairs [see Bodnar, Col. 1, lines 6-32]. Thus, the combination of Ostrovski and Bodnar would train the actor network to predict the best action to take based on maximizing return distribution of each possible action as predicted by the critic network);
wherein the first model is different than the second model (Bodnar discloses that reinforcement learning uses an actor network and a critic network, with the actor network being trained to predict the probabilities of actions based on a current state, and the critic network being trained to predict the return for the state-action pairs [see Bodnar, Col. 1, lines 6-32]).
Regarding claim 9, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 8 and further teaches:
wherein the device is an autonomous driving robot (A system of this aspect may be comprised in a control system for an autonomous or semi-autonomous vehicle (which may be a land, sea or air vehicle), in a control system for a robotic agent, in a control system for a mechanical agent, or in a control system for an electronic agent. [see Ostrovski, para. 12]);
wherein the first model and the second model are configured to predict the action of the device and the reward, respectively (Ostrovski discloses the reinforcement learning agent predicting the action of the device and a quantile network to predict the reward [see Ostrovski, para. 5 and 58]), based on a position of an obstacle around the autonomous driving robot (the observations may include, for example, one or more of images, object position data, and sensor data to capture observations as the agent as it interacts with the environment, for example sensor data from an image, distance, or position sensor or from an actuator. [see Ostrovski, para. 32]; The observations may also include, for example, sensed electronic signals such as motor current or a temperature signal; and/or image or video data for example from a camera or a LID AR sensor, e.g., data from sensors of the agent or data from sensors that are located separately from the agent in the environment. [see Ostrovski, para. 35]), a path through which the autonomous driving robot is to move (the agent may be a robot interacting with the environment to accomplish a specific task, e.g., to locate an object of interest in the environment or to move an object of interest to a specified location in the environment or to navigate to a specified destination in the environment; or the agent may be an autonomous or semi-autonomous land or air or sea vehicle navigating through the environment. [see Ostrovski, para. 31]), and a velocity of the autonomous driving robot (the observations may include data characterizing the current state of the robot, e.g., one or more of: joint position, joint velocity, joint force, torque or acceleration, for example gravity-compensated torque feedback, and global or relative pose of an item held by the robot. [see Ostrovski, para. 33]; the observations may similarly include one or more of the position, linear or angular velocity, force, torque or acceleration, and global or relative pose of one or more parts of the agent. The observations may be defined in 1, 2 or 3 dimensions, and may be absolute and/or relative observations. [see Ostrovski, para. 34]).
Regarding claim 10, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the learning model learns the distribution of rewards by iteratively estimating a reward according to the action of the device for the given situation (The reinforcement learning system described in this specification includes a quantile function neural network that implicitly models the quantile function of the probability distribution over possible returns that would be received if an agent performs a particular action in response to an observation. [see Ostrovski, para. 15]; The system 100 is configured to train the quantile function network 112 over multiple training iterations using reinforcement learning techniques. The system 100 trains the quantile function network 112 by iteratively (i.e., at each training iteration) adjusting the current values of the quantile function network parameters. By training the quantile function network 112, the system 100 may cause the quantile function network 112 to generate outputs that result in the selection of actions 102 to be performed by the agent 104 which increase a cumulative measure of reward received by the system 100. [see Ostrovski, para. 60]),
wherein each iteration comprises learning an episode that represents a movement of the device from a start position to a goal position and updating the learning model, the device being a robot configured to drive from the start position to the goal position (A system of this aspect may be comprised in a control system for an autonomous or semi-autonomous vehicle (which may be a land, sea or air vehicle), in a control system for a robotic agent, in a control system for a mechanical agent, or in a control system for an electronic agent. [see Ostrovski, para. 12] The system 100 is configured to train the quantile function network 112 over multiple training iterations using reinforcement learning techniques. The system 100 trains the quantile function network 112 by iteratively (i.e., at each training iteration) adjusting the current values of the quantile function network parameters. By training the quantile function network 112, the system 100 may cause the quantile function network 112 to generate outputs that result in the selection of actions 102 to be performed by the agent 104 which increase a cumulative measure of reward received by the system 100. [see Ostrovski, para. 60]; At each time step, the system 100 may receive a reward 110 based on the current state of the environment 106 and the action 102 of the agent 104 at the time step. In general, the reward 110 is a numerical value. The reward 110 can be based on any event or aspect of the environment 106. For example, the reward 110 may indicate whether the agent 104 has accomplished a task (e.g., navigating to a target location in the environment 106) or the progress of the agent 104 towards accomplishing a task. [see Ostrovski, para. 30]);
wherein, when the episode starts, the first risk-measure parameter is sampled to obtain a sampled risk-measure parameter and the sampled risk-measure parameter is fixed until the episode ends (Ostrovski discloses using an observation-action-probability tuple as input [see Ostrovski, para. 49] and sampling probability values for a current and next observation [see Ostrovski, para. 10]. In sampling the probability value, implementing a risk-sensitive policy using a risk measure function [see Ostrovski, para. 58] that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Thus, the risk-measure probability is sampled for a current observation tuple, and fixed for that tuple until the episode ends when the next observation tuple is used with it's corresponding risk-measure probability).
Regarding claim 11, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 10 and further teaches:
wherein updating of the learning model is performed using the sampled risk-measure parameter, or performed by resampling the first risk-measure parameter to obtain a resampled risk-measure parameter and using the resampled risk-measure parameter, the sampled risk-measure parameter being stored in a buffer (Ostrovski discloses that the network can be updated using as many samples as desired per training iteration [see Ostrovski, para. 17]. Ostrovski further discloses using an observation-action-probability tuple as input [see Ostrovski, para. 49] and sampling probability values for a current and next observation [see Ostrovski, para. 10]. In sampling the probability value, implementing a risk-sensitive policy using a risk measure function [see Ostrovski, para. 58] that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Thus, the risk-measure probability is sampled for a current observation tuple, and fixed for that tuple until the episode ends when the next episode begins and a new observation tuple is used with a new sampled risk-measure probability. Bodnar discloses a replay buffer 110 which can be used to store data associated with episodes from the offline episode database [see Bodnar, Col. 7, lines 48–52]. Thus, it would have been obvious to store the risk-measure parameter for an episode in the replay buffer 110 along with the other associated episode data).
Regarding claim 13, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the device is an autonomous driving robot, and (A system of this aspect may be comprised in a control system for an autonomous or semi-autonomous vehicle (which may be a land, sea or air vehicle), in a control system for a robotic agent, in a control system for a mechanical agent, or in a control system for an electronic agent. [see Ostrovski, para. 12]);
wherein the value input through the user terminal is set as a value of the first risk-measure parameter while the autonomous driving robot is autonomously driving in the environment or before the autonomous driving robot is autonomously driving in the environment (Ostrovski discloses that the device can be an autonomous driving robot [see Ostrovski, para. 12], and sampling probability values during a current observation [see Ostrovski, para. 10]. In sampling the probability value, the network distorts the probability value by the risk-measure [see Ostrovski, para. 7]. Bodnar discloses using "out-of-band" signals such as user preferences to determine the risk-measure [see Bodnar, Col. 2, lines 34–48]. Thus, the second risk-measure parameter can be set while the robot is driving by sampling the transformed risk-measure probability at the time of observation based on user preferences. Additionally, the user could explicitly input their preferences which can be used exclusively for the risk-measure [see Bodnar, Col. 13, lines 16-20]. Thus, the user preference set for the risk-measure before the robot is driving is used as the risk-measure)
Regarding claim 14, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
A non-transitory computer-readable record medium storing computer-executable instructions that, when executed by a processor, cause the processor to (Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus [see Ostrovski, para. 84]; The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. [see Ostrovski, para. 85]).
Regarding claim 15, claim 15 contains substantially similar limitations to those found in claim 1. Therefore it is rejected for the same reason as claim 1 above. Additionally, the combination of Ostrovski and Bodnar further teaches:
A computer system comprising: memory storing computer-executable instructions; and (According to another aspect there is provided a system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the respective operations of a method according to any aspect or implementation described herein. [see Ostrovski, para. 12]) at least one processor configured to execute the computer-executable instructions such that the at least one processor is configured to (Computers suitable for the execution of a computer program can be based on general or special purpose micro processors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. [see Ostrovski, para. 89]).
Regarding claim 16, Ostrovski teaches a method of training a model used to determine an action of a device for a situation, the method comprising (FIG. 5 is a flow diagram of an example process for selecting an action to be performed by an agent using a quantile function network. [see Ostrovski, para. 24; FIG. 5]; FIG. 6 is a flow diagram of an example process for training a quantile function network. [see Ostrovski, para. 25; FIG. 6]):
controlling the device in the environment based on the value of the risk-measure parameter (Ostrovski discloses a method of selecting an action for an agent interacting with an environment [see Ostrovski, para. 5], using an observation-action-probability tuple as input to the quantile network [see Ostrovski, para. 49], that implements a risk-sensitive action selection policy using a risk measure function [see Ostrovski, para. 58], that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Ostrovski further discloses the agent performing the selected action [see Ostrovski, para. 8]);
training, by a processor (Ostrovski discloses training the quantile function network [see Ostrovski, para. 25; FIG. 6], which can be done by a central processing unit executing a computer program [see Ostrovski, para. 89]), the model to learn a distribution of rewards according to the action of the device for the situation using the first risk-measure parameter (Ostrovski discloses a reinforcement learning system that learns a probability distribution over possible returns if an agent performs a particular action in response to an observation [see Ostrovski, para. 15], using an observation-action-probability tuple as input to the quantile network [see Ostrovski, para. 49], that implements a risk-sensitive action selection policy using a risk measure function [see Ostrovski, para. 58], that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Thus, the system can be used when it has finished learning);
wherein the model learns, using a quantile regression method, the distribution of rewards obtainable according to the action of the device by learning values of the rewards (Ostrovski discloses that quantile regression is used for the quantile network [see Ostrovski, para. 11] that learns the distribution of rewards obtainable according to an action for a given situation [see Ostrovski, para. 5]) corresponding to first parameter values belonging to a first range and sampling a second risk-measure parameter that belongs to a second range corresponding to the first range (Ostrovski discloses using a first set of probability values and a second set of probability values [see Ostrovski, para. 79] such that the first and second probability values are randomly sampled within the range [0,1] [see Ostrovski, para. 78], then distorted by the risk measure function [see Ostrovski, para. 59], and equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]) and learning a value of a reward corresponding to the second risk-measure parameter within the distribution of rewards (Ostrovski discloses that the quantile network learns the distribution of rewards obtainable according to an action for a given situation [see Ostrovski, para. 5]. Thus, when using the observation-action-probability tuple of the second probability value tuple [see Ostrovski, para. 78], wherein the probability value is a transformed risk-measure probability [see Ostrovski, para. 70] it would learn the corresponding value of a reward)
However, Ostrovski fails to teach setting a first risk-measure parameter as a value requested by a user, the value requested by the user being input through a user terminal of a user device or a user interface of the device in consideration of an environment in which the device is controlled; and such that the first risk-measure parameter for the model is able to be re-set to a different value via the user interface based on a characteristic of the environment, without requiring the model to be trained.
In the same field of endeavor, Bodnar teaches:
setting a first risk-measure parameter as a value requested by a user, the value requested by the user being input through a user terminal of a user device or a user interface of the device in consideration of an environment in which the device is controlled (Bodnar discloses a plurality of "out-of-band" signals, including user preferences and environmental attributes, that can be used to determine the value of the desired risk-measure [see Bodnar, Col. 10, line 61-Col. 11, line 5], that the user preference may be explicitly input/selected by a user [see Bodnar, Col. 13, lines 16-20], and user interface input devices that allow user interaction with the computing device [see Bodnar, Col. 17, lines 39-42 and FIG. 7]. Thus, it would have been obvious that the user could set their preference based on the environmental attributes, such that the only signal needed is the user preference);
such that the first risk-measure parameter for the model is able to be re-set to a different value via the user interface based on a characteristic of the environment, without requiring the model to be trained (Bodnar discloses a plurality of "out-of-band" signals, including user preferences and environmental attributes, that can be used to determine the value of the desired risk-measure [see Bodnar, Col. 10, line 61-Col. 11, line 5], that the user preference may be explicitly input/selected by a user [see Bodnar, Col. 13, lines 16-20], and user interface input devices that allow user interaction with the computing device [see Bodnar, Col. 17, lines 39-42 and FIG. 7]. Thus, if the user inputs different preferences or environmental attributes change, the value of the risk-measure can be changed accordingly within the spectrum. Further, because the risk-measure value is on constrained to a spectrum, and the model was trained on the spectrum, any value derived from the out-of-band signals within the spectrum can be used without retraining the model).
It would have been obvious to one of ordinary skill, in the art at the time before the effective filing date of the invention to incorporate setting a first risk-measure parameter as a value requested by a user, the value requested by the user being input through a user terminal of a user device or a user interface of the device in consideration of an environment in which the device is controlled; and such that the first risk-measure parameter for the model is able to be re-set to a different value via the user interface based on a characteristic of the environment, without requiring the model to be trained as suggested in Bodnar into Ostrovski because both methods perform machine learning to control a device in a given environment (see Ostrovski, Abstract; see Bodnar, Abstract). Incorporating the teaching of Bodnar into Ostrovski would allow out-of-band signals aside from input to the network to determine how conservatively or risk-seeking the device should act [see Bodnar, Col. 2, lines 34-48].
Regarding claim 18, the combination of Ostrovski and Bodnar as applied in claim 16 above teaches all the limitations of claim 16 and further teaches:
wherein a minimum value among the first parameter values corresponds to a minimum value among values of the rewards and a maximum value among the first parameter values corresponds to a maximum value among the values of the rewards (In an implementation each of the probability values is transformed by a distortion risk measure function prior to being processed by the quantile function network. In an implementation the distortion risk measure function is a non-decreasing function mapping a domain [0,1] to a range [0,1]; the distortion risk measure function maps the point 0 in the domain to the point 0 in the range; and the distortion risk measure function maps the point 1 in the domain to the point 1 in the range. [see Ostrovski, para. 7]).
Regarding claim 19, the combination of Ostrovski and Bodnar as applied in claim 16 above teaches all the limitations of claim 16 and further teaches:
wherein the learning model further comprises: (The reinforcement learning system described in this specification includes a quantile function neural network that implicitly models the quantile function of the probability distribution over possible returns that would be received if an agent performs a particular action in response to an observation. [see Ostrovski, para. 15; FIG. 1]);
a first model configured to predict the action of the device for the situation (By training the quantile function network 112, the system 100 may cause the quantile function network 112 to generate outputs that result in the selection of actions 102 to be performed by the agent 104 which increase a cumulative measure of reward received by the system 100. [see Ostrovski, para. 60]);
a second model configured to predict a reward according to the action (The sampling engine 118 may generate probability values 116 which are sampled from a uniform probability distribution over the interval [0,1]. In this case, the quantile values generated by the quantile function network 112 by processing action-observation-probability value tuples can be understood to represent return values that are randomly sampled from the return distribution that would result from the agent performing the action in response to the observation. Therefore, the system 100 selects the actions 102 to be performed by the agent based on respective measures of central tendency (e.g., means) of the return distribution corresponding to each possible action. [see Ostrovski, para. 58]);
wherein each of the first model and the second model is trained using the first risk-measure parameter (Ostrovski discloses training the quantile network and the agent [see Ostrovski, para. 60] using an observation-action-probability tuple as input [see Ostrovski, para. 49], and implementing a risk-sensitive policy using a risk measure function [see Ostrovski, para. 58] that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]);
wherein the training comprises training the first model to predict an action that maximizes the reward predicted from the second model as a next action of the device (The system 100 selects an action 102 to be performed by the agent 104 at the time step based on the measures of central tendency 128 corresponding to the actions. In some implementations, the system 100 selects an action having a highest corresponding measure of central tendency 128 from amongst all the actions in the set of actions that can be performed by the agent 104. [see Ostrovski, para. 55]; The sampling engine 118 may generate probability values 116 which are sampled from a uniform probability distribution over the interval [0,1]. In this case, the quantile values generated by the quantile function network 112 by processing action-observation-probability value tuples can be understood to represent return values that are randomly sampled from the return distribution that would result from the agent performing the action in response to the observation. Therefore, the system 100 selects the actions 102 to be performed by the agent based on respective measures of central tendency (e.g., means) of the return distribution corresponding to each possible action. [see Ostrovski, para. 58]).
Regarding claim 20, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 2 and further teaches:
wherein the device is an autonomous driving robot (A system of this aspect may be comprised in a control system for an autonomous or semi-autonomous vehicle (which may be a land, sea or air vehicle), in a control system for a robotic agent, in a control system for a mechanical agent, or in a control system for an electronic agent. [see Ostrovski, para. 12]), and the determining of the action of the device further comprises (a method of selecting an action to be performed by a reinforcement learning agent interacting with an environment [see Ostrovski, para. 5]):
selecting, as the action, an action that causes the device to operate in a more risk-seeking manner as the first risk-measure parameter becomes a more risk-seeking value (Ostrovski discloses selecting an action of the device according to a risk-sensitive action policy that can be made more risk-averse or more risk-seeking depending on the risk-measure function [see Ostrovski, para. 18]. Bodnar discloses value of the risk-measure parameter on a spectrum ranging from risk-averse to risk-seeking behavior [see Bodnar, Col. 10, lines 61–63]. Thus, as the set value of the risk-measure becomes a more risk-seeking value a more risk-seeking action is chosen);
selecting, as the action, an action that causes the device to operate in a more risk-averse manner as the first risk-measure parameter becomes a more risk-averse value. (Ostrovski discloses selecting an action of the device according to a risk-sensitive action policy that can be made more risk-averse or more risk-seeking depending on the risk-measure function [see Ostrovski, para. 18]. Bodnar discloses value of the risk-measure parameter on a spectrum ranging from risk-averse to risk-seeking behavior [see Bodnar, Col. 10, lines 61–63]. Thus, as the set value of the risk-measure becomes a more risk-averse value a more risk-averse action is chosen).
Regarding claim 21, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
receiving a second value input through the user terminal, the value input through the user terminal being a first value and the second value being different from the first value; re-setting the first risk-measure parameter based on the second value (Bodnar discloses a plurality of "out-of-band" signals, including user preferences and environmental attributes, that can be used to determine the value of the desired risk-measure [see Bodnar, Col. 10, line 61-Col. 11, line 5], that the user preference may be explicitly input/selected by a user [see Bodnar, Col. 13, lines 16-20], and user interface input devices that allow user interaction with the computing device [see Bodnar, Col. 17, lines 39-42 and FIG. 7]. Bodnar further explains that the critic network takes in current state data, and determining state-action returns based on state data for each iteration [see Bodnar, Col. 1, lines 25-47 and FIG. 3], thus, if the user inputs different preferences or environmental attributes change, the value of the risk-measure can be changed accordingly by the risk distortion module. It would have been obvious to one of ordinary skill in the art before the effective filing date to receive a second input at the user terminal, different than the first input at the user terminal, that can be used by the risk distortion module of the critic network to set a new value of the risk-measure parameter).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Ostrovski et al. (US 2020/0364557 A1, first cited in office action mailed 02/26/2025), hereinafter Ostrovski, in view of Bodnar et al. (US 11,571,809 B1, first cited in office action mailed 08/15/2025), hereinafter Bodnar, as applied in claim 1 above, and further in view of Samuelson (Safety-Aware Optimal Control of Stochastic Systems Using Conditional Value-at-Risk, first cited in office action mailed 02/26/2025), hereinafter Samuelson.
Regarding claim 12, the combination of Ostrovski and Bodnar as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the first risk-measure parameter is a parameter representing a conditional value-at-risk (CVaR) risk measure (Ostrovski discloses using an observation-action-probability tuple as input [see Ostrovski, para. 49], and implementing a risk-sensitive policy using a risk measure function [see Ostrovski, para. 58] that distorts the probability distribution sample resulting in an equivalent to sampling a transformed risk-measure probability distribution [see Ostrovski, para. 70]. Ostrovski further discloses using CVaR as a risk-measure function [see Ostrovski, para. 65; FIG. 4 elements 414 and 416]. Thus, the transformed risk-measure probability sample can represent a CVaR risk-measure).
However, the combination of Ostrovski and Bodnar fails to teach a conditional value-at-risk (CVaR) risk measure that is a number within a range greater than 0 and less than or equal to 1, or a power-law risk measure that is a number within the range less than zero.
In the same field of endeavor, Samuelson teaches:
a conditional value-at-risk (CVaR) risk measure that is a number within a range greater than 0 and less than or equal to 1 (First, we introduce a novel measure of safety risk by using CVaR and the distance between the system state and a desired set A for safety. This safety risk measure represents the conditional expectation of the distance between the state and A within the (1—α) worst-case quantile of an associated safety loss distribution, where α ∈ (0, 1). [see Samuelson, pg. 1, Sect. 1, para. 3]).
It would have been obvious to one of ordinary skill, in the art at the time before the effective filing date of the invention to incorporate a conditional value-at-risk (CVaR) risk measure that is a number within a range greater than 0 and less than or equal to 1 as suggested in Samuelson into the combination of Ostrovski and Bodnar because both methods use CVaR risk measures (see Ostrovski, FIG. 4; see Samuelson, pg. 1, Section 1. Introduction, para. 3). Incorporating the teaching of Samuelson into the combination of Ostrovski and Bodnar would give an important advantage of CVaR over the value-at-risk (VaR) or the chance constraints in that CVaR takes into account the possibility of tail events in which safety losses exceed VaR while VaR is incapable of distinguishing situations beyond VaR (see Samuelson, pg. 1, Section 1. Introduction, para. 3).
Response to Arguments
Applicant’s arguments, filed 02/02/2026, traversing the rejections of claims 6 and 7 under 35 U.S.C. 112(b) have been fully considered and are persuasive, the rejections of claims 6 and 7 under 35 U.S.C. 112(b) are respectfully withdrawn.
Applicant’s arguments, filed 02/02/2026, traversing the rejections of claims 1-20 under 35 U.S.C. 103 have been fully considered and are not persuasive. Applicant argues with respect to claim 1, that Ostrovski cannot describe a distribution of rewards that learns values of the rewards with both first parameters values in a first range and a sampled risk-measure parameter, and both a parameter value and a risk-measure parameter are required to obtain a value of a reward and multiple parameter values and a risk measure parameters are utilized to obtain a distribution of rewards; argues with respect to claim 8, that neither Ostrovski nor Bodnar describe the first model being different than the second model; argues with respect to claim 10, neither Ostrovski nor Bodnar describe iterations that include episodes representing a movement of a device from a start position to a goal position where a robot is configured to drive from the start position to the goal position; Examiner respectfully disagrees.
With respect to Applicant’s arguments addressing claim 1, Applicant argues Ostrovski describes the following claim limitation recited in claim 1: “wherein the learning model has learned, using a quantile regression method, the distribution of rewards obtainable according to the action of the device for the given situation by learning values of the rewards corresponding to first parameter values belonging to a first range and sampling a second risk-measure parameter that belongs to a second range corresponding to the first range and learning a value of a reward corresponding to the second risk-measure parameter within the distribution of rewards”. In other words the claims recites two parts: 1) the model using quantile regression learns a distribution of rewards obtainable according to the action of the device by learning values of the rewards corresponding to first parameter values belonging to a first range, and 2) the model learns a value of a reward corresponding to a sampled second risk-measure parameter within the distribution of rewards, the second risk-measure parameter belonging to a second range corresponding to the first range. As described above in the 35 U.S.C. 103 section, the quantile regression model of Ostrovksi learns a distribution of rewards according to the action of the device for a given situation [see Ostrovski, para. 5 and 11] by using probability values sampled with the range [0,1] and then distorted by a risk-measure function, such that the values are equivalent to sampling the risk-measure probability distribution [see Ostrovski, para. 59, 70, and 78-79]. Further, Ostrovski describes learning a value of a reward according to a sampled risk-measure parameter (the risk-measure parameter being equivalently sampled as described with the first parameter values) by using an observation-action-probability tuple [see Ostrovski, para. 5, 59, 70, and 78-79]. Ostrovski further describes this is done during learning from experience tuples with two sets of probability values, one for the current observation-action-probability tuple, and a second for the next observation-argmax action-probability tuple, to calculate a temporal difference used to update the quantile network [see Ostrovski, para. 78-80]. Thus, Ostrovski describes using first risk-parameter values to learn the distribution of rewards, and learning the value of a reward corresponding to a risk-measure parameter within the distribution of rewards, as part of training the quantile network.
Applicant further argues that the risk measure function includes only a probability value and a hyper-parameter, and that only the probability value is sampled thus the function cannot describe a distribution of rewards that learns values of the rewards with both first parameter values in a first range and a sampled risk-measure parameter. However, with respect to applicant’s argument that the model learns the distribution of rewards according to both first parameter values in a first range and a sampled risk-measure parameter and that a result of a function cannot describe a value that is sampled and used with a first parameter value to obtain a reward; while this interpretation was previously discussed in the interview on 01/14/2026, the claim language has not been amended to reflect this interpretation, thus, the interpretation is not required by the claim language. If this interpretation is desired, applicant is encouraged to amend the claim language to reflect that. Further, Ostrovski states that transforming the probability values via the distortion risk measure function is “equivalent to sampling the probability values from a transformed probability distribution defined by the composition of the original probability distribution and the distortion risk measure” [see Ostrovski, para. 70], thus, the claim limitation would be rendered obvious in view of Ostrovski.
Applicant argues that claims 15 and 16 recite similar elements as claim 1 and cannot be rendered obvious for similar reasons, and that claims 17, 19, and 20 cannot be rendered obvious by virtue of their dependence from claim 16. For similar reasons as claim 1 above, applicant’s arguments with respect to claims 15-17 and 19-20 are not persuasive.
With respect to applicant’s arguments addressing claim 8 that neither Ostrovski nor Bodnar describe using both a first and second model that are different, while it was previously discussed in the interview on 01/14/2026 that Ostrovski a quantile network that achieves the goal of both the first and second model and thus can act as both the first and second model, Bodnar as cited in the 35 U.S.C. 103 section above describes learning with both an actor network and a critic network [see Bodnar, Col. 1, lines 6-32], thus, the combination of Ostrovski and Bodnar would render obvious using a first and second model that are different from each other.
With respect to applicant’s arguments addressing claim 10 that neither Ostrovski nor Bodnar describe a robot configured to drive from the start position to the goal position, however, as cited above in the 35 U.S.C. 103 section, Ostrovski describes providing a reward at a time step based on the agent navigating to a target location [see Ostrovski, para. 30] and that the agent is an autonomous or vehicle [see Ostrovski para. 12].
For at least the aforementioned reasons, the rejections of claims 1-21 under 35 U.S.C. 103 are respectfully maintained.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Dabney, Will, et al. ("Implicit quantile networks for distributional reinforcement learning." International conference on machine learning. PMLR, 2018.) discloses a quantile network that learns using a risk-measure parameter to predict an action for an agent to take in the current state and to predict a reward for the agent to take that action.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAKE BREEN whose telephone number is (571)272-0456. The examiner can normally be reached Monday - Friday, 7:00 AM - 3:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.T.B./Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143