DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the application filed on 09/15/2023. Claims 1-14 are presented in the case. Claims 1, 13 and 14 are independent claims.
Priority
Applicant's claim for the benefit of a Singaporean Patent Application No. 10202102725Q, filed on March 17, 2021 is acknowledged.
Information Disclosure Statement
The information disclosure statement submitted on 09/15/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1-13 are directed to a method and claim 14 is directed to a system. Therefore, the claims are eligible under Step 1 for being directed to a process and a machine respectively.
Independent claims 1 and 14:
Step 2A Prong 1:
Claims recite:
training a deep reinforcement learning model for autonomous control of a machine - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation to perform the training of a deep reinforcement learning model.
minimizing a loss function of the policy network - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation to minimize a loss function.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
the model being configured to output, by a policy network, an agent action in response to input of state information and a value function, the agent action representing a control signal for the machine - the steps recited at a high level of generality, and amounts to mere data inputting and outputting, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
wherein the loss function of the policy network comprises an autonomous guidance component and a human guidance component - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
wherein the autonomous guidance component is zero when the state information is indicative of input of a human input signal at the machine - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
A system for training a deep reinforcement learning model for autonomous control of a machine, the system comprising: storage; and
at least one processor in communication with the storage; wherein the storage comprises machine-readable instructions for causing the at least one processor to execute a method according to claim 1 - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
the model being configured to output, by a policy network, an agent action in response to input of state information and a value function, the agent action representing a control signal for the machine - the steps recited at a high level of generality, and amounts to mere data inputting and outputting, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
wherein the loss function of the policy network comprises an autonomous guidance component and a human guidance component - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
wherein the autonomous guidance component is zero when the state information is indicative of input of a human input signal at the machine - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
A system for training a deep reinforcement learning model for autonomous control of a machine, the system comprising: storage; and
at least one processor in communication with the storage; wherein the storage comprises machine-readable instructions for causing the at least one processor to execute a method according to claim 1 - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 2:
Step 2A Prong 1: The claim recites the abstract ideas of claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the model has an actor-critic architecture comprising an actor part and a critic part, and wherein the actor part comprises the policy network - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein the model has an actor-critic architecture comprising an actor part and a critic part, and wherein the actor part comprises the policy network - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 3:
Step 2A Prong 1: The claim recites the abstract ideas of claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the critic part comprises at least one value network configured to output the value function - the steps recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein the critic part comprises at least one value network configured to output the value function - the steps recited at a high level of generality, and amounts to mere data outputting, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 4:
Step 2A Prong 1:
Claim recites:
wherein the at least one value network is configured to estimate the value function based on the Bellman equation - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation to estimate the value function based on the Bellman equation.
Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claims do not provide a practical application and is not considered to be significantly more. As such, the claims are ineligible.
Dependent claim 5:
Step 2A Prong 1:
Claim recites:
wherein the critic part comprises a first value network paired with a second value network, each value network having the same architecture, for reducing or preventing overestimation - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation for reducing or preventing overestimation.
Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claims do not provide a practical application and is not considered to be significantly more. As such, the claims are ineligible.
Dependent claim 6:
Step 2A Prong 1: The claim recites the abstract ideas of claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein each value network is coupled to a target value network - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein each value network is coupled to a target value network - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 7:
Step 2A Prong 1: The claim recites the abstract ideas of claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the policy network is coupled to a target policy network - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein the policy network is coupled to a target policy network - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 8:
Step 2A Prong 1: The claim recites the abstract ideas of claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the deep reinforcement learning model comprises a priority experience replay buffer for storing, for a series of time points: the state information; the agent action; a reward value; and an indicator as to whether a human input signal is received - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein the deep reinforcement learning model comprises a priority experience replay buffer for storing, for a series of time points: the state information; the agent action; a reward value; and an indicator as to whether a human input signal is received - These limitations amount to components of a general purpose computer that applies a judicial exception, by use of conventional computer functions (see MPEP § 2106.05(b)).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 9:
Step 2A Prong 1: The claim recites the abstract ideas of claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the machine is an autonomous vehicle - the step recited at a high level of generality, and amounts to merely indicating a field of use or technological environment in which the judicial exception is performed (see MPEP § 2106.05(h)).
Dependent claim 10:
Step 2A Prong 1:
Claims recite:
wherein the loss function includes an adaptively assigned weighting factor applied to the human guidance component - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses a mathematical concept of a mathematical calculation to apply an adaptively assigned weighting factor the human guidance component.
Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claims do not provide a practical application and is not considered to be significantly more. As such, the claims are ineligible.
Dependent claim 11:
Step 2A Prong 1: The claim recites the abstract ideas of claim 10.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the weighting factor comprises a temporal decay factor - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein the weighting factor comprises a temporal decay factor - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Dependent claim 12:
Step 2A Prong 1: The claim recites the abstract ideas of claim 10.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
wherein the weighting factor comprises an evaluation metric for evaluating a trustworthiness of the human guidance component - the step recited at a high level of generality, and amounts to selecting a particular data source or type of data to be manipulated, which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
wherein the weighting factor comprises an evaluation metric for evaluating a trustworthiness of the human guidance component - viewed individually or in combination, describes selecting a particular data source or type of data to be manipulated similar to selecting information, based on types of information and availability of information in a power-grid environment, for collection, analysis and display described in MPEP § 2106.05(g).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Independent claim 13:
Step 2A Prong 1:
Claims recite:
determining, by the trained deep reinforcement learning model in response to input of the state information, an agent action indicative of a control signal - Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating data and selecting data based on judgement, which is observing, evaluating and judging that is practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: This judicial exception is not integrated into a practical application because they recite the additional elements:
obtaining parameters of a trained deep reinforcement learning model trained by a method according to claim 1 - the steps recited at a high level of generality, and amounts to mere data gathering which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g));
receiving state information indicative of an environment of the machine - the steps recited at a high level of generality, and amounts to mere data gathering which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g));
transmitting the control signal to the machine - the steps recited at a high level of generality, and amounts to mere data transmission which is well known which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g)).
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to the abstract idea.
Step 2B: The claims do not include additional elements that amount to significantly more than the judicial exception.
The additional elements:
obtaining parameters of a trained deep reinforcement learning model trained by a method according to claim 1 - the steps recited at a high level of generality, and amounts to mere data gathering which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g));
receiving state information indicative of an environment of the machine - the steps recited at a high level of generality, and amounts to mere data gathering which is a form of insignificant extra-solution activity (see MPEP § 2106.05(g));
transmitting the control signal to the machine - which is a well-understood, routine, conventional activity similar to receiving or transmitting data over a network described in MPEP 2106.05(d)(II).
Accordingly, these additional elements do not amount to significantly more than the judicial exception. As such, the claims are ineligible.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-3, 6-9 and 13-14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Palanisamy et al. (hereinafter Palanisamy), US 20200033868 A1.
Regarding independent claim 1, Palanisamy teaches a method of training a deep reinforcement learning model for autonomous control of a machine ([0009] In one embodiment, each of the driving policy learner modules comprises a Deep Reinforcement Learning (DRL) algorithm that can process input information from at least some of the driving experiences to learn and generate an output comprising: a set of parameters representing a policy that are developed through DRL, and wherein each policy is processible by at least one of the driver agents to generate an action for controlling the vehicle), the model being configured to output, by a policy network, an agent action in response to input of state information and a value function ([0075] The driving environment processor on the vehicle receives the information about the environment acquired using the variety of sensors on the vehicle as well as from other infrastructure-based information about the environment (e.g., from satellites/V2X etc), processes it and provides it as the “observation” input to the driver agent process. In the cases when the driving environments (e.g., in simulated driving environments) is fully observable, or assuming that they are fully observable improves the performance of the driving agents, the information about the environment can be provided as the “state” input to the agent. As used herein, the term “action (A),” when used with reference to a driving experience, can refer to the action performed by the autonomous driver agent which can include lower level control signals like steering, throttle, brake values or higher-level driving decisions like “accelerate by x.y”, “make a left lane change”, “stop in z meters.”; [0109] Each policy prescribes a distribution over a space of actions for any given state. The DRL algorithm 132 processes input information from driving experiences 122 (gathered by the driver agents 116-1 . . . 116-n from several driving environments 114) to generate an output that optimizes the expected discounted future rewards for each driver agent 116-1 . . . 116-n. The DRL algorithm 132 outputs parameters representing a policy (e.g., new policy parameters for a new policy or updated policy parameters for an existing policy). Depending on the implementation, the policy parameters can one or more of: (1) estimated (or predicted) values of state/action/advantage as determined by a state/action/advantage value function (i.e., estimate of how good it is to be in this state; estimate of how good an action is in this state; or estimate of an advantage of taking some action in this state); or (2) a policy distribution. The state/action/advantage value function(s) are used by the DRL algorithm 132 to produce policies (or parameters for policies) which are eventually used by the driving agents 116. The value functions are more like what the learners have learnt from their vast experiences collected by the driving agents from several different environments over a long period of time. The value functions are like the understanding of the world (driving environments). As such, unique policies 118 can be generated for each driver agent 116 to optimize performance of that driver agent 116 while operating in a certain driving environment and driving scenario and following that particular policy), the agent action representing a control signal for the machine ([0048] Control signals 72 (e.g., steering torque or angle signals used to generate corresponding steering torque or angle commands, and brake/throttle control signals used to generate acceleration commands) are sent to the actuator system 90, which processes the control signals 72 to generate the appropriate commands to control various vehicle systems and subsystems), the method comprising:
minimizing a loss function of the policy network ([0097] Reinforcement learning agents interact with an environment by receiving an observation that characterizes the current state of the environment, and in response, performing an action. Reinforcement learning (RL) can be used by an agent to learn to control a vehicle from sensor outputs. Reinforcement learning differs from supervised learning in that correct input-output pairs are not presented, but instead a machine (software agent) learns to take actions in some environment to maximize some form of reward or minimize a cost. Taking an action moves the environment/system from one state to another; [0112] A loss function is a function that maps an event or values of one or more variables onto a real number intuitively representing some “cost” associated with the event. Loss functions are used to measure the inconsistency between a predicted value and an actual value, or the inconsistency between a predicted value and a target value. Based on a metric implemented using the loss function, the loss function processes a batch of inputs (e.g., all of the learning targets from the learning target module 138, and all of the predictions from the DRL algorithm 132) to compute an overall output loss. As such, the overall output loss combines the losses for all the outputs of the DRL algorithm 132. When the DRL algorithm is an actor-critic based reinforcement learning algorithm, in which the critic predicts the state/action/advantage value function and the actor produces a policy distribution, the loss is the overall combined loss for both the actor and the critic (for the batch of inputs));
wherein the loss function of the policy network comprises an autonomous guidance component and a human guidance component ([0003] Vehicle automation has been categorized into numerical levels ranging from Zero, corresponding to no automation with full human control, to Five, corresponding to full automation with no human control; [0039] the autonomous vehicle 10 can be, for example, a Level Four or Level Five automation system. A Level Four system indicates “high automation”, referring to the driving mode-specific performance by an automated driving system of all aspects of the dynamic driving task, even if a human driver does not respond appropriately to a request to intervene. A Level Five system indicates “full automation”, referring to the full-time performance by an automated driving system of all aspects of the dynamic driving task under all roadway and environmental conditions that can be managed by a human driver); and
wherein the autonomous guidance component is zero when the state information is indicative of input of a human input signal at the machine ([0003] Vehicle automation has been categorized into numerical levels ranging from Zero, corresponding to no automation with full human control, to Five, corresponding to full automation with no human control).
Regarding dependent claim 2, Palanisamy teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Palanisamy further teaches wherein the model has an actor-critic architecture comprising an actor part and a critic part, and wherein the actor part comprises the policy network ([0016] In one embodiment, each of the driving policy learner modules comprises: a Deep Reinforcement Learning (DRL) algorithm that is configured to: process input information from at least some of the driving experiences to learn and generate an output comprising: a set of parameters representing a policy that are developed through DRL. Each policy is processible by at least one of the driver agents to generate an action for controlling the vehicle. Each DRL algorithm comprises: a policy-gradient-based reinforcement learning algorithm; or a value-based reinforcement learning algorithm or an actor-critic based reinforcement learning algorithm. The output of the DRL algorithm comprises one or more of: (1) estimated values of state/action/advantage as determined by a state/action/advantage value function; and (2) a policy distribution; [0112] When the DRL algorithm is an actor-critic based reinforcement learning algorithm, in which the critic predicts the state/action/advantage value function and the actor produces a policy distribution, the loss is the overall combined loss for both the actor and the critic (for the batch of inputs)).
Regarding dependent claim 3, Palanisamy teaches all the limitations as set forth in the rejection of claim 2 that is incorporated. Palanisamy further teaches wherein the critic part comprises at least one value network configured to output the value function ([0112] When the DRL algorithm is an actor-critic based reinforcement learning algorithm, in which the critic predicts the state/action/advantage value function and the actor produces a policy distribution).
Regarding dependent claim 6, Palanisamy teaches all the limitations as set forth in the rejection of claim 3 that is incorporated. Palanisamy further teaches wherein each value network is coupled to a target value network ([0017] In one embodiment, each of the driving policy learner modules further comprises: a learning target module configured to process trajectory steps of a driver agent within a driving environment to compute desired learning targets that are desired to be achieved, wherein each trajectory step comprises: a state, an observation, an action, a reward, a next-state and a next-observation, and wherein each learning target represents a result of an action that is desired for a given driving experience. In one embodiment, each of the learning targets comprises at least one of: a value target that comprises: an estimated value of a state/action/advantage to be achieved; and a policy objective to be achieved).
Regarding dependent claim 7, Palanisamy teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Palanisamy further teaches wherein the policy network is coupled to a target policy network ([0017] In one embodiment, each of the driving policy learner modules further comprises: a learning target module configured to process trajectory steps of a driver agent within a driving environment to compute desired learning targets that are desired to be achieved, wherein each trajectory step comprises: a state, an observation, an action, a reward, a next-state and a next-observation, and wherein each learning target represents a result of an action that is desired for a given driving experience. In one embodiment, each of the learning targets comprises at least one of: a value target that comprises: an estimated value of a state/action/advantage to be achieved; and a policy objective to be achieved).
Regarding dependent claim 8, Palanisamy teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Palanisamy further teaches wherein the deep reinforcement learning model comprises a priority experience replay buffer for storing, for a series of time points: the state information; the agent action; a reward value; and an indicator as to whether a human input signal is received ([0100] Experience replay is another technique used to store the agent's experiences at each time step, et=(st, at, rt, st+1) in a dataset D=e1, . . . , eN. This dataset D can be pooled over many episodes into replay memory. Here, s denotes the sequence, a denotes the action, and r denotes the reward for a specific timestep; [0087] the driving experiences collected by each driver agent 116-1 . . . 116-n can be stored in priority order (e.g., in an order that is ranked based on novelty/priority of each driving experience as determined by a prioritization algorithm 134 of the driving policy generation module 130). For example, the driving policy generation module 130 can update the relative priority/novelty/impact/effectiveness 126 of the driving experiences 124 in the experience memory 120, and then rank the driving experiences in a priority order. In one embodiment, when a driver agent 116 acquires a driving experience, it adds its own estimate of the priority as the Instance information (I) as described. The driving policy learner module(s) 131, which have access to much more information through the pooled experience memory 120, can update a value of priority/novelty/impact/effectiveness so that driving experiences with higher novelty/impact/effectiveness/priority are retrieved more often when they are sampled from the experience memory 120).
Regarding dependent claim 9, Palanisamy teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Palanisamy further teaches wherein the machine is an autonomous vehicle ([0006] System, methods and controller are provided for controlling an autonomous vehicle; [0039] In various embodiments, the vehicle 10 is an autonomous vehicle and an autonomous driving system (ADS) is incorporated into the autonomous vehicle 10 (hereinafter referred to as the autonomous vehicle 10) that intelligently controls the vehicle 10).
Regarding independent claim 13, Palanisamy teaches a method for autonomous control of a machine ([0006] System, methods and controller are provided for controlling an autonomous vehicle), comprising:
obtaining parameters of a trained deep reinforcement learning model trained by a method according to claim 1 ([0006] processing, at the one or more driver agents, received parameters for at least one candidate policy; Note: See rejection of claim 1 above for a trained deep reinforcement learning model trained by a method according to claim 1);
receiving state information indicative of an environment of the machine ([0075] Each driving environment processor 114-1 to 114-n can process sensor information that describes a particular driving environment. The sensor information can be acquired using the vehicle's on-board sensors including but not limited to cameras, radars, lidars, V2X communication and other sensors described herein. Driver agents 116-1 . . . . 116-n are artificial intelligence based autonomous driver agents. Each of the driver agents 116-1 . . . . 116-n can gather different driving experiences from different driving environments observed by the driving environment processors 114-1 to 114-n. In one embodiment, each driving experience can be represented in a large, multi-dimensional tensor that includes information from a particular driving environment at a particular time. Each experience includes: state (S), observation (O), action (A), reward (R), next state (S{circumflex over ( )}′), next observation (O{circumflex over ( )}′), goal (G), and instance information (I));
determining, by the trained deep reinforcement learning model in response to input of the state information, an agent action indicative of a control signal ([0075] As used herein, the term “action (A),” when used with reference to a driving experience, can refer to the action performed by the autonomous driver agent which can include lower level control signals like steering, throttle, brake values or higher-level driving decisions like “accelerate by x.y”, “make a left lane change”, “stop in z meters.” As used herein, the term “reward (R),” when used with reference to a driving experience, can refer to a signal that signifies how desirable the autonomous driver agent's performed action (A) is at some given time and environment conditions; [0076] Each policy 118 can process state (S) of the driving environment (as observed by a corresponding driving environment processor 114), and generate actions (A) that are used to control a particular AV that is operating in that state (S) of the driving environment); and
transmitting the control signal to the machine ([0076] Each low-level controller 120-1 . . . 120-n of FIG. 5 processes the action (or control signals 72 of FIG. 3) to generate signals or commands that control the actuators (actuator devices 42 a-42 n of FIG. 1) in accordance with the action (or control signals 72 of FIG. 3) to schedule and execute one or more control actions to be performed to automate driving tasks).
Regarding independent claim 14, it is a system claim that corresponding to the method of claim 1. Therefore, it is rejected for the same reason as claim 1 above.
Palanisamy further teaches the system (Fig. 1) comprising:
Storage (Fig. 1, 46; [0052]); and
at least one processor in communication with the storage (Fig. 1, 44; [0052]);
wherein the storage comprises machine-readable instructions for causing the at least one processor to execute a method according to claim 1 ([0053]).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Palanisamy as applied in claim 3, in view of SHALEV-SHWARTZ et al. (hereinafter SHALEV-SHWARTZ), US 20190333381 A1.
Regarding dependent claim 4, Palanisamy teaches all the limitations as set forth in the rejection of claim 3 that is incorporated. Palanisamy does not explicitly teach wherein the at least one value network is configured to estimate the value function based on the Bellman equation.
However, in the same field of endeavor, SHALEV-SHWARTZ teaches wherein the at least one value network is configured to estimate the value function based on the Bellman equation ([0206] The system may also be trained through value based learning (learning Q or V functions). Suppose a good approximation can be learned to the optimal value function V*. An optimal policy may be constructed (e.g., by relying on the Bellman equation); [0321] Many RL algorithms approximate the V function or the Q function in one way or another. Value iteration algorithms, e.g., the Q learning algorithm, may rely on the fact that the V and Q functions of the optimal policy may be fixed points of some operators derived from Bellman's equation. Actor-critic policy iteration algorithms aim to learn a policy in an iterative way, where at iteration t, the “critic” estimates Qπ t and based on this estimate, the “actor” improves the policy).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of V and Q functions of the optimal policy being fixed points of some operators derived from Bellman's equation as suggested in SHALEV-SHWARTZ into Palanisamy’s system because both of these systems are addressing autonomous vehicle navigation using reinforcement learning techniques. This modification would have been motivated by the desire of a fully autonomous vehicle that is capable of navigating on roadways (SHALEV-SHWARTZ, [0003]).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Palanisamy as applied in claim 3, in view of KARTAL et al. (hereinafter KARTAL), US 20200143206 A1.
Regarding dependent claim 5, Palanisamy teaches all the limitations as set forth in the rejection of claim 3 that is incorporated. Palanisamy does not explicitly teach wherein the critic part comprises a first value network paired with a second value network, each value network having the same architecture, for reducing or preventing overestimation.
However, in the same field of endeavor, KARTAL teaches wherein the critic part comprises a first value network paired with a second value network, each value network having the same architecture, for reducing or preventing overestimation ([0071] A standard reinforcement learning setting comprises an agent interacting in an environment over a discrete number of steps. At time t the agent in state st takes an action at and receives a reward rt. The discounted return is defined as:
R t:∞=Σt=1 ∞γt r t,
the state-value function is the expected return (sum of discounted rewards) from state s following a policy π after taking action a from state s: π(a|s):
V π(s)=Figure US20200143206A1-20200507-P00001[R t:∞ |s t =s,π],
and the action-value function is the expected return following policy π after taking action a from state s:
Q π(s,a)=Figure US20200143206A1-20200507-P00001[R t:∞ |s t =s,a t =a,π]; [0074] The policy and the value function are updated after every tmax actions or when a terminal state is reached. One softmax output may be used for the policy π(at|st;θ) head and one linear output for the value function V(st;θv) head, with all non-output layers shared (see FIGS. 7A and 7B). FIGS. 7A and 7B illustrate, in component diagrams, examples of a neural network architecture of an A3C 700 and A3C-TP 750 actor-critic neural network architecture, in accordance with some embodiments. The convolutional neural networks (CNNs) 710 shown represent convolutional neural network layers, and FC 720 represent a fully-connected layer. FIG. 7A shows a standard actor-critic neural network architecture 700 for A3C with policy (actor) 740 and value (critic) 730 heads. FIG. 7B shows the neural network architecture 750 of A3C-TP is identical to that of A3C, except with the additional auxiliary TP head 750).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of an algorithm A3C (Asynchronous Advantage Actor-Critic) employing a parallelized asynchronous training scheme as suggested in KARTAL into Palanisamy’s system because both of these systems are addressing deep reinforcement learning. This modification would have been motivated by the desire of distributed RL methods enabling efficient learning with better exploration (KARTAL, [0011]).
Claims 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Palanisamy as applied in claim 1, in view of WU et al. (hereinafter WU), US 20180322391 A1.
Regarding dependent claim 10, Palanisamy teaches all the limitations as set forth in the rejection of claim 1 that is incorporated. Palanisamy does not explicitly teach wherein the loss function includes an adaptively assigned weighting factor applied to the human guidance component.
However, in the same field of endeavor, WU teaches wherein the loss function includes an adaptively assigned weighting factor applied to the human guidance component ([0073] In one example non-limiting implementation, as part of the forward pass of training DNN 200, the system scales the loss value by some factor S (see FIG. 6 block 680 and FIG. 6A); [0139] Sometimes, weight decay as shown in FIG. 13D is applied to the loss added to the loss function itself because this may avoid the network learning weights that are too big that can result in overfit).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of selecting scaling factor S and for adjusting the weight update as suggested in WU into Palanisamy’s system because both of these systems are addressing deep reinforcement learning. This modification would have been motivated by the desire to improve applying reduced precision to machine learning and training (WU, [0011]).
Regarding dependent claim 11, the combination of Palanisamy and WU teaches all the limitations as set forth in the rejection of claim 10 that is incorporated. WU teaches wherein the weighting factor comprises a temporal decay factor ([0139] Sometimes, weight decay as shown in FIG. 13D is applied to the loss added to the loss function itself because this may avoid the network learning weights that are too big that can result in overfit. Thus, some neural network frameworks and/or learning algorithms punish the weights through the weight decay parameter in order to prevent the weights from getting too large. If there are many weights that are allowed to vary in value widely, this results in a complex function that is difficult to analyze to determine what may be wrong. Increasing the weight decay parameter has the effect of dumbing down the training to avoid overfit errors. This adapts the network to work well on the data it has seen while still allowing it to work well on data that it has not yet seen. In this example, the forward propagation, the backward propagation, and the weight updates occur with the appropriate parameter(s) and/or the weight gradients being modified or adjusted to compensate for the S scaling that was performed on the loss factor before the gradients were calculated for the iteration).
Regarding dependent claim 12, the combination of Palanisamy and WU teaches all the limitations as set forth in the rejection of claim 10 that is incorporated. WU teaches wherein the weighting factor comprises an evaluation metric for evaluating a trustworthiness of the human guidance component ([0132] As an example, if S=1000, then the weight gradients will be one thousand times larger than they would have otherwise have been and using such weight gradients to update the weights will result in the weight update that is one thousand times larger than it should be. It is useful to address this in order to prevent the neural training from failing or becoming inaccurate. Additionally, since the neural network training is iterative, inaccuracies at an earlier iteration have the capacity to affect all succeeding iterations, and systematically introduced errors may be compounded).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Applicant is required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action.
HOEL et al. (US 20230142461 A1) discloses a method of controlling an autonomous vehicle using a reinforcement learning, RL, agent.
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMY P HOANG whose telephone number is (469)295-9134. The examiner can normally be reached M-TH 8:30-5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER WELCH can be reached at 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AMY P HOANG/Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143