DETAILED ACTION
Claims 1-20 are presented for examination.
This Office Action is in response to submission of documents on June 4, 2026.
Rejection of claims 1-20 under 35 U.S.C. 101 for being directed to unpatentable subject matter are withdrawn.
Rejection of claims 1-20 under 35 U.S.C. 103 as being obvious over Grimm in view of Nagy is withdrawn.
Rejection of claims 1-4, 6-11, 13-19, and 20 are rejected under 35 U.S.C. 103 as being obvious over Grimm in view of Akash and Li.
Rejection of claims 5, 12, and 19 under 35 U.S.C. 103 as being obvious over Grimm in view of Akash, Li, and Nagy.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see pages 9-11 of Response, filed June 4, 2026, with respect to the rejection under 35 U.S.C. 101 have been fully considered and are persuasive. Accordingly, the rejection of claims 1-20 under 35 U.S.C. 101 has been withdrawn.
Applicant’s arguments, see pages 11-13 of Response, filed June 4, 2026, with respect to the rejection under 35 U.S.C. 103 have been fully considered and are persuasive. Accordingly, the rejection of claims 1-20 under 35 U.S.C. 103 has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Akash, Li, and Nagy.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
The independent claims recite:
executing, by the processor, a simulation of a driving environment, wherein, in the simulation:
the simulated vehicle receives an action of the joint policy;
(ii) an intervention action of the adaptive HMI system updates an acceptance state of the simulated human driver;
(iii) the acceptance state modulates a transition probability of a distraction state of the simulated human driver;
(iv) the distraction state determines a vehicle action applied to the simulated vehicle; and
(v) the simulation outputs the observations;
The limitation, as currently presented, recites a simulation that receives, as input, an action of the joint policy and provides, as output, observations. However, the Specification does not disclose such a process. Instead, referring to FIG. 1, when training the one or more policies, input from environment 12 is received by the simulated human 14 and adaptive HMI 18. “The one or more policies 134 are learned from observations of the world based on collected experience and are gradually improved through maximization of a total reward for each roll-out generated by the one or more policies 134 as it is trained. Observations are potentially corrupted measurements of the state. The model of the human is expressed as the MDP, which specifies the probability of transitioning from one state to another, along with any rewards (e.g., speed preferences) or penalties (e.g., collisions) collected along the way.” Spec. at [0030].
Based on the disclosure and FIG. 1, the process for training the joint policy requires receiving observations, evaluating the observations via the one or more policies to determine one or more actions, and outputting the actions. The joint policy can then be optimized to maximize rewards from actions.
Thus, because the Specification does not appear to disclose using received actions to determine an observation, as claimed, claims 1-20 are rejected under 35 U.S.C. 102(a).
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Independent claims 1, 8, and 15 recite “observations,” but it is unclear what is being observed and/or what entity is performing the observations. For the purposes of interpretation, Examiner interprets “observations” to mean “input from the current state of a simulated environment,” as is inferred by environment 12 and the arrows originating from that component.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-20 are rejected under 35 U.S.C. 103 as being obvious over Grimm, et al., (WIPO Pub. No. 2020/060478, hereinafter “Grimm”) in view of Akash, et al., (“Toward Adaptive Trust Calibration for Level 2 Driving Automation,” hereinafter “Akash”) and Li, et al., (“Online Markov Decision Processes with Time-varying Transition Probabilities and Rewards,” hereinafter “Li”).
Claim 1
Grimm discloses:
A method comprising:
A virtual traffic agent can for example be a car, truck, bus, bike or motor bike. Once a virtual traffic agent has been trained in a way that replicates human driving behavior for a predetermined geographical area, the goal is to inject one or more trained virtual traffic agents into a simulation environment representing the predetermined geographical area where they can interact, cooperate with and challenge an autonomous vehicle system controlling an autonomous vehicle under test. Grimm at [0040].
Setting, by a processor, parameters of rewards and a Markov Decision Process (MDP) of a
Reinforcement learning uses the formal framework of Markov Decision Process (MDP) to define the interaction between a learning agent and its environment in terms of states, actions and rewards. Grimm at [0066].
“Reinforcement learning” requires the use of one or more processors and “defin[ing] the interaction” is analogous to “setting parameters.”
training, by the processor via the execution of the simulation, the
The presently disclosed methods and systems of training a traffic agent 20 are based on reinforcement learning methods. Reinforcement learning is a category of machine learning in which a machine (agent) learns a policy specifying which action to take in any situation (state), in order to maximize the expected reward according to a reward function. Grimm at [0046].
evaluating output of the trained
When the traffic agent 20 is in the evaluation phase, the evaluator 155 will proceed to judge the performance of the traffic agent 20. This is done through a comprehensive set of predefined mathematical indicators. For example, if the predefined mathematical indicators are indicative of driver behavior patterns, the evaluator 155 will compare each of the mathematical indicators associated with the traffic agent 20 and compare them with the mathematical indicators obtained from real world driving behavior data collected from real drivers or vehicle trajectories extracted from GPS data or video data. Grimm at [0065].
Grimm does not appear to disclose:
the joint policy comprising a function mapping observations to actions of a simulated human driver of a simulated vehicle and intervention actions of an adaptive human-machine interface (HMI) system,
executing, by the processor, a simulation of a driving environment, wherein, in the simulation: (i) the simulated vehicle receives an action of the joint policy; (ii) an intervention action of the adaptive HMI system updates an acceptance state of the simulated human driver; (iii) the acceptance state modulates a transition probability of a distraction state of the simulated human driver; (iv) the distraction state determines a vehicle action applied to the simulated vehicle; and (v) the simulation outputs the observations;
Akash teaches, which is analogous art, discloses:
the joint policy comprising a function mapping observations to actions of a simulated human driver of a simulated vehicle and intervention actions of an adaptive human-machine interface (HMI) system,
In this paper, we present 1) an interaction model of human trust workload dynamics and 2) optimization of the system transparency level based on the estimated human state in a hands-off SAE (Society of Automotive Engineers) level 2 driving context. The driving automation was chosen as it is a promising application of action automation-based systems. Akash at pg. 2, col. 1.
executing, by the processor, a simulation of a driving environment, wherein, in the simulation:
We develop a probabilistic model of the user trust and workload dynamics using human subject data collected using a driving simulator for urban driving scenes. We then optimize system behavior of dynamically varying automation transparency to achieve a better trust-workload tradeoff considering the automation performance. Akash at pg. 2, cols. 1-2.
(i) the simulated vehicle receives an action of the joint policy;
An extension of HMMs, partially observable Markov decision processes (POMDPs), provide a framework that does account for these actions and enables the design and synthesis of a policy to choose optimal actions based on a desired reward function. Akash at pg. 2, col. 2.
(ii) an intervention action of the adaptive HMI system updates an acceptance state of the simulated human driver;
Finally, we assume that human trust and workload are influenced by the characteristics of the automation—reliability and transparency—as well as that of the environment—i.e. scene complexity. The model structure based on these assumptions is illustrated in Figure 3. Table 1 summarizes the definitions of the model variables. Akash at pg. 3, col. 2.
See pg. 3, col. 2 (Figure 3):
PNG
media_image1.png
198
306
media_image1.png
Greyscale
S’T and S’W are analogous to “updated acceptance states.”
aT is analogous to the “intervention action of the adaptive HMI system.”
See also pg. 4, Table 1, which lists the definitions for the states of Figure 3:
PNG
media_image2.png
260
684
media_image2.png
Greyscale
(iv) the distraction state determines a vehicle action applied to the simulated vehicle;
However, while more information is typically communicated to the human to achieve greater transparency, it often results in increased cognitive workload [24, 56] and can distract the human from the most critical information [7] as well as sacrifice one of the primary benefits of the automation, i.e., reduction of human workload. Therefore, a tradeoff between increased trust and increased workload exists when considering increased transparency [6]. Akash at pg. 2, col. 1.
and (v) the simulation outputs the observations;
Given the need to strike a balance between model fidelity and complexity, we assume that any coupled interactions between these particular states and observations can be captured through the coupled interaction between trust and workload; doing so facilitates parameterization of the model as described later in this section. Akash at pg. 3, col. 2.
Akash is analogous art to the claimed invention because both are directed to policy training of an HMI system that performs actions while in operation. It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to combine Grimm and Akash to result in a system that trains a policy by simulating real world driving conditions. Motivation to combine includes development of an accurate warning system for a driver without requiring real environment conditions. Using a live driver in an actual real world situation poses potential dangers to the driver and to others that is prevented when simulations are employed to determine appropriate policies and rewards for a driver based on distraction.
While Akash teaches an acceptance state and a distraction state, the reference does not appear to disclose:
(iii) the
Li, which is analogous art, discloses:
(iii) the
In this paper, we focus on designing online algorithms under both time-varying transition probabilities and rewards, aiming at good dynamic regrets. Li at pg. 2, col. 1.
As indicated in the Specification, the changes (i.e., “modulation” of transition probabilities is dependent on time in an acceptance state: “From the above equations, one can see that if the simulated human driver 14 accepts an alert from the adaptive HMI system 18, the acceptance state will be set to 1. As a result, the transition probabilities in Equations (4) and (5) are modulated. This modulation remains in effect for at least N time steps. In the framework, at every timestep t, first c, is updated, followed by the acceptance state it, and then finally the distraction variable d.” Spec. at [0038].
Li is analogous art to the claimed invention because it is directed to a Markov Decision Process that modulates transition probabilities depending on time spent in a state. It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to combine Li with Grimm and Akash to result in a simulation and training of a policy that includes changing a transition probability to a distracted state based on time spent in an acceptance state. Motivation to combine includes utilizing a more realistic modeling of the real distraction tendencies of a driver, thus improving a system that is deployed to provide alerts to a driver when distraction is likely. Such a system would improve road safety.
Claim 2
Grimm discloses:
wherein the driving environment is a simulated road environment.
The road infrastructure database 121 includes a road infrastructure generator 122 which generates road infrastructure elements such as basic rules and balanced elements in order to provide an unbiased training. Grimm at [0049].
Claim 3
Grimm discloses:
wherein actions of the at least one policy includes human initiated vehicle actions by the
For example, if an aggressive driver behavior pattern is to be emulated, some indicators of aggression, both qualitatively and quantitatively, may be used. Grimm at [0060].
Claim 4
Grimm discloses:
wherein the human-initiated vehicle actions include speeding up the simulated vehicle, slowing down the simulated vehicle, causing the simulated vehicle to move left, causing the simulated vehicle to move right, and maintaining the speed of the simulated vehicle.
For example, if the traffic agent is a simulated vehicle navigating through a simulation environment, the action to be selected may be a driving controller 160 or control inputs to control the simulated vehicle. The action selected can be from one of the choices: keep, accelerate, decelerate, drive left or drive right. Grimm at [0053].
Claim 6
Grimm discloses:
wherein the parameters of the rewards of the at least one policy include cautiousness exhibited by the simulated human driver, a likelihood of the simulated human driver becoming distracted and attentive, and
Qualitatively, the perceived level of attention, i.e. potential impairment of driver, distraction, asleep, or refusing to cooperate with other road users may be used as reward parameters 140. Grimm at [0060].
a willingness of the
In another example, a traffic agent 20 can also be trained in other driving competencies that emulate real-world drivers, for example, cooperative driving behavior or defensive driving behavior that allow traffic agents to communicate with each other by giving way to each other or to inhibit letting the other road user on a highway entry ramp. Grimm at [0048].
The traffic agent learning system 100 may be trained through exposure to various navigational states, and when the system applies the policy, it provides a reward based on a reward function that is designed to reward desired navigational or driver behavior. Grimm at [0046].
“Desired navigational or driver behavior” is analogous to “cautiousness,” “willingness,” and distraction level of a driver.
Claim 7
Grimm discloses:
wherein the at least one policy is one of: a joint policy modeling actions of the simulated human driver and the adaptive HMI system; and separate policies that separately model actions of the simulated human driver and the adaptive HMI system.
FIG. 5 illustrates an abstract and representative example of a neural network architecture employed by an embodiment of the invention. It illustrates the relationship between the traffic agent 20, the simulation environment 400 and the evaluator 155. Grimm at [0065].
As illustrated in FIG. 5, the traffic agent 20 (analogous to the “simulated human driver”) and the simulation environment 400 (analogous to the HMI) are part of the same modeling system.
FIG. 7 illustrates a graphical diagram of some of the components of the traffic agent learning system 100. Particularly, it illustrates the relationship between the traffic agent 20, simulation environment 400, the reward generator 153, reward parameters 140 and the neural network model 151. Grimm at [0067].
Claim 8-11 and 13-14
Claims 8-14 recite:
a processor; and
a memory in communication with the processor, the memory storing instructions
for a method that is substantially the same as the method disclosed in claims 1-7.
Accordingly, for at least the same reasons and based on the same prior art as claims 1-7, claims 8-14 are rejected under 35 U.S.C. 103 as being obvious over Grimm in view of Akash and Li.
Claim 15-18 and 20
Claims 15-20 recite:
a non-transitory computer-readable medium storing instructions
to perform a method that is substantially the same as the method disclosed in claim 1-6.
Accordingly, for at least the same reasons and based on the same prior art as 1-6, claims 15-20 are rejected under 35 U.S.C. 103 as being obvious over Grimm in view of Akash and Li.
Claims 5, 12, and 19 are rejected under 35 U.S.C. 103 as being obvious over Grimm in view of Akash, Li, and further in view of Nagy, et al., (U.S. Pat. No. 10,449,957, hereinafter "Nagy").
Claim 5, 12, and 19
Grimm, Akash, and Li do not appear to disclose:
wherein the intervention actions include providing an alert to the simulated human driver and not providing the alert to the simulated human driver.
Nagy, which is analogous art, discloses:
wherein the intervention actions include providing an alert to the simulated human driver and not providing the alert to the simulated human driver.
The human machine interface (HMI) 22 provides an interface between the autonomous vehicle control system 10 and the driver. The HMI 22 is electrically coupled to the electronic controller 12 and receives input from the driver, receives information from the electronic controller 12, and provides feedback (e.g., audio, visual, haptic. or a combination thereof) to the driver based on the received information. Nagy at col. 6, lines 45-53.
Nagy is analogous art to the claimed invention because both are directed to alerts provided by an HMI based on human behavior. It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to combine the alert actions of Nagy with the system disclosed in the other references to result in a system that provides alerts when an MDP indicates that the driver is likely in a distracted state. Motivation to combine includes the inclusion of alerts to an HM, which improves the likelihood of transitioning to an acceptance state thereby improving the safety of the driver and other drivers when the system is deployed.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Misu, et al., (U.S. Pat. No. 11,498,591): Discloses determining driver trust in an automated driving system and adjusting the system based on driver trust.
Drescher, et al. (U.S. Pat. No. 9,542,848): Discloses adjusting an alarm of a vehicle based on learning habits and reactions of drivers.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Communication
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSEPH MORRIS whose telephone number is (703)756-5735. The examiner can normally be reached M-F 8:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ryan Pitaro can be reached at (571) 272-4071. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JOSEPH MORRIS
Examiner
Art Unit 2188
/JOSEPH P MORRIS/Examiner, Art Unit 2188
/RYAN F PITARO/Supervisory Patent Examiner, Art Unit 2188