DETAILED ACTION
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 5/18/2026 has been entered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This action is in response to the amendment filed on 5--/18/2026 for application 17/951,001. Applicant’s amendment filed on 5/18/2026 has been entered. Claim 1 – 12, 14 and 16 – 22 are pending and have been examined.
Claim 1, 6, 16, 20 are amended.
Claim 13 is canceled.
Claim 21 – 22 are new.
Claim rejection under 35 U.S.C. 101 has been withdrawn in light of the applicants amendment.
Response to Argument
Applicant’s arguments with respect to claim rejection under 35 U.S.C. 102 and 103 section have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Objection
Claim 14 is objected to because of the following informalities: Claim 14 is depending on Claim 13 which is canceled. Appropriate correction is required. For the examination purpose, Claim 14 is interpreted as depending on Claim 1.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 1 – 5, 14, 17 – 22 are rejected under 35 U.S.C. 103 as being unpatentable over Sivaraman et al., (hereinafter Sivaraman), US20210397885 in view of Weinstein-Raun, US20180079423.
Regarding Claim 1, Sivaraman discloses: A method performed by one or more computers, the method comprising: obtaining context data characterizing an environment, the context data comprising data characterizing a plurality of agents in the environment (0030 – 0040, “sensor data may include, without limitation, sensor data from any of the sensors 102 of the vehicle 800 (and/or other vehicles (agents in the environment) or objects, such as robotic devices, VR systems, AR systems, etc.,”, “In some examples, the sensor data may be generated by one or more forward-facing sensors, side-view sensors, interior sensors, and/or rear-view sensors of the vehicle 800 and/or other machine type. This sensor data may be useful for identifying, detecting, classifying, and/or tracking movement of objects (context data that characterize agents in the environment) around the vehicle 800 and/or other machines within the environment”, “The classifier 108 may generate raw classification outputs using the sensor data as input”);
generating, based on processing the context data using a neural network having a softmax output layer, a respective predicted multi-category classification for each of one or more target agents of the plurality of agents in the environment, the respective predicted multi-category classification comprising a respective score generated by the softmax output layer for each of a plurality of categories (refer to the mapping above & , 0028, “The classifier 108 may be a multiclass classifier that uses one or more CNNs (neural network) ----or other deep neural network (DNN) and/or machine learning models-to process the inputs.”; 0031 – 0040, “The classifier 108 may generate raw classification outputs using the sensor data as input”, “The softmax output 210 may be generated by the softmax function 114, and the softmax output 210 may include a confidence score for each class the multiclass classifier is trained to identify”; in this case the categories as claimed are the classes identified by the multi-class classifier)
controlling one or both of steering or braking of a vehicle based on the respective predicted multi-state classification generated for each of the one or more target agents (refer to the mapping above & 0173, “AEB systems detect an impending forward collision with another vehicle or other object, and may automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter.”).
Sivaraman does not explicitly teach:
including an active participant category that indicates that the target agent is an active participant of the traffic in the environment and two or more stationary non-participant categories that represent different categories of stationary non-participant of traffic in the environment
Weinstein-Raun, in the same field of endeavor, explicitly teach:
including an active participant category that indicates that the target agent is an active participant of the traffic in the environment and two or more stationary non-participant categories that represent different categories of stationary non-participant of traffic in the environment (Fig. 7 & 0089 – 0100, “dynamic Bayesian network 700 includes various nodes 702-736 (states/categories)”, “in various embodiments, the intermediate nodes 726-734 include: (i) a blocked state 726 representing whether the target vehicle is blocked from movement; (ii) a pulled over state 728 representing whether the target vehicle is pulled over ( e.g., in certain embodiments, along the roadway); (iii) a motion state 730 pertaining to motion of the target vehicle ( e.g., as to a magnitude and direction of movement of the target vehicle, in certain embodiments); (iv) a passable state 732 (e.g., as to whether the vehicle 10 is able to successfully maneuver around the target vehicle, in certain embodiments); and (vi) an apparent activity state 734 (e.g., as to an indication of whether the target vehicle is active or inactive)”; 0073, “change the current path to drive around the target vehicle if the target vehicle is inactive”; 0086, “the target vehicle comprises a vehicle that is stopped within or proximate the same lane … contact would be likely if the target vehicle were to remain stopped and the vehicle 10 were to maintain its current path”; i.e., the system evaluates each of the target agents the probability/likelihood on each of the stationary non-participant state (blocked for movement, pulled over, passable state) and moving participant state (motion state) that the targe agent is likely in and use these information to determine/control the policy of the vehicle.);
Sivaraman and Weinstein-Raun both teach the use of neural network multi-class classifier on the vehicle environment sensing/detection for autonomous driving and are analogous. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention with a reasonable likelihood of success to further include the detection/determination of the traffic participation status of surrounding vehicles for the maneuver policy of the autonomous vehicle taught by Weinstein-Raun into the system of Sivaraman to achieve the claimed teaching. One of the ordinary skill in the art would have motivated to make this modification to improve movement of autonomous vehicle in response to another vehicle (Weinstein-Raun, 0002 – 0004).
Regarding Claim 2, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 1. Weinstein-Raun further teach: the two or more stationary nonparticipant categories comprise two or more of: a first stationary non-participant category that indicates that the target agent is a double parked vehicle in the environment, a second stationary non-participant category that indicates that the target agent is a parked vehicle in the environment, a third stationary non-participant category that indicates that the target agent is a pulled over vehicle in the environment, or a fourth stationary non-participant category that indicates that the target agent is a stalled vehicle in the environment (refer to the mapping in Claim 1, the system identifies the possibility of at least the following two non-participant states/categories that the target vehicle is pulled over, the target vehicle is blocked (stalled) ).
The reason for combination is same as Claim 1
Regarding Claim 3, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 2. Weinstein-Raun further teach: the respective score for each of the plurality of categories represents a likelihood that the target agent is in the category (refer to the mapping in Claim 1 & Sivaraman, 0031 – 0040, the confidence score is referring to the likelihood of each states of the target vehicle/agent).
Regarding Claim 4, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 1. The combination further teach: the context data characterizing the environment comprises data characterizing a plurality of road features in the environment and data characterizing one or more traffic light signals in the environment (refer to the mapping in Claim 1 & Sivaraman, 0091, “information about objects and status of objects as perceived by the controller … one or more objects (e.g., a street sign, caution sign, traffic light changing, etc.)”; 0183, “surrounding environment information ( e.g., intersection information, vehicle information, road information, etc.), and/or other information.”; i.e., the system uses road feature information and traffic light information to make decisions).
Regarding Claim 5, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 4. The combination further teach: the data characterizing each of the plurality of road features comprises one or more road feature vectors characterizing the road feature; the data characterizing each of the plurality of agents comprises one or more agent vectors characterizing the agent; and the data characterizing each of the plurality of traffic light signals comprises one or more traffic light vectors characterizing the traffic light signal (refer to the mapping in Claim 1, the system of Sivaraman uses computerized machine learning model to process the input/sensor/vehicle data for inference. One of skilled in the art would recognize that a collection of values is a vector in machine learning field. Such understanding can be easily found online for example “vector is a general term with many uses. In this case, think of it as a list of values or a row in a table.”, StackOverflow, “What is vector in terms of machine learning”).
Regarding Claim 14, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 13. The combination further teach: each of the plurality of agents is an agent in a vicinity of the vehicle in the environment, and the context data comprises data generated from data captured by one or more sensors of the vehicle (refer to the mapping in Claim 1, & Sivaraman, 0030 – 0040, other vehicles are in the vicinity of the vehicle and context data are collected by forward-facing sensors, side-view sensors, rear-view sensors of the vehicle 800 and/or other machine type).
Regarding Claim 17 – 19 and 21, these are the corresponding system claim of Claim 1 – 3 & 14. Sivaraman further teach: one or more computers, and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations (Sivaraman, 0197, “computer-storage media … computer-readable instructions,”). These claims are rejected with same reason.
Regarding Claim 20 and 22, these are the corresponding computer-readable storage media claim of Claim 17 & 21. These claims rejected with same reason.
Claim(s) 6 – 12 are rejected under 35 U.S.C. 103 as being unpatentable over Sivaraman et al., (hereinafter Sivaraman), US20210397885 in view of Weinstein-Raun, US20180079423 as applied to claim 5 above, and further in view of Gao et al., (hereinafter Gao), “VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation”.
Regarding Claim 6, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 5. The combination further teach: processing the context data using the neural network (refer to the mapping in Claim 1 & Sivaraman teaches using neural network to perform classification)
The combination does not explicitly teach:
generating, from the context data, (i) a respective road feature embedding for each of the plurality of road features, (ii) a respective agent embedding for each of the plurality of agents that characterizes a state of the agent at a current time point, and (iii) a respective traffic light embedding for each of the plurality of traffic light signals that characterizes a state of the traffic light signal at the current time point; and for each of the one or more target agents of the plurality of agents in the environment: generating, from the respective road feature embeddings, the respective agent embeddings, and the respective traffic light embeddings, (i) agent interaction embeddings characterizing the states of other agents in the environment relative to the target agent, and (ii) road feature interaction embeddings characterizing the plurality of road features in the environment relative to the target agent.
Gao, in the same field of endeavor, explicitly teach:
generating, from the context data, (i) a respective road feature embedding for each of the plurality of road features, (ii) a respective agent embedding for each of the plurality of agents that characterizes a state of the agent at a current time point, and (iii) a respective traffic light embedding for each of the plurality of traffic light signals that characterizes a state of the traffic light signal at the current time point (Gao, sec. 3.1, “Most of the annotations from an HD map are in the form of splines (e.g. lanes), closed shape (e.g. regions of intersections) and points (e.g. traffic lights), with additional attribute information such as the semantic labels of the annotations and their current states (e.g. color of the traffic light, speed limit of the road). For agents, their trajectories are in the form of directed splines with respect to time. All of these elements can be approximated as sequences of vectors”, “Our polyline subgraph network … embedding the ordering information into vectors”); and
for each of the one or more target agents of the plurality of agents in the environment: generating, from the respective road feature embeddings, the respective agent embeddings, and the respective traffic light embeddings, (i) agent interaction embeddings characterizing the states of other agents in the environment relative to the target agent, and (ii) road feature interaction embeddings characterizing the plurality of road features in the environment relative to the target agent (Gao, fig. 2 & sec. 1, the global interaction graph captures interaction embeddings between agents (the green blocks) and the interaction (road feature interaction embeddings) between each road feature embeddings/blue blocks and green blocks).
Sivaraman and Weinstein-Raun combination and Gao both teach multi agent behavior prediction learning model for the autonomous vehicle and are analogous. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention with a reasonable likelihood of success to further apply the VectorNet of Gao’s teaching to the system of Sivaraman and Weinstein-Raun combination to achieve the claimed teaching. One of the ordinary skill in the art would have motivated to make this modification for the model “outperforms the state of the art” (Gao, abs.).
Regarding Claim 7, Sivaraman, Weinstein-Raun and Gao combination teaches all the limitation of Claim 6. The combination further teach: generating the respective road feature embedding for each of the plurality of road features comprises generating a respective polyline that represent the road feature (Gao, sec. 1, “geographic extent of the road features … can be closely approximated as polylines”).
The reason for combination is same as Claim 6.
Regarding Claim 8, Sivaraman, Weinstein-Raun and Gao combination teaches all the limitation of Claim 6. The combination further teach: generating the respective road feature embedding for each of the plurality of road features comprises generating a respective polyline embedding for the respective polyline that represents the road feature using a road feature encoder neural network (Gao, sec. 3.1, “lanes … regions … intersections (road feature) … All of these elements can be approximated as sequences of vectors”, “we treat each vector vi belonging to a polyline Pj as a node in the graph with node features”; sec. 3.2, “Function genc(.) transforms the individual node features”, “genc(.) is a multi-layer perceptron (MLP)”; i.e., use MLP, which is a neural network, to transform/encode road feature into embeddings).
The reason for combination is same as Claim 6
Regarding Claim 9, Sivaraman, Weinstein-Raun and Gao combination teaches all the limitation of Claim 6. The combination further teach: generating the respective agent embedding and the respective traffic light embedding comprises: processing the one or more agent vectors using an agent encoder neural network to generate the respective agent embedding; and processing the one or more traffic light vectors using a traffic light encoder neural network to generate the respective traffic light embedding (Gao, sec. 3.1, “traffic lights … agents … All of these elements can be approximated as sequences of vectors”, “we treat each vector vi belonging to a polyline Pj as a node in the graph with node features”; sec. 3.2, “Function genc(.) transforms the individual node features”, “genc(.) is a multi-layer perceptron (MLP)”; i.e., use MLP, which is a neural network, to transform/encode traffic light and agent into embeddings).
The reason for combination is same as Claim 6.
Regarding Claim 10, Sivaraman, Weinstein-Raun and Gao combination teaches all the limitation of Claim 9. The combination further teach: the agent encoder neural network and the traffic light encoder neural network are each a respective multi-layer perceptron or a recurrent neural network (Gao, sec. 3.2, “genc(.) is a multi-layer perceptron (MLP)”).
The reason for combination is same as Claim 6.
Regarding Claim 11, Sivaraman, Weinstein-Raun and Gao combination teaches all the limitation of Claim 6. The combination further teach: in generating the agent interaction embeddings and the road feature interaction embeddings comprises: processing the respective road feature embeddings, the respective agent embeddings, and the respective traffic light embeddings using a self-attention neural network to generate the agent interaction embeddings and the road feature interaction embeddings (Gao, sec. 3.3, “We now consider modeling the high-order interactions on the polyline node features {p1, p2, …, pP} with a global interaction graph”, “Our graph network is implemented as a self-attention operation”).
The reason for combination is same as Claim 6
Regarding Claim 12, Sivaraman, Weinstein-Raun and Gao combination teaches all the limitation of Claim 6. The combination further teach: generating the predicted classification for the target agent comprises: processing the respective agent interaction embeddings and the respective road feature interaction embeddings using an output neural network to generate the predicted classification (Gao, sec. 1, “trajectories of other moving agents are propagated to the target agent node through the GNN. We can then take the output node feature corresponding to the target agent to decode its future trajectories”; sec. 3.3, “We now consider modeling the high-order interactions on the polyline node features {p1, p2, …, pP} with a global interaction graph”, “We then decode the future trajectories from the nodes corresponding the moving agents”, “we use an MLP as the decoder function”; i.e., use MLP, which is a neural network, to decode/process the graphical neural network, which models the interactions between nodes, to get the predicted result. Examiner notes that even though Gao’s final output is value instead of class as claimed, Weinstein-Raun teaches predicting the classes representing different state of the agent. The combination renders obviousness of the claimed limitation. ).
The reason for combination is same as Claim 6.
Claim(s) 16 is rejected under 35 U.S.C. 103 as being unpatentable over Sivaraman et al., (hereinafter Sivaraman), US20210397885 in view of Weinstein-Raun, US20180079423 as applied to claim 1 above, and further in view of Pink, US20140336912.
Regarding Claim 16, Sivaraman and Weinstein-Raun combination teaches all the limitation of Claim 1. The combination further teach: obtaining context data characterizing the environment comprises obtaining data indicating that the one or more target agents are stationary at a current time point (refer to the mapping in Claim 1 & Weinstein-Raun, 0092, “percept nodes included … an observed movement state 708 as to whether or not the target vehicle is currently moving”),
The combination does not explicitly teach:
wherein the generating is performed only for target agents that are indicated as stationary at the current time point
Pink, explicitly teach:
wherein the generating is performed only for target agents that are indicated as stationary at the current time point (Pink, 0009, “upon detection of a stopped vehicle, and decision data being formed based on the starting intention criterion, which include a pieces of information on whether or not the stopped vehicle will start”; i.e., Pink teaches an intention prediction model that focus only on the stopped/stationary vehicle).
Sivaraman and Weinstein-Raun combination and Pink both teach multi agent behavior prediction learning model for the autonomous vehicle and are analogous. It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention with a reasonable likelihood of success to further apply Pink’s control logic that applies specifically to the intention of stopped vehicle to the system of Sivaraman and Weinstein-Raun combination to achieve the claimed teaching. One of the ordinary skill in the art would have motivated to make this modification to reduce the cost of time and energy (Pink, 0005).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: Caruana, “MultiTask Learning” which teaches performing different type of machine learning tasks using common semantic/embedding layer/data with different output model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHIEN MING CHOU whose telephone number is (571)272-9354. The examiner can normally be reached Monday- Friday 9 am - 5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, HITESH PATEL can be reached on 571-270-5442. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHIEN MING CHOU/Examiner, Art Unit 3667
/Hitesh Patel/Supervisory Patent Examiner, Art Unit 3667
6/23/26