Prosecution Insights
Last updated: October 02, 2026
Application No. 18/991,070

VEHICLE GUIDANCE BY A MULTIMODAL LARGE LANGUAGE MODEL

Final Rejection §103
Filed
Dec 20, 2024
Examiner
PARK, KYLE S
Art Unit
3666
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
Zoox Inc.
OA Round
2 (Final)
66%
Grant Probability
Favorable
3-4
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
101 granted / 154 resolved
+13.6% vs TC avg
Strong +34% interview lift
Without
With
+33.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
9 currently pending
Career history
177
Total Applications
across all art units

Statute-Specific Performance

§101
25.3%
-14.7% vs TC avg
§103
40.2%
+0.2% vs TC avg
§102
7.9%
-32.1% vs TC avg
§112
25.2%
-14.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 154 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims This Final action is in response to the applicant’s amendment/response of April 30, 2026. Claims 1-20 are pending and have been considered as follows. Information Disclosure Statement The information disclosure statement (IDS) submitted on April 30, 2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. The information disclosure statement filed April 30, 2026 fails to comply with 37 CFR 1.98(a)(2), which requires a legible copy of each cited foreign patent document; each non-patent literature publication or that portion which caused it to be listed; and all other information or that portion which caused it to be listed. It has been placed in the application file, but the information referred to therein has not been considered. The legible copy of cited foreign patent documents, “CN117407694” and “JP2024043564” are unavailable. Response to Arguments Applicant’s arguments/amendments with respect to the objection to the claims have been fully considered and are persuasive. Therefore, the objection to the claims has been withdrawn. Applicant’s arguments/amendments with respect to the interpretation of claims under 35 USC §112(f) have been fully considered and are persuasive. Therefore, the interpretation under 35 USC §112(f) as presented in the Office Action of March 19, 2026 has been withdrawn. However, the Examiner notes that under MPEP 2181(I)(A), “device for”/”device configured to” is interpreted under 35 USC §112(f). Therefore, new interpretation under 35 USC §112(f), regarding “vehicle computing device” in claims 1, 6, and 17, is presented below based on the amendments to the claims presented in the Amendment of 30 April 2026. Applicant’s arguments/amendments with respect to the rejection of claims under 35 USC § 101 have been fully considered and are persuasive. Therefore, the rejection of claims under 35 USC § 101 has been withdrawn. Applicant’s arguments/amendments with respect to the rejection of claims under 35 USC § 103 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “vehicle computing device” in claims 1, 6, and 17, described in Applicant’s specification as “The vehicle computing device(s) 404 may include one or more processors 416 and memory 418 communicatively coupled with the one or more processors 416.” [0076]. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 2, 5-8, 11, and 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al., US 2025/0200283 A1, hereinafter referred to as Wu, in view of Buyval et al., US 2025/0236314 A1, hereinafter referred to as Buyval, and further in view of Alsharif, US 2025/0206332 A1, hereinafter referred to as Alsharif, respectively. As to claim 1, Wu teaches a system comprising: one or more processors (see at least Claim 17 regarding one or more processors, Wu); and one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising (see at least paragraph 426 regarding one or more non-transitory computer-readable storage media, Wu): receiving data associated with an autonomous vehicle (see at least paragraph 41 regarding sensor data 104 (or other raw data captured or representative of an environment) is obtained with respect to a specific environment 102. See also at least FIG. 3 and paragraphs 77-85 regarding sensor data from the capture device 202 (along with potentially other observations) is provided to a perception module 302. A perception module 302 in at least one embodiment can perform tasks including those discussed in more detail elsewhere herein, such as to extract features from the sensor data and attempt to identify objects in the environment, as well as to determine relevant information about those objects. … The sensor or perception data can be interpreted and/or correlated in the cross-attention layer(s) of one or more trained models, Wu); retrieving, based at least in part on receiving the request, map data from a database associated with the autonomous vehicle, the map data describing a region of the environment within a threshold distance from the autonomous vehicle (see at least FIG. 3 and paragraphs 81-82 regarding a mapping module 308 can access map data stored to a map repository 310 or other such location. The map repository 310 may be available on the vehicle or accessible over a wireless data connection, for example, where relevant map data can be pre-fetched by the vehicle based on a current and/or anticipated location of the vehicle, such as for a given minimum distance of the vehicle or along a current navigation route. Such an alignment module 312 can attempt to “align” the perception data and the mapping data to provide a more accurate and reliable interpretation of the location and surroundings of the ego vehicle. See also at least paragraphs 347-351 regarding server(s) 1678 may receive, over network(s) 1690 and from vehicles, image data representative of images showing unexpected or changed road conditions, such as recently commenced road-work. In at least one embodiment, server(s) 1678 may transmit, over network(s) 1690 and to vehicles, neural networks 1692, updated or otherwise, and/or map information 1694, including, without limitation, information regarding traffic and road conditions, Wu); inputting the data and the map data into a multimodal large language model (MLLM) (see at least FIG. 3 and paragraphs 77-89 regarding a trained language model 314 can take both perception data and aligned map data as input, Wu); receiving, from the MLLM, text indicating a solution for the event (see at least FIG. 3 and paragraphs 77-101 regarding a language model, or other language-based generative model such as an LLM, can take both aligned map data and detection result data as input, and can produce an updated description of the surrounding environment. In at least one embodiment, this description can be in the format of a tokenized description, or string of text-based tokens, in a domain-specific language, such as RTL. The geolocation of the ego car can also be updated based in part on the updated information. Being able to predict future states of the environment, including dynamic objects, allows for predictions of undesirable actions or states, as may relate to collisions or other operational states that may fall below a minimum safety threshold or otherwise fail to meet at least one requirement of safe operation. The availability and understanding of semantic information and relationships can also help a trained model to make more accurate predictions or inferences than may be possible using other approaches or components. Instead of only sending the tokenized description of the environment to a downstream component to attempt to perform an operation such as collision avoidance, the model can infer or predict a potential collision and send a warning or other indication of the potential undesirable action. This warning may be sent in the same tokenized description, or as a separate tokenized description of higher importance that may be sent to a different system, or directly to a specific system, Wu); and transmitting, by the remote computing device, the solution to the autonomous vehicle (see at least paragraphs 347-351 regarding server(s) 1678 may transmit a signal to vehicle 1600 instructing a fail-safe computer of vehicle 1600 to assume control, notify passengers, and complete a safe parking maneuver, Wu). Wu does not explicitly teach wherein the solution is configured to cause a vehicle computing device of the autonomous vehicle to determine a trajectory for controlling the autonomous vehicle in the environment. However, such matter is taught by Buyval (see at least FIGS. 2-3 and paragraphs 35-37 regarding an alternative driving scenario 300 where pedestrian 302 is near autonomous vehicle 304 equipped with camera 305 is depicted, according to an embodiment. Crosswalk 306 is in front of autonomous vehicle 304 and there is again a possibility that pedestrian 302 will enter the crosswalk. This time pedestrian 302 is in motion towards crosswalk 306. Camera image data records movement of pedestrian 302 and this image data is passed to the LVLM. A prompt is also passed to the LVLM, such as “do you see anything dangerous?” The LVLM outputs descriptions that become inputs for the MPPI controller's description parser. The description parser adjusts the cost model and dynamic model accordingly. Possible trajectories 320-325 are considered by the MPPI controller. As a result of the extra cost added as a result of pedestrian 302, trajectory 320 now becomes the most costly path. Trajectories 321-325 are progressively less costly than 320, but the combination of a pedestrian in a crosswalk is enough to make paths that cross into neighboring lane 320 too expensive to be chosen by the MPPI controller. In this scenario, the lowest cost path is stopping at imaginary line 330 and allowing pedestrian 302 to cross at crosswalk 306. This results because the LVLM has interpreted camera data and indicates that there is a pedestrian who is about to cross the road at a crosswalk, leading to an increase in the track cost of the neighboring lane. Therefore, from the optimizer's perspective, stopping is now the optimal solution. See also at least Claim 1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Buyval which teaches causing a vehicle computing device of the autonomous vehicle to determine a trajectory for controlling the autonomous vehicle in the environment with the system of Wu as both systems are directed to a system and method for receiving the environmental information and determining the guidance for the vehicle based on the environmental information using the language model, and one of ordinary skill in the art would have recognized the established utility of causing a vehicle computing device of the autonomous vehicle to determine a trajectory for controlling the autonomous vehicle in the environment and would have predictably applied it to improve the system of Wu. Wu, as modified by Buyval, does not explicitly teach receiving, by a remote computing device, a request from the autonomous vehicle to assist with an event in an environment. However, such matter is taught by Alsharif (see at least Abstract. See also at least paragraph 27 regarding upon detection of the unexpected issue, the vehicle can generate a question based on the unexpected issue and transmit the question along with sensor data as part of a request for assistance to the remote computing system. In response to receiving the request, the remote computing system can use a VLM to answer the question, which can then be provided to the vehicle for use to overcome the issue. See also at least paragraphs 109-136 regarding a vehicle requesting remote assistance). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Alsharif which teaches receiving, by a remote computing device, a request from the autonomous vehicle to assist with an event in an environment with the system of Wu, as modified by Buyval, as both systems are directed to a system and method for receiving the environmental information and determining the guidance for the vehicle based on the environmental information using the language model, and one of ordinary skill in the art would have recognized the established utility of receiving, by a remote computing device, a request from the autonomous vehicle to assist with an event in an environment and would have predictably applied it to improve the system of Wu as modified by Buyval. As to claim 2, Wu teaches wherein the data comprises a first portion structured as image data, and the operations further comprising: inputting the image data into an image model (see at least FIG. 3 and paragraphs 77-82 regarding sensor data from the capture device 202 (along with potentially other observations) is provided to a perception module 302, Wu); inputting the map data into a scene model (see at least FIG. 3 and paragraphs 77-82 regarding at least a selected portion of the perception data from the perception module 302, and the geolocation and/or local map data from the mapping module 308, can be provided as input to an alignment module 312, Wu); receiving first output data from the image model (see at least FIG. 3 and paragraphs 77-82 regarding a perception module 302 in at least one embodiment can perform tasks including those discussed in more detail elsewhere herein, such as to extract features from the sensor data and attempt to identify objects in the environment, as well as to determine relevant information about those objects. Feature extraction or feature inference may be performed by an encoder in at least one embodiment to extract and encode features that may be relatively low-level and may not have a clear sematic meaning attached. The features may be used to generate a relatively universal and/or generic representation of the sensor data. The sensor or perception data can be interpreted and/or correlated in the cross-attention layer(s) of one or more trained models, Wu); receiving second output data from the scene model (see at least FIG. 3 and paragraphs 77-82 regarding an alignment module 312 (and similarly a localization module as mentioned above) can be a dedicated or stand-alone component or process in at least one embodiment, or can be part of a system component that also supports other functionalities, among other such options. Such an alignment module 312 can attempt to “align” the perception data and the mapping data to provide a more accurate and reliable interpretation of the location and surroundings of the ego vehicle, Wu); and inputting, as input data, the first output data and the second output data into the MLLM (see at least FIG. 3 and paragraphs 77-89 regarding a trained language model 314 can take both perception data and aligned map data as input, Wu). As to claim 5, Wu teaches wherein the vehicle computing device having first computational resources (see at least FIG. 3 and paragraphs 267-270 regarding vehicle 1600 may include any number of SoCs 1604. In at least one embodiment, each of SoCs 1604 may include, without limitation, central processing units (“CPU(s)”) 1606, graphics processing units (“GPU(s)”) 1608, processor(s) 1610, cache(s) 1612, accelerator(s) 1614, data store(s) 1616, and/or other components and features not illustrated, Wu), and the MLLM utilizes second computational resources of the remote computing device, the second computational resources greater than the first computational resources (see at least FIG. 3 and paragraphs 81 and 267-270 regarding the map repository 310 may be available on the vehicle or accessible over a wireless data connection. SoC(s) 1604 may be combined in a system (e.g., system of vehicle 1600) with a High Definition (“HD”) map 1622 which may obtain map refreshes and/or updates via network interface 1624 from one or more servers (not shown in FIG. 16C), Wu). As to claim 6, Wu teaches one or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations comprising (see at least paragraph 426 regarding one or more non-transitory computer-readable storage media, Wu): inputting first data associated with a first data format and second data associated with a second data format into a large language model (LLM), the first data associated with a first source and the second data associated with a second source different from the first source (see at least FIG. 3 and paragraphs 77-89 regarding sensor data from the capture device 202 (along with potentially other observations) is provided to a perception module 302. A perception module 302 in at least one embodiment can perform tasks including those discussed in more detail elsewhere herein, such as to extract features from the sensor data and attempt to identify objects in the environment, as well as to determine relevant information about those objects. … The sensor or perception data can be interpreted and/or correlated in the cross-attention layer(s) of one or more trained models. a mapping module 308 can access map data stored to a map repository 310 or other such location. The map repository 310 may be available on the vehicle or accessible over a wireless data connection, for example, where relevant map data can be pre-fetched by the vehicle based on a current and/or anticipated location of the vehicle, such as for a given minimum distance of the vehicle or along a current navigation route. Such an alignment module 312 can attempt to “align” the perception data and the mapping data to provide a more accurate and reliable interpretation of the location and surroundings of the ego vehicle. A trained language model 314 can take both perception data and aligned map data as input, Wu); receiving, from the LLM and based at least in part on the first data and the second data, a solution for the vehicle relative to the event (see at least FIG. 3 and paragraphs 77-101 regarding a language model, or other language-based generative model such as an LLM, can take both aligned map data and detection result data as input, and can produce an updated description of the surrounding environment. In at least one embodiment, this description can be in the format of a tokenized description, or string of text-based tokens, in a domain-specific language, such as RTL. The geolocation of the ego car can also be updated based in part on the updated information. Being able to predict future states of the environment, including dynamic objects, allows for predictions of undesirable actions or states, as may relate to collisions or other operational states that may fall below a minimum safety threshold or otherwise fail to meet at least one requirement of safe operation. The availability and understanding of semantic information and relationships can also help a trained model to make more accurate predictions or inferences than may be possible using other approaches or components. Instead of only sending the tokenized description of the environment to a downstream component to attempt to perform an operation such as collision avoidance, the model can infer or predict a potential collision and send a warning or other indication of the potential undesirable action. This warning may be sent in the same tokenized description, or as a separate tokenized description of higher importance that may be sent to a different system, or directly to a specific system, Wu); and transmitting, by the remote computing device, the solution to the vehicle (see at least paragraphs 347-351 regarding server(s) 1678 may transmit a signal to vehicle 1600 instructing a fail-safe computer of vehicle 1600 to assume control, notify passengers, and complete a safe parking maneuver, Wu). Wu does not explicitly teach the solution configured to cause a vehicle computing device of the vehicle to determine a trajectory for controlling the vehicle in the environment. However, such matter is taught by Buyval (see at least FIGS. 2-3 and paragraphs 35-37 regarding an alternative driving scenario 300 where pedestrian 302 is near autonomous vehicle 304 equipped with camera 305 is depicted, according to an embodiment. Crosswalk 306 is in front of autonomous vehicle 304 and there is again a possibility that pedestrian 302 will enter the crosswalk. This time pedestrian 302 is in motion towards crosswalk 306. Camera image data records movement of pedestrian 302 and this image data is passed to the LVLM. A prompt is also passed to the LVLM, such as “do you see anything dangerous?” The LVLM outputs descriptions that become inputs for the MPPI controller's description parser. The description parser adjusts the cost model and dynamic model accordingly. Possible trajectories 320-325 are considered by the MPPI controller. As a result of the extra cost added as a result of pedestrian 302, trajectory 320 now becomes the most costly path. Trajectories 321-325 are progressively less costly than 320, but the combination of a pedestrian in a crosswalk is enough to make paths that cross into neighboring lane 320 too expensive to be chosen by the MPPI controller. In this scenario, the lowest cost path is stopping at imaginary line 330 and allowing pedestrian 302 to cross at crosswalk 306. This results because the LVLM has interpreted camera data and indicates that there is a pedestrian who is about to cross the road at a crosswalk, leading to an increase in the track cost of the neighboring lane. Therefore, from the optimizer's perspective, stopping is now the optimal solution. See also at least Claim 1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Buyval which teaches causing a vehicle computing device of the vehicle to determine a trajectory for controlling the vehicle in the environment with the system of Wu as both systems are directed to a system and method for receiving the environmental information and determining the guidance for the vehicle based on the environmental information using the language model, and one of ordinary skill in the art would have recognized the established utility of causing a vehicle computing device of the vehicle to determine a trajectory for controlling the vehicle in the environment and would have predictably applied it to improve the system of Wu. Wu, as modified by Buyval, does not explicitly teach receiving, by a remote computing device, a request for assistance from a vehicle indicating an event in an environment. However, such matter is taught by Alsharif (see at least Abstract. See also at least paragraph 27 regarding upon detection of the unexpected issue, the vehicle can generate a question based on the unexpected issue and transmit the question along with sensor data as part of a request for assistance to the remote computing system. In response to receiving the request, the remote computing system can use a VLM to answer the question, which can then be provided to the vehicle for use to overcome the issue. See also at least paragraphs 109-136 regarding a vehicle requesting remote assistance). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Alsharif which teaches receiving, by a remote computing device, a request for assistance from a vehicle indicating an event in an environment with the system of Wu, as modified by Buyval, as both systems are directed to a system and method for receiving the environmental information and determining the guidance for the vehicle based on the environmental information using the language model, and one of ordinary skill in the art would have recognized the established utility of receiving, by a remote computing device, a request for assistance from a vehicle indicating an event in an environment and would have predictably applied it to improve the system of Wu as modified by Buyval. As to claim 7, Wu teaches wherein the first data comprises sensor data associated with a sensor of the vehicle and the second data comprises map data (see at least FIG. 3 and paragraphs 77-89 regarding sensor data from the capture device 202 (along with potentially other observations) is provided to a perception module 302. A perception module 302 in at least one embodiment can perform tasks including those discussed in more detail elsewhere herein, such as to extract features from the sensor data and attempt to identify objects in the environment, as well as to determine relevant information about those objects. … The sensor or perception data can be interpreted and/or correlated in the cross-attention layer(s) of one or more trained models. a mapping module 308 can access map data stored to a map repository 310 or other such location. The map repository 310 may be available on the vehicle or accessible over a wireless data connection, for example, where relevant map data can be pre-fetched by the vehicle based on a current and/or anticipated location of the vehicle, such as for a given minimum distance of the vehicle or along a current navigation route. Such an alignment module 312 can attempt to “align” the perception data and the mapping data to provide a more accurate and reliable interpretation of the location and surroundings of the ego vehicle, Wu). As to claim 8, Wu teaches inputting the first data into a first machine learned model and the second data into a second machine learned model different from the first machine learned model (see at least FIG. 3 and paragraphs 77-82 regarding sensor data from the capture device 202 (along with potentially other observations) is provided to a perception module 302. At least a selected portion of the perception data from the perception module 302, and the geolocation and/or local map data from the mapping module 308, can be provided as input to an alignment module 312, Wu); receiving first output data from the first machine learned model and second output data from the second machine learned model (see at least FIG. 3 and paragraphs 77-82 regarding a perception module 302 in at least one embodiment can perform tasks including those discussed in more detail elsewhere herein, such as to extract features from the sensor data and attempt to identify objects in the environment, as well as to determine relevant information about those objects. Feature extraction or feature inference may be performed by an encoder in at least one embodiment to extract and encode features that may be relatively low-level and may not have a clear sematic meaning attached. The features may be used to generate a relatively universal and/or generic representation of the sensor data. The sensor or perception data can be interpreted and/or correlated in the cross-attention layer(s) of one or more trained models. An alignment module 312 (and similarly a localization module as mentioned above) can be a dedicated or stand-alone component or process in at least one embodiment, or can be part of a system component that also supports other functionalities, among other such options. Such an alignment module 312 can attempt to “align” the perception data and the mapping data to provide a more accurate and reliable interpretation of the location and surroundings of the ego vehicle, Wu); and inputting, as input data, the first output data and the second output data into the LLM (see at least FIG. 3 and paragraphs 77-89 regarding a trained language model 314 can take both perception data and aligned map data as input, Wu). As to claim 11, Examiner notes claim 11 recites similar limitations to claim 5 and is rejected under the same rational. As to claim 16, Wu does not explicitly teach wherein the solution represents a token or a waypoint for the vehicle to navigate relative to the event. However, such matter is taught by Buyval (see at least FIGS. 2-3 and paragraphs 35-37 regarding an alternative driving scenario 300 where pedestrian 302 is near autonomous vehicle 304 equipped with camera 305 is depicted, according to an embodiment. Crosswalk 306 is in front of autonomous vehicle 304 and there is again a possibility that pedestrian 302 will enter the crosswalk. This time pedestrian 302 is in motion towards crosswalk 306. Camera image data records movement of pedestrian 302 and this image data is passed to the LVLM. A prompt is also passed to the LVLM, such as “do you see anything dangerous?” The LVLM outputs descriptions that become inputs for the MPPI controller's description parser. The description parser adjusts the cost model and dynamic model accordingly. Possible trajectories 320-325 are considered by the MPPI controller. As a result of the extra cost added as a result of pedestrian 302, trajectory 320 now becomes the most costly path. Trajectories 321-325 are progressively less costly than 320, but the combination of a pedestrian in a crosswalk is enough to make paths that cross into neighboring lane 320 too expensive to be chosen by the MPPI controller. In this scenario, the lowest cost path is stopping at imaginary line 330 and allowing pedestrian 302 to cross at crosswalk 306. This results because the LVLM has interpreted camera data and indicates that there is a pedestrian who is about to cross the road at a crosswalk, leading to an increase in the track cost of the neighboring lane. Therefore, from the optimizer's perspective, stopping is now the optimal solution. See also at least Claim 1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Buyval which teaches wherein the solution represents a token or a waypoint for the vehicle to navigate relative to the event with the system of Wu as both systems are directed to a system and method for receiving the environmental information and determining the guidance for the vehicle based on the environmental information using the language model, and one of ordinary skill in the art would have recognized the established utility of having wherein the solution represents a token or a waypoint for the vehicle to navigate relative to the event and would have predictably applied it to improve the system of Wu. As to claim 17, Examiner notes claim 17 recites similar limitations to claim 6 and is rejected under the same rational. As to claim 18, Examiner notes claim 18 recites similar limitations to claim 8 and is rejected under the same rational. Claim(s) 3, 4, 9, 10, 14, 15, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al., US 2025/0200283 A1, hereinafter referred to as Wu, in view of Buyval et al., US 2025/0236314 A1, hereinafter referred to as Buyval, in view of Alsharif, US 2025/0206332 A1, hereinafter referred to as Alsharif, and further in view of Yin et al., US 2025/0342346 A1, hereinafter referred to as Yin, respectively. As to claim 3, Wu, as modified by Buyval and Alsharif, does not explicitly teach inputting the first output data into a first projector; receiving, from the first projector, a first common representation; inputting the second output data into a second projector; receiving from the second projector, a second common representation; or inputting the first common representation and the second common representation into the MLLM. However, Yin teaches inputting the first output data into a first projector (see at least FIG. 1B and paragraphs 57-58 regarding the multi-modal inputs 102 includes one or more of text 102A, image 102B, video 102C, or audio 102D. … The image token sequence 114B is projected into a textual embedding space corresponding to the tokenizer 110A through a projector 112B. That is, the projector 112B maps image representations (e.g., output from the image encoder 110B) into an embedding space compatible with textual representations (e.g., output from the tokenizer 110A)); receiving, from the first projector, a first common representation (see at least FIG. 1B and paragraphs 57-58); inputting the second output data into a second projector (see at least FIG. 1B and paragraphs 57-58 regarding the multi-modal inputs 102 includes one or more of text 102A, image 102B, video 102C, or audio 102D. … The image token sequence 114B is projected into a textual embedding space corresponding to the tokenizer 110A through a projector 112B. That is, the projector 112B maps image representations (e.g., output from the image encoder 110B) into an embedding space compatible with textual representations (e.g., output from the tokenizer 110A). Similarly, the video encoder 110C encodes an input video 102C to generate a video token sequence 114C. The video token sequence 114C includes one or more video tokens. The video token sequence 114C is projected into the textual embedding space corresponding to the tokenizer 110A through a projector 112C); receiving from the second projector, a second common representation (see at least FIG. 1B and paragraphs 57-58); and inputting the first common representation and the second common representation into the MLLM (see at least FIG. 1B and paragraphs 57-58 regarding a textual embedding input is formed based on the tokens aligned in the textual embedding space, which serves as input to the LLM 120. The LLM 120 then generates a corresponding textual embedding output). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Yin which teaches inputting the first output data into a first projector; receiving, from the first projector, a first common representation; inputting the second output data into a second projector; receiving from the second projector, a second common representation; and inputting the first common representation and the second common representation into the MLLM with the system of Wu, as modified by Buyval and Alsharif, as both systems are directed to a system and method for controlling execution of a task based on analyzed input information using the language model, and one of ordinary skill in the art would have recognized the established utility of inputting the first output data into a first projector; receiving, from the first projector, a first common representation; inputting the second output data into a second projector; receiving from the second projector, a second common representation; and inputting the first common representation and the second common representation into the MLLM and would have predictably applied it to improve the system of Wu as modified by Buyval and Alsharif. As to claim 4, Wu, as modified by Buyval and Alsharif, does not explicitly teach receiving input text indicating a condition for the MLLM to consider during processing; or inputting, into the MLLM, the input text. However, Yin teaches receiving input text indicating a condition for the MLLM to consider during processing (see at least FIG. 1B and paragraphs 42 and 57-58 regarding a tokenizer configured to encode text input to generate a text token sequence that includes one or more second textual tokens in the textual embedding space. The tokenizer 110A encodes an input text 102A to generate a text token sequence 114A. The text token sequence 114A includes one or more text tokens. For example, the text token sequence 114A is represented by a high-dimensional embedding consisting of one or more vectors, with each vector corresponding to a text token); and inputting, into the MLLM, the input text (see at least FIG. 1B and paragraphs 57-58). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Yin which teaches receiving input text indicating a condition for the MLLM to consider during processing; and inputting, into the MLLM, the input text with the system of Wu, as modified by Buyval and Alsharif, as both systems are directed to a system and method for controlling execution of a task based on analyzed input information using the language model, and one of ordinary skill in the art would have recognized the established utility of receiving input text indicating a condition for the MLLM to consider during processing; and inputting, into the MLLM, the input text and would have predictably applied it to improve the system of Wu as modified by Buyval and Alsharif. As to claim 9, Examiner notes claim 9 recites similar limitations to claim 3 and is rejected under the same rational. As to claim 10, Examiner notes claim 10 recites similar limitations to claim 4 and is rejected under the same rational. As to claim 14, Wu, as modified by Buyval and Alsharif, does not explicitly teach receiving text associated with a user input from a user interface; determining a token to represent the text; or inputting the token into the LLM. However, Yin teaches receiving text associated with a user input from a user interface (see at least paragraphs 48-49 regarding the multi-modal inputs 102 include inputs of various modalities, such as, text, image, video, and audio. In at least one embodiment, the multi-modal inputs 102 include user inputs, outputs from the system 100 based on previous multi-modal inputs 102, or a combination thereof. See also at least FIG. 2A and paragraphs 80-81. See also at least paragraphs 57-59); determining a token to represent the text (see at least paragraphs 48-49. See also at least FIG. 1B and paragraphs 57-59 regarding the tokenizer 110A encodes an input text 102A to generate a text token sequence 114A. See also at least FIG. 2A and paragraphs 80-81); and inputting the token into the LLM (see at least paragraphs 48-49. See also at least FIG. 1B and paragraphs 57-59 regarding a textual embedding input is formed based on the tokens aligned in the textual embedding space, which serves as input to the LLM 120. A text output 104A is generated based on one or more generation text token 124A. See also at least FIG. 2A and paragraphs 80-81). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Yin which teaches receiving text associated with a user input from a user interface; determining a token to represent the text; and inputting the token into the LLM with the system of Wu, as modified by Buyval and Alsharif, as both systems are directed to a system and method for controlling execution of a task based on analyzed input information using the language model, and one of ordinary skill in the art would have recognized the established utility of receiving text associated with a user input from a user interface; determining a token to represent the text; and inputting the token into the LLM and would have predictably applied it to improve the system of Wu as modified by Buyval and Alsharif. As to claim 15, Wu, as modified by Buyval and Alsharif, does not explicitly teach wherein the first data in the first data format is received from a tokenizer, the second data in the second data format is received from a multilayer perceptron. However, Yin teaches wherein the first data in the first data format is received from a tokenizer (see at least FIG. 1B and paragraphs 57-58 regarding the tokenizer 110A encodes an input text 102A to generate a text token sequence 114A), the second data in the second data format is received from a multilayer perceptron (see at least FIG. 1B and paragraphs 57-58. See also at least paragraphs 160-167 regarding A deep neural network (DNN) model includes multiple layers of many connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained with enormous amounts of input data to quickly solve complex problems with high accuracy. In one example, a first layer of the DNN model breaks down an input image of an automobile into various sections and looks for basic patterns such as lines and angles. The second layer assembles the lines to look for higher level patterns such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the final few layers generate a label for the input image, identifying the model of a specific automobile brand. Images generated applying one or more of the techniques disclosed herein may be used to train, test, or certify DNNs used to recognize objects and environments in the real world. Such images may include scenes of roadways, factories, buildings, urban settings, rural settings, humans, animals, and any other physical object or real-world setting. Such images may be used to train, test, or certify DNNs that are employed in machines or robots to manipulate, handle, or modify physical objects in the real world). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Yin which teaches wherein the first data in the first data format is received from a tokenizer, the second data in the second data format is received from a multilayer perceptron with the system of Wu, as modified by Buyval and Alsharif, as both systems are directed to a system and method for controlling execution of a task based on analyzed input information using the language model, and one of ordinary skill in the art would have recognized the established utility of having wherein the first data in the first data format is received from a tokenizer, the second data in the second data format is received from a multilayer perceptron and would have predictably applied it to improve the system of Wu as modified by Buyval and Alsharif. As to claim 19, Examiner notes claim 19 recites similar limitations to claim 3 and is rejected under the same rational. As to claim 20, Examiner notes claim 20 recites similar limitations to claim 4 and is rejected under the same rational. Claim(s) 12 is rejected under 35 U.S.C. 103 as being unpatentable over Wu et al., US 2025/0200283 A1, hereinafter referred to as Wu, in view of Buyval et al., US 2025/0236314 A1, hereinafter referred to as Buyval, in view of Alsharif, US 2025/0206332 A1, hereinafter referred to as Alsharif, and further in view of Gupta et al., US 2024/0411808 A1, hereinafter referred to as Gupta, respectively. As to claim 12, Wu, as modified by Buyval and Alsharif, does not explicitly teach transmitting the solution to a user interface associated with an operator; receiving, from the user interface, an input comprising a suggested command for the vehicle to execute; or transmitting the input to the vehicle. However, Gupta teaches transmitting the solution to a user interface associated with an operator (see at least paragraphs 46-49 regarding a user provides a voice prompt to the voice assistance system, by saying “Hey Mercedes, suggest a national park to drive to from Sunnyvale”. The voice assistance system or user prompt processing system may generate a prompt message (e.g., a digital message) that is based on the voice prompt, and communicate the prompt message to generative AI system. The generative AI system may generate an AI system response that identifies some of the most famous or visited national parks around Sunnyvale (e.g., “Yosemite National Park is a great choice for a drive from Sunnyvale. It's about a three-hour drive, and you'll get to see some of the most stunning scenery in the United States.”). The generative AI system may send the AI system response to the user prompt processing system or voice assistance system. The voice assistance system may outputs a voice response that reads out aloud. This may a response such as, for example: “Yosemite National Park is a great choice for a drive from Sunnyvale. It's about a three-hour drive, and you'll get to see some of the most stunning scenery in the United States.”); receiving, from the user interface, an input comprising a suggested command for the vehicle to execute; and transmitting the input to the vehicle (see at least paragraphs 46-49 regarding the head unit of the vehicle may display Yosemite National Park as Point of Interest in the head unit of the vehicle. The vehicle may start route guidance to Yosemite National Park if the user taps (or otherwise selects) on Yosemite National Park). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Gupta which teaches transmitting the solution to a user interface associated with an operator; receiving, from the user interface, an input comprising a suggested command for the vehicle to execute; and transmitting the input to the vehicle with the system of Wu, as modified by Buyval and Alsharif, as both systems are directed to a system and method for controlling execution of a task based on analyzed input information using the language model, and one of ordinary skill in the art would have recognized the established utility of transmitting the solution to a user interface associated with an operator; receiving, from the user interface, an input comprising a suggested command for the vehicle to execute; and transmitting the input to the vehicle and would have predictably applied it to improve the system of Wu as modified by Buyval and Alsharif. Claim(s) 13 is rejected under 35 U.S.C. 103 as being unpatentable over Wu et al., US 2025/0200283 A1, hereinafter referred to as Wu, in view of Buyval et al., US 2025/0236314 A1, hereinafter referred to as Buyval, in view of Alsharif, US 2025/0206332 A1, hereinafter referred to as Alsharif, and further in view of Zahid et al., US 2025/0296586 A1, hereinafter referred to as Zahid, respectively. As to claim 13, Wu, as modified by Buyval and Alsharif, does not explicitly teach wherein the LLM is trained based at least in part on log data received from an additional vehicle and associated solution data associated with an operator. However, such matter is taught by Zahid (see at least paragraph 110 regarding training an artificial intelligence (AI) model based on sensor data received from a plurality of vehicles and actions performed by the plurality of vehicles while travelling along a route in 244E. See also at least paragraph 245 regarding the AI model 1010 may be a generative AI model such as a large language model (LLM), a multi-modal LLM, a transformative neural network, or the like, which is capable of creating custom images of a road such as a road 1020 based on image data captured by vehicles that travel along the road 1020). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the system of Zahid which teaches wherein the LLM is trained based at least in part on log data received from an additional vehicle and associated solution data associated with an operator with the system of Wu, as modified by Buyval and Alsharif, as both systems are directed to a system and method for controlling execution of a task based on analyzed input information using the language model, and one of ordinary skill in the art would have recognized the established utility of having wherein the LLM is trained based at least in part on log data received from an additional vehicle and associated solution data associated with an operator and would have predictably applied it to improve the system of Wu as modified by Buyval and Alsharif. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLE S. PARK whose telephone number is (571)272-3151. The examiner can normally be reached Mon-Thurs 9:00AM-5:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Anne M ANTONUCCI can be reached at (313)446-6519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.S.P./Examiner, Art Unit 3666 /ANNE MARIE ANTONUCCI/Supervisory Patent Examiner, Art Unit 3666
Read full office action

Prosecution Timeline

Dec 20, 2024
Application Filed
Mar 19, 2026
Non-Final Rejection mailed — §103
Apr 08, 2026
Examiner Interview Summary
Apr 08, 2026
Applicant Interview (Telephonic)
Apr 30, 2026
Response Filed
Aug 17, 2026
Final Rejection mailed — §103
Sep 22, 2026
Applicant Interview (Telephonic)
Sep 22, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12736368
HIGH DEFINITION MAPPING FOR AUTONOMOUS SYSTEMS AND APPLICATIONS
4y 4m to grant Granted Sep 15, 2026
Patent 12735032
VEHICLE CONTROL DEVICE, VEHICLE CONTROL METHOD, AND STORAGE MEDIUM
1y 7m to grant Granted Sep 15, 2026
Patent 12729021
AERONAUTICAL OBSTACLE KNOWLEDGE NETWORK
3y 11m to grant Granted Sep 08, 2026
Patent 12718637
VEHICLE DIAGNOSTIC METHOD BASED ON MODE $06 AND IN-USE MONITOR PERFORMANCE RATIO DATA
2y 2m to grant Granted Aug 25, 2026
Patent 12694723
INFORMATION COLLECTION SYSTEM AND INFORMATION COLLECTION METHOD
2y 10m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
66%
Grant Probability
99%
With Interview (+33.5%)
2y 8m (~11m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 154 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month