Prosecution Insights
Last updated: October 02, 2026
Application No. 18/955,615

MULTI-MODAL LARGE LANGUAGE MODEL WITH TOKENIZED OBJECT-LEVEL KNOWLEDGE FOR AUTONOMOUS DRIVING

Non-Final OA §101§102
Filed
Nov 21, 2024
Priority
May 08, 2024 — provisional 63/644,021
Examiner
NGUYEN, THUY-VI THI
Art Unit
3668
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
52%
Grant Probability
Moderate
1-2
OA Rounds
1y 9m
Est. Remaining
63%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
406 granted / 787 resolved
At TC average
Moderate +11% lift
Without
With
+11.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
19 currently pending
Career history
806
Total Applications
across all art units

Statute-Specific Performance

§101
21.2%
-18.8% vs TC avg
§103
35.3%
-4.7% vs TC avg
§102
17.8%
-22.2% vs TC avg
§112
21.8%
-18.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 787 resolved cases

Office Action

§101 §102
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This is in response to Applicant’s communication filed on 5/8/24, wherein: Claims 1-34 are currently pending; Claims 27-28 overcome the prior of record. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 23-28 are directed towards “A machine readable readable medium” having stored thereon a set of instructions. These claims are not found in one of the four statutory categories of invention. The broadest reasonable interpretation of a claim drawn to a computer or machine readable medium typically covers forms of non transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of "computer readable medium" and therefore is directed to non-statutory subject matter. See in In re Nuijten, 500 F.3d 1346, 1356-57. The Examiner suggests in this situation to overcoming this rejection under 35 US.C. § 101, the following approach. A claim drawn to such a computer readable medium that covers both transitory and non-transitory embodiments may be amended to narrow the claim to cover only statutory embodiments to avoid a rejection under 35 US.C. § 101 by adding the limitation "non-transitory" to the preamble of the claim. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-4, 8, 12-15, 19, 23-26, 29-32 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by GOPALKRISHMA ET AL (US 12,511,876). Herein after GOPALKRISHMA. As for claim 1, GOPALKRISHMA discloses a processor comprising: one or more arithmetic logic units (ALUs) configured to perform, using one or more neural networks, motion planning for an autonomous vehicle {see at least figures 1-3, 6, at least col. 1, lines 38-57}, the one or more neural networks comprising: a tokenizer configured to: identify, based on visual data obtained from an environment of the autonomous device, a plurality of objects in the environment; and generate, for the plurality of identified objects, a plurality of object level visual tokens in a latent token embedding space {see at least figures 2-3, tokenizer 211, 213; figures 4-6; col 1, lines 11-30; col. 6, lines 20-60; and at least figure 4-5, col. 10, lines 56-67, col. 11, lines 11-15} an adapter configured to align the object level visual tokens to a text embedding space to generate aligned tokens corresponding to the object level visual tokens {see at least figures 4- 5, and col. 11, lines 15-45} a large language model (LLM) configured to: determine, based on the aligned tokens, critical objects from the plurality of objects identified in the environment; and generate, based on the critical objects, a planning output for the autonomous vehicle {see at least figure 5-6, col. 11, lines 40-53 and claim 1}. As for claim 2, GOPALKRISHMA discloses , wherein the tokenizer is further configured to generate, based on the visual data, a plurality of scene level tokens in the latent token embedding space, each scene level token of the plurality of scene level tokens providing scene information {see at least figures 2, 4-5, at least col. 6, lines 20-60; col. 11, lines 1-30}; wherein the adapter is further configured to align the scene level tokens to the text embedding space to generate aligned tokens corresponding to the scene level tokens; and wherein the aligned tokens comprise the aligned tokens corresponding to the object level visual tokens and the aligned tokens corresponding to the scene level tokens {see at least figures 4-5, col. 11, lines 22-53}. As for claim 3, GOPALKRISHMA discloses wherein the tokenizer is further configured to generate a traffic agent token based on past state history from one or more traffic agents; and wherein generating the planning output for the autonomous vehicle is further based on the aligned tokens and the traffic agent token {see at least figures 6, at least col. 11, lines 1-67, col. 12, lines 1-25}. As for claim 4, GOPALKRISHMA discloses wherein the visual data comprises at least one of: multi-view video frames; or high-definition (HD) maps, wherein the tokenizer is configured to identify the plurality of objects in the environment based on the visual data and symbolic representations that are obtained from the visual data {see at least figures 6, abstract, col. 3, lines 10-33; col. 11, lines 1-67, col. 12, lines 1-25}. As for claim 8, GOPALKRISHMA discloses wherein the tokenizer comprises: a first querying transformer for extracting features of object tracking; a second querying transformer for extracting features of map elements; and a third querying transformer for extracting features of object motion {see at least col. 5, lines 65-67; col. 6, lines 1-29; col. 6, lines 38-61} As for claims 12-15, 19, 23-26, 29-32, the limitations of these claims have been noted in the rejection above. They are therefore rejected for the same reason sets forth above. Claims 5-7, 9-11, 16-18, 20-22, 33-34 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Jang (US 2025/0145174): The method of learning a driving-decision algorithm for a vehicle includes: receiving a dataset related to driving from a database in which a driving scenario is stored; obtaining a first output by inputting the dataset into a first large language model (LLM), wherein the first output includes an inference of the first LLM based on the dataset; based on the first output, training a decision-trainer and a behavior-and-trajectory planner. Li et al (US 2023/0237773): Vision-language models are configured to match an image with a proper caption. Vision-language pre-training (VLP) has been used to improve performance of downstream vision and language tasks by pretraining models on large-scale image-text pairs. Wu (US 2023/0135659): techniques to facilitate financial natural language processing (NLP) training and tasks, such as sentiment analysis, machine reading comprehension, question answering, and causal inferencing. Chaudhury et al (US 2022/0172080): A computer-implemented method is provided for learning multimodal feature matching. The method includes training an image encoder to obtain encoded images. The method further includes training a common classifier on the encoded images by using labeled images. The method also includes training a text encoder while keeping the common classifier in a fixed configuration by using learned text embeddings and corresponding labels for the learned text embeddings. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Kira Nguyen whose telephone number is (571)270-1614. The examiner can normally be reached on Monday to Friday 9:00-5:00 ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Khoi Tran can be reached on 571-272-6919. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KIRA NGUYEN/Primary Examiner, Art Unit 3656
Read full office action

Prosecution Timeline

Nov 21, 2024
Application Filed
Aug 19, 2026
Non-Final Rejection mailed — §101, §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12746692
INFORMATION PROCESSING SYSTEM, INFORMATION PROCESSING METHOD, AND PROGRAM
2y 6m to grant Granted Sep 29, 2026
Patent 12746674
TASK EXECUTION METHOD AND APPARATUS, DEVICE, AND COMPUTER MEDIUM
1y 10m to grant Granted Sep 29, 2026
Patent 12741366
Robot System for Lead-Through Programming
2y 0m to grant Granted Sep 22, 2026
Patent 12722285
MANIPULATOR ROBOT
2y 4m to grant Granted Sep 01, 2026
Patent 12702502
SYSTEMS AND METHODS FOR DYNAMIC ADJUSTMENTS BASED ON LOAD INPUTS FOR ROBOTIC SYSTEMS
2y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
52%
Grant Probability
63%
With Interview (+11.4%)
3y 8m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 787 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month