Prosecution Insights
Last updated: October 02, 2026
Application No. 18/938,399

SURROUND VIEW VISUALIZATION USING VISION LANGUAGE MODELS

Non-Final OA §103
Filed
Nov 06, 2024
Examiner
PATEL, JITESH
Art Unit
2612
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
324 granted / 411 resolved
+16.8% vs TC avg
Moderate +12% lift
Without
With
+12.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
24 currently pending
Career history
425
Total Applications
across all art units

Statute-Specific Performance

§101
6.5%
-33.5% vs TC avg
§103
62.0%
+22.0% vs TC avg
§102
2.3%
-37.7% vs TC avg
§112
18.5%
-21.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 411 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 9-10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Gopalkrishna et al (US 20240354336 A1) in view of Tao et al (US 20260091792 A1). Regarding claim 1, Gopalkrishna discloses one or more processors comprising processing circuitry (Gopalkrishna [0037], “hardware processor subsystem can include one or more data processing elements (e.g., logic circuits, processing circuits”) to: prompt a vision-language model (VLM) to generate, based at least on image data, one or more responses indicating a selection of at least one environment visualization technique from a plurality of supported environment visualization techniques (Gopalkrishna [0019], “using advanced vision-language models (VLM) responsive to user queries (prompt)”; [0042], “input query … accepting image submissions (query/prompt a vision-language model (VLM), based at least on image data)”; [0076], “a ranked list of relevant concepts can be presented to a user … interface may include visualization tools, such as graphs or heatmaps (exemplary environment visualization techniques) … refinement can be performed in real-time during use by responsiveness to user input (responses indicating a selection of at least one environment visualization technique from a plurality of supported environment visualization techniques)”; [0079], “camera 902 can capture nuanced details for the accurate semantic interpretation of visual content (image data).”); generate a visualization of at least a portion of the environment using the at least one environment visualization technique to process the image data based at least on the selection by the VLM (Gopalkrishna [0042], “input query … accepting image submissions (query/prompt a vision-language model (VLM)”)[0077], “final presentation may involve sophisticated display algorithms that organize the images in a manner that is most conducive to the user's needs, including for example, being grouped by concept relevance (generate a visualization … using the at least one environment visualization technique to process the image data) … specified by the user (based at least on the selection by the VLM)”); and cause presentation of the visualization of at least the portion of the environment on a display (Gopalkrishna [0077], “a final curated set of images, which have been semantically refined and verified through iterative user feedback, can be presented to the user”). Gopalkrishna does not disclose (highlighted missing features) generated using one or more cameras of an ego-machine in an environment of at least the portion of the environment on a display associated with the ego-machine However, Tao discloses (highlighted missing features) using one or more cameras of an ego-machine in an environment (Tao [0045], “At 402, images or image data is received by the system. The images can be generated from one or more image sensors (e.g., cameras) mounted to a vehicle (using one or more cameras of an ego-machine in an environment) as the vehicle is driving through an environment”; [0048], “fig. 4 … ego vehicle”) of at least the portion of the environment on a display associated with the ego-machine (Tao figs. 2, 4; [0033], “a display device 232 … for displaying information to a user (display associated with the ego-machine)”; [0045], “vehicle is driving through an environment”; [0048], “fig. 4 … ego vehicle”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Gopalkrishna with Tao to include functionality for implementing the VLM system in a vehicle. This would have enhanced Gopalkrishna by extending features to an additional application. Regarding claim 9, Gopalkrishna in view of Tao discloses the one or more processors of claim 1, wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations (Gopalkrishna [0053], “applying machine learning and deep learning techniques”); a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more multi-model language models; a system implementing one or more vision language models (VLMs); a system implementing one or more multi-modal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Gopalkrishna [0083], “supporting cloud-based operations”). Claim 10 recites a method which corresponds to the function performed by the one or more processors of claim 1. As such, the mapping and rejection of claim 1 above is considered applicable to the method of claim 10. Claim 17 recites a method which corresponds to the function performed by the one or more processors of claim 9. As such, the mapping and rejection of claim 9 above is considered applicable to the method of claim 17. Claims 2 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Gopalkrishna in view of Tao and further view of Xue et al (US 20240160917 A1). Regarding claim 2, Gopalkrishna in view of Tao discloses the one or more processors of claim 1, but does not disclose wherein at least one of the plurality of supported environment visualization techniques comprises modeling at least a portion of the environment, wherein the selection by the VLM represents a determination to generate the visualization based at least on modeling the environment using a detected 3D surface topology in the environment. However, Xue discloses wherein at least one of the plurality of supported environment visualization techniques comprises modeling at least a portion of the environment, wherein the selection by the VLM represents a determination to generate the visualization based at least on modeling the environment using a detected 3D surface topology in the environment (Xue [0020], “an improved 3D visual recognition model … A vision language model (VLM) … may be used for generating representations of image and text. The features from 3D point cloud may be aligned to the vision/language feature space”; [0047], “each 3D model may represent a 3D object, e.g., using a collection of points in 3D space, connected by various geometric entities such as triangles, lines, curved surfaces, (reads on 3D surface topology in an environment) etc.”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Gopalkrishna further with Xue to model an environment using 3D surface geometric data/topology. This would have been done to accurately represent an environment in a realistic manner. Claim 11 recites a method which corresponds to the function performed by the one or more processors of claim 2. As such, the mapping and rejection of claim 2 above is considered applicable to the method of claim 11. Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Gopalkrishna in view of Tao and further view of Ren et al (US 20260100062 A1). Regarding claim 3, Gopalkrishna in view of Tao discloses the one or more processors of claim 1, but does not disclose wherein at least one of the plurality of supported environment visualization techniques comprises a configuration of one or more parameters for the environment visualization technique, wherein the selection by the VLM represents a determination to configure a size parameter, of the one or more parameters, of one or more voxels of a signed distance function modeling the environment. However, Ren discloses wherein at least one of the plurality of supported environment visualization techniques comprises a configuration of one or more parameters for the environment visualization technique, wherein the selection by the VLM represents a determination to configure a size parameter, of the one or more parameters, of one or more voxels of a signed distance function modeling the environment (Ren [0070], “a first prompt of the VLM 410 generates three different outputs”; [0071], “adopt a 3D voxel-based representation for the map of size L×W×H−W and L (configure a size parameter) expand as the robot explores more areas … volumetric truncated signed distance function (TSDF) fusion is applied to update (1) occupancy of the voxels (voxels of a signed distance function modeling the environment)”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Gopalkrishna further with Ren to use size data of a volumetric truncated signed distance function (TSDF) voxels when mapping an area of an environment. This would have been done to accurate generate and utilize accurate data related to the environment in an error free manner. Claim 12 recites a method which corresponds to the function performed by the one or more processors of claim 3. As such, the mapping and rejection of claim 3 above is considered applicable to the method of claim 12. Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Gopalkrishna in view of Tao and further view of Park et al (US 20250356671 A1). Regarding claim 5, Gopalkrishna in view of Tao discloses the one or more processors of claim 1, but does not disclose wherein the selection by the VLM represents a determination to generate the visualization based at least on modeling the environment using a 3D bowl. However, Park discloses the selection by the VLM represents a determination to generate the visualization based at least on modeling the environment using a 3D bowl (Park [0052], “Large-scale vision-language models”; [0054], “The produced segmentation map 205 from the method 220 labels the ball with multiple personalized semantic labels, including … “bowl,” … In FIG. 2, the method 230 shows the disclosed model (modeling the environment using a 3D bowl)”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Gopalkrishna further with Park to model an environment comprises various objects. This would have been done to accurately represent an environment. Claim 14 recites a method which corresponds to the function performed by the one or more processors of claim 5. As such, the mapping and rejection of claim 5 above is considered applicable to the method of claim 14. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Gopalkrishna in view of Tao and further view of Xue et al (US 20240160917 A1). Regarding claim 8, Gopalkrishna in view of Tao discloses the one or more processors of claim 1, but does not disclose wherein the at least one of the plurality of supported environment visualization techniques comprises a configuration of one or more parameters for the environment visualization technique, wherein the selection by the VLM comprises a configuration that represents at least one of a position or orientation of a virtual camera associated with rendering the visualization. However, Xue discloses the at least one of the plurality of supported environment visualization techniques comprises a configuration of one or more parameters for the environment visualization technique, wherein the selection by the VLM comprises a configuration that represents at least one of a position or orientation of a virtual camera associated with rendering the visualization (Xue [0020], “vision language model (VLM) that is pre-trained … the 3D visual recognition framework to leverage the abundant semantics captured in the vision/language feature spaces”; [0053], “a plurality of image candidates having different viewpoints … virtual camera may be controlled by the image generator (parameters) to provide different perspectives and angles (position/orientation) of the 3D object (a configuration that represents at least one of a position or orientation of a virtual camera associated with rendering the visualization)”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Gopalkrishna further with Xue to utilize a virtual camera position/orientation. This would have been done to provide an ability to view objects from multiple desired perspectives. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Mason et al (US 20250191387 A1) in view of Frazzoli et al (US 20170277195 A1). Regarding claim 18, Mason discloses a system (Mason [0004], “system”) comprising: one or more processors to generate, within a simulation rendered using one or more light transport simulation algorithms, a visualization of at least a portion of a simulated environment using an environment visualization pipeline selected based at least on prompting a vision-language model (VLM) (Mason [0007], “at least one processor”; [0050], “performance of 3D algorithms … efficiently representing three-dimensional space … ray tracing (using one or more light transport simulation algorithms, a visualization of at least a portion of a simulated environment)”; [0115], “a visual language model (VLM)”; [0096], “users and applications can query this spatial data structure 1518 by passing the input language query 1519 or visual query 152 (prompting a VLM)”; [0118], “corresponding to the process of FIG. 22. That is, the same ingestion pipeline (exemplary environment visualization pipeline selected based at least on prompting a vision-language model)”). Mason does not disclose an environment around a simulated ego-machine However, Frazzoli discloses an environment around a simulated ego-machine (Frazzoli [0036], “commands sent to the functional devices of the ego vehicle … The simulator process also contains models of other objects such as other vehicles, cyclists, and pedestrians and can predict their trajectories in a similar way.”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Mason with Frazzoli to include functionality for implementing the system in a simulated vehicle environment. This would have enhanced Mason by extending features to an additional application. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Mason in view of Frazzoli and further view of Yerli (US 20240362518 A1). Regarding claim 19, Mason in view of Frazzoli discloses the system of claim 18, wherein the simulation is generated, at least in part, using one or more content creation applications of a three-dimensional (3D) content collaboration platform for 3D assets. However, Yerli discloses the simulation is generated, at least in part, using one or more content creation applications of a three-dimensional (3D) content collaboration platform for 3D assets (Yerli [0039], “the environment 100 may utilize to collaborate for creating contents … for example, an augmented reality environment, a virtual reality environment, and 2-dimensional (2-D) or 3-dimensional (3-D) simulated environments”; [0044], “one or more users may paint specific areas in a computing environment for creating contents (e.g. …3D curves (3D assets) that can define a specific area for creation or refinement”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Mason further with Yerli to utilize a collaborative content creation system to generate 3D environments. This would have enhanced Mason by allowing multiple users to collaborate and create a variety 3D environments. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Mason in view of Frazzoli and further view of Yerli and further view of Regarding claim 20, Mason in view of Frazzoli and further view of Yerli discloses the system of claim 19, but does not disclose wherein the simulated environment is represented in at least one content creation application of the one or more content creation applications using an OpenUSD format. However, Becke discloses the simulated environment is represented in at least one content creation application of the one or more content creation applications using an OpenUSD format (Becke [0062], “environment may be encapsulated and/or encoded via an interchangeable format such as a universal scene description (USD) format (e.g., OpenUSD)”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Mason further with Becke to generate content by using an OpenUSD format. This would have been done easily create simulated environments by using a standardized creating format. Allowable Subject Matter Claims 4 and 6-7 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 4, none of the prior art or record, alone or in combination, disclose the claim as recited. In particular, none of the prior art or record, alone or in combination, disclose, “the selection by the VLM represents a determination to configure a threshold parameter, of the one or more parameters, that designates one or more threshold distances encoded by a signed distance function modeling the environment” Claim 13 is allowable similar to claim 4 for reciting similar subject matter as claim 4. Regarding claim 6, none of the prior art or record, alone or in combination, disclose the claim as recited. In particular, none of the prior art or record, alone or in combination, disclose, “the selection by the VLM designates one or more parameters of a 3D bowl modeling the environment”. Claim 15 is allowable similar to claim 4 for reciting similar subject matter as claim 6. Regarding claim 7, none of the prior art or record, alone or in combination, disclose the claim as recited. In particular, none of the prior art or record, alone or in combination, disclose, “the selection by the VLM represents a configuration of one or more stitching parameters, of the one or more parameters, associated with stitching overlapping frames of the image data applied to the VLM”. Claim 16 is allowable similar to claim 4 for reciting similar subject matter as claim 7. Conclusion See the notice of references cited (PTO-892) for prior art made of record, including art that is not relied upon but considered pertinent to applicant's disclosure. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JITESH PATEL whose telephone number is (571)270-3313. The examiner can normally be reached 8am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said A. Broome can be reached at (571) 272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JITESH PATEL/Primary Examiner, Art Unit 2612
Read full office action

Prosecution Timeline

Nov 06, 2024
Application Filed
Jun 30, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12710687
ELECTRONIC DEVICE AND IMAGE MAPPING METHOD
2y 6m to grant Granted Aug 18, 2026
Patent 12708445
MOBILE VIRTUAL REALITY SYSTEM FOR SURGICAL ROBOTIC SYSTEMS
2y 0m to grant Granted Aug 18, 2026
Patent 12704936
USER INTERFACE ELEMENTS FOR FACILITATING DIRECT-TOUCH AND INDIRECT HAND INTERACTIONS WITH A USER INTERFACE PRESENTED WITHIN AN ARTIFICIAL-REALITY ENVIRONMENT, AND SYSTEMS AND METHODS OF USE THEREOF
2y 6m to grant Granted Aug 11, 2026
Patent 12693736
AUGMENTED REALITY AND SCREEN IMAGE RENDERING COORDINATION
2y 4m to grant Granted Jul 28, 2026
Patent 12694580
GENERATIVE AI TECHNIQUES FOR ADAPTING STYLE OF A SPACE
2y 1m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
91%
With Interview (+12.3%)
2y 2m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 411 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month