Prosecution Insights
Last updated: October 02, 2026
Application No. 18/949,003

PROMPT-BASED THREE-DIMENSIONAL CONTENT GENERATION USING AN ARTIFICIAL INTELLIGENCE PIPELINE

Non-Final OA §103
Filed
Nov 15, 2024
Examiner
GE, JIN
Art Unit
2619
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
440 granted / 552 resolved
+17.7% vs TC avg
Strong +19% interview lift
Without
With
+18.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
25 currently pending
Career history
572
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
62.0%
+22.0% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 552 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Election/Restrictions Claims 10-20 are withdrawn from further consideration pursuant to 37 CFR 1.142(b) as being drawn to a nonelected Species, there being no allowable generic or linking claim. Election was made without traverse in the reply filed on 07/06/2026. Applicant’s election without traverse of claims1-9 in the reply filed on 07/06/2026 is acknowledged. Response to Amendment This is in response to applicant’s amendment/response filed on 07/06/2026, which has been entered and made of record. Claims 21-31 have been added. Claims 1-9 and 21-31 are pending in the application. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2, 7-9, 21-22, and 27-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2026/0024273 to Tatarchenko et al. in view of U.S. US2026/0004193 to Lee et al. Regarding claim 1, Tatarchenko et al. teach at least one processor (Fig 1, par 0033), comprising: one or more circuits to: generate, based at least on an input text prompt, one or more layout images for a three-dimensional (3D) scene (par 0035-0038, “The at least one text prompt 106 comprises a textual description Y of a three-dimensional layout of the scene. The at least one text prompt comprises a description of a style of the scene”, also par 0090); determine, for at least a portion of objects depicted in the one or more layout images, individual object location information (par 0005-0006, “producing the bounding box, for example, comprises determining a box center of the bounding box depending on the description of the position and determining a box orientation of the bounding box depending on the description of the orientation”, par 0037-0038, “An exemplary description of an exemplary three-dimensional layout of an exemplary scene is: “A double bed is positioned in the middle of the room, a bit nearer to the tip left wall, set a ta right angle. In close proximity of the top left corner, a nightstand is situated set at a right angle. Another nightstand can be found placed near the bottom left corner, also set at a right angel. In the bottom left corner, a wardrobe is positioned, with no particular orientation. Lastly, a shelf is set in the top right corner, with no rotation.””, Figs 2-3, par 0046-0057, “FIG. 2 schematically depicts the exemplary three-dimensional layout 200. The exemplary three-dimensional layout 200 comprises a bounding box 202 for the double bed positioned in the middle of the layout 200. … FIG. 3 schematically depicts an exemplary three-dimensional scene 300 assembled depending on the three-dimensional layout 200. The exemplary three-dimensional scene 300 comprises three-dimensional model 302 for the double bed positioned in the middle of the scene 300. “, also par 0091-0096, “e exemplary description of the layout comprises a description of a position of the objects 502, 504, 508, 512, 514 in the scene 300, 1100 in the two-dimensional perspective. The exemplary description comprises a description of an orientation of the objects 502, 504, 508, 512, 514 in the exemplary scenes 300, 1100 in the two-dimensional perspective”); generate, for at least the portion of objects depicted in the one or more layout images, one or more individual object textual descriptions (par 0037-0038, “An exemplary description of an exemplary three-dimensional layout of an exemplary scene is: “A double bed is positioned in the middle of the room, a bit nearer to the tip left wall, set a ta right angle. In close proximity of the top left corner, a nightstand is situated set at a right angle. Another nightstand can be found placed near the bottom left corner, also set at a right angel. In the bottom left corner, a wardrobe is positioned, with no particular orientation. Lastly, a shelf is set in the top right corner, with no rotation.””, par 0120, “the three-dimensional bounding boxes 202, 204, 208, 212, 214 are generated for the objects 502, 504, 508, 512, 514 in the scene 300, 1100 depending on the exemplary description of the position and the exemplary description of the orientation”, also par 0090-0095); obtain, based at least on the individual object textual descriptions, one or more individual 3D representations for at least the portion of objects (Fig 3, par 0052-0057, “FIG. 3 schematically depicts an exemplary three-dimensional scene 300 assembled depending on the three-dimensional layout 200. The exemplary three-dimensional scene 300 comprises three-dimensional model 302 for the double bed positioned in the middle of the scene 300….”, par 0125-0128, “ Assembling the scene depending on the layout for example comprises retrieving a three-dimensional model of the at least one object from a database that comprises three-dimensional models of objects …. For example, three-dimensional models of the objects 502, 504, 508, 512, 514 are retrieved from the database for assembling the exemplary scene 300”, also claim 5); and generate a representation of the 3D scene based on the one or more individual 3D representations and individual object location information (Fig 3, par 0052-0057, “FIG. 3 schematically depicts an exemplary three-dimensional scene 300 assembled depending on the three-dimensional layout 200. The exemplary three-dimensional scene 300 comprises three-dimensional model 302 for the double bed positioned in the middle of the scene 300….”, par 0137, “Rendering the digital image from the three-dimensional Gaussian Splatting representation may comprise providing a viewpoint, and rendering a view of the scene from the viewpoint”, also see Fig 7, par 0066-0071). But Tatarchenko et al. keep silent for teaching generate, based at least on the individual object textual descriptions, one or more individual 3D representations for at least the portion of objects. In related endeavor, Lee et al. teach generate, based at least on the individual object textual descriptions, one or more individual 3D representations for at least the portion of objects (par 0008, “some embodiments of the present disclosure may aim to provide a method and a system for generating intended high-resolution and precise 3D shape contents only by inputting text through a machine learning model trained to restore 3D contents with a small amount of data processing “, par 0153-0154, “the language model of the computing system 1000 may organize and provide the text prompt generated based on the context of the conversation to be confirmed before the 3D content is generated”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Tatarchenko et al. to include generate, based at least on the individual object textual descriptions, one or more individual 3D representations for at least the portion of objects as taught by Lee et al. to automatically generating high-resolution and detailed 3D contents through a text description input by users in a natural language using machine learning models. Regarding claim 2, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, and Tatarchenko et al. further teach wherein the individual object location information includes at least one of an object size, an object position, or an object pose (par 0005, “ the at least one text prompt may comprise a description of a position of at least one object in the scene in a two-dimensional perspective, and in that the at least one text prompt comprises a description of an orientation of the at least one object in the scene in a two-dimensional perspective, wherein generating the layout comprises producing a three-dimensional bounding box for the at least one object in the scene depending on the description of the position and the description of the orientation”, par 0095, “The exemplary description of the layout comprises a description of a position of the objects 502, 504, 508, 512, 514 in the scene 300, 1100 in the two-dimensional perspective. The exemplary description comprises a description of an orientation of the objects 502, 504, 508, 512, 514 in the exemplary scenes 300, 1100 in the two-dimensional perspective”). Regarding claim 7, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, and Tatarchenko et al. further teach wherein the one or more individual 3D representations for the portion of objects include a plurality of individual 3D representations depicting the individual object from a plurality of perspectives (par 0005, par 0094-0095, “The exemplary description of the layout comprises a description of a position of the objects 502, 504, 508, 512, 514 in the scene 300, 1100 in the two-dimensional perspective. The exemplary description comprises a description of an orientation of the objects 502, 504, 508, 512, 514 in the exemplary scenes 300, 1100 in the two-dimensional perspective.“). Regarding claim 8, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, and Tatarchenko et al. further teach wherein the one or more layout images corresponds to a two-dimensional (2D) representation of the 3D scene (par 0015, “an example embodiment of the present invention may comprise providing three different viewpoints, and determining for the three viewpoints, the synthetic digital image showing the scene from the respective viewpoint. This uses three viewpoints provided by the three-dimensional Gaussian Splatting”, par 0005, par 0094-0095, “The exemplary description of the layout comprises a description of a position of the objects 502, 504, 508, 512, 514 in the scene 300, 1100 in the two-dimensional perspective. The exemplary description comprises a description of an orientation of the objects 502, 504, 508, 512, 514 in the exemplary scenes 300, 1100 in the two-dimensional perspective.“, par 0064-0066, “FIG. 5 schematically depicts a first exemplary digital image 500 comprising a view of a synthetic three-dimensional scene from a first viewpoint rendered from the exemplary three-dimensional Gaussian Splatting representation 400”). Regarding claim 9, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, and further teach wherein the at least one processor is comprised in at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs); a system for performing operations for a conversational AI application; a system for performing operations for a generative AI application; a system for performing operations using a language model; a system for performing one or more operations using a large language model (LLM); a system for performing one or more operations using a vision language model (VLM); a system for performing one or more operations using a multi-modal language model; a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for performing one or more generative content operations using a language model; a system for synthetic data generation; a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources (Tatarchenko et al.: par 0002, par 0102, Lee et al.: abstract, par 0088). Regarding claims 21-22 and 27-28, the method claims 21-22 and 27-28 are similar in scope to claims 1-2 and 7-8 and are rejected under the same rational. Regarding claims 29-30, the system claims 29-30 are similar in scope to claims 1 and 2+7 and are rejected under the same rational. Claim(s) 3, 23, and 31 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2026/0024273 to Tatarchenko et al. in view of U.S. US2026/0004193 to Lee et al., further in view of U.S. PGPubs 2026/0105581 to Kumaravel et al. Regarding claim 3, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, but keep silent for teaching wherein the one or more circuits are further to: generate an asset list for the one or more layout images; and provide the asset list to a visual model used to generate the individual object location information. In related endeavor, Kumaravel et al. teach wherein the one or more circuits are further to: generate an asset list for the one or more layout images (par 0003, “a computer-implemented method that can include receiving a first input image depicting a first object. The method can also include receiving a second input image depicting a second object. The method can also include obtaining a layout specifying a first location of the first object and a second location of the second object”, par 0053, “FIG. 6 shows example pairs of input images having respective objects that are deemed compatible by object compatibility checking 404 of image generation pipeline 400. Input image 602 shows a table 604 and input image 606 shows a sofa 608. The table and sofa can be paired together during object pairing for subsequent generation of an image that includes both the table and the sofa. Input image 610 shows a chair 612 and input image 614 shows a bookshelf 616. The chair and bookshelf can be paired together during object pairing for subsequent generation of an image that includes both the chair and the bookshelf”); and provide the asset list to a visual model used to generate the individual object location information (par 0057, “Layout and canvas generation 408 of image generation pipeline 400 can involve generating layouts for placement of the objects in realistic scenes. The layouts can specify locations and sizes of the respective objects, and can be generated using a number of techniques discussed more below. Once the layouts are determined, two-dimensional canvases can be generated by placing the objects in a two-dimensional image at the locations specified by the respective layouts. FIG. 8 illustrates an example canvas 802 with table 604 and sofa 608, an example canvas 804 with chair 612 and bookshelf 616, and an example canvas 806 with bed 620 and lamp 624”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Tatarchenko et al. as modified by Lee et al. to include wherein the one or more circuits are further to: generate an asset list for the one or more layout images; and provide the asset list to a visual model used to generate the individual object location information as taught by Kumaravel et al. to use a generative image model to generate a new natural like image from existing images of objects through generating the depicted objects in the specified locations in a layout with accurately convey. Regarding claim 23, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 21, the system claim 23 is similar in scope to claim 3 and are rejected under the same rational. Regarding claim 31, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 29, the system claim 31 is similar in scope to claim 3 and are rejected under the same rational. Claim(s) 4 and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2026/0024273 to Tatarchenko et al. in view of U.S. US2026/0004193 to Lee et al., further in view of U.S. PGPubs 2025/0078346 to Couleaud et al.. Regarding claim 4, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, but keep silent for teaching wherein the one or more circuits are further to: generate a cropped image for an individual object of the portion of objects within the one or more layout images; and provide, to a vision language model, at least one of the one or more layout images or the cropped image to generate the individual object textual description. In related endeavor, Couleaud et al. teach wherein the one or more circuits are further to: generate a cropped image for an individual object of the portion of objects within the one or more layout images (par 0042, “As shown in FIG. 1, the image processing system may perform segmentation 114 of image 110 to obtain images 116, 118, and 120 (e.g., different identified portions, segments, and/or objects of image 210 of FIG. 2), which may respectively correspond to images 216, 218, and 220. Images 216, 218, and 220 may each comprise a depiction of a respective object of the plurality of objects 211, 213 and 215 of image 210 ….the image processing system may perform image segmentation 114 (e.g., semantic segmentation and/or instance segmentation) on image 210 to identify, localize, distinguish, and/or extract the different objects, and/or different types or classes of the objects, or portions thereof, of image 210”); and provide, to a vision language model, at least one of the one or more layout images or the cropped image to generate the individual object textual description (par 0050, “as shown in FIG. 1, images 116, 118 and 120 may be input (e.g., sequentially or in parallel) to trained machine learning model 126 (e.g., an image-to-text machine learning model). In some embodiments, machine learning model 126 may comprise a contrastive language-image pretraining (CLIP) model. Model 126 may generate a set of textual prompts 128, 130, and 132 corresponding to (e.g., describing or interpreting) the plurality of images 116, 118 and 120, receptively (which in turn may depict or represent objects 211, 213 and 215 of FIG. 2). In some embodiments, textual prompts 128, 130, and 132 of FIG. 1 may respectively correspond to textual prompts 228, 230 and 232 of FIG. 2”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Tatarchenko et al. as modified by Lee et al. to include wherein the one or more circuits are further to: generate a cropped image for an individual object of the portion of objects within the one or more layout images; and provide, to a vision language model, at least one of the one or more layout images or the cropped image to generate the individual object textual description as taught by Couleaud et al. to provide a guided attention mechanism to selectively attends to different regions of the text in order to generate images that match the textual description more closely. Regarding claim 24, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 21, the system claim 24 is similar in scope to claim 4 and are rejected under the same rational. Claim(s) 5 and 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2026/0024273 to Tatarchenko et al. in view of U.S. US2026/0004193 to Lee et al., further in view of U.S. PGPubs 2024/0236443 to Hings et al. Regarding claim 5, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, but keep silent for teaching wherein the 3D scene is converted to a different file format than at least one format of the one or more individual 3D representations. In related endeavor, Hinds et al. teach wherein the 3D scene is converted to a different file format than at least one format of the one or more individual 3D representations (par 0030, “ cause the at least one processor to parse a scene file to extract scene file data, sending code configured to cause the at least one processor to send the scene file data to a converter, and translating code configured to cause the at least one processor to translate, by the converter and based on a metadata framework, the scene file data from a first scene graph format to a second scene graph format compatible with a renderer interface, and the metadata framework includes an organization of metadata into at least one of systems and subsystems of the systems comprising collections of information common across a plurality of scene graph formats, and the metadata is specified in an Immersive Technologies Media Format”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Tatarchenko et al. as modified by Lee et al. to include wherein the 3D scene is converted to a different file format than at least one format of the one or more individual 3D representations as taught by Hinds et al. to employ media delivery systems and architectures that reformat the media from an input or network “ingest” media format to a distribution media format. Regarding claim 25, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 21, the system claim 25 is similar in scope to claim 5 and are rejected under the same rational. Claim(s) 6 and 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2026/0024273 to Tatarchenko et al. in view of U.S. US2026/0004193 to Lee et al., further in view of U.S. PGPubs 2025/0342628 to Zhao et al. Regarding claim 6, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 1, but keep silent for teaching wherein the one or more circuits are further to: receive the input text prompt; and modify the input text prompt based at least on one or more preferences for the one or more layout images. In related endeavor, Zhang et al. teach wherein the one or more circuits are further to: receive the input text prompt; and modify the input text prompt based at least on one or more preferences for the one or more layout images (Fig 4, par 076-0083, “As further illustrated, the text-to-image editing system 106 includes example input-output pairs 408a-408c in the in-context examples 402. In other words, FIG. 4 illustrates each in-context example of the in-context examples 402 having an input-output pair. In one or more embodiments the text-to-image editing system 106 uses each input-output pair to indicate, to the large language model 404, an example of an output that is to be generated from a corresponding input. In particular, the text-to-image editing system 106 uses each input-output pair to instruct the large language model 404 regarding the formatting or language of the output that is to be used based on the corresponding input. ….. as FIG. 4 illustrates, the prompt 406 instructs the large language model 404 when to generate natural language text output indicating that the whole image (rather than any particular segment of the image) is to be generated. Further, the example input-output pair 408c provides an example in which the example output indicates that the whole image is targeted for modification based on the included example input.). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Tatarchenko et al. as modified by Lee et al. to include wherein the one or more circuits are further to: receive the input text prompt; and modify the input text prompt based at least on one or more preferences for the one or more layout images as taught by Zhang et al. to use a model implementing artificial intelligence to generate a modified version of a digital image having edited content based on a natural language description of the modification. Regarding claim 26, Tatarchenko et al. as modified by Lee et al. teach all the limitation of claim 21, the system claim 26 is similar in scope to claim 6 and are rejected under the same rational. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jin Ge whose telephone number is (571)272-5556. The examiner can normally be reached 8:00 to 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571)272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. JIN . GE Examiner Art Unit 2619 /JIN GE/ Primary Examiner, Art Unit 2619
Read full office action

Prosecution Timeline

Nov 15, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749262
THREE-DIMENSIONAL MODEL GENERATION METHOD, THREE-DIMENSIONAL MODEL GENERATION DEVICE, AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 4m to grant Granted Sep 29, 2026
Patent 12748521
Image Processing Method, Electronic Device, and Non-Transitory Readable Storage Medium
2y 3m to grant Granted Sep 29, 2026
Patent 12743815
ON COMPRESSION OF A MESH WITH MULTIPLE TEXTURE MAPS
2y 1m to grant Granted Sep 22, 2026
Patent 12731176
Method for creating digital art from photos and videos of coins and various methods of presenting the art to be viewed.
3y 8m to grant Granted Sep 08, 2026
Patent 12731352
Video System with Scene-Based Object Insertion Feature
3y 0m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
98%
With Interview (+18.8%)
2y 6m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 552 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month