Prosecution Insights
Last updated: October 02, 2026
Application No. 18/790,535

AUTOMATIC ANNOTATION OF THREE-DIMENSIONAL SHAPE DATA FOR TRAINING TEXT TO 3D GENERATIVE AI SYSTEMS AND APPLICATIONS

Final Rejection §103
Filed
Jul 31, 2024
Examiner
SALVUCCI, MATTHEW D
Art Unit
2613
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Final)
72%
Grant Probability
Favorable
4-5
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
357 granted / 494 resolved
+10.3% vs TC avg
Strong +27% interview lift
Without
With
+27.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
27 currently pending
Career history
512
Total Applications
across all art units

Statute-Specific Performance

§101
4.6%
-35.4% vs TC avg
§103
62.9%
+22.9% vs TC avg
§102
16.1%
-23.9% vs TC avg
§112
14.0%
-26.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 494 resolved cases

Office Action

§103
DETAILED ACTION Status of Claims Applicant's amendments filed on 29 July 2026 have been entered. Claims 1-7, 9-17, and 19 have been amended. Claim 8 has been canceled. Claim 21 has been added. Claims 1-7 and 9-21 are still pending in this application, with claims 1, 10, and 19 being independent. Response to Arguments Applicant's arguments filed 29 have been fully considered but they are not persuasive. With respect to independent claims 1 and 10, Applicant argues for their allowance. It is noted that these claims are in fact allowed. No mention of any rejection of independent claim 19 is presented. This claim stands rejected for the reasons below. Reasons for Allowance Claims 1-7 and 9-18 are allowed. The following is an examiner’s statement of reasons for allowance: Claim 1 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a method comprising: generating third data that associates the one or more shapes with the one or more first captions and the one or more second captions, as presented in the environment of the remaining limitations of claim 1. It is noted that the closest prior art, Menashof et al. (US Pub. 2025/0148753), hereinafter Menashof, shows generating, based at least on first data representative of one or more poses, one or more images to represent one or more shapes associated with the one or more poses from one or more perspectives of one or more canonical viewpoints; generating, based at least on one or more language models processing second data associated with the one or more images: one or more first captions describing the one or more shapes; and one or more second captions describing the one or more shapes. However, Menashof fails to disclose or suggest the one or more first captions associated with one or more first numbers of words that are less than a threshold number of words, the one or more second captions being associated with one or more second numbers of words that are equal to or greater than the threshold number of words; and generating third data that associates the one or more shapes with the one or more first captions and the one or more second captions. Claim 10 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a system comprising: generate third data that associates the shape with the first caption and the second caption, as presented in the environment of the remaining limitations of claim 10. It is noted that the closest prior art, Menashof, shows one or more processors to: generate first data representing a format for generating captions associated with at least a shape; generate second data associated with one or more images of the shape. However, Menashof fails to disclose or suggest the format indicating to generate at least a first caption that includes a first length and a second caption that includes a second length that is greater than the first length; generate, using one or more language models and based at least on the first data and the second data: the first caption that describes the shape using the first length; and the second caption that describes the shape using the second length; and generate third data that associates the shape with the first caption and the second caption. The remaining claims 2-7, 9, and 11-18 depend from one of the above independent claims, either directly or indirectly, and are accordingly allowable. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Menashof et al. (US Pub. 2025/0148753), hereinafter Menashof, in view of Greasley (US Pub. 2021/0397842). Regarding claim 19, Menashof discloses one or more processors (Fig. 7) comprising processing circuitry to: receive first data representative of one or more poses associated with one or more shapes (Paragraph [0024]: the encoder extracts every tenth frame of the video and causes a first generative machine model to generate a descriptor for the extracted frames (e.g., a natural language description of the soccer ball and the location). In this example, the extracted frames are selected based on an interval of time (e.g., every tenth frame), although, as described in greater detail below, other algorithms can be used for selecting frames of the video to be extracted; Paragraph [0071]: the system implementing the method 300 generates descriptors based at least in part on the pivot images and/or video. In one example, an LLM or other machine learning model generates descriptors based at least in part on the pivot images. The descriptors, in an embodiment, include natural language descriptions of the pivot images. In other embodiments, the descriptors include structured data (e.g., a JavaScript Object Notation [JSON] file) that indicates attributes (e.g., location, movement, size, shape, etc.) of objects within the pivot images. For example, the descriptors include text-based data (e.g., data, pseudo-code, source code, etc.) that is provided to a generative model of an encoder to enable the generative model to reconstruct the video. In an embodiment, descriptors are generated for the video and/or portions thereof); generate, based at least on the one or more poses, one or more images to represent the one or more shapes (Paragraph [0005]: the encoder includes a large language model (LLM) or other natural language model to generate descriptors of the pivot images. For example, the descriptors can include natural language descriptions of the objects, backgrounds, interactions, and concepts included in the pivot images; Paragraph [0024]: the encoder extracts every tenth frame of the video and causes a first generative machine model to generate a descriptor for the extracted frames (e.g., a natural language description of the soccer ball and the location). In this example, the extracted frames are selected based on an interval of time (e.g., every tenth frame), although, as described in greater detail below, other algorithms can be used for selecting frames of the video to be extracted; Paragraph [0071]: the system implementing the method 300 generates descriptors based at least in part on the pivot images and/or video. In one example, an LLM or other machine learning model generates descriptors based at least in part on the pivot images. The descriptors, in an embodiment, include natural language descriptions of the pivot images. In other embodiments, the descriptors include structured data (e.g., a JavaScript Object Notation [JSON] file) that indicates attributes (e.g., location, movement, size, shape, etc.) of objects within the pivot images. For example, the descriptors include text-based data (e.g., data, pseudo-code, source code, etc.) that is provided to a generative model of an encoder to enable the generative model to reconstruct the video. In an embodiment, descriptors are generated for the video and/or portions thereof); generate, based at least on one or more language models processing second data associated with one or more images, one or more captions describing the one or more shapes (Paragraph [0005]: the encoder includes a large language model (LLM) or other natural language model to generate descriptors of the pivot images. For example, the descriptors can include natural language descriptions of the objects, backgrounds, interactions, and concepts included in the pivot images; Paragraph [0024]: the encoder extracts every tenth frame of the video and causes a first generative machine model to generate a descriptor for the extracted frames (e.g., a natural language description of the soccer ball and the location). In this example, the extracted frames are selected based on an interval of time (e.g., every tenth frame), although, as described in greater detail below, other algorithms can be used for selecting frames of the video to be extracted; Paragraph [0071]: the system implementing the method 300 generates descriptors based at least in part on the pivot images and/or video. In one example, an LLM or other machine learning model generates descriptors based at least in part on the pivot images. The descriptors, in an embodiment, include natural language descriptions of the pivot images. In other embodiments, the descriptors include structured data (e.g., a JavaScript Object Notation [JSON] file) that indicates attributes (e.g., location, movement, size, shape, etc.) of objects within the pivot images. For example, the descriptors include text-based data (e.g., data, pseudo-code, source code, etc.) that is provided to a generative model of an encoder to enable the generative model to reconstruct the video. In an embodiment, descriptors are generated for the video and/or portions thereof). Menashof does not explicitly disclose the one or more shapes from one or more canonical viewpoints; and generate third data that associates the one or more shapes with the one or more captions. However, Greasley teaches generating a text description for images (Paragraphs [0038]-[0039]), further comprising: the one or more shapes from one or more canonical viewpoints (Paragraph [0014]: the XR system may detect head movement and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. As another example, the XR system may detect movement of the electronic device presenting the XR environment (e.g., a mobile phone, a tablet, a laptop, or the like) and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some situations (e.g., for accessibility reasons), the XR system may adjust characteristic(s) of graphical content in the XR environment in response to representations of physical motions; Paragraph [0026]: the environment is an XR environment, the objects 112 are XR representations of objective-effectuators. In some implementations, the objective-effectuators model characters from fictional materials such as movies, video games, comics, and novels. For example, the character object 112a may represent and model the behavior of a character from a fictional comic, and the character object 112b represents and models the behavior of a character from a fictional video game. In some implementations, the XR environment includes XR representations of objective-effectuators that represent and model the behavior of characters from different fictional materials (e.g., from different movies/games/comics/novels). In various implementations, the objective-effectuators represent and model the behavior of physical elements (e.g., tangible objects). For example, in some implementations, the objective-effectuators model the behavior of equipment (e.g., machinery such as planes, tanks, robots, cars, etc.). In the example of FIG. 1, the robot object 112c represents and models the behavior of a robot, and the drone object 112d represents and models the behavior of a drone); and generate third data that associates the one or more shapes with the one or more captions (Paragraph [0038]: if the user 106 indicates a preference for a short listening experience, the electronic device 102 and/or the controller 104 may generate a terse text description, e.g., with a low word count, for example, less than a threshold number of words. On the other hand, if the user 106 indicates a preference for a longer listening experience, the electronic device 102 and/or the controller 104 may generate a more verbose text description, e.g., with a high word count, for example, greater than a threshold number of words; Paragraph [0039]: if the user 106 is moving at a low velocity or is stationary, the electronic device 102 and/or the controller 104 may generate a text description with a high word count (e.g., greater than a threshold number of words). If the user 106 is moving at a high velocity, the electronic device 102 and/or the controller 104 may generate a text description with a low word count (e.g., less than a threshold number of words) or may increase the speed (e.g., words per minute) of the narration to keep pace with the movement of the user 106; Paragraph [0053]: length of the text description may be determined based on the velocity of the user. For example, if the user is stationary or moving at a low velocity, the audio output generator 224 may generate a verbose text description with a high word count (e.g., greater than a threshold number of words). If the user is moving at a high velocity, the audio output generator 224 may generate a terse text description with a low word count (e.g., less than a threshold number of words). In some implementations, the audio output generator 224 selects a depth to which the ontology 226 is traversed). Greasley teaches that this will allow for user preference of descriptions to be realized (Paragraph [0038]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Menashof with the features of above as taught by Greasley so as to allow for user preference of descriptions to be realized as presented by Greasley. Regarding claim 20, Menashof, in view of Greasley teaches the one or more processors of claim 19, Menashof discloses wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs) (Paragraph [0005]: the encoder includes a large language model (LLM) or other natural language model to generate descriptors of the pivot images); a system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Allowable Subject Matter Claim 21 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim 21 would be allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest one or more processors comprising processing circuitry wherein: the one or more captions include at least the one or more first captions that includes the one or more first lengths and the one or more second captions that includes the one or more second lengths, as presented in the environment of the remaining limitations of claim 19. It is noted that the closest prior art, Menashof, shows obtain fourth data representative of a format for generating captions associated with the one or more shapes. However, Menashof fails to disclose or suggest the format indicating to generate one or more first captions that include one or more first lengths and one or more second captions that include one or more second lengths that is greater than the one or more first lengths, wherein: the one or more captions are further generated based at least on the fourth data; and the one or more captions include at least the one or more first captions that includes the one or more first lengths and the one or more second captions that includes the one or more second lengths. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW D SALVUCCI whose telephone number is (571)270-5748. The examiner can normally be reached M-F: 7:30-4:00PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, XIAO WU can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW SALVUCCI/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Jul 31, 2024
Application Filed
Apr 02, 2026
Non-Final Rejection mailed — §103
Jun 03, 2026
Applicant Interview (Telephonic)
Jun 09, 2026
Non-Final Rejection mailed — §103
Jul 29, 2026
Response Filed
Aug 11, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737976
METHOD FOR CONSTRUCTING STRUCTURAL SEMANTIC MAP UNDER UNDERGROUND WEAK-LIGHT AND LOW-TEXTURE ENVIRONMENT
2y 2m to grant Granted Sep 15, 2026
Patent 12731340
METHOD FOR INTRAOPERATIVE DISPLAY FOR SURGICAL SYSTEMS
4y 6m to grant Granted Sep 08, 2026
Patent 12729975
DISPLAY CONTROL APPARATUS, DISPLAY DEVICE, AND DISPLAY CONTROL METHOD
1y 11m to grant Granted Sep 08, 2026
Patent 12731222
IMAGE PROCESSING DEVICE AND OPERATING METHOD THEREOF
1y 10m to grant Granted Sep 08, 2026
Patent 12711643
TARGET DIGITAL TWIN MODEL GENERATION SYSTEM, CONTROL SYSTEM FOR ROBOT, VIRTUAL SHOP GENERATION SYSTEM, TARGET DIGITAL TWIN MODEL GENERATION METHOD, CONTROL METHOD FOR ROBOT, AND VIRTUAL SHOP GENERATION METHOD
2y 4m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+27.4%)
2y 11m (~9m remaining)
Median Time to Grant
High
PTA Risk
Based on 494 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month