Prosecution Insights
Last updated: October 01, 2026
Application No. 18/634,134

SYNTHETIC DATA GENERATION USING VIEWPOINT AUGMENTATION FOR AUTONOMOUS SYSTEMS AND APPLICATIONS

Final Rejection §103
Filed
Apr 12, 2024
Priority
Apr 14, 2023 — provisional 63/459,355
Examiner
SHIN, SOO JUNG
Art Unit
2667
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
87%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
547 granted / 628 resolved
+25.1% vs TC avg
Strong +16% interview lift
Without
With
+16.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
27 currently pending
Career history
646
Total Applications
across all art units

Statute-Specific Performance

§101
8.3%
-31.7% vs TC avg
§103
38.3%
-1.7% vs TC avg
§102
18.3%
-21.7% vs TC avg
§112
25.7%
-14.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 628 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Response to Amendment The amendment filed on August 7, 2026 has been entered. The amendment of claims 1, 2, 10, 12, 14, 17, and 18 and cancellation of claim 7 have been acknowledged. In view of the amendment, the 35 U.S.C. 102(a)(1) and Double Patenting rejections have been withdrawn. Response to Arguments Applicant’s arguments filed on August 7, 2026, with respect to the amended claims, have been fully considered but are moot because the arguments rely on newly added and/or amended claim limitations. The examiner has revised the rejections to match the new claim limitations. Claim Rejections - 35 USC § 103 Claim(s) 1-4, 8-14, and 16-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (“3D Hierarchical Refinement and Augmentation for Unsupervised Learning of Depth and Pose From Monocular Video,” IEEE Transactions on Circuits and Systems for Video Technology, Vol. 33, No. 4, April 2023, published 19 October 2022), in view of Duluk Jr. et al. (US 2004/0130552 A1), hereinafter referred to as Wang and Duluk Jr., respectively. Regarding claim 1, Wang teaches a processor comprising: one or more circuits (Wang pg. 1782: “All the experiments for 2 PoseNets are implemented on a NVIDIA RTX 2080Ti GPU … implemented on NVIDIA RTX 3090 GPU”) to: generate, using a machine-learning model and based at least on a first image of a set of sequential images corresponding to a first viewpoint, a first transformed image corresponding to a second viewpoint (Wang Fig. 1: “the process of pose estimation refinement … given two adjacent frames Xt and Xt+1, with their corresponding camera view Pt and Pt+1, we utilize our 3D hierarchical refinement method in Sec. III-B to synthesize transitional camera view Pmwarp”; Wang pg. 1778-1779, §III-B: “Our method has multiple PoseNets, NP1, NP2, …, NPM, with the same network structure to refine the pose estimation … we use the image warping based on 2D-3D transformation to construct the intermediate image between Xt and Xt+1 to make next residual pose estimation easier”; Wang pg. 1782 left column: “11 long driving stereo sequences with available ground truth trajectories … using sequences 00-08 for training and sequences 09-10 for testing”; Wang pg. 1782 right column: “A snippet of two sequential images is used to estimate both forward pose and backward pose between two frames. The losses are used in two-directional transformation and image reconstruction”); and update one or more parameters of the machine-learning model based at least on a loss determined according to the first transformed image and a second image of the set of the sequential images (Wang Fig. 1 & pg. 1778-1779, §III-B discussed above, Wang uses a hierarchical refinement method; also see Wang pg. 1778 right column: “Through the image reconstruction loss LR, geometry consistency loss LGC, and depth smooth loss Lsmooth, the multi-scale depths and multi-layer poses are trained together … we propose the data augmentation loss Laug based on 2D-3D transformation … The overall training loss is L … where α1, α2, α3, and α4 are hyperparameters”; Wang Fig. 2: see “Refine”). However, Wang does not appear to explicitly teach generating a 3D mesh representative of the first image and generating a transformed image based at least on a rendering of the mesh. Pertaining to the same field of endeavor, Duluk Jr. teaches generating a 3D mesh and transformed image based on the rendering of the 3D mesh (Duluk Jr. ¶¶0016: “Generation of pictures or images, is commonly called rendering. Generally, in three-dimensional (3D) computer graphics, geometry that represents surfaces (or volumes) of objects in a scene is translated into pixels (picture elements)”; Duluk Jr. ¶¶0017: “In a 3D animation, a sequence of images is displayed … allows a user to change his viewpoint or change the geometry in real-time, thereby requiring the rendering system to create new images on-the-fly in real-time”; Duluk Jr. ¶¶0018: “each renderable object generally has its own local object coordinate system, and therefore needs to be translated (or transformed) from object coordinates to pixel display coordinates. Conceptually, this is a 4-step process: 1) translation (including scaling for size enlargement or shrink) from object coordinates to world coordinates, which is the coordinate system for the entire scene; 2) translation from world coordinates to eye coordinates, based on the viewing point of the scene; 3) translation from eye coordinates to perspective translated eye coordinates, where perspective scaling (farther objects appear smaller) has been performed; and 4) translation from perspective translated eye coordinates to pixel coordinates, also called screen coordinates”; Duluk Jr. ¶¶0091: “FIG. 41 is a diagrammatic illustration showing the manner in which the user specified point is adjusted to the rendered point in the Geometry Unit”; Duluk Jr. ¶¶0740: “3D application programs generally render many vertices with the same state information S3. If fact, most APIs require the state information S3 to be constant for all the vertices in a polygon mesh (or, line strips, triangle strips, etc.)”). Wang and Duluk Jr. are considered to be analogous art because they are directed 3D image processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the 3D hierarchical refinement and augmentation for unsupervised learning of depth and pose from videos (as taught by Wang) to generate a transformed image based on the rendering of the 3D mesh (as taught by Duluk Jr.) because the combination can create new 3D animated images on the fly in real-time (Duluk Jr. ¶¶0016). Regarding claim 2, Wang, in view of Duluk Jr., teaches the processor of claim 1, wherein the one or more circuits are to: identify a respective depth map associated with each image of the set of sequential images (Wang Fig. 2: “The DepthNet estimates depth maps at four scales … reconstruction loss is generated by all-scale depth maps and all-level poses”; Wang pg. 1780 left column: “to find the corresponding depth in depth map Dt for Dt+1”; Wang Fig. 6); and update the one or more parameters of the machine-learning model further based at least on a second loss determined according to depth values of one or more mesh faces of an output of the machine-learning model and a respective depth map associated with the second image (Wang Figs. 1-2 & pg. 1778-1779, §III-B discussed above; Wang 1781 left column: “more than one pixel in the original images may be projected to the same grid in the augmented images. To solve the collision problem, pixels with minimum depth in original images Xt are selected to be displayed on the augmented image Xaugt”; Wang Fig. 3). Regarding claim 3, Wang, in view of Duluk Jr., teaches the processor of claim 2, wherein the one or more circuits are to update the one or more parameters of the machine-learning model further based at least on a third loss determined according to an estimated depth map of the output of the machine-learning model and a respective depth map associated with the first image (Wang Figs. 1-2 & pg. 1778-1779, §III-B discussed above; Wang eqs. (1), (14)-(16), (19)). Regarding claim 4, Wang, in view of Duluk Jr., teaches the processor of claim 1, wherein the one or more circuits are to generate at least one mask for at least one image of the set of sequential images (Wang pg. 1779 right column: “Loss Functions with Masks”; Wang pg. 1780 left column: “the depth inconsistency map is used to generate the occlusion weight mask … The binary auto-mask [14] is also used to filter the objects which are static relative to the camera and textureless regions … The final masked image reconstruction loss of Dn and Tm”; Wang Fig. 4). Regarding claim 8, Wang, in view of Duluk Jr., teaches the processor of claim 1, wherein the loss comprises one or more of a L1 loss, a structural similarity (SSIM) loss, or a minimal loss (Wang pg. 1778-1779, §III-B discussed above; also see Wang eq. (8), (12)-(16) & pg. 1780). Regarding claim 9, Wang, in view of Duluk Jr., teaches the processor of claim 1, wherein the one or more circuits are to execute the machine-learning model to generate a set of transformed images corresponding to at least the second viewpoint (Wang Figs. 1-2 discussed above; also see Wang Figs. 5-6). Regarding claim 10, Wang, in view of Duluk Jr., teaches the processor of claim 1, wherein the one or more circuits are to update one or more second parameters of a second machine-learning model using the first transformed image (Wang Fig. 2 & pg. 1778-1779, §III-B discussed above – multiple parameters are refined using the transformed images). Regarding claim 11, Wang, in view of Duluk Jr., teaches the processor of claim 1, wherein the processor is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational Al operations; a system for performing generative Al operations using a large language model (LLM); a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Wang Abstract: “autonomous robots and autonomous driving”; Wang pg. 1779 left column: “virtual intermediate view”; Wang pg. 1780 right column: “autonomous driving scenario”). Regarding claim 12, Wang teaches a system comprising: one or more processors (Wang pg. 1782 discussed above) to: identify a first set of images corresponding to a first viewpoint (Wang Fig. 1 discussed above); generate, using a machine-learning model and based at least on the first set of images, a second set of images corresponding to the first set of images and at least one second viewpoint (Wang Fig. 1, pg. 1778-1779, §III-B, pg. 1782 discussed above); and update one or more parameters of a second machine-learning model using a dataset comprising the second set of images (Wang Figs. 1-2, & pg. 1778-1779, §III discussed above, also see Wang Fig. 5 describing pose augmentation learning). However, Wang does not appear to explicitly teach generating a 3D mesh representative of each image of the first set of images and generating a second set of images based at least on a rendering of each mesh. Pertaining to the same field of endeavor, Duluk Jr. teaches generating a 3D mesh representative of each image of the first set of images and generating a second set of images based at least on a rendering of each mesh (Duluk Jr. ¶¶0016- ¶¶0018, ¶¶0091 & ¶¶0740 discussed above). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the 3D hierarchical refinement and augmentation for unsupervised learning of depth and pose from videos (as taught by Wang) to generate a transformed image based on the rendering of the 3D mesh (as taught by Duluk Jr.) because the combination can create new 3D animated images on the fly in real-time (Duluk Jr. ¶¶0016). Regarding claim 13, Wang, in view of Duluk Jr., teaches the system of claim 12, wherein the one or more processors are to iteratively execute the machine-learning model using a first image of the first set of images as input to generate a plurality of images included in the second set of images, each of the plurality of images corresponding to a respective viewpoint different from the first viewpoint (Wang Figs. 1-2, 5 & 1778-1780, §III-B discussed above). Regarding claim 14, Wang, in view of Duluk Jr., teaches the system of claim 13, wherein the one or more processors are to execute the machine-learning model further using at least an indication of the at least one second viewpoint (Wang Figs. 1-2, 4-5 & 1778-1780, §III-B discussed above – the final augmented image is based on the mask H; also see Wang Fig. 6). Claim 16 is rejected using the same rationale as applied to claim 11 discussed above. Regarding claim 17, Wang, in view of Duluk Jr., teaches that the processor and system perform a method comprising the processes described in claims 1 and 7. Therefore, claim 17 is rejected using the same rationale as applied to claims 1 and 12 discussed above. Claim 18 is rejected using the same rationale as applied to claim 2 discussed above. Claim 19 is rejected using the same rationale as applied to claim 3 discussed above. Claim 20 is rejected using the same rationale as applied to claim 4 discussed above. Claim(s) 5, 6, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (IEEE TCSVT, Vol. 33, No. 4, April 2023, published 19 October 2022), in view of Duluk Jr. et al. (US 2004/0130552 A1), and further in view of Seo et al. (US 2020/0090322 A1), hereinafter referred to as Wang, Duluk Jr., and Seo, respectively. Regarding claim 5, Wang, in view of Duluk Jr., teaches the processor of claim 4, wherein the one or more circuits are to generate the at least one mask using a second machine-learning model updated to predict the at least one mask to correspond to one or more objects (Wang pg. 1779-1780 & Fig. 4 discussed above). However, Wang, in view of Duluk Jr., does not appear to explicitly teach that the one or more objects are proximate to a device that captured the at least one image. Pertaining to the same field of endeavor, Seo teaches that the one or more objects are proximate to a device that captured the at least one image (Seo ¶¶0041: “a region-based mask may be created that associates masks with regions in the image based on the importance of the respective regions”; Seo Fig. 2 & ¶¶0046: “the machine learning model(s) 104 may use image data representative of input image 210A as input and may output an image mask 210B including the image blindness regions 212 and 214”; Seo Fig. 5 & ¶¶0072: “the ground truth data 404 may include annotations for blindness region(s) 406 (e.g., blindness region 522), blindness classification(s) 408 (e.g., blocked area), and blindness attribute(s) 410 (e.g., object (e.g., pedestrian) in proximity, during the day) to train the machine learning model(s) 104 to recognize and classify sensor blindness based on the regions and associated causes thereof”). Wang, in view of Duluk Jr., and Seo are considered to be analogous art because they are directed to neural networks for augmenting image data. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the 3D hierarchical refinement and augmentation for unsupervised learning of depth and pose from videos (as taught by Wang, in view of Duluk Jr.) to detect and estimate objects that are proximate to the image capture device (as taught by Seo) because the combination allows the machine learning model to recognize a road region or less important regions (e.g., sky) even when an input image includes variations in color or positioning (Seo ¶¶0066). Regarding claim 6, Wang, in view of Duluk Jr., teaches the processor of claim 4, but does not appear to explicitly teach that the one or more circuits are to generate the at least one mask using a second machine-learning model updated to predict the at least one mask to correspond to a sky depicted in the at least one image. Pertaining to the same field of endeavor, Seo teaches that the one or more circuits are to generate the at least one mask using a second machine-learning model updated to predict the at least one mask to correspond to a sky depicted in the at least one image (Seo Figs. 2, 5 & ¶¶0041, ¶¶0046, ¶¶0072 discussed above; further see Seo ¶¶0073: “The sky region may be annotated in training image 560 to indicate the sky near the horizon in the image 560 … The labeling may also train the machine learning model(s) to learn that the sky region—when blocked or blurred—is not as important of a region for determining usability of sensor data”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the 3D hierarchical refinement and augmentation for unsupervised learning of depth and pose from videos (as taught by Wang, in view of Duluk Jr.) to detect and estimate sky regions (as taught by Seo) because the combination allows the machine learning model to recognize a road region or less important regions (e.g., sky) even when an input image includes variations in color or positioning (Seo ¶¶0066). Regarding claim 15, Wang, in view of Duluk Jr. teaches the system of claim 12, but does not appear to explicitly teach that the second machine-learning model comprises a segmentation model. Pertaining to the same field of endeavor, Seo teaches that the second machine-learning model comprises a segmentation model (Seo Fig. 5 & ¶0071: “the machine learning model(s) 104 may be trained to predict potential blindness regions 108 as well as blindness classification(s) 110 and/or blindness attribute(s) 112 associated therewith” – different classified regions are segmented). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the 3D hierarchical refinement and augmentation for unsupervised learning of depth and pose from videos (as taught by Wang) to use a segmentation model (as taught by Seo) because the combination allows the machine learning model to ignore ego-vehicle regions and further include contextual information (Seo ¶¶0071). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOO J SHIN whose telephone number is (571)272-9753. The examiner can normally be reached M-F; 10-6. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571)272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Soo Shin/Primary Examiner, Art Unit 2667 571-272-9753 soo.shin@uspto.gov
Read full office action

Prosecution Timeline

Apr 12, 2024
Application Filed
May 08, 2026
Non-Final Rejection mailed — §103
Aug 06, 2026
Applicant Interview (Telephonic)
Aug 06, 2026
Examiner Interview Summary
Aug 07, 2026
Response Filed
Sep 02, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738049
ALGORITHM AND METHOD FOR DYNAMICALLY VARYING QUANTIZATION PRECISION OF DEEP LEARNING NETWORK
3y 2m to grant Granted Sep 15, 2026
Patent 12737862
METHOD AND SYSTEM FOR COMPUTER-AIDED ANEURYSM TRIAGE
2y 4m to grant Granted Sep 15, 2026
Patent 12725295
ADJACENT ITEM FILTERING FOR ACCURATE COMPARTMENT CONTENT MAPPING
2y 8m to grant Granted Sep 01, 2026
Patent 12725264
MASKING A DETECTED OBJECT IN A VIDEO STREAM
2y 5m to grant Granted Sep 01, 2026
Patent 12705917
AMBIGUITY RESOLUTION FOR OBJECT SELECTION AND FASTER APPLICATION LOADING FOR CLUTTERED SCENARIOS
3y 1m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+16.2%)
2y 2m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 628 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month