Prosecution Insights
Last updated: August 17, 2026
Application No. 18/497,940

TECHNIQUES FOR RECONSTRUCTING DIFFERENT THREE-DIMENSIONAL SCENES USING THE SAME TRAINED MACHINE LEARNING MODEL

Non-Final OA §103
Filed
Oct 30, 2023
Priority
Nov 15, 2022 — provisional 63/383,880
Examiner
NGUYEN, ANH TUAN V
Art Unit
2619
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
72%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
361 granted / 501 resolved
+10.1% vs TC avg
Strong +20% interview lift
Without
With
+19.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
23 currently pending
Career history
538
Total Applications
across all art units

Statute-Specific Performance

§101
9.2%
-30.8% vs TC avg
§103
69.3%
+29.3% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
12.5%
-27.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 501 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/29/2026 has been entered. Claims 1, 3-5, 11, 13-14, and 20 were amended. Claim 21 was added. Claims 1, 3-11, and 13-21 are pending in the application. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3-4, 7, 11, 13, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borisov (US 2017/0064287) in view of Ramirez de Chanlatte et al. (US 2023/0147722) and Steedly et al. (US 2010/0238164). Regarding claim 1, Borisov teaches/suggests: A computer-implemented method for generating three-dimensional (3D) representations of scenes, the method comprising: mapping a first red, green, blue, and depth (RGBD) image associated with both a first scene and a first viewpoint to a first surface representation of at least a first portion of the first scene (Borisov [0013]-[0014] “ building a surface representation from a point cloud, usually as a set of triangles … generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” [0005] “The source data for a scene reconstruction algorithm is a set of pairs of RGB and depth images” [The texture corresponding to a first camera position meets the first surface representation.]), mapping a second RGBD image associated with both the first scene and a second viewpoint to a second surface representation of at least a second portion of the first scene (Borisov [0013]-[0014] “building a surface representation from a point cloud, usually as a set of triangles … generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” [0005] “The source data for a scene reconstruction algorithm is a set of pairs of RGB and depth images” [The texture corresponding to a second camera position meets the second surface representation.]), aggregating at least the first surface representation and the second surface representation in a 3D space to generate a first fused surface representation of the first scene (Borisov [0024] “Blend several RGB images to create a seamless texture”); and mapping the first fused surface representation of the first scene to a 3D representation of the first scene (Borisov [0014] “generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” [0005] “The output of the algorithm is a 3D model of a scene consisting of a mesh and a texture”). Borisov does not teach/suggest using a geometry encoder. Nor does Borisov teach/suggest: wherein the mapping the first RGBD image using the geometry encoder is based at least on a first plurality of input vectors that are determined based on the first viewpoint and a first depth image included in the first RGBD image; Ramirez de Chanlatte, however, teaches/suggests using a geometry encoder (Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 leverages the depth map of the real 2D image 302 to learn shape features (e.g., elements of a visible object shape) of the digital object portrayed in the real 2D image 302. For example, the 3D-object-reconstruction-machine-learning model 207 encodes shape features (e.g., a wing shape, a fuselage shape, and a tail shape of an airplane) into a shape feature encoding—also referred to as a predicted latent shape feature vector”). Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the first and second mappings of Borisov to include the geometry encoder of Ramirez de Chanlatte for machine learning. As such, Borisov as modified by Ramirez de Chanlatte teaches/suggests: wherein the mapping the first RGBD image using the geometry encoder is based at least on a first plurality of input vectors that are determined based on the first viewpoint and a first depth image included in the first RGBD image (Borisov [0013]-[0014] “building a surface representation from a point cloud, usually as a set of triangles … generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” [0005] “The source data for a scene reconstruction algorithm is a set of pairs of RGB and depth images” Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 leverages the depth map of the real 2D image 302 to learn shape features (e.g., elements of a visible object shape) of the digital object portrayed in the real 2D image 302. For example, the 3D-object-reconstruction-machine-learning model 207 encodes shape features (e.g., a wing shape, a fuselage shape, and a tail shape of an airplane) into a shape feature encoding—also referred to as a predicted latent shape feature vector” [0032] “a depth map can include depth information derived from … multiple images from different viewpoints”); Borisov and Ramirez de Chanlatte are silent regarding: wherein at least part of the first portion of the first scene overlaps with at least part of the second portion of the first scene; Steedly, however, teaches/suggests: wherein at least part of the first portion of the first scene overlaps with at least part of the second portion of the first scene (Steedly [0084] “a geometric proxy generation module 520 operates to automatically generate a 3D model of the scene using any of a number of conventional techniques. For example, capturing overlapping images of a scene from two or more slightly different perspectives allows conventional stereo imaging techniques to be used”); Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the images of Borisov as modified by Ramirez de Chanlatte to be overlapped as taught/suggested by Steedly for stereo imaging. Regarding claim 3, Borisov as modified by Ramirez de Chanlatte and Steedly teaches/suggests: The computer-implemented method of claim 1, wherein the geometry encoder generates a first geometric surface representation of the at least first portion of the first scene (Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 leverages the depth map of the real 2D image 302 to learn shape features (e.g., elements of a visible object shape) of the digital object portrayed in the real 2D image 302. For example, the 3D-object-reconstruction-machine-learning model 207 encodes shape features (e.g., a wing shape, a fuselage shape, and a tail shape of an airplane) into a shape feature encoding—also referred to as a predicted latent shape feature vector”). The same rationale to combine as set forth in the rejection of claim 1 is incorporated herein. Regarding claim 4, Borisov as modified by Ramirez de Chanlatte and Steedly teaches/suggests: determining a second plurality of input vectors based on a third viewpoint and a third RGBD image that is associated with both a second scene and the third viewpoint (Borisov [0005] “The source data for a scene reconstruction algorithm is a set of pairs of RGB and depth images” [Reconstructing another scene meets the second scene.] Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 leverages the depth map of the real 2D image 302 to learn shape features (e.g., elements of a visible object shape) of the digital object portrayed in the real 2D image 302. For example, the 3D-object-reconstruction-machine-learning model 207 encodes shape features (e.g., a wing shape, a fuselage shape, and a tail shape of an airplane) into a shape feature encoding—also referred to as a predicted latent shape feature vector” [0032] “a depth map can include depth information derived from … multiple images from different viewpoints”); and executing the geometry encoder on the second plurality of input vectors to generate a second geometric surface representation of at least a portion of the second scene (Borisov [0013]-[0014] “building a surface representation from a point cloud, usually as a set of triangles … generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 leverages the depth map of the real 2D image 302 to learn shape features (e.g., elements of a visible object shape) of the digital object portrayed in the real 2D image 302. For example, the 3D-object-reconstruction-machine-learning model 207 encodes shape features (e.g., a wing shape, a fuselage shape, and a tail shape of an airplane) into a shape feature encoding—also referred to as a predicted latent shape feature vector”). The same rationale to combine as set forth in the rejection of claim 1 is incorporated herein. Regarding claim 7, Borisov as modified by Ramirez de Chanlatte and Steedly teaches/suggests: The computer-implemented method of claim 1, wherein the first surface representation comprises a geometric surface representation of the at least first portion of the first scene and a texture surface representation of the at least first portion of the first scene (Borisov [0005] “The output of the algorithm is a 3D model of a scene consisting of a mesh and a texture”). The mesh meets the geometric surface representation. Claims 11 and 13 recite limitation(s) similar in scope to those of claims 1 and 3, respectively, and are rejected for the same reason(s). Borisov as modified by Ramirez de Chanlatte and Steedly further teaches/suggests one or more non-transitory computer readable media including instructions (Borisov [0005] “3d scene reconstruction from RGBD-camera images (especially using IOS Ipad device with attached Structure Sensor)”). Regarding claim 19, Borisov as modified by Ramirez de Chanlatte and Steedly teaches/suggests: The one or more non-transitory computer readable media of claim 11, wherein the first viewpoint is specified by at least one of a rotation matrix, a 3D translation, or an intrinsic matrix associated with a camera (Borisov [0008] “Intrinsic parameters: parameters of RGB camera that define how 3D point map to pixels in an image generated by the camera”). Claim 20 recites limitation(s) similar in scope to those of claim 1, and is rejected for the same reason(s). Borisov as modified by Ramirez de Chanlatte and Steedly further teaches/suggests one or more memories storing instructions; and one or more processors coupled to the one or more memories (Borisov [0005] “3d scene reconstruction from RGBD-camera images (especially using IOS Ipad device with attached Structure Sensor)”). Claim(s) 5-6 and 14-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borisov (US 2017/0064287) in view of Ramirez de Chanlatte et al. (US 2023/0147722) and Steedly et al. (US 2010/0238164) as applied to claims 1 and 11 above, and further in view of Wang et al. (US 2023/0071559). Regarding claim 5, Borisov as modified by Ramirez de Chanlatte and Steedly does not teach/suggest: The computer-implemented method of claim 1, wherein mapping the first RGBD image further comprises executing a trained texture encoder on a first red, green, and blue (RGB) image included in the first RGBD image to generate a first plurality of texture feature vectors associated with a first plurality of pixels included in the first RGB image. Wang, however, teaches/suggests executing a trained texture encoder (Wang [0067] “The first step of the present neural rendering is to assign to each 3D point p.sub.i a feature vector f.sub.i that encodes the object's appearance, its alpha matte, and its contextual information”). Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the first and second mappings of Borisov as modified by Ramirez de Chanlatte and Steedly to include the texture encoder of Wang for machine learning. As such, Borisov as modified by Ramirez de Chanlatte, Steedly, and Wang teaches/suggests a trained texture encoder on a first red, green, and blue (RGB) image included in the first RGBD image to generate a first plurality of texture feature vectors associated with a first plurality of pixels included in the first RGB image (Borisov [0013]-[0014] “building a surface representation from a point cloud, usually as a set of triangles … generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” [0005] “The source data for a scene reconstruction algorithm is a set of pairs of RGB and depth images” Wang [0067] “The first step of the present neural rendering is to assign to each 3D point p.sub.i a feature vector f.sub.i that encodes the object's appearance, its alpha matte, and its contextual information”). Regarding claim 6, Borisov as modified by Ramirez de Chanlatte, Steedly, and Wang teaches/suggests: The computer-implemented method of claim 5, further comprising projecting the first plurality of texture feature vectors onto a first plurality of 3D surface points to generate a first texture surface representation of the at least first portion of the first scene (Wang [0067] “When rendering a novel target viewpoint V, all points and thus their features are projected onto a view-dependent feature map M.sub.q”). The same rationale to combine as set forth in the rejection of claim 5 is incorporated herein. Claim 14 recites limitation(s) similar in scope to those of claim 5, and is rejected for the same reason(s). Regarding claim 15, Borisov as modified by Ramirez de Chanlatte, Steedly, and Wang teaches/suggests: The one or more non-transitory computer readable media of claim 14, further comprising: executing the trained texture encoder on a second RGB image associated with a second scene to generate a second plurality of texture feature vectors associated with a second plurality of pixels included in the second RGB image (Borisov [0013]-[0014] “building a surface representation from a point cloud, usually as a set of triangles … generated a seamless texture map for a surface from images from multiple frames, corresponding to different camera positions in space” [0005] “The source data for a scene reconstruction algorithm is a set of pairs of RGB and depth images” [Reconstructing another scene meets the second scene.] Wang [0067] “The first step of the present neural rendering is to assign to each 3D point p.sub.i a feature vector f.sub.i that encodes the object's appearance, its alpha matte, and its contextual information”); and projecting the second plurality of texture feature vectors onto a second plurality of 3D surface points to generate a second texture surface representation of at least a portion of the second scene (Wang [0067] “When rendering a novel target viewpoint V, all points and thus their features are projected onto a view-dependent feature map M.sub.q”). The same rationale to combine as set forth in the rejection of claim 5 is incorporated herein. Regarding claim 16, Borisov as modified by Ramirez de Chanlatte, Steedly, and Wang teaches/suggests: The one or more non-transitory computer readable media of claim 11, wherein the second surface representation comprises a plurality 3D surface points that are associated with a plurality of geometry feature vectors (Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 leverages the depth map of the real 2D image 302 to learn shape features (e.g., elements of a visible object shape) of the digital object portrayed in the real 2D image 302. For example, the 3D-object-reconstruction-machine-learning model 207 encodes shape features (e.g., a wing shape, a fuselage shape, and a tail shape of an airplane) into a shape feature encoding—also referred to as a predicted latent shape feature vector”) and a plurality of texture feature vectors (Wang [0067] “The first step of the present neural rendering is to assign to each 3D point p.sub.i a feature vector f.sub.i that encodes the object's appearance, its alpha matte, and its contextual information”). The same rationales to combine as set forth in the rejection of claims 1 and 5 are incorporated herein. Claim(s) 8 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borisov (US 2017/0064287) in view of Ramirez de Chanlatte et al. (US 2023/0147722) and Steedly et al. (US 2010/0238164) as applied to claims 1 and 11 above, and further in view of Shen et al. (US 2022/0392162). Regarding claim 8, Borisov as modified by Ramirez de Chanlatte and Steedly teaches/suggests: The computer-implemented method of claim 1, wherein mapping the first fused surface representation comprises: executing a trained geometry decoder on the plurality of geometry input vectors to generate a plurality of signed distance function values (Ramirez de Chanlatte [0064] “the 3D-object-reconstruction-machine-learning model 207 uses the shape feature encoding based on the depth map to assign predicted SDF values to query points sampled near an object surface of the digital object”). The same rationale to combine as set forth in the rejection of claim 1 is incorporated herein. Borisov as modified by Ramirez de Chanlatte and Steedly does not teach/suggest: performing one or more interpolation operations on the first fused surface representation to generate a plurality of geometry input vectors; Shen, in view of Borisov, teaches/suggests: performing one or more interpolation operations on the first fused surface representation to generate a plurality of geometry input vectors (Borisov [0024] “Blend several RGB images to create a seamless texture” Shen [0026] “The machine learning model(s) may then be used to generate a feature vector F.sub.vol(v, x) for a grid vertex v∈custom-character.sup.3 via trilinear interpolation”); Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the geometry encoding of Borisov as modified by Ramirez de Chanlatte and Steedly to include the interpolation of Shen to create a seamless mesh. Claim 17 recites limitation(s) similar in scope to those of claim 8, and is rejected for the same reason(s). Claim(s) 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borisov (US 2017/0064287) in view of Ramirez de Chanlatte et al. (US 2023/0147722) and Steedly et al. (US 2010/0238164) as applied to claim 1 above, and further in view of Sugano et al. (US 2020/0410754). Regarding claim 21, Borisov as modified by Ramirez de Chanlatte and Steedly does not teach/suggest: The computer-implemented method of claim 1, wherein the first viewpoint associated with the first RGBD image is specified via camera metadata associated with the first RGBD image. Sugano, however, teaches/suggests the first viewpoint associated with the first RGBD image is specified via camera metadata associated with the first RGBD image (Sugano [0161] “performs an encoding process based on a predetermined encoding method with respect to metadata including the 2-dimensional image data and the depth image data of a plurality of viewpoints ... a camera parameter of each viewpoint”). Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the viewpoints of Borisov as modified by Ramirez de Chanlatte and Steedly to be included in the metadata of Sugano for encoding. Allowable Subject Matter Claims 9-10 and 18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The limitations “generating a first plurality of texture input vectors based on the first fused surface representation and a first plurality of signed distance function (SDF) values generated by a trained geometry decoder” and “executing a trained texture decoder on the first plurality of texture input vectors to generate a first plurality of radiance values,” taken as a whole, render the claims patentably distinct over the prior art. Response to Arguments Applicant's arguments filed on 04/29/2026 have been fully considered but they are moot in view of the new ground(s) of rejection set forth in this Office action. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: US 2021/0174513 – depth completion US 2022/0237879 – joint training US 2024/0070884 – depth autoencoder Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANH-TUAN V NGUYEN whose telephone number is 571-270-7513. The examiner can normally be reached on M-F 9AM-5PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JASON CHAN can be reached on 571-272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANH-TUAN V NGUYEN/ Primary Examiner, Art Unit 2619
Read full office action

Prosecution Timeline

Show 2 earlier events
Oct 31, 2025
Response Filed
Nov 13, 2025
Examiner Interview Summary
Nov 13, 2025
Applicant Interview (Telephonic)
Dec 30, 2025
Final Rejection mailed — §103
Feb 26, 2026
Response after Non-Final Action
Apr 29, 2026
Request for Continued Examination
May 04, 2026
Response after Non-Final Action
Jun 11, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12657824
SYSTEMS AND METHODS FOR FACE ASSET CREATION AND MODELS FROM ONE OR MORE IMAGES
2y 8m to grant Granted Jun 16, 2026
Patent 12626456
ELECTRONIC DEVICE FOR DISPLAYING VIRTUAL OBJECT AND OPERATION METHOD THEREOF
2y 5m to grant Granted May 12, 2026
Patent 12614358
AUGMENTED REALITY ENVIRONMENT MELDING
3y 1m to grant Granted Apr 28, 2026
Patent 12608856
SYSTEM FOR AND METHOD OF GRAPHICALLY REPRESENTING INFORMATION
3y 5m to grant Granted Apr 21, 2026
Patent 12591359
ELECTRONIC DEVICE COMPRISING DISPLAY THAT OPTIMALLY DISPLAY CONTENT WITH RESPECT TO CAMERA HOLE, AND METHOD FOR CONTROLLING DISPLAY THEREOF
3y 2m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
72%
Grant Probability
92%
With Interview (+19.7%)
2y 10m (~1m remaining)
Median Time to Grant
High
PTA Risk
Based on 501 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month