Prosecution Insights
Last updated: August 17, 2026
Application No. 18/915,547

DYNAMIC NOVEL VIEW RECONSTRUCTION BASED ON FLOW REMATCHING

Non-Final OA §103
Filed
Oct 15, 2024
Priority
Sep 30, 2024 — GR 20240100666
Examiner
DU, HAIXIA
Art Unit
2611
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
86%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
490 granted / 567 resolved
+24.4% vs TC avg
Strong +18% interview lift
Without
With
+17.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
11 currently pending
Career history
582
Total Applications
across all art units

Statute-Specific Performance

§101
10.3%
-29.7% vs TC avg
§103
52.0%
+12.0% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
20.5%
-19.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 567 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are present for examination. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 4, 7-9, 11, 12, and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over US Patent Publication No. 20240257309 A1 to Holland in view of US Patent Publication No. 20250193363 A1 to Liu et al. Regarding claim 1, Holland discloses One or more processors comprising processing circuitry to (Holland, para. [0041], systems include powerful processors, para. [0067], the image processor may include one or more processors): cause an image rendering model to generate an estimated image of a scene based at least on a plurality of images of the scene, at least one image of the plurality of images associated with at least one of a different time or a different view (Holland, para. [0052], disclosing generating a combined image based on two or more input images captured by image sensors of separate devices, the image captured devices may have relative positioning (or pose) varies between images captured at different times, the system can generate an image with a novel viewpoint from the input images and the localization images, para. [0059], disclosing capturing images and/or videos that include multiple images in a particular sequence, para. [0133], disclosing the viewpoint synthesis engine can be configured to obtain the first image, the second image, and the localization information to generate an image with a novel viewpoint, the novel viewpoint can be different from the viewpoint of the first image or the viewpoint of the second image, para. [0134], disclosing the viewpoint synthesis engine can be implemented with a GAN, a generator G for a GAN can receive a set of images of an area captured from different viewpoints and generate an output image of the area from a novel viewpoint). However, Holland does not expressly disclose update the image rendering model based at least on the estimated image, the plurality of images, and one or more criteria for motion associated with the estimated image. On the other hand, Liu discloses update the image rendering model based at least on the estimated image, the plurality of images, and one or more criteria for motion associated with the estimated image (Liu, para. [0047], disclosing after generating the rendered image, the three-dimensional scene model can be trained based on the differences between the rendered image and the image samples, the network parameters of the three-dimensional scene model can be iteratively adjusted until the difference between the rendered image and the image samples satisfy a preset requirement that may include the differences between the rendered image and the image samples being less than a difference threshold, indicating the rendered image can correspond to the estimated image, the image samples can correspond to the plurality of images, and the difference threshold between rendered image and image samples can correspond to one or more criteria for motion associated with the rendered image as the estimated image as the difference between rendered image and image samples can correspond to motion associated with the rendered image as the estimated image). Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Holland and Liu. The suggestion/motivation would have been to allow the three-dimensional scene model to reconstruct the three-dimensional structure of the target scene, as suggested by Liu (see Liu, para. [0049]). Regarding claim 4, Holland in view of Liu discloses the one or more processors of claim 1, wherein the one or more criteria for motion comprise at least one of: (i) a rigidity constraint that limits changes in shape of objects over time or (ii) a continuity constraint that assumes smooth transitions in object motion within the scene (Liu, para. [0047], disclosing after generating the rendered image, the three-dimensional scene model can be trained based on the differences between the rendered image and the image samples, the network parameters of the three-dimensional scene model can be iteratively adjusted until the difference between the rendered image and the image samples satisfy a preset requirement that may include the differences between the rendered image and the image samples being less than a difference threshold, indicating the rendered image can correspond to the estimated image, the image samples can correspond to the plurality of images, and the difference threshold between rendered image and image samples can correspond to one or more criteria for motion associated with the rendered image as the estimated image as the difference between rendered image and image samples can correspond to motion associated with the rendered image as the estimated image, and the difference can correspond to a rigidity constraint that limits changes in shape of objects over time, because limiting the differences will limit changes in shape of objects over time). Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Holland and Liu. The suggestion/motivation would have been to allow the three-dimensional scene model to reconstruct the three-dimensional structure of the target scene, as suggested by Liu (see Liu, para. [0049]). Regarding claim 7, Holland in view of Liu discloses the one or more processors of claim 1, wherein the image rendering model comprises at least one of: (i) a Gaussian splatting model or (ii) a neural radiance field (NeRF) model (Liu, para. [0027], disclosing using the neural radiance field module of the three-dimensional scene model to process the compressed image feature and obtain a rendered image). Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Holland and Liu. The suggestion/motivation would have been to allow the three-dimensional scene model to reconstruct the three-dimensional structure of the target scene, as suggested by Liu (see Liu, para. [0049]). Regarding claim 8, Holland in view of Liu discloses the one or more processors of claim 1, wherein the plurality of images of the scene comprises a plurality of multi-view images, wherein the plurality of multi-view images corresponds to a plurality of different viewpoints captured at a plurality of different time points (Holland, para. [0052], disclosing generating a combined image based on two or more input images captured by image sensors of separate devices, the image captured devices may have relative positioning (or pose) varies between images captured at different times, the system can generate an image with a novel viewpoint from the input images and the localization images, para. [0059], disclosing capturing images and/or videos that include multiple images in a particular sequence, para. [0133], disclosing the viewpoint synthesis engine can be configured to obtain the first image, the second image, and the localization information to generate an image with a novel viewpoint, the novel viewpoint can be different from the viewpoint of the first image or the viewpoint of the second image, indicating the captured images can correspond to a plurality of multi-view images captured at different pose/positions (different viewpoints) and different times (time points)). Regarding claim 9, Holland in view of Liu discloses the one or more processors of claim 1, wherein the processing circuitry is to: apply a scene reconstruction using the image rendering model to render the estimated image for one or more viewpoints and one or more temporal intervals based on the plurality of images of the scene (Holland, para. [0133], disclosing the viewpoint synthesis engine can be configured to obtain the first image, the second image, and the localization information to generate an image with a novel viewpoint, the novel viewpoint can be different from the viewpoint of the first image or the viewpoint of the second image, para. [0134], disclosing the viewpoint synthesis engine can be implemented with a GAN, para. [0135], disclosing during inference, the viewpoint synthesis engine can generate images from novel viewpoint based on the first image and the second image, indicating the viewpoint synthesis engine can correspond to the image rendering model to generate images from novel viewpoint during inference as scene reconstruction, the generated images from novel viewpoint can correspond to the estimated image for one viewpoint and one temporal interval based on the input images as the plurality of images of the scene). Regarding claim 11, Holland in view of Liu discloses the one or more processors of claim 1, wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more multi-model language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Holland, para. [0075], disclosing an XR system that can mixed virtual content appears to be at a location in the environment, para. [0077], disclosing the XR system includes remote control, indicating the XR system can correspond to a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content). Regarding claim 12, it recites similar limitations of claim 1 discussed above but in a system form. The rationale of claim 1 rejection is applied to reject claim 12. Regarding claim 17, it recites similar limitations of claim 7 discussed above but in a system form. The rationale of claim 7 rejection is applied to reject claim 17. Regarding claim 18, it recites similar limitations of claim 8 discussed above but in a system form. The rationale of claim 8 rejection is applied to reject claim 18. Regarding claim 19, it recites similar limitations of claim 9 discussed above but in a system form. The rationale of claim 9 rejection is applied to reject claim 19. Regarding claim 20, it recites similar limitations of claim 1 discussed above but in a method form. The rationale of claim 1 rejection is applied to reject claim 20. Claim(s) 2, 3, 13, and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Holland in view of Liu as applied to claim 1 above, and further in view of US Patent Publication No. 20200090345 A1 to Krebs et al. Regarding claim 2, Holland in view of Liu discloses the one or more processors of claim 1. However, Holland or Liu does not expressly disclose wherein the one or more criteria for motion comprise a velocity field representing a plurality of deformations in space over time. On the other hand, Krebs discloses the one or more criteria for motion comprise a velocity field representing a plurality of deformations in space over time (Krebs, para. [0009], disclosing a predicted dense velocity field to provide estimated velocities at each pixel in the input frame, para. [0013], disclosing minimizing a loss function that compares the final frame with a predicted final frame generated from each previous frame by warping each previous frame based on a sum of the dense velocity field generated for that frame and the dense velocity fields generated for all intermediate frames between that frame and the final frame, indicating the velocity field can correspond to the one or more criteria for motion (minimizing the loss function), the velocity field representing pixel changes from a previous frame as a plurality of deformations in space over time). Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Holland in view of Liu with Krebs. The suggestion/motivation would have been to provide performing motion synthesis from an input image to generate a synthetic sequence of images, as suggested by Krebs (see Krebs, para. [0018]). Regarding claim 3, Holland in view of Liu discloses the one or more processors of claim 1. However, Holland or Liu does not expressly disclose wherein updating the image rendering model further comprises rematching a velocity field generated from the estimated image with a prior velocity field corresponding with the scene. On the other hand, Krebs discloses updating the image rendering model further comprises rematching a velocity field generated from the estimated image with a prior velocity field corresponding with the scene (Krebs, para. [0009], disclosing a predicted dense velocity field to provide estimated velocities at each pixel in the input frame, para. [0013], disclosing training the deep neural network to minimize a loss function that compares the final frame with a predicted final frame generated from each previous frame by warping each previous frame based on a sum of the dense velocity field generated for that frame and the dense velocity fields generated for all intermediate frames between that frame and the final frame, indicating training the deep neural network to minimize the loss function can correspond to update the deep neural network as the image rendering model by rematching the velocity field from the predicted final frame (corresponding to a velocity field generated from the estimated image) with a velocity field of a pervious frame as the prior velocity field corresponding to the scene so that the loss function can be minimized). Before the invention was effectively filed, it would have been obvious for a person skilled in the art to combine Holland in view of Liu with Krebs. The suggestion/motivation would have been to provide performing motion synthesis from an input image to generate a synthetic sequence of images, as suggested by Krebs (see Krebs, para. [0018]). Regarding claim 13, it recites similar limitations of claim 2 discussed above but in a system form. The rationale of claim 2 rejection is applied to reject claim 13. Regarding claim 14, it recites similar limitations of claim 3 discussed above but in a system form. The rationale of claim 3 rejection is applied to reject claim 14. Allowable Subject Matter Claims 5, 6, 10, 15, and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 5, none of the prior art references on the record discloses wherein the one or more criteria for motion correspond to a machine learning (ML) model updated to generate one or more deformations of the estimated image based on parameters derived from one or more historical scenes. Claim 6 depends from claim 5 with additional limitations. Claim 15 recites similar limitations discussed above with respect to claim 5. Claim 16 depends from claim 15 with additional limitations. Regarding claim 10, Holland in view of Liu discloses the one or more processors of claim 1, wherein updating the image rendering model comprises minimizing a reconstruction loss, and wherein the reconstruction loss corresponds to a measure of discrepancy between the estimated image and the plurality of images of the scene (Liu, para. [0047], disclosing after generating the rendered image, the three-dimensional scene model can be trained based on the differences between the rendered image and the image samples, the network parameters of the three-dimensional scene model can be iteratively adjusted until the difference between the rendered image and the image samples satisfy a preset requirement that may include the differences between the rendered image and the image samples being less than a difference threshold, indicating the difference can correspond to the reconstruction loss). However, none of the prior art references on the record discloses wherein updating the image rendering model comprises minimizing a reconstruction loss and a rematch loss and wherein the rematch loss corresponds to a measure of deviation between the estimated image and the one or more criteria for motion associated with an image flow. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US Patent Publication No. 20210326583 A1 to Donatsch et al., which discloses training a machine learning model to predict user expression using a plurality of images and a plurality of values for a movement metric. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAIXIA DU whose telephone number is (571)270-5646. The examiner can normally be reached Monday - Friday 8:00 am-4:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at 571-272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HAIXIA DU/Primary Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Oct 15, 2024
Application Filed
Jul 16, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694480
HIGH-RESOLUTION MULTIVIEW-CONSISTENT RENDERING AND ALPHA MATTING FROM SPARSE VIEWS
2y 10m to grant Granted Jul 28, 2026
Patent 12688638
ELECTRONIC DEVICE FOR GENERATING THREE-DIMENSIONAL PHOTO BASED ON IMAGES ACQUIRED FROM PLURALITY OF CAMERAS, AND METHOD THEREFOR
2y 5m to grant Granted Jul 21, 2026
Patent 12688652
SURFACE MESH SELF-INTERSECTION DETECTION
2y 5m to grant Granted Jul 21, 2026
Patent 12685620
IMAGE-BASED LONGITUDINAL ANALYSIS
1y 11m to grant Granted Jul 21, 2026
Patent 12682584
AUGMENTED REALITY DEVICE OPERATION WITH ROBOTIC TOTAL STATION
2y 2m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
86%
Grant Probability
99%
With Interview (+17.9%)
2y 3m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 567 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month