Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 1/16/25 is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 6-10, 12 and 14-20 are rejected under 35 U.S.C. 103 as being unpatentable over Song et al. (US 2026/0065425 A1).
As to Claim 1, Song teaches A method for reducing noise in synthesized images generated by a view synthesis network using an image generation network (Song, [0005, 0070]), the method comprising:
capturing images of one or more target objects or sites using cameras; estimating camera poses for the captured images (Song discloses “At operation 2005, the system obtains a first training set including a first training image and a second training image, where the second training image depicts an object from the first training image from a different view…The first training image and the second training image include the same object having different views or poses” in [0192]; pose in [0030].);
for each target object or site, dividing training samples comprising image-camera pose pairs into a first training set and a second training set (Song discloses a first training set in [0192] and a second training set in [0196]);
providing camera poses associated with the first training set to a view synthesis network, which synthesizes views from the camera poses to obtain a first set of synthesized images (Song discloses “At operation 2010, the system trains, using the first training set during a first training stage, an image generation model to generate a synthetic image that preserves an identity…” in [0193]);
comparing the first set of synthesized images with images in the first training set to obtain a first comparison result (Song discloses “Training component 1040 computes an identity preserving loss based on the preliminary output and the second training image” in [0121], see also [0238]. It is obvious that the loss function is calculated by comparing the output image and input image.);
inputting the first comparison result to the view synthesis network (Song discloses “Training component 1040 updates parameters of the image generation model 1025 based on the identity preserving loss” in [0121], see also [0238]);
providing camera poses associated with the second training set to the view synthesis network, which synthesizes views from the camera poses to obtain a second set of synthesized images; providing the second set of synthesized images to an image generation network to generate noise-reduced images (Song discloses “obtaining a second training set including a training foreground image, a training background image, and a ground-truth composite image that depicts an object from the training foreground image in a scene from the training background image; and training, using the second training set during a second training stage, the image generation model to generate a composite image based on an input foreground image and an input background image, wherein the composite image depicts an object from the input foreground image in a scene from the input background image” in [0005]);
comparing the noise-reduced images with images in the second training set to obtain a second comparison result; and inputting the second comparison result to the image generation network, wherein the view synthesis network has not been trained on images associated with the second training set (Song discloses “Training component 1040 computes a compositing loss based on the preliminary composite output and the ground-truth composite image. Training component 1040 updates parameters of the image generation model 1025 based on the compositing loss” in [0123], see also [0143] and Fig 4 & 27. It is obvious that the loss function is calculated by comparing the output image and input image.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Song with estimated pose for the captured image and comparing output image with input image to calculate a loss function during a generative network.
As to Claim 2, the modified Song teaches The method of claim 1, further comprising using the noise-reduced images as additional training samples for the view synthesis network (Song, Fig 13.)
As to Claim 3, the modified Song teaches The method of claim 1, wherein the synthesized views from randomly sampled camera poses are input into the image generation network before being used as additional training samples (Song discloses “In some cases, video segmentation datasets are used in the second training stage. Referring to FIG. 25, each training pair comes from one video with instance-level segmentation labels. Two distinct frames are randomly sampled (e.g., two frames are randomly sampled from frames 2500)” in [0228].)
As to Claim 4, the modified Song teaches The method of claim 3, further comprising using at least some of the randomly sampled camera poses to train view synthesis networks to enable the view synthesis networks to learn additional 3D structures consistently (Song discloses “In some cases, video segmentation datasets are used in the second training stage. Referring to FIG. 25, each training pair comes from one video with instance-level segmentation labels. Two distinct frames are randomly sampled (e.g., two frames are randomly sampled from frames 2500); one frame serves as the target image 2515, while an object 2505 is extracted from the other frame as the augmented input (i.e., augmented object 2510).” in [0228].)
As to Claim 6, the modified Song teaches The method of claim 4, further comprising sharing one or more image generation network parameters across two or more of the view synthesis networks to improve a noise reduction performance based on accumulated training data (Song discloses “In some examples, training component 1040 freezes an image encoder 1030 of the image generation model 1025 during the second training stage. In some examples, training component 1040 trains the image encoder 1030 of the image generation model 1025 during the first training stage” in [0124]; “The second training stage 1200 includes a process of taking the learned image encoder 1215 from the first training stage 1100 (see FIG. 11) and freezing the backbone of image encoder 1215” in [0140], see also [0239, 0242].)
As to Claim 7, the modified Song teaches The method of claim 6, wherein the accumulated training data uses training data from two or more target objects or sites (Song, Fig 4-8, 26-27.)
As to Claim 8, the modified Song teaches The method of claim 6, wherein the image generation network learns from training data of different target objects or sites (Song, Fig 4-8, 26-27.)
As to Claim 9, the modified Song teaches The method of claim 6, wherein the one or more shared image generation networks use training data from multiple target objects or sites to enhance noise reduction performance across different view synthesis networks (Song, Fig 4-8, 26-27.)
As to Claim 10, the modified Song teaches The method of claim 9, wherein each of the different view synthesis networks learns its own target site or object (Song discloses “image generation model 1025 comprises parameters stored in the at least one memory and is trained to encode a foreground image to obtain a foreground embedding that preserves an identity of an object in the foreground image and generates a composite image based on a background image and the foreground embedding, wherein the composite image depicts the object from the foreground image within a scene from the background image” in [0116], see also Fig 4-8, 26-27.)
As to Claim 12, the modified Song teaches The method of claim 1, further comprising optimizing the image generation network by comparing synthesized images with real images from the second training set using an objective function (Song discloses “For example, during the training process, an algorithm adjusts machine learning parameters to minimize an error or loss between predicted outputs and actual targets according to optimization techniques like gradient descent, stochastic gradient descent, or other optimization algorithms [112” in [0112]; compositing loss in [0123, 0143, 0207] and identity preserving loss in [0121, 0201-0203].)
Claim 14 recites similar limitations as claim 1 but in a computer readable medium form. Therefore, the same rationale used for claim 1 is applied.
Claim 15 is rejected based upon similar rationale as Claim 2.
Claim 16 is rejected based upon similar rationale as Claim 3.
Claim 17 is rejected based upon similar rationale as Claim 4.
Claim 18 is rejected based upon similar rationale as Claim 6.
Claim 19 is rejected based upon similar rationale as Claim 9.
Claim 20 is rejected based upon similar rationale as Claim 10.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Song et al. (US 2026/0065425 A1) in view of Lukac et al. (US 2022/0114365 A1).
As to Claim 5, the modified Song teaches The method of claim 1. The combination of Lukac further teaches wherein estimating the camera poses for the captured images comprises using at least one of a structure-from-motion algorithm or a point-matching algorithm, wherein the point-matching algorithm is used in response to capturing depth data with 3D sensors to enhance a camera pose estimation (Lukac discloses “Further, the input images can be localized to the terrain model using Structure-from-Motion. Global bundle adjustment can then be used to refine camera parameters belonging to the images and three-dimensional points” in [0049]; “Pairs of cameras can be selected that have at least 30 corresponding three-dimensional points in a Structure-from-Motion reconstruction. For instance, for each pair, a camera pose and depth map can be used to un-project image pixels into a dense three-dimensional model” in [0052]; “wherein estimating the camera pose further comprises: matching two-dimensional points of the candidate renders in relation to a three-dimensional model using rendered camera parameters and a depth map related to the candidate renders” in claim 14.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Song with the teaching of Lukac so as to estimate camera pose comprising matching 2D points in relation to a 3D model using camera parameters and a depth map.
Claims 11 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Song et al. (US 2026/0065425 A1) in view of Roessle et al. (GANeRF: Leveraging Discriminators to Optimize Neural Radiance Fields, cited in IDS).
As to Claim 11, the modified Song teaches The method of claim 10. The combination of Roessle further teaches wherein the view synthesis networks generate views by casting rays from a camera center in a view direction and sampling points along the rays (Roessle,
PNG
media_image1.png
328
726
media_image1.png
Greyscale
)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Song with the teaching of Roessle so that NeRF models represent a scene by providing a density and an RGB color for each point in 3D space, the color depending additionally on the viewing direction to account for view-dependent effects.
As to Claim 13, the modified Song teaches The method of claim 1. The combination of Roessle further teaches comprising calculating pixel values of the view by integrating colors and densities of sampled points (Roessle,
PNG
media_image1.png
328
726
media_image1.png
Greyscale
)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Song with the teaching of Roessle so that NeRF models represent a scene by providing a density and an RGB color for each point in 3D space, the color depending additionally on the viewing direction to account for view-dependent effects.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIMING HE whose telephone number is (571)270-1221. The examiner can normally be reached on Monday-Friday, 8:30am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached on 571-272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WEIMING HE/
Primary Examiner, Art Unit 2611