Prosecution Insights
Last updated: October 02, 2026
Application No. 18/709,218

High Dynamic Range View Synthesis from Noisy Raw Images

Final Rejection §103
Filed
May 10, 2024
Priority
Nov 15, 2021 — provisional 63/279,363 +1 more
Examiner
BAYNES, SAMUEL DAVID
Art Unit
2665
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Final)
90%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
9 granted / 10 resolved
+28.0% vs TC avg
Strong +17% interview lift
Without
With
+16.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
16 currently pending
Career history
22
Total Applications
across all art units

Statute-Specific Performance

§101
10.9%
-29.1% vs TC avg
§103
59.4%
+19.4% vs TC avg
§102
6.9%
-33.1% vs TC avg
§112
18.8%
-21.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 10 resolved cases

Office Action

§103
DETAILED ACTION The amendment filed on 06/26/2026 has been entered for consideration. Applicant amended claims 1, 6-7, 10, and 14-17, and canceled claim 20. Claims 1-19 remain pending. The new grounds of rejection in the office action have been necessitated by applicants’ amendments. Therefore, this action is made FINAL. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 07/28/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Amendment The amendment filed on 06/26/2026 has been entered. Applicant amended claims 1, 6-7, 10, and 14-17, and canceled claim 20. Claims 1-19 remain pending. Applicant’s amendments to the Claims have overcome each and every Objection and 112(b) rejection previously set forth in the Non-Final Office Action mailed on 03/30/2026. Accordingly, the Claim Objections and 112(b) Rejections has been withdrawn. Amendment to Claim 7 remedies the Examiner’s inability to perform prior art search. Support for the amendments can be found in at least paragraphs 10, 26, 30, 39- 40, 56, 81 - 82, 102, 104, 108, 124- 126, 133, 137, 149, & 168 and FIGS. 3 - 5 and 8 of the written description and drawings as originally filed. Claim 17 was amended to include the subject matter of claim 20, which the Examiner previously indicated as allowable subject matter. Accordingly, the 35 U.S.C. 103 rejection for claim 17 is withdrawn, and the 35 U.S.C. 103 rejections for claims 18-19 are also withdrawn by virtue of their dependency on independent claim 17. Response to Arguments Applicant’s arguments, see pages 8-11 of the applicant’s remarks, filed 06/26/2026 and consistent with the arguments expressed during the course of an interview with the Examiner held on 05/28/2026, with respect to claim 1 as amended to narrow the scope of the view rendering as “wherein the view rendering is a rendering of a high dynamic range view, and descriptive of one or more predicted color values and one or more predicted volume density values," have been fully considered and are persuasive. The rejection under 35 USC 103 of claim 1 as amended has been withdrawn. Regarding claim 1 as amended, Applicant argues that the prior art references of Mildenhall (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”) and Lim (US 20130322752 Al) fail to teach or render obvious the amended limitation requiring the view rendering of a high dynamic range (HDR) view. Specifically, Applicant argues that (1) Lim is directed to image processing and noise reduction rather than neural radiance field based view synthesis; (2) paragraph [0247] of Lim merely concerns “stitching images based on feature detection configured to detect the locations of corners of objects in the image frame”; therefore, Lim fails to teach rendering, let alone "the view rendering is a rendering of a high dynamic range view"; (3) the cited Lim passages paragraph [0218], which states that acquired image data may undergo significant processing before appearing as a finished image, paragraph [0353] of Lim, which discusses raw image data being received by sensors and the potential nonlinearity, and the previously discussed paragraph [0247] fail to teach view synthesis rendering of images and is instead focused on reducing noise in an image using a sparse filter and a noise threshold, as further reflected in the Abstract. Therefore, Applicant argues the Lim reference fails to teach view rendering, let alone “view rendering is a rendering of a high dynamic range view” as recited in amended claim 1, and subsequently, it does not cure the deficiencies of Mildenhall. Applicant’s arguments are found to be persuasive with respect to the amened limitation. Although paragraph [0247] of Lim is not limited to panoramic stitching and expressly identifies certain high dynamic range (HDR) imaging algorithms in addition to stitching, Lim’s HDR teachings concern processing or combining captured image data rather than generating an HDR view using a neural radiance field. Mildenhall teaches neural radiance field view-based synthesis, while Lim teaches raw, noisy, and HDR-related image processing. However, the cited teachings do not sufficiently teach or suggest modifying Mildenhall such that the view rendering generated by the neural radiance field itself is a rendering of a high dynamic range view. Accordingly, the combination does not adequately address the newly amended limitation. Applicant also argues that a person of ordinary skill in the art would not have sought out Lim to cure the deficiencies of Mildenhall because Lim is directed to a different technical problem and has a different solution, namely image processing, feature detection, stitching, and noise reduction rather than neural radiance field based view rendering. Applicant’s arguments that a person of ordinary skill in the art would not seek out Lim to cure the deficiencies of Mildenhall have been fully considered but they are not persuasive in their entirety. A prior art reference is analogous if it is either in the field of the inventor’s endeavor or, if not, then be reasonably pertinent to the particular problem with which the inventor was concerned, in order to be relied upon as a basis for rejection of the claimed invention. See In re Oetiker, 977 F.2d 1443, 24 USPQ2d 1443 (Fed. Cir. 1992); MPEP § 2141.01(a). Mildenhall, Lim, and the instant application broadly concern digital image processing. Further, Lim teaches the acquisition and processing of digital image data, including raw and noisy sensor image data, sensor linearization, and HDR image processing (Lim ¶ [0218], ¶ [0247], ¶ [0353]). These teachings are reasonably pertinent to Mildenhall’s use of captured image data to train a neural radiance field and compare rendered image values with corresponding observed image values. Thus, the fact that Lim addresses a different specific image processing problem does not establish that its teachings would not have reasonably been considered by a person of ordinary skill in the art Nevertheless, Applicant’s arguments are persuasive with respect to the amended limitation requiring the view rendering to be a rendering of a high dynamic range view. Although Lim is properly considered analogous art, the cited teachings of Mildenhall and Lim do not sufficiently teach or suggest modifying Mildenhall such that the neural radiance field model would generate a rendering of a high dynamic range view in the manner claimed. Accordingly, the claim 1 rejection under 35 U.S.C. 103 based on Mildenhall in view of Lim is withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view newly found prior art, namely Onzon et al. (US 20220269910 A1) and “ADOP: Approximate Differentiable One-Pixel Point Rendering” by Rückert et al., while remaining to primarily rely on the teachings of Mildenhall ("NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis"), as seen below. Although Examiner maintains Lim is analogous art, Examiner has withdrawn Lim as a reference in the new grounds of rejection. In summary, Rückert teaches a differentiable neural rendering system that generates novel views from calibrated camera images and scene geometry and generates high dynamic range (HDR) output while accounting for exposure and camera response, and Onzon teaches using raw linear HDR sensor images in the training pipeline, including sensor-noisy image data. Thus, these two prior art references address the discrepancies in the teachings of Mildenhall and further establish an obviousness 35 U.S.C. 103 rejection for the respective amended claims, as seen below. Claim 7, which was previously unable to be interpreted and examined for prior art due to the 35 U.S.C. 112(b) issue cited in the Office Action mailed on 03/30/2026. Claim 7 is now rejected under 35 U.S.C. 103 as being obvious under Mildenhall in view of Onzon et al., in further view of Rückert et al., and in further view of newly cited Lehtinen et al. (“Noise2Noise: Learning Image Restoration without Clean Data”; copy provided by examiner) prior art reference, as seen in the Claims Rejection section below. Additionally, "Neural Camera Simulators” by Ouyang et al. is introduced to teach discrepancies of Mildenhall, Onzon, and Rückert where Lim was previously relied upon in claims 6 and 9. Please refer to the Claim Rejections section below to see the new ground(s) for rejection. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Examiner notes, Onzon et al. (US 20220269910 A1) qualifies as prior art under 35 U.S.C. 102(a)(2) because it claims benefit of provisional application no. 63/175,505 and its filling date of 04/15/2021. See MPEP 2154.01(b). Claim(s) 1-2, 4-5, and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Mildenhall (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”; examiner relied on a more easily readable copy of the article than provided by applicant, updated version of reference provided by examiner Office Action mailed on 03/30/2026) in view of Onzon et al. (US 20220269910 A1) and Rückert et al. (“ADOP: Approximate Differentiable One-Pixel Point Rendering”; copy provided by examiner). Regarding Claim 1, Mildenhall teaches: (Original) A computing system, the system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising (Mildenhall teaches a system that can use a NVIDIA V100 GPU (Abstract; p. 8, section 5.3, last paragraph; p. 17-18 Annex A). NVIDIA V100 GPU is known by one of ordinary skill to comprise one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations.): obtaining a training dataset, wherein the training dataset comprises a plurality of three-dimensional positions, a plurality of two-dimensional view directions, and a plurality of (Abstract "…input is a single continuous 5D coordinate (spatial location (x,y,z) and viewing direction (θ,φ))"; see FIG. 1-3 and their corresponding descriptions; see p. 14. section 7, first paragraph; Mildenhall further teaches using pixels representing images throughout the model, including in the training data set and implementation (e.g. p. 6, section 4, paragraph 1 "Rendering a view from our continuous neural radiance field requires estimating this integral C(r) for a camera ray traced through each pixel of the desired virtual camera.", and p. 9, section 5.3, paragraph 1 and p. 9, section 6, paragraph 1, teach the implementation and use of synthetic renderings of objects using pixels). Under the broadest reasonable interpretation of the claim and to a person of ordinary skill in the art, the use of pixels representative of an image constitutes the use of digital images, i.e. "a plurality of bits structured in a format".); processing a first three-dimensional position of the plurality of three-dimensional positions and a first two-dimensional view direction of the plurality of two-dimensional view directions with a neural radiance field model to generate a view rendering (See p. 5, FIG. 2 (found below)), PNG media_image1.png 574 930 media_image1.png Greyscale wherein the neural radiance field model comprises one or more multi-layer perceptrons (p. 1, section 1, second paragraph "Our method optimizes a deep fully-connected neural network without any convolutional layers (often referred to as a multilayer perceptron or MLP) to represent this function by regressing from a single 5D coordinate (x,y,z,θ,φ) to a single volume density and view-dependent RGB color."; p. 2, section 1, bullet point from last paragraph "An approach for representing continuous scenes with complex geometry and materials as 5D neural radiance fields, parameterized as basic MLP networks."), and wherein the view rendering is (FIG. 2 and description (seen above); p. 5, section 3, paragraph 2 (seen below)); PNG media_image2.png 326 936 media_image2.png Greyscale evaluating a loss function that evaluates a difference between the view rendering and a first image of the plurality of (p. 2, section 1, last paragraph "…our technical contributions are…A differentiable rendering procedure based on classical volume rendering techniques, which we use to optimize these representations from standard RGB images. This includes a hierarchical sampling strategy to allocate the MLP’s capacity towards space with visible scene content."; P. 9, section 5.3, paragraph 1 "Our loss is simply the total squared error between the rendered and true pixel colors for both the coarse and fine renderings" and equation 6 includes loss function to measure the difference between the rendered pixel color and the ground truth pixel colors (i.e. minimized error).), wherein the first image is associated with at least one of the first three-dimensional position or the first two- dimensional view direction (p. 9, section 5.3, paragraph 1 "…a dataset of captured RGB images of the scene, the corresponding camera poses and intrinsic parameters, and scene bounds (we use ground truth camera poses, intrinsics, and bounds for synthetic data, and use the COLMAP structure-from-motion package [39] to estimate these parameters for real data)."); and adjusting one or more parameters of the neural radiance field model based at least in part on the loss function (Mildenhall teaches a neural network parameterized by a fully connected MLP (p. 2, section 1, paragraph 5 "…representing continuous scenes with complex geometry and materials as 5D neural radiance fields, parameterized as basic MLP networks"). Such MLP are understood in the art to include learnable parameters (e.g., weights and biases). Mildenhall further teaches, defining and minimizing an error between rendered and ground truth pixel colors as a loss function (p. 2, section 1, first paragraph "…we can use gradient descent to optimize this model by minimizing the error between each observed image and the corresponding views rendered from our representation."; see also p. 9, section 5.3, paragraph 1; p.5-6, section 4, paragraphs 1-2), and optimizing the network parameters over 100-300k iterations using an optimizer (p. 9, section 5.3, paragraph 2), thereby iteratively adjusting one or more parameters of the neural radiance field model based at least in part on the loss function.). Mildenhall fails to explicitly disclose: (1) raw noisy images, wherein the plurality of raw noisy images comprise a plurality of high dynamic range images comprising a plurality of unprocessed bits structured in a raw format (2) wherein the view rendering is a rendering of a high dynamic range view. In a related art, Onzon teaches (1) forming a training dataset using raw linear HDR images output by an HDR image sensor ([0197]-[0198] “Real life data is collected by the HDR sensor, and the HDR sensor can output: a single raw linear HDR image, for example by setting the HDR sensor to produce a linear HDR image when collecting HDR data for the training dataset”; [0204] “Either the raw linear HDR image or corresponding “n” raw linear LDR images may be used to create the training dataset.”; [0206] Either way, the linear HDR image for the training dataset is formed, either as a direct output of the HDR sensor, or as a fusion of n LDR images outputted by the HDR sensor.), wherein the raw sensor image data comprise digitized pre-ISP pixel values represented in a raw format (see [0014], [0119]-[0122]). Onzon further teaches that the HDR training-set images contain sensor noise ([0229], [0235]). However, Onzon does not directly train its neural network using the raw noisy HDR images. Rather, Onzon derives simulated raw LDR images from the HDR images and trains its auto-exposure network using the simulated LDR images ([0214]-[0219], [0227]-[0230]). Thus, Onzon does not teach evaluating a NeRF rendering loss directly against the raw noisy HDR training image as claimed. In a related art, Rückert teaches (2) a differentiable neural novel view rendering pipeline in which a deep neural renderer reconstructs an HDR image, which may subsequently be converted to LDR by a differentiable tone mapping/camera model (see p. 3 section 3, including third paragraph of left column “The neural renderer (Figure 2 middle) takes the multi resolution neural image to produce a single HDR output image.”, and Figure 2, as seen below). Thus, Rückert expressly teaches that its neural rendering system handles varying exposure and generates high dynamic range output. PNG media_image3.png 251 320 media_image3.png Greyscale However, Rückert does not teach training its neural renderer directly using the claimed raw noisy HDR images comprising unprocessed bits structured in a raw format. Mildenhall, Onzon, and Rückert collectively fail to expressly teach training Mildenhall’s neural radiance field by evaluating the rendering loss directly against the raw noisy HDR images themselves. Mildenhall trains using standard RGB images, while Onzon forms a raw linear HDR training dataset but derives simulated raw LDR images from the HDR training dataset for actual neural network training. The references are analogous art because they concern neural processing of image data and are reasonably pertinent to the representation and use of image data for learned image generation. See MPEP § 2141.01(a). In particular, Onzon addresses raw/HDR training image representation, while Rückert concerns neural novel view synthesis and expressly discusses and NeRF and Mildenhall’s work in its Related Work section (p. 2, right column sub section C). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to use Onzon’s raw linear HDR training images as Mildenhall’s observed training images, rather than reducing the HDR information to LDR, to retain the greater dynamic range information available in the HDR sensor data and avoid unnecessary HDR/LDR conversion, particularly because Onzon expressly teaches storing the training dataset directly as linear HDR images to avoid generating/converting HDR imagery during training and the associated processing time required to do so (Onzon [0201]-[0202]). It would further have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to render the resulting scene representation in HDR as taught by Rückert, which applies HDR representation to neural novel view synthesis and teaches that accounting for camera exposure and response characteristics enables high dynamic range output (see Rückert, Abstract; Fig. 1; Section 1 Introduction). Regarding Claim 2, Mildenhall, Onzon, and Rückert teach the computing system of claim 1. Mildenhall further teaches: wherein the operations further comprise: processing the view rendering with a color correction model to generate a color corrected rendering (See p. 1-2, paragraph 2). Regarding Claim 4, Mildenhall, Onzon, and Rückert teach the computing system of claim 1, including: evaluating the loss function that evaluates the difference between the view rendering and the first image of the plurality of raw noisy images Mildenhall further teaches: mosaic masking being comprised in evaluating the loss function by sampling pixels from the dataset of a frame and rendering true pixel colors, thereby altering pixels of a portion of a frame, a known form of “mosaic masking”) (Mildenhall p. 9, section 5.3; Mildenhall teaches the system uses a NVIDIA V100 that uses frames when rendering p. 18, section A, subsection Rendering Details). Regarding Claim 5 Mildenhall, Onzon, and Rückert teach the computing system of claim 1, including: evaluating the loss function that evaluates the difference between the view rendering and the first image of the plurality of raw noisy images The system taught by Mildenhall, Onzon, and Rückert in claim 1 fails to explicitly disclose exposure adjustments being comprised in the evaluation of the loss function. However, Onzon further teaches: adjusting the exposure of training image data by scaling a linear HDR image according to a base exposure and a randomly varied exposure shift during training (Onzon, [0219]-[0221]; see also [0224]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to further apply the exposure adjustment teachings of Onzon to the image data used in the combined Mildenhall, Onzon, and Rückert system’s loss evaluation to account for varying exposure conditions and provide a more consistent comparison between rendered observed image data, thereby improving training across differently exposed images. Regarding Claim 8, Mildenhall, Onzon, and Rückert teach the computing system of claim 1. Mildenhall further teaches: wherein the first image comprises a real-world photon signal data generated by a camera, and wherein the view rendering comprises predicted photon signal data (see FIG. 2’s description “We synthesize images by sampling 5D coordinates (location and viewing direction) along camera rays (a), feeding those locations into an MLP to produce a color and volume density (b), and using volume rendering techniques to composite these values into an image (c). This rendering function is differentiable, so we can optimize our scene representation by mini-mizing the residual between synthesized and ground truth observed images.”. Camera rays are known in the art to be equivalent to photons). Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Mildenhall (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”) in view of Onzon et al. (US 20220269910 A1), in further view of Rückert et al. (“ADOP: Approximate Differentiable One-Pixel Point Rendering”), and in further view of Li (“A reweighted L2 method for image restoration with Poisson and mixed Poisson-Gaussian noise”; copy provided by examiner with Office Action mailed on 03/30/2026). Regarding Claim 3, Mildenhall, Onzon, and Rückert teach the computing system of claim 1, including the loss function. Mildenhall, Onzon, and Rückert fail to explicitly disclose: wherein the loss function comprises a reweighted L2 loss. In a related art, Li teaches: a reweighted L2 loss, referred to as reweighted L2 “fidelity”, for noise related image restoration (Abstract) that iteratively estimates noise variance (p. 2, section 1, 2nd to last paragraph). It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon and Rückert, to incorporate the teachings of Li in order to more efficiently estimate noise variance. The teachings lie in the same field of endeavor as the instant application of image processing with an aim at improving image quality. Claims 6 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Mildenhall (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”) in view of Onzon et al. (US 20220269910 A1), in further view of Rückert et al. (“ADOP: Approximate Differentiable One-Pixel Point Rendering”), and in further view of Ouyang et al. (“Neural Camera Simulators”; copy provided by examiner). Regarding Claim 6, Mildenhall, Onzon, and Rückert teach the computing system of claim 1. Mildenhall further teaches: wherein the operations further comprising: obtaining an input view direction and an input position (Abstract "..input is a single continuous 5D coordinate (spatial location (x,y,z) and viewing direction (θ,φ))"); processing the input view direction and the input position with the neural radiance field model to generate predicted (p. 1-2, paragraph 2); and processing the predicted (Abstract "We describe how to effectively optimize neural radiance fields to render photorealistic novel views"; FIG. 1 and caption "we show two novel views rendered from our optimized NeRF representation." Under the broadest reasonable interpretation of the claim, examiner interprets the novel view output by Mildenhall’s NeRF model to reasonable correspond to the claimed output.). Mildenhall, Onzon, and Rückert fail to explicitly disclose: generating predicted quad bayer filter data for use in subsequent generation of output view rendering. In a related art, Ouyang teaches: unpacking Bayer raw image data into four RGGB channels for neural processing (p. 5 right column section 4.2), thereby providing separate red, first-green, second-green, and blue raw filter data. Ouyang is reasonably related to the other references because each concerns neural processing of camera image data, and Ouyang specifically teaches representing and processing raw sensor image data for use by a neural image processing model. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon, and Rückert, to further incorporate the teachings of Ouyang by unpacking Bayer raw image data into four RGGB channels, because Onzon already teaches Bayer pattern sampling of raw image data (Onzon [0040], [0227]) and Ouyang teaches that separating the Bayer pattern into RGGB channels facilitates neural processing of the raw sensor data (Ouyang p. 5 right column section 4.2). Regarding Claim 9, Mildenhall, Onzon, and Rückert teach the computing system of claim 1, including the plurality of raw noisy images. Mildenhall further teaches: (Currently Amended) The computing system of claim 1, wherein the plurality of (p. 8-9, section 5.3, paragraph 1 "…a dataset of captured RGB images of the scene, the corresponding camera poses and intrinsic parameters, and scene bound"). Mildenhall fails to explicitly disclose: raw noisy images and using red-green-green-blue datasets. However, Mildenhall, Onzon, and Rückert teach the plurality of raw noisy images in claim 1. In a related art, Ouyang teaches using RGGB datasets for neural processing of raw sensor data (p. 5 right column section 4.2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon and Rückert, to incorporate the teachings of Ouyang to more accurately model the raw image data captured by the camera sensor. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Mildenhall (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”) in view of Onzon et al. (US 20220269910 A1), in further view of Rückert et al. (“ADOP: Approximate Differentiable One-Pixel Point Rendering”), and in further view of Lehtinen et al. (“Noise2Noise: Learning Image Restoration without Clean Data”; copy provided by examiner). Regarding claim 7, Mildenhall, Onzon, and Rückert teach the computing system of claim 1, including a loss function. Mildenhall, Onzon, and Rückert fail to explicitly disclose: wherein the loss function comprises a stop gradient. In a related art, Lehtinen teaches: an HDR loss in which the gradient of denominator is treated as zero during optimization (p. 6 “High dynamic range (HDR)” subsection found on left and right column), thereby preventing gradients from propagating through that portion of the loss. Thus, treating the denominator gradient as zero corresponds to the claimed “stop gradient,” because that term is excluded from contributing gradient information during backpropagation. Lehtinen is related to Mildenhall, Onzon, and Rückert because each concerns neural processing of image data, particularly under noisy or HDR imaging conditions. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined neural rendering system taught by Mildenhall, Onzon, and Rückert to incorporate Lehtinen’s stop gradient loss technique to prevent the loss normalization term from influencing parameter updates and thereby improve stability when training with HDR image data. Claim(s) 10-16 are rejected under 35 U.S.C. 103 as being unpatentable over Mildenhall (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”) in view of Onzon et al. (US 20220269910 A1), in further view of Rückert et al. (“ADOP: Approximate Differentiable One-Pixel Point Rendering”), and in further view of Chen (“Learning to See in the Dark”; copy provided by examiner with Office Action mailed on 03/30/2026). Regarding Claim 10, Mildenhall teaches: (Original) A computer-implemented method for view rendering (Abstract "We present a method that achieves state-of-the-art results for synthesizing novel views of complex scenes…"; p. 17-18 Annex A), the method comprising: obtaining, by a computing system comprising one or more processors (Mildenhall teaches a system that can use a NVIDIA V100 GPU (Abstract; p. 8, section 5.3, last paragraph; p. 17-18 Annex A). NVIDIA V100 GPU is known by one of ordinary skill to comprise one or more processors that cause the computing system to perform operations.), an input two- dimensional view direction and an input three-dimensional position associated with an environment (Abstract "Our algorithm represents a scene … whose input is a single continuous 5D coordinate (spatial location (x,y,z) and viewing direction (θ,φ))…"; see FIGS. 1-3 and their corresponding descriptions) obtaining, by the computing system, a neural radiance field model, (p. 1, section 1, paragraph 2) wherein the neural radiance field model was trained on a training dataset, wherein the training dataset comprises a plurality of (Mildenhall teaches training a neural radiance field model using a set of images of a scene, including camera poses, directions, and spatial locations (Abstract, lines 1-4 and 10-12), which serve as training data for optimizing the model (p. 8-9, section 5.3).) processing, by the computing system, the input two-dimensional view direction and the input three-dimensional position with the neural radiance field model to generate prediction data, wherein the prediction data comprises one or more predicted density values and one or more predicted color values (Abstract, lines 4-10; p. 1-2, section 1, paragraph 2; FIG. 2 and p. 5, section 3, paragraphs 1-2) processing, by the computing system, the prediction data with an image (FIG. 2 and p. 4-5, section 3, paragraphs 1-3; FIG. 3 and p. 5-6, paragraphs 1-2). Mildenhall fails to explicitly disclose: (1) raw noisy images, wherein the plurality of raw noisy images comprise a plurality of high dynamic range images comprising a plurality of unprocessed bits structured in a raw format (2) wherein the view rendering is a rendering of a high dynamic range view, and (3) the use of an image augmentation block to generate predicted view rendering. In a related art, Onzon teaches (1) forming a training dataset using raw linear HDR images output by an HDR image sensor ([0197]-[0198] “Real life data is collected by the HDR sensor, and the HDR sensor can output: a single raw linear HDR image, for example by setting the HDR sensor to produce a linear HDR image when collecting HDR data for the training dataset”; [0204] “Either the raw linear HDR image or corresponding “n” raw linear LDR images may be used to create the training dataset.”; [0206] Either way, the linear HDR image for the training dataset is formed, either as a direct output of the HDR sensor, or as a fusion of n LDR images outputted by the HDR sensor.), wherein the raw sensor image data comprise digitized pre-ISP pixel values represented in a raw format (see [0014], [0119]-[0122]). Onzon further teaches that the HDR training-set images contain sensor noise ([0229], [0235]). However, Onzon does not directly train its neural network using the raw noisy HDR images. Rather, Onzon derives simulated raw LDR images from the HDR images and trains its auto-exposure network using the simulated LDR images ([0214]-[0219], [0227]-[0230]). Thus, Onzon does not teach evaluating a NeRF rendering loss directly against the raw noisy HDR training image as claimed. In a related art, Rückert teaches (2) a differentiable neural novel view rendering pipeline in which a deep neural renderer reconstructs an HDR image, which may subsequently be converted to LDR by a differentiable tone mapping/camera model (see p. 3 section 3, including third paragraph of left column “The neural renderer (Figure 2 middle) takes the multi resolution neural image to produce a single HDR output image.”, and Figure 2, as seen below). Thus, Rückert expressly teaches that its neural rendering system handles varying exposure and generates high dynamic range output. PNG media_image3.png 251 320 media_image3.png Greyscale However, Rückert does not teach training its neural renderer directly using the claimed raw noisy HDR images comprising unprocessed bits structured in a raw format. Mildenhall, Onzon, and Rückert collectively fail to expressly teach training Mildenhall’s neural radiance field by evaluating the rendering loss directly against the raw noisy HDR images themselves. Mildenhall trains using standard RGB images, while Onzon forms a raw linear HDR training dataset but derives simulated raw LDR images from the HDR training dataset for actual neural network training. The references are analogous art because they concern neural processing of image data and are reasonably pertinent to the representation and use of image data for learned image generation. See MPEP § 2141.01(a). In particular, Onzon addresses raw/HDR training image representation, while Rückert concerns neural novel view synthesis and expressly discusses and NeRF and Mildenhall’s work in its Related Work section (p. 2, right column sub section C). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to use Onzon’s raw linear HDR training images as Mildenhall’s observed training images, rather than reducing the HDR information to LDR, to retain the greater dynamic range information available in the HDR sensor data and avoid unnecessary HDR/LDR conversion, particularly because Onzon expressly teaches storing the training dataset directly as linear HDR images to avoid generating/converting HDR imagery during training and the associated processing time required to do so (Onzon [0201]-[0202]). It would further have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to render the resulting scene representation in HDR as taught by Rückert, which applies HDR representation to neural novel view synthesis and teaches that accounting for camera exposure and response characteristics enables high dynamic range output (see Rückert, Abstract; Fig. 1; Section 1 Introduction). Mildenhall, Onzon, and Rückert collectively fail to expressly teach: (3) the use of an image augmentation block to generate predicted view rendering. In a related art, Chen teaches: applying data augmentation to training images, including cropping or patch-based processing (p. 5, section 4.2, first paragraph), which constitutes an image augmentation block in the training pipeline. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon and Rückert, to incorporate the teachings of Chen in order to increase the effectiveness and efficiency of training the model and predicting view rendering when a variety of unknown types of images (e.g. raw images) are provided to the model. All references and the instant application all lie in the same field of endeavor of image processing with a specific aim of improving image quality, specifically relating to colors. Also, Mildenhall, Chen, and the instant application use “Adam optimizers” for training the network (Mildenhall p. 9, section 5.3, last paragraph; Chen, p. 5, section 4.2, first paragraph). Regarding Claim 11, Mildenhall, Onzon, Rückert, and Chen teach the computer-implemented method of claim 10, including processing the prediction data with an augmentation block to generate predicted view rendering, descriptive of a predicted scene. Mildenhall, Onzon, and Rückert fail to explicitly teach: wherein the image augmentation block adjusts a focus of the prediction data. However, Chen and Onzon each teach: well-known image adjustment techniques for focusing. For example, Chen teaches camera settings like focus, and focal length can be adjusted to maximize the quality of images (Chen, p. 3, section 3, paragraph 5) and teaches there are a variety of deblurring techniques found in prior art (Chen, Abstract), while Onzon teaches altering a path of light rays from a scene to be captured in order to focus a captured scene and produce a raw LDR image (Onzon FIG. 3A-1, [0145], and [0160]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon, Rückert, and Chen, to incorporate the further focus techniques of Chen and/or Onzon to increase the accuracy of the model and corresponding predicted view rendering by adjusting focus of the prediction data throughout the training pipeline, specifically by use of the modified image augmentation block taught collectively by Mildenhall, Onzon, Rückert, and Chen in claim 10. Regarding Claim 12, Mildenhall, Onzon, Rückert, and Chen teach the computer-implemented method of claim 10, including processing the prediction data with an augmentation block to generate predicted view rendering, descriptive of a predicted scene. The method taught by Mildenhall, Onzon, Rückert, and Chen, in claim 10, fails to explicitly disclose wherein the image augmentation block adjusts an exposure level of the prediction data. However, Onzon further teaches: adjusting an exposure level of image data by scaling the image according to a base exposure and an exposure shift, including an exposure value predicted by a neural network ([0219]-[0221], [0224]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon, Rückert, and Chen, to further apply Onzon’s exposure adjustment techniques to the combined system and augmentation block taught by Mioldenhall, Onzon, Rückert, and Chen to account for varying image exposure levels and improve consistency of the image data used by the model. Regarding Claim 13, Mildenhall, Onzon, Rückert, and Chen teach the computer-implemented method of claim 10, including processing the prediction data with an augmentation block to generate predicted view rendering, descriptive of a predicted scene. The method taught by Mildenhall, Onzon, Rückert, and Chen, in claim 10, fails to explicitly teach: wherein the image augmentation block adjusts a tone-mapping of the prediction data. However, Rückert further teaches: adjusting tone-mapping of neural-rendered image data using a learnable differentiable tone-mapping operator applied to the HDR output of the neural renderer (see p. 3 FIG. 2, and p. 6 section “V. Differentiable Tone Mapping”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon, Rückert, and Chen, to further apply Rückert’s learnable tone-mapping operation to the combined system and image augmentation block previously taught by Mildenhall, Onzon, Rückert, and Chen in order to convert the HDR rendered output into visually usable display representation. Regarding Claim 14, Mildenhall, Onzon, Rückert, and Chen teach the computer-implemented method of claim 10, including a plurality of noisy input dataset. Mildenhall further teaches: (Mildenhall teaches sampling camera rays, which are known in the art to be photons (see FIG. 2’s description)). While Mildenhall doesn’t by itself explicitly disclose: wherein each noisy input dataset the plurality of noisy input datasets, Mildenhall, as previously modified by Onzon, Rückert, and Chen, teach these principles in claim 10. Refer back to claim 10 for further details. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Mildenhall, previously modified by Onzon, Rückert, and Chen, to incorporate the further focus techniques of Mildenhall in order to more accurately capture location and viewing direction. Regarding Claim 15, Mildenhall, Onzon, Rückert, and Chen teach the computer-implemented method of claim 10, including a plurality of noisy input dataset. Mildenhall further teaches: wherein each (Examiner notes, signal data is known in the art as data related to the visual representation of an image (e.g. pixel color represented at a location). Mildenhall teaches using an input dataset of a plurality input datasets as a "sparse set of input views" (Abstract). Mildenhall further teaches input data sets comprise of data associated with RGB images (i.e. "data associated with at least one of a red value, a green value, or a blue value") (p. 8, section 5.3 "We optimize a separate neural continuous volume representation network for each scene. This requires only a dataset of captured RGB images of the scene…"; p. 2, section 1, last paragraph, lines 9-21); thus, Mildenhall also teaches the use of signal data.) Mildenhall fails to explicitly disclose: noisy input dataset. However, Mildenhall, as previously modified by Onzon, Rückert, and Chen (in claim 10), does teach a noisy input dataset. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate the teachings of input datasets associated with color values, discussed in the previous paragraph of the present office action and taught by Mildenhall, to the teachings of Mildenhall, previously modified by Onzon, Rückert, and Chen, to increase the accuracy of the system by accounting for color values. Regarding Claim 16, Mildenhall, Onzon, Rückert, and Chen teach the computer-implemented method of claim 10, including a plurality of noisy input datasets. Onzon further teaches: mosaicked raw image data by Bayer pattern sampling a linear HDR image to form a simulated raw image ([0227]), and adding sensor noise to the simulate raw image ([0229]), and subsequently demosaicing the raw image in the ISP ([0180]; see also [0243]). Thus, Mildenhall, as previously modified by Onzon, Rückert, and Chen teach and/or render obvious wherein each noisy input dataset of the plurality of noisy input datasets (detailed in claim 10) comprises one or more noisy mosaicked linear raw images. Allowable Subject Matter Claims 17-19 are allowed. None of the prior art either alone or in combination disclose, teach or suggest the operations recited in claim 17. Accordingly, claims 18-19 are allowed by virtue of their dependency on claim 17. Prior Art Not Relied Upon The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure, are: Kalantari (US 11094043 B2) which teaches systems and methods for generating high dynamic range (HDR) images using a neural network, including capturing multiple images at different exposure levels, processing image data captured in RAW format, while accounting for noise in low-exposure input images, and training the neural network using corresponding HDR ground-truth images to generate reconstructed HDR images as output. Çoğalan et al. (“Deep Joint Deinterlacing and Denoising for Single Shot Dual-IOS HDR Reconstruction”; copy provided by examiner) which teaches neural network based HDR image reconstruction from noisy RAW dual-IOS Bayer image data, wherein trained deep convolutional neural networks process the RAW image data prior to demosaicking to reconstruct low and high exposure images that are subsequently merged to generate low noise HDR images. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMUEL DAVID BAYNES whose telephone number is (571)272-0607. The examiner can normally be reached Monday - Friday 8:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen R Koziol can be reached at (408)918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SDB/ Samuel D. Baynes Examiner, Art Unit 2665 /Stephen R Koziol/Supervisory Patent Examiner, Art Unit 2665
Read full office action

Prosecution Timeline

May 10, 2024
Application Filed
Mar 31, 2026
Non-Final Rejection mailed — §103
May 28, 2026
Examiner Interview Summary
May 28, 2026
Applicant Interview (Telephonic)
Jun 26, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12685438
METHOD AND APPARATUS FOR DETECTING PENETRATION DEPTH OF RIBOFLAVIN IN CORNEA
2y 2m to grant Granted Jul 21, 2026
Patent 12688557
EXTENDED U-NET FOR MULTI-INFORMATION EXTRACTION AND APPLICATION METHOD THEREOF IN LOW-DOSE X-RAY IMAGING
1y 11m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
90%
Grant Probability
99%
With Interview (+16.7%)
2y 5m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 10 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month