DETAILED ACTION
Response to Amendment
Applicant’s amendments filed on 27 January 2026 have been entered. Claims 1, 10, 13, 14, 16, 18 and 19 have been amended. Claims 1-20 are still pending in this application, with claims 1, 10 and 16 being independent.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 4, and 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over in Mehta et al. (US 20220292341 A1), referred herein as Mehta view of Rematas et al. (US 20230281913 A1), referred herein as Rematas, Martin et al. (US 20240005590 A1), referred herein as Martin and Kim et al. (US 20240185505 A1), referred herein as Kim.
Regarding Claim 1, Mehta in view of Rematas and Martin teaches a method, comprising:
obtaining a dataset for training a neural Mehta [0008] training the modulator network and the synthesizer network based on the loss function; [0082] the system includes a neural network to receive at least a visual signal and encode visual signals (e.g., images, video, shapes, and 3D scenes) which are implicit representations for signals that are encoded in weights of the neural network; [0144] In some embodiments, training data includes pairs of {(x.sub.j, y.sub.i), y.sub.ij}, where x.sub.j are image coordinates, y.sub.i s the lighting direction corresponding to image i and y.sub.ij is the radiance at x.sub.j in image I; [0123] reconstruction from sparse samples (e.g., light-fields, compression); FIG.7: 700);
Mehta teaches a neural network, but does not explicitly teaches training a neural radiance field (NeRF). However, Rematas teaches
teaches training a neural radiance field (NeRF) (Rematas [0096] FIG. 9 includes renderings from a DS-NeRF model, a Mip-NeRF model). Mehta in view of Tancik further teaches
inputting, to the NeRF during a first training iteration, a first positional encoding of the dataset (Mehta [0005] The modulator network generates modulation parameters to modulate amplitude, phase, and/or frequency of intermediate layers of the synthesizer network; [0120] during each training iteration, the output of the model is compared to the known annotation information in the training data; [0073] network 415 can adjust the frequencies, phase, and/or amplitudes of the sinusoid (e.g., sine) in the synthesizer network 425 to match a target signal; [0145] ReLU-MLP with positional encoding (PE) which uses basic conditioning-by-concatenation) are implemented for comparison, where [x.sub.j, y.sub.i] is the input),
the first positional encoding associated with a first length of a visible frequency Mehta [0072] the synthesizer network 425 operates at the tile level and includes K hidden layers having hidden features h.sub.1, . . . , h.sub.K; Rematas [0006] the pre-trained semantic segmentation model can determine one or more pixels are associated with a sky in the environment and masks the one or more pixels); and
Mehta teaches a frequency spectrum (see [0094]), but does not teach a frequency band. Martin teaches a visible frequency band (Martin [0048] The hyper-parameter m controls the number of frequency bands (and therefore the highest frequency) used in the mapping).
Mehta in view of Rematas and Martin further teaches
inputting, to the NeRF during a second training iteration, a second positional encoding of the dataset, the second positional encoding associated with a second length of the visible frequency band based at least on applying a second frequency mask to the second positional encoding (Mehta [0120] during each training iteration, the output of the model is compared to the known annotation information in the training data… the parameters of the model are updated accordingly, and a new set of predictions are made during the next iteration; [0073] network 415 can adjust the frequencies, phase, and/or amplitudes of the sinusoid (e.g., sine) in the synthesizer network 425 to match a target signal; [0072] the synthesizer network 425 operates at the tile level and includes K hidden layers having hidden features h.sub.1, . . . , h.sub.K; Rematas [0129] often tied to semantics, the system can run a pre-trained semantic segmentation model on every image, and can then mask pixels of people; Martin introducing a parameter α that windows the frequency bands of the positional encoding).
In view of Kim, the prior art further teaches
the second positional encoding including at least one signal that was excluded from the first positional encoding (Kim [0038] The light field network F.sub.θ 306 then takes as input positionally encoded 4D ray coordinates γ(r) and outputs color c (integrated radiance) 308 along each ray; [0041] After embedding, positional encoding 316 is applied to the N-dimensional feature produced by the embedding network 314 and then the positionally encoded N-dimensional feature is feed into the light field MLP F.sub.θ 318; [0045] A positional encoding 336 is then applied to the re-parameterized or transformed ray coordinates. The positionally encoded ray coordinates are then passed into the light field network F.sub.θ 338 to produce a color 340 for the ray). The output of embedding network 314 or 334 is a unique signal that was excluded from the positional encoding 304.
Rematas discloses a radiance fields for three-dimensional reconstruction and novel view synthesis in large-scale environments. Rematas is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mehta to incorporate the teachings of Rematas, and apply the NeRF model of Rematas to the local neural implicit functions, as taught by Mehta, and pre-trained semantic segmentation model with making pixel of Rematas to generate replace the hidden layers having hidden features, as taught by Mehta.
Doing so would be able to leverage a plurality of panoramic images and corresponding lidar data to accurately learn a large outdoor environment to then generate view synthesis outputs and three-dimensional reconstruction outputs.
Martin discloses techniques of image synthesis using a neural radiance field (NeRF) includes generating a deformation model of movement experienced by a subject in a non-rigidly deforming scene. Martin is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mehta to incorporate the teachings of Martin, and apply the method of training iteration includes a particular frequency into the modulator network including multiple multi-layer perceptron (MLP) layers with rectified linear unit (ReLU) activation.
Doing so would minimize the error between each observed image and the corresponding views rendered from our representation.
Kim discloses a computing system may access a set of training images for training a neural light field network for view synthesis. Kim is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mehta to incorporate the teachings of Kim, and apply the positional encodings with feature-based or transformation-based embedding into the modulator network including multiple multi-layer perceptron (MLP) layers with rectified linear unit (ReLU) activation.
Doing so would provide an improved method or technique for view synthesis that may render an image with far fewer network evaluations than existing view synthesis approaches (e.g., NeRF-based approaches), while still maintaining a small memory footprint, and also be able to better represent complex light matter interactions or view dependent effects, such as light reflections and refractions.
Regarding Claim 3, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1, and further teaches wherein the first length of the visible frequency band is at least one of shorter or associated with a lower frequency than the second length of the visible frequency band, and wherein the first training iteration occurring before the second training iteration (Martin [0023] generating the deformation model includes applying a positional encoding to a position coordinate within the scene to produce a periodic function of position, the periodic function having a frequency that increases with training iteration for the MLP).
Regarding Claim 4, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1, and further teaches further comprising inputting, to the NeRF during a third training iteration that is subsequent to both the first training iteration and the second training iteration, a third positional encoding of the dataset, the third positional encoding associated with a full length of the visible frequency band (Martin [During training, [00004].. where t is the current training iteration, and N is a hyper-parameter for when α should reach the maximum number of frequencies m. [0051] The simplest version of the deformation uses a translational vector field V: (x, ω.sub.i).fwdarw.t, defining the deformation as T(x, ω.sub.i)=x+V(x, ω.sub.i). This formulation is sufficient to represent all continuous deformations).
Regarding Claim 8, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1, and further teaches wherein a first value associated with the first frequency mask includes a linear relationship with a second value associated with the second frequency mask, the linear relationship being based at least on the first training iteration of the NeRF and the second training iteration of the NeRF (Mehta [0058] A rectified linear activation function may be used as a default activation function for many types of neural networks. Using a rectified linear activation function may enable the use of stochastic gradient descent with backpropagation of errors to train deep neural networks. The rectified linear activation function may operate similar to a linear function, but it may enable complex relationships in the data to be learned.; Rematas [0006] the pre-trained semantic segmentation model can determine one or more pixels are associated with a sky in the environment and masks the one or more pixels).
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over in Mehta et al. (US 20220292341 A1), referred herein as Mehta view of Rematas et al. (US 20230281913 A1), referred herein as Rematas, Martin et al. (US 20240005590 A1), referred herein as Martin, Kim et al. (US 20240185505 A1), referred herein as Kim, and Tancik et al. (Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains, 2020), referred herein as Tancik.
Regarding Claim 2, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1, and further teaches training iteration being associated with frequency regularization (Martin [0048] Coarse-to-fine data 145 represents a coarse-to-fine deformation regularization… This function projects a positional vector x∈custom-character in a high dimensional space using a set of sine and cosine functions of increasing frequencies).
The prior arts does not teach all the claimed limitation herein. However, Tancik teaches further comprising obtaining, based at least on the first training iteration and the second training iteration, a trained NeRF configured to generate the few-shot neural rendering while minimizing at least one of overfitting or underfitting in the 3D scene (Tancik pp. 4; 4.2: We show in Section 5 that a Fourier feature input mapping can be tuned to lie between these “underfitting’ and “overfitting” extremes, enabling both fast convergence and low test error).
Tancik discloses a method using Fourier feature mapping to transform the effective NTK into a stationary kernel with a tunable bandwidth. Tancik is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mehta to incorporate the teachings of Tancik, and apply the method of training MLPs (4 layers, 1024 channels, ReLU activations) to fit a bandlimited noise signal (c = 8; n = 32) using Fourier feature mappings with different p values into the, as taught by Tancik to the modulator network that generates modulation parameters to modulate amplitude, phase, and/or frequency of intermediate layers of the synthesizer network.
Doing so would greatly improve the performance of MLPs for low-dimensional regression tasks relevant to the computer vision and graphics communities.
Claim(s) 9-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over in Mehta et al. (US 20220292341 A1), referred herein as Mehta view of Rematas et al. (US 20230281913 A1), referred herein as Rematas, Martin et al. (US 20240005590 A1), referred herein as Martin, Kim et al. (US 20240185505 A1), referred herein as Kim and Xian et al. (US 11748940 B1), referred herein as Xian.
Regarding Claim 9, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1, but does not teach the claimed limitation herein. However, Xian teaches further comprising, during at least one of the first training iteration or the second training iteration, causing a masking of one or more density scores located within a threshold proximity of an origin of a ray to reduce an occlusion in the 3D scene (Xian col 9, ln 5-7: the system may define the scene depth of a ray using accumulative depth values along the ray modulated with the transmittance and the volume density; col 9, ll52-57: The system may render photorealistic views with correctly filled dis-occluded contents compared to view synthesis with per-frame depth-based warping. The system may fill in the dis-occluded content implicitly in the 3D space and may produce significantly fewer artifacts than the traditional methods; col 12, ln 22-27: The system may estimate the scene depth from the learned volume density of the scene and measure its difference from the input depth d.sub.t. In particular embodiments, the system may measure the distance where the accumulated transmittance T becomes less than a certain threshold).
Xian discloses a computing system may determine a view position, a view direction, and a time with respect to a scene. Xian is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mehta to incorporate the teachings of Xian, and apply the method of constraining the dis-occluded contents by propagating the color and volume density across time to the modulator network that generates modulation parameters to modulate amplitude, phase, and/or frequency of intermediate layers of the synthesizer network.
Doing so would generate a free-viewpoint video rendering experience on various casual videos (e.g., by smartphones) while preserving motion and texture details for conveying a vivid sense of 3D.
Regarding Claims 10-13, Mehta in view of Rematas, Martin, Kim and Xian teaches a system (Mehta [0005] a signal processing apparatus trained using machine learning techniques to create a continuous representation of the signal), comprising:
determine, using the NeRF, a first density score representing a predicted volumetric density at a first point of multiple points disposed along a ray cast from an origin associated with a camera (Kim [0027] For each sample point 104, 5D coordinates (X,Y,Z,θ,ϕ) are provided as inputs to the function F.sub.θ, which gives density (σ) and radiance (L.sub.e) at each sample point in a volume) and through a pixel of the image an origin associated with a camera and through a pixel of the image (Xian col 2, ln 1-5: The space-time NeRF framework may use a continuous volume rendering method which allows the color of a pixel to be determined by integrating the radiance as modulated by the volume density along the camera ray).
The metes and bounds of the rest limitations of the claims substantially correspond to the limitations set forth in claims 1, 3, 8 and 9; thus they are rejected on similar grounds and rationale as their corresponding limitations.
Regarding Claim 14, Mehta in view of Rematas, Martin, Kim and Xian teaches the system of claim 10, and further teaches wherein the first density score is distinguishable from a second density score associated with a second point of the multiple points disposed along the ray, the one or more processors further configured to cause the second density score to be included based at least on a second distance between the origin and the second point meeting or exceeding the threshold (Xian col 2, ll. 43-49: wherein the first density score is distinguishable from a second density score associated with a second point of the multiple points disposed along the ray, the one or more processors further configured to cause the second density score to be included based at least on a second distance between the origin and the second point meeting or exceeding the threshold).
Regarding Claim 15, Mehta in view of Rematas, Martin, Kim and Xian teaches the system of claim 10, and further teaches wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing one or more simulation operations;
a system for performing one or more digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing one or more deep learning operations;
a system implemented using an edge device;
a system for performing one or more generative AI operations;
a system for performing operations using a large language model;
a system implemented using a robot;
a system for performing one or more generative AI operations;
a system for performing operations using a large language model;
a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content (Mehta [0046] a synthesizer network 330 comprising a plurality of synthesizer layers, wherein the synthesizer network 330 represents a continuous function of a signal parameter of the digital signal);
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.
Regarding Claims 16-20, Mehta in view of Rematas, Martin, Kim and Xian teaches a processor (Mehta [0005] a signal processing apparatus trained using machine learning techniques to create a continuous representation of the signal).
The metes and bounds of the limitations of the claims substantially correspond to the limitations set forth in claims 1, 3, 9 and 15; thus they are rejected on similar grounds and rationale as their corresponding limitations.
Allowable Subject Matter
Claim(s) 5-7 is/are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding Claim 5, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1, but neither Mehta, Rematas, Martin or Tanci teaches based at least on a difference between the first length of the visible frequency band and the second length of the visible frequency band, the first positional encoding includes a first signal of the dataset while excluding a second signal of the dataset. Therefore, the claim 5 in the context of the claim 1 as a whole would be allowable if rewritten in independent form.
Regarding Claim 6, Mehta in view of Rematas, Martin and Kim teaches the method of claim 5. Therefore the claim 6 is allowable by virtue of its dependency of claim 5.
Regarding Claim 7, Mehta in view of Rematas, Martin and Kim teaches the method of claim 1. Xian (US 11748940 B1) further teaches
the first signal associated with a first frequency that is less than a threshold; and the second signal associated with a second frequency that is greater than the threshold (col16, ln52 - col17, ln8: spatial frequency components higher than a threshold spatial frequency, spatial frequency components lower than a threshold spatial frequency, temporal frequency components higher than a temporal frequency threshold, temporal frequency components lower than a temporal frequency threshold, spatiotemporal power components falling with a particular spatiotemporal frequency range, etc.).
However, the combination of the prior art does not teach the first positional encoding includes a first signal of the dataset while excluding a second signal of the dataset, and the second positional encoding includes the first signal and the second signal. Therefore, the claim 7 in the context of the claim 1 as a whole would be allowable if rewritten in independent form.
Response to Arguments
Applicant’s arguments, see page 8, filed on 27 January 2026, with respect to the rejection(s) of claim(s) 1 under 103 rejection have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Kim.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Samantha (Yuehan) Wang whose telephone number is (571)270-5011. The examiner can normally be reached Monday-Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached on (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Samantha (YUEHAN) WANG/
Primary Examiner
Art Unit 2617