Prosecution Insights
Last updated: October 02, 2026
Application No. 18/516,625

TRANSFORMERS AS NEURAL RENDERERS

Non-Final OA §103
Filed
Nov 21, 2023
Priority
Dec 09, 2022 — provisional 63/431,620
Examiner
MCCULLEY, RYAN D
Art Unit
2611
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
70%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
357 granted / 509 resolved
+8.1% vs TC avg
Strong +28% interview lift
Without
With
+27.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
24 currently pending
Career history
534
Total Applications
across all art units

Statute-Specific Performance

§101
7.9%
-32.1% vs TC avg
§103
56.5%
+16.5% vs TC avg
§102
14.6%
-25.4% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 509 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 22 June 2026 has been entered. Response to Arguments Applicant’s arguments filed 22 June 2026 have been fully considered but they are not persuasive. Applicant argues “according to Mildenhall, the continuous scene representation is a neural radiance field implemented by an MLP … Instead of using a continuous scene representation over 3-D space, Suhail uses epipolar geometry to aggregate information from multiple views … combining Suhail and Mildenhall as proposed would render Suhail inoperable … the MLP of Mildenhall, which does not aggregate information related to multiple epipolar points, could not be used to implement either 2-MLP or 1-MLP [of Suhail]” (Remarks, pgs. 9-10). The Examiner respectfully disagrees. In the rejection below, Suhail is cited to describe sampling locations along rays, using a neural network to obtain feature values of samples, and pooling the feature values to obtain a color value. However, Suhail does not specifically describe the object representation being 3D, and the object representation is not described as being bounded. Mildenhall remedies these deficiencies by showing that it is known to represent an object three-dimensionally, and that such a 3D object representation can be bounded to improve ray sampling. When Suhail is modified using these concepts, the modified Suhail would teach that the object is a 3D object instead of a 2D object, and the 3D object is bounded for ray sampling purposes. Such modifications of Suhail would improve realism and efficiency. In the modified Suhail of the rejection below, the MLPs are not being replaced by any MLPs of Mildenhall. Instead, only the 2D representation of Suhail is being replaced by a bounded 3D representation. The test for obviousness is not whether the features of a secondary reference may be bodily incorporated into the structure of the primary reference. Rather, the test is what the combined teachings of the references would have suggested to those of ordinary skill in the art. The Examiner maintains that after reading Mildenhall, one having ordinary skill in the art would have found it obvious to use a bounded 3D object representation as the object representation in Suhail. Any remaining arguments are considered moot based on the foregoing. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6 and 8-11 are rejected under 35 U.S.C. 103 as being unpatentable over Suhail et al. (“Light Field Neural Rendering”; hereinafter “Suhail”) in view of Mildenhall et al. (“NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”; hereinafter “Mildenhall”), and further in view of Adamkiewicz et al. (“Vision-Only Robot Navigation in a Neural Radiance World”; hereinafter “Adamkiewicz”). Regarding claim 1, Suhail discloses A method comprising: sampling a plurality of samples along a ray intersecting a representation of at least one object (“sample a sequence of P points … along the ray,” pg. 8271, sec. 3.2; Fig. 2 illustrates the ray intersecting a representation of an object); using at least one neural network (Fig. 2 illustrates multiple Transformer and MLP neural networks) to obtain a set of feature values based at least in part on the plurality of samples (“sample a sequence of P points … along the ray … features associated to the epipolar points and target ray,” pg. 8271, sec. 3.2); using the set of feature values to obtain at least one pooled feature value; using the at least one pooled feature value (“aggregating features,” pg. 8271, sec. 3.2) to obtain a color value (“predicts the target ray color by aggregating features,” pg. 8271, sec. 3.2). Suhail does not disclose the representation of the object being a three-dimensional representation or the plurality of samples to be sampled within a bounded region comprising the three-dimensional representation. In the same art of computer graphics, Mildenhall teaches sampling a three-dimensional representation of an object (see Fig. 2) and the plurality of samples to be sampled within a bounded region comprising the three-dimensional representation (“camera ray r(t) = o + td with near and far bounds tn and tf,” pg. 101, sec. 4). Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the 3D object and bounded sampling of Mildenhall to Suhail. The motivation would have been “to render high-resolution photorealistic novel views of real objects” (Mildenhall, pg. 100, col. 1, para. 1). The combination of Suhail and Mildenhall does not disclose causing a device to move based at least in part on a path of motion determined based at least in part on the color value. In the same art of neural rendering, Adamkiewicz teaches causing a device to move based at least in part on a path of motion determined based at least in part on the color value (“navigating a robot through a 3D environment represented as a NeRF using only an onboard RGB camera for localization … a trajectory optimization algorithm that avoids collisions with high-density regions in the NeRF … an optimization based filtering method to estimate 6DoF pose and velocities for the robot in the NeRF given only an onboard RGB camera,” abstract; Fig. 2 illustrates using a color output of a neural renderer to update a trajectory). Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Adamkiewicz to the combination of Suhail and Mildenhall. The motivation would have been “Planning and executing a trajectory with onboard sensors is a fundamental building block of many robotic applications” (Adamkiewicz, pg. 1, sec. I, para. 1). Regarding claim 2, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious dividing the bounded region into subregions, and obtaining a sample of the plurality of samples along the ray from each of the subregions (“camera ray r(t) = o + td with near and far bounds tn and tf … we use a stratified sampling approach where we partition [tn, tf] into N evenly spaced bins and then draw one sample uniformly at random from within each bin,” Mildenhall, pg. 101, sec. 4; see claim 1 for motivation to combine). Regarding claim 3, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious encoding the plurality of samples with positional information indicating a position of each of the plurality of samples with respect to at least one other of the plurality of samples, the at least one neural network to use the positional information to obtain the set of feature values (“positional encoding … We aggregate the P outputs corresponding to the epipolar points … to obtain the reference view features,” Suhail, pg. 8272, sec. 3.3.1). Regarding claim 4, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the at least one neural network comprises a transformer encoder to obtain the set of feature values (“Epipolar feature transformer,” Suhail, pg. 8272, sec. 3.3.1) Regarding claim 5, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious calculating the at least one pooled feature value based at least in part on the set of feature values (“computes a feature representation per reference view by aggregating features associated to the epipolar points and target ray,” Suhail, pg. 8271, sec. 3.2). Regarding claim 6, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the set of feature values comprises a plurality of values for a plurality of features, the at least one pooled feature value comprises a corresponding pooled feature value for each of the plurality of features, at least a portion of the plurality of values are associated with each of the plurality of features (Suhail, Fig. 2 illustrates these properties of the features and values), and for each of the plurality of features, the corresponding pooled feature value is calculated as a maximum or an average of those of the plurality of values in the portion associated with the feature (“The aggregation is a weighted average,” Suhail, pg. 8272, sec. 3.3.1). Regarding claim 8, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the at least one neural network comprises a multilayer perceptron that uses the at least one pooled feature value to obtain the color value (Suhail, Fig. 2 illustrates the feature aggregation used as input to an MLP to predict color). Regarding claim 9, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the at least one neural network uses a photometric loss function to obtain the color value (“During training, we minimize the L2 loss between the observed and predicted colors,” Suhail, pg. 8273, sec. 3.4; “a NeRF-based photometric loss,” Adamkiewicz, pg. 2, col. 1; see claim 1 for motivation to combine). Regarding claim 10, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious determining the path of motion based at least in part on the color value (Adamkiewicz, Fig. 2 illustrates using a color output of a neural renderer to update a trajectory; see claim 1 for motivation to combine). Regarding claim 11, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious casting the ray through the three-dimensional representation (Mildenhall, Fig. 2; see claim 1 for motivation to combine). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Suhail, Mildenhall, and Adamkiewicz, and further in view of Hori et al. (US 2024/0046085; hereinafter “Hori”). Regarding claim 7, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the at least one neural network comprises a transformer encoder (“Epipolar feature transformer,” Suhail, pg. 8272, sec. 3.3.1) comprising at least one layer comprising self-attention functionality (“self-attention transformer,” Suhail, pg. 8272, sec. 3.3.1). The combination of Suhail, Mildenhall, and Adamkiewicz does not disclose that the self-attention functionality determines a set of dependency metrics based at least in part on a set of parameter values associated with the plurality of samples and feed forward functionality that outputs one or more feature values based at least in part on the set of dependency metrics. In the same art of extracting visual features using machine learning, Hori teaches self-attention functionality that determines a set of dependency metrics based at least in part on a set of parameter values associated with the set of samples and feed forward functionality that outputs one or more feature values based at least in part on the set of dependency metrics (“The self-attention layer 410 [of Fig. 4] extracts temporal dependency … the feed-forward layers 412 are applied in a point-wise manner. The encoded representations for audio and visual features are obtained,” para. 73). Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Hori to the combination of Suhail, Mildenhall, and Adamkiewicz. The motivation would have been to optimize the feature extraction. Claims 12-20 and 24-28 are rejected under 35 U.S.C. 103 as being unpatentable over Suhail in view of Mildenhall. Regarding claim 12, Suhail discloses A system comprising: at least one processor; and memory storing instructions that when executed by the at least one processor cause the system to (these are considered inherent aspects of a computer-based neural rendering architecture): perform one or more transformer models (“transformer-based model,” abstract) to obtain a set of feature values based at least in part on a plurality of locations along a ray intersecting a representation of an object (“computes a feature representation per reference view by aggregating features associated to the epipolar points and target ray,” pg. 8271, sec. 3.2); perform one or more machine learning processes to obtain a color value based at least in part on the set of feature values (“predicts the target ray color by aggregating features associated to each reference view,” pg. 8271, sec. 3.2); and generate a view of the object using the color value (“synthesize novel views of a scene,” pg. 8271, sec. 3). Suhail does not disclose the representation of the object being a three-dimensional representation or the plurality of locations to be positioned within a bounded region comprising the three-dimensional representation. In the same art of computer graphics, Mildenhall teaches sampling a three-dimensional representation of an object (see Fig. 2) and the plurality of locations to be positioned within a bounded region comprising the three-dimensional representation (“camera ray r(t) = o + td with near and far bounds tn and tf,” pg. 101, sec. 4). Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the 3D object and bounded sampling of Mildenhall to Suhail. The motivation would have been “to render high-resolution photorealistic novel views of real objects” (Mildenhall, pg. 100, col. 1, para. 1). Regarding claim 13, the combination of Suhail and Mildenhall renders obvious obtain the plurality of locations by dividing a portion of the ray within the bounded region into subregions and obtain a sample from within each subregion (“camera ray r(t) = o + td with near and far bounds tn and tf … we use a stratified sampling approach where we partition [tn, tf] into N evenly spaced bins and then draw one sample uniformly at random from within each bin,” Mildenhall, pg. 101, sec. 4; see claim 12 for motivation to combine). Regarding claim 14, the combination of Suhail and Mildenhall renders obvious wherein the one or more transformer models obtain the set of feature values using positional encodings obtained based at least in part on the plurality of locations (“positional encoding … We aggregate the P outputs corresponding to the epipolar points … to obtain the reference view features,” Suhail, pg. 8272, sec. 3.3.1). Regarding claim 15, the combination of Suhail and Mildenhall renders obvious wherein the one or more transformer models comprises a transformer encoder (“Epipolar feature transformer,” Suhail, pg. 8272, sec. 3.3.1). Regarding claim 16, the combination of Suhail and Mildenhall renders obvious aggregate at least a portion of the set of feature values to obtain an aggregated feature value, the one or more machine learning processes to obtain the color value based at least in part on the aggregated feature value (“predicts the target ray color by aggregating features,” Suhail, pg. 8271, sec. 3.2). Regarding claim 17, the combination of Suhail and Mildenhall renders obvious wherein each of the set of feature values is associated with a feature in a set of features (see Suhail, Fig. 2), and the portion of the set of feature values are aggregated by at least one of selecting a maximum one of the set of feature values for each feature within the set of features or calculating an average of those of the set of feature values associated with each feature within the set of features (“The aggregation is a weighted average,” Suhail, pg. 8272, sec. 3.3.1). Regarding claim 18, the combination of Suhail and Mildenhall does not specifically recite wherein the one or more transformer models comprises a transformer encoder usable for natural language processing. The Examiner previously took Official Notice that both the concepts and the advantages of using a transformer encoder for natural language processing were well known and expected in the art before the effective filing date of the claimed invention. Since Applicant did not traverse the Official Notice, it is now taken as Applicant Admitted Prior Art. It would have been obvious before the effective filing date of the claimed invention to use a transformer encoder usable for natural language processing in the combination of Suhail and Mildenhall in order to increase accuracy and efficiency of the machine learning model. Regarding claim 19, the combination of Suhail and Mildenhall renders obvious wherein the one or more machine learning processes comprise a multilayer perceptron to obtain the color value (Suhail, Fig. 2 illustrates the feature aggregation used as input to an MLP to predict color). Regarding claim 20, the combination of Suhail and Mildenhall renders obvious wherein the one or more machine learning processes uses a photometric loss function when obtaining the color value (“During training, we minimize the L2 loss between the observed and predicted colors,” Suhail, pg. 8273, sec. 3.4). Regarding claim 24, it is rejected using the same citations and rationales described in the rejection of claim 12. Regarding claim 25, the combination of Suhail and Mildenhall renders obvious cast the ray through the three-dimensional representation of the obstacle and between a virtual image capture device and a focal point of the virtual image capture device (Mildenhall, Fig. 2; see claim 12 for motivation to combine). Regarding claim 26, the combination of Suhail and Mildenhall renders obvious determine bounds of the bounded region along the ray, and obtain the plurality of samples by dividing a portion of the ray within the bounds of the bounded region into subregions and obtain a sample from each of the subregions (“camera ray r(t) = o + td with near and far bounds tn and tf … we use a stratified sampling approach where we partition [tn, tf] into N evenly spaced bins and then draw one sample uniformly at random from within each bin,” Mildenhall, pg. 101, sec. 4; see claim 12 for motivation to combine). Regarding claim 27, the combination of Suhail and Mildenhall renders obvious wherein a multilayer perceptron is to obtain the at least one color value (Suhail, Fig. 2 illustrates the feature aggregation used as input to an MLP to predict color), and the at least one transformer-based machine learning model comprises a transformer encoder to obtain the set of feature values (“Epipolar feature transformer,” Suhail, pg. 8272, sec. 3.3.1). Regarding claim 28, the combination of Suhail and Mildenhall renders obvious use a photometric loss function to supervise performing the at least one transformer-based machine learning model and determining the at least one color value (“During training, we minimize the L2 loss between the observed and predicted colors,” Suhail, pg. 8273, sec. 3.4). Claims 21-23 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Suhail and Mildenhall, and further in view of Adamkiewicz. Regarding claim 21, the combination of Suhail and Mildenhall does not disclose a device, wherein the instructions, when executed by the at least one processor, cause the at least one processor to determine a path of motion based at least in part on the color value, and instruct the device to move based at least in part on the path of motion. In the same art of neural rendering, Adamkiewicz teaches a device, wherein the instructions, when executed by the at least one processor, cause the at least one processor to determine a path of motion based at least in part on the color value, and instruct the device to move based at least in part on the path of motion (“navigating a robot through a 3D environment represented as a NeRF using only an onboard RGB camera for localization … a trajectory optimization algorithm that avoids collisions with high-density regions in the NeRF … an optimization based filtering method to estimate 6DoF pose and velocities for the robot in the NeRF given only an onboard RGB camera,” abstract; Fig. 2 illustrates using a color output of a neural renderer to update a trajectory). Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Adamkiewicz to the combination of Suhail and Mildenhall. The motivation would have been “Planning and executing a trajectory with onboard sensors is a fundamental building block of many robotic applications” (Adamkiewicz, pg. 1, sec. I, para. 1). Regarding claim 22, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the device is an autonomous machine or a semi-autonomous machine (“autonomous driving or drone flight,” Adamkiewicz, pg. 1, sec. I; see claim 21 for motivation to combine). Regarding claim 23, the combination of Suhail, Mildenhall, and Adamkiewicz renders obvious wherein the device is an autonomous vehicle (“autonomous driving or drone flight,” Adamkiewicz, pg. 1, sec. I; see claim 21 for motivation to combine), and the instructions, when executed by the at least one processor, cause the at least one processor to: obtain image data captured by at least one image capture device associated with the autonomous vehicle, generate the three-dimensional representation of the object based on the image data (“Camera Images” of Adamkiewicz, Fig. 2; see claim 21 for motivation to combine; “captured RGB images of the scene,” Mildenhall, pg. 102, sec. 5.2; see claim 12 for motivation to combine); and cast the ray through the three-dimensional representation of the object (Mildenhall, Fig. 2; see claim 12 for motivation to combine). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ryan McCulley whose telephone number is (571)270-3754. The examiner can normally be reached Monday through Friday, 8:00am - 4:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RYAN MCCULLEY/Primary Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Show 2 earlier events
Jan 15, 2026
Applicant Interview (Telephonic)
Jan 15, 2026
Examiner Interview Summary
Jan 30, 2026
Response Filed
Apr 21, 2026
Final Rejection mailed — §103
Jun 22, 2026
Response after Non-Final Action
Jul 21, 2026
Request for Continued Examination
Jul 23, 2026
Response after Non-Final Action
Aug 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749259
Image Generation System with Controllable Scene Lighting
3y 1m to grant Granted Sep 29, 2026
Patent 12737955
RIGID BODY SIMULATION BY RELATIVE SLEEPING
2y 10m to grant Granted Sep 15, 2026
Patent 12731326
GEMSTONE CUT ANALYSIS
2y 3m to grant Granted Sep 08, 2026
Patent 12718467
DEDICATED RAY MEMORY FOR RAY TRACING IN GRAPHICS SYSTEMS
1y 11m to grant Granted Aug 25, 2026
Patent 12711695
IMAGE RENDERING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM
2y 2m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
70%
Grant Probability
98%
With Interview (+27.9%)
2y 6m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 509 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month