Prosecution Insights
Last updated: October 02, 2026
Application No. 18/532,392

Generation of Video for a Location Via a Generative Machine-Learned Model

Final Rejection §103§112
Filed
Dec 07, 2023
Examiner
CHEN, ALAN S
Art Unit
Tech Center
Assignee
Google LLC
OA Round
2 (Final)
91%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 91% — above average
91%
Career Allowance Rate
1048 granted / 1152 resolved
+31.0% vs TC avg
Moderate +7% lift
Without
With
+6.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
33 currently pending
Career history
1170
Total Applications
across all art units

Statute-Specific Performance

§101
12.8%
-27.2% vs TC avg
§103
22.7%
-17.3% vs TC avg
§102
36.3%
-3.7% vs TC avg
§112
20.5%
-19.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1152 resolved cases

Office Action

§103 §112
CTNF 18/532,392 CTNF 79889 Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Drawings 06-22-07 AIA The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: reference character "3000" appearing in FIG. 3 as the outer system-level label enclosing the generative video application 3100 . Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification 07-29 AIA The disclosure is objected to because of the following informalities: paragraph [0162] recites "The computing device 50 can be a user computing device or a server computing device," but the figure being described is FIG. 5C, whose computing device is identified as "5300" (see [0163]); "50" appears to be a typographical error for "5300" the specification uses both "Wi-Fi" ([0056]) and "WiFi" ([0085]) for the same wireless networking standard, contrary to the consistent-terminology requirement of 37 CFR 1.71(a) reference character "334" (navigation application 334) is referred to in [0099] and [0117] but is not expressly introduced or described in parallel to the introduction of related elements 134 ([0079]) and 332 ([0091]) . Appropriate correction is required. Claim Objections 07-29-01 AIA Claim 12 is objected to because of the following informalities: the claim recites "...the video depicts the one or more object included in the scene"; the noun should be plural ("objects") to maintain consistency with the immediately preceding limitation "one or more objects to be included in the scene." Appropriate correction is required. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 8 and 12 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Regarding claim 8: Claim 8 recites “implementing one or more large language models to determine a plurality of variables based on the query.” The limitation introduces the term “a plurality of variables” as a new claim element that is not further defined, referenced, or used anywhere in the claim. Specifically, the “plurality of variables” determined by the one or more large language models bears no express or implied relationship to any other claim element, neither the “conditioning parameters,” the “generative machine-learned model,” the “video,” nor any recited operation of the computer platform. The “variables” are introduced and then abandoned, leaving it entirely unclear how their determination affects the scope of the claimed invention or contributes to generating the video. A person of ordinary skill in the art would not be able to determine, with reasonable certainty, what the metes and bounds of claim 8 are with respect to this limitation. See Nautilus, Inc. v. Biosig Instruments, Inc., 572 U.S. 898 (2014). Regarding claim 12: Claim 12 recites “wherein the query comprises a text query that specifies one or more objects to be included in the scene and wherein the video depicts the one or more object included in the scene.” The phrase “the one or more object” in the second clause is grammatically inconsistent with “one or more objects” in the first clause, “object” (singular) is used where “objects” (plural) was previously introduced. It is unclear whether this inconsistency narrows the claim scope to a single object or whether it is a drafting error intended to refer to the same antecedent “one or more objects.” This grammatical ambiguity renders the metes and bounds of claim 12 uncertain. See MPEP § 2173.05(e). Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-20-02-aia AIA This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 07-21-aia AIA Claims 1-4, 9-11, 13, 15-17 and 20 are r ejected under 35 USC 103 as being unpatentable over B lock-NeRF: Scalable Large Scene Neural View Synthesis to Tancik et al. (hereinafter Tancik) in view of NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections to Martin-Brualla et al. (hereinafter Martin-Brualla). P er claim 1, Tancik discloses A computer platform for generating a video (Tancik: Abstract and Section 1…Tancik describes a computer system that reconstructs a city-scale environment using neural radiance fields and renders movies of the environment from arbitrary camera trajectories, which constitutes a computer platform that generates video under BRI, “We present Block-NeRF, a variant of Neural Radiance Fields that can represent large-scale environments… capable of rendering an entire neighborhood of San Francisco; Fig.1…”Video results can be found on the project website”) , comprising: one or more memories configured to store instructions (Tancik: Section 5 and Supplement (“The architectural and optimization specifics are provided in the supplement”)…Tancik describes implementing the Block-NeRF pipeline on a computer system that stores model weights and rendering code; a PHOSITA would understand such a pipeline to require one or more memories storing executable instructions, “Each Block-NeRF was trained on data from 38 to 48 different data collection runs”) ; and one or more processors configured to execute the instructions to perform operations (Tancik: Fig. 3 and Section 1…Tancik's pipeline executes two MLPs (f σ , f c ) plus a visibility MLP f v on each ray sample to render each pixel; running such MLPs requires one or more processors executing instructions, “The first MLP f σ predicts the density σ for a position x in space. The network also outputs a feature vector that is concatenated with viewing direction d, the exposure level, and an appearance embedding. These are fed into a second MLP f c that outputs the color for the point”) , the operations comprising: receiving a query from a user relating to a location (Tancik: Fig. 2…Tancik renders a target view in response to a target viewpoint (location) supplied to the system, which constitutes a query from a user relating to a location since the viewpoint identifies the geographic location at which the scene is to be rendered, “To render a target view in the scene, the visibility maps are computed for all of the NeRFs within a given radius”; Section 4.3.1…”we only consider Block-NeRFs that are within a set radius of the target viewpoint”) ; in response to receiving the query, generating conditioning parameters based at least in part on the query… (Tancik: Section 4.2.1…Tancik optimizes per-image appearance embedding vectors that capture varying environmental conditions (weather, lighting) and then applies/interpolates those embeddings to the rendering at the target location; this constitutes generating conditioning parameters that provide values for one or more conditions associated with the scene to be rendered at the location, “we follow NeRF-W [40] and use Generative Latent Optimization [5] to optimize per-image appearance embedding vectors… This allows the NeRF to explain away several appearance-changing conditions, such as varying weather and lighting”; Section 4.2.3 (Exposure Input)…Tancik further conditions the model on a numerical exposure value that the user can alter at inference, which constitutes a value for a lighting condition derived in response to user input, “we find that feeding the camera exposure information to the appearance prediction part of the model allows the NeRF to compensate for the visual differences; Section 5.2…”the exposure input marginally improves the reconstruction, but more importantly provides us with the ability to change the exposure during inference”) ; generating, using a generative machine-learned model, the video, wherein the video depicts the scene at the location and with the values for the one or more conditions (Tancik: Section 1, Section 4 and Fig. 1…Tancik's Block-NeRF is a generative neural-network-based scene representation (an extension of NeRF) and renders frames along a camera trajectory at the target location with the appearance embedding controlling weather/lighting, which constitutes generating a video depicting the scene at the location with the values for the conditions, “scene conditioned NeRFs are also capable of changing environmental lighting conditions such as camera exposure, weather, or time of day, which can be used to further augment simulation scenarios”; Fig. 4 …The same scene rendered under different appearance codes shows the location with varying lighting and weather conditions specified by the conditioning parameters, “The appearance codes allow the model to represent different lighting and weather conditions”) ; and providing the video for presentation to the user (Tancik: Abstract and Fig. 1…Tancik provides the rendered video to the user via the project website and as the output of the rendering pipeline, which reads on providing the video for presentation, “Video results can be found on the project website waymo.com/research/block-nerf”) . Tancik does not expressly disclose, but combined with Martin-Brualla does teach: conditioning parameters provide values for one or more conditions associated with a scene to be rendered at the location (Martin-Brualla: Section 1 and Section 4.1 (Latent Appearance Modeling)…Martin-Brualla learns a per-image appearance embedding that explicitly models exposure, lighting, weather, and post-processing as a latent-space condition controlling the rendered appearance of a landmark scene; this is conditioning parameters providing values for one or more conditions associated with a scene to be rendered at the location, “we model per-image appearance variations such as exposure, lighting, weather, and post-processing in a learned low-dimensional latent space… The learned latent space provides control of the appearance of output renderings”). Tancik and Martin-Brualla are analogous art because they are both within the same field of endeavor, specifically neural radiance field (NeRF) scene reconstruction and novel view synthesis of real-world environments captured from photographs. They address the same problem solving area of producing photorealistic renderings of a scene from a sparse set of captured images under varying environmental conditions such as weather, time of day, and lighting. Tancik explicitly cites Martin-Brualla’s appearance embedding technique as the basis for its conditioning mechanism (Tancik: Section 2.2 and Section 4.2.1…”We incorporate techniques from NeRF in the Wild (NeRF-W) [40], which adds a latent code per training image to handle inconsistent scene appearance when applying NeRF to landmarks from the Photo Tourism dataset.”), which is the core subject of Martin-Brualla (Martin-Brualla: Section 1). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to combine Tancik’s per-location, large-scale NeRF rendering pipeline with Martin-Brualla’s per-image latent appearance embedding controlling weather/lighting in order to render videos of a queried location under explicit, user-selectable weather and lighting conditions. Tancik already incorporates Martin-Brualla’s appearance-embedding architecture (Tancik: Section 2.2 — “We incorporate techniques from NeRF in the Wild (NeRF-W) [40]”), so applying its full conditioning capability to render videos of a chosen location under specified conditions involves only routine combination of known elements yielding predictable results, with no unexpected interaction or new functionality. The suggestion/motivation for doing so would have been Tancik’s explicit statement that scene-conditioned NeRFs can change environmental lighting conditions such as exposure, weather and time of day to augment simulation scenarios (Tancik: Section 1), combined with Martin-Brualla’s seminal teaching that a learned latent appearance embedding provides controllable interpolation over exposure, lighting, weather and post-processing (Martin-Brualla: Section 1). A PHOSITA would have been motivated to apply Martin-Brualla’s explicit weather/lighting controls within Tancik’s per-location, multi-block rendering pipeline to provide users with on-demand renderings of a queried location at a selected weather/lighting state. Per claim 2, Tancik combined with Martin-Brualla discloses claim 1, Tancik further teaches the generative machine-learned model comprises a neural radiance field (NeRF) (Tancik: Abstract and Section 3.1…Tancik's Block-NeRF is explicitly an extension of NeRF (a neural radiance field), so the generative machine-learned model literally comprises a neural radiance field, “We present Block-NeRF, a variant of Neural Radiance Fields; Section 3…”We build upon NeRF [42] and its extension mip-NeRF [3]”). Per claim 3, Tancik combined with Martin-Brualla discloses claim 1, Tancik further teaches generating the conditioning parameters comprises retrieving current values for the one or more conditions at the location (Tancik: Section 5.1…Tancik's collection vehicles capture sensor data including a scalar exposure value at the time and location of capture, and the system retrieves these current exposure values to condition the rendering; under BRI, retrieving a sensor-derived exposure value at the location constitutes retrieving current values for the one or more conditions at the location, “Each camera captures images at 10 Hz and stores a scalar exposure value…”; Section 4.2.3…”We find that feeding the camera exposure information to the appearance prediction part of the model allows the NeRF to compensate for the visual differences”). Per claim 4, Tancik combined with Martin-Brualla discloses claim 1, Tancik further teaches generating, using the generative machine-learned model, the video comprises conditioning the generative machine-learned model with the conditioning parameters (Tancik: Fig. 3 and Section 4.2…Tancik feeds the appearance embedding and exposure inputs directly into the color-predicting MLP f c , thereby conditioning the generative model on those parameters before rendering each frame of the output video, this constitutes conditioning the generative machine-learned model with the conditioning parameters, “The network also outputs a feature vector that is concatenated with viewing direction d, the exposure level, and an appearance embedding. These are fed into a second MLP f c that outputs the color for the point”). Per claim 9, Tancik combined with Martin-Brualla discloses claim 1, Tancik further teaches generating the conditioning parameters comprises predicting future values for the one or more conditions based on current values for the one or more conditions at the location and/or based on historical values for the one or more conditions at the location (Tancik: Section 4.2.1 and Fig. 4…Tancik interpolates between appearance embeddings observed at different points in time across the training collection (for example, day vs. night, clear vs. cloudy) to synthesize new appearance conditions at the same location; under BRI, generating a new condition value at the location from historical condition values observed there constitutes predicting future values for the one or more conditions based on historical values for the one or more conditions at the location, “We can additionally manipulate these appearance embeddings to interpolate between different conditions observed in the training data (such as cloudy versus clear skies, or day and night)”). Per claim 10, Tancik combined with Martin-Brualla discloses claim 1, Tancik further teaches generating, using the generative machine-learned model, the video comprises: generating a series of camera poses based at least in part on the query; and rendering, respectively from the series of camera poses, a series of images of the scene at the location and with the values for the one or more conditions (Tancik: Section 5.1…Tancik renders a video by tracing camera rays from each successive camera pose along a trajectory and applying volume rendering to produce an image at each pose; the trajectory of camera poses is constructed based on the user-supplied target view, and the resulting per-pose images form the rendered video at the location with the conditioned appearance, “the vehicle pose is known and all cameras are calibrated. Using this information, we calculate the corresponding camera ray origins and directions in a common coordinate system; Section 3.1…”NeRF randomly samples distances {t i } N i=0 along the ray and passes the points r(t i ) and direction d through its MLPs to calculate σ i and c i ”). Per claim 11, Tancik combined with Martin-Brualla discloses claim 1, Tancik further teaches the computer platform comprises a database configured to store a plurality of generative machine-learned models respectively associated with a plurality of different locations; and generating, using the generative machine-learned model, the video comprises retrieving, from among the plurality of generative machine-learned models, the generative machine-learned model associated with the location (Tancik: Abstract…Tancik decomposes the city-scale environment into a plurality of independently-trained Block-NeRFs that are each associated with a respective geographic block (i.e., a respective location) and stored in the system, and dynamically retrieves the subset of Block-NeRFs whose geographic origin is within a set radius of the target viewpoint for rendering; under BRI this is a database storing a plurality of generative ML models for different locations and retrieving the one associated with the queried location, “we demonstrate that when scaling NeRF to render city-scale scenes spanning multiple blocks, it is vital to decompose the scene into individually trained NeRFs”; Section 1…”To compute a target view, only a subset of the Block-NeRFs are rendered and then composited based on their geographic location compared to the camera”; Section 4.3.1…“we only consider Block-NeRFs that are within a set radius of the target viewpoint”). Per claim 13, Tancik combined with Martin-Brualla discloses claim 1, Tancik in combination with Martin-Brualla further teaches the generative machine-learned model has been trained on a training dataset comprising a plurality of reference images of the location, and the training dataset comprises values for the one or more conditions for at least some of the plurality of reference images (Tancik: Section 5.1 (Alamo Square Dataset)…Tancik trains each Block-NeRF on millions of reference images captured at the corresponding location, with each image accompanied by a scalar exposure value (a value for a lighting condition) and an implicitly-trained appearance embedding reflecting per-image weather/lighting; this constitutes training on a plurality of reference images of the location plus values for the one or more conditions for at least some of them, “Each camera captures images at 10 Hz and stores a scalar exposure value … Each Block-NeRF is trained on between 64,575 to 108,216 images… The overall dataset is composed of 13.4 h of driving time sourced from 1,330 different data collection runs, with a total of 2,818,745 training images”); further, Martin-Brualla teaches that the per-image latent embedding paired with each landmark photograph encodes the weather/lighting condition of that photograph (Martin-Brualla: Section 4.1 (Latent Appearance Modeling) and Fig. 3…Martin-Brualla explicitly pairs each landmark training image with a learned appearance embedding capturing exposure, lighting and weather of that image, which is the value of the condition associated with that reference image, “we optimize an appearance embedding for each input image, thereby granting NeRF-W the flexibility to explain away photometric and environmental variations between images by learning a shared appearance representation across the entire photo collection”). The rational to combine Tancik with Martin-Brualla is the same as provided for the parent claim. Claims 15 and 16 are substantially similar in scope and spirit as claims 1 and 2, respectively. Therefore, the rejections of claims 1 and 2 are applied accordingly. Per claim 17, Tancik combined with Martin-Brualla discloses claim 15. Tancik further teaches generating the conditioning parameters comprises: retrieving current values for the one or more conditions at the location, or predicting future values for the one or more conditions based on current values for the one or more conditions at the location and/or based on historical values for the one or more conditions at the location (Tancik: Section 5.1…Tancik teaches by retrieving the captured per-image exposure value (a current value of a lighting condition at the location), and additionally teaches predicting future values for one or more conditions based on current values for one or more conditions at the location based on historical values by interpolating between historically-observed appearance embeddings at the same location to synthesize new condition values, “Each camera captures images at 10 Hz and stores a scalar exposure value”; Section 4.2.1…”We can additionally manipulate these appearance embeddings to interpolate between different conditions observed in the training data (such as cloudy versus clear skies, or day and night)”). Claim 20 is substantially similar in scope and spirit as claim 1. Therefore, the rejection of claim 1 is applied accordingly. As to the recited non-transitory CRM, Tancik's Block-NeRF rendering pipeline necessarily executes on a computer system whose model weights and rendering code are stored on non-transitory storage media (e.g., disk, ROM, SSD) so that the trained parameters of each Block-NeRF persist between training runs and inference invocations; a PHOSITA would understand the trained-then-loaded-at-inference workflow described by Tancik to require non-transitory storage (Tancik: Section 1, Section 4 and Section 5…Tancik's per-block independence and ability to update or expand the environment without retraining presupposes that the parameters of each Block-NeRF persist between sessions, which under BRI requires storage on a non-transitory computer-readable medium, “Modeling these Block-NeRFs independently allows for maximum flexibility, scales up to arbitrarily large environments and provides the ability to update or introduce new regions in a piecewise manner without retraining the entire environment”) . 07-21-aia AIA Claim s 5-8, 18 and 19 are rejected under 35 USC 103 as being unpatentable over Tancik in view of Martin-Brualla and further in view of BERT for Joint Intent Classification and Slot Filling to Chen et al. (hereinafter Chen) . Per claims 5, 6 and 18, 19, Tancik combined with Martin-Brualla discloses claim 1 and claim 15, respectively. Tancik combined with Martin-Brualla does not expressly disclose but with Chen does teach: Claims 5 and 18… extracting the values for the one or more conditions from the query (Chen: Section 1, Section 3.2 and Table 1…Chen teaches using a BERT-based joint intent-classification / slot-filling model to identify slot values literally present in a natural-language user query (e.g., extracting “movie” and “Steven Spielberg” as slot values from the query “Find me a movie by Steven Spielberg”); applied to the query of claims 1 and 15, this slot-extraction constitutes extracting the values for the one or more conditions from the query, “Intent classification focuses on predicting the intent of the query, while slot filling extracts semantic concepts. Table 1 shows an example of intent classification and slot filling for user query “Find me a movie by Steven Spielberg”… genre = movie, directed by = Steven Spielberg”). Claims 6 and 19… inferring the values for the one or more conditions from the query (Chen: Section 4.4 (Ablation Analysis and Case Study)…Chen's joint BERT classifier predicts a structured intent label and slot labels even for queries whose values are not literally present (e.g., disambiguating phrases by leveraging pre-trained world knowledge); inferring missing or implied values from the surface form of the query constitutes inferring the values for the one or more conditions from the query under BRI, “joint BERT correctly predicts the slot labels and intent because “mother joan of the angels” is a movie entry in Wikipedia. The BERT model was pre-trained partly on Wikipedia and possibly learned this information for this rare phrase”). Claim 7…Tancik combined with Martin-Brualla and further with Chen discloses claim 6, Chen further teaches inferring the values for the one or more conditions comprises providing the query to a sequence processing model, wherein the sequence processing model is configured to output the values for the one or more conditions in response to the query (Chen: Section 3.1 and Fig. 1…Chen's BERT model is a sequence-processing model: its input is a token sequence x = (x 1 … x T ) and its output is a per-token slot-label sequence y s plus an intent label; providing the user query to such a sequence-processing model to obtain the slot values directly constitutes the limitation under BRI, “Given an input token sequence x = (x 1 , …, x T ), the output of BERT is H = (h 1 , …, h T )”; , Section 3.2…”For slot filling, we feed the final hidden states of other tokens h 2 , …, h T into a softmax layer to classify over the slot filling labels”). Claim 8… implementing one or more large language models to determine a plurality of variables based on the query (Chen: Abstract, Section 1…Chen implements BERT, a pre-trained large-scale bidirectional transformer language model trained on BooksCorpus (800M words) plus English Wikipedia (2,500M words), to predict a plurality of variables (the intent label plus a sequence of slot labels) from the input user query; under BRI this constitutes implementing one or more large language models to determine a plurality of variables based on the query, “Bidirectional Encoder Representations from Transformers (BERT) (Devlin et al., 2018), was proposed and has created state-of-the-art models for a wide variety of NLP tasks… BERT is pre-trained on BooksCorpus (800M words) (Zhu et al., 2015) and English Wikipedia (2,500M words)”; Equations (1)-(3) and Table 2…The variables determined for each query (intent y i plus per-token slot variables y s n ) are produced jointly by softmax classification heads on top of the BERT encoder, being “plurality of variables based on the query” under BRI, “y i = softmax(W i h 1 + b i )… y s n = softmax(W s h n + b s ), n ∈ 1… N”). As it pertains to claims 5-8, 18 and 19, Tancik, Martin-Brualla and Chen are analogous art because they are each from the broader field of machine-learning-driven user-facing systems and address the shared problem of converting user inputs into structured parameters consumable by a downstream model. Tancik’s rendering pipeline already accepts a structured input (target viewpoint plus a numerical appearance/exposure parameter), and Chen teaches that a BERT-based joint intent-classification / slot-filling model is the state-of-the-art way to convert a natural-language user query into exactly that kind of structured (intent, slots) tuple. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to interpose Chen’s BERT-based natural-language understanding front-end between a user’s natural-language query and Tancik’s/Martin-Brualla’s appearance-conditioned rendering pipeline so that users could request rendered videos of a queried location at a selected weather/lighting state by typing natural-language descriptions instead of having to supply numerical exposure values and viewpoint coordinates directly. This is a routine combination of a known natural language understanding front-end with a known generative-rendering back-end and yields predictable results: the same renderings Tancik/Martin-Brualla already produce, but addressable by natural language. The suggestion/motivation for doing so would have been Chen’s explicit statement that BERT enables state-of-the-art intent classification and slot filling for goal-oriented dialogue systems (Chen: Section 1), combined with the long-recognized desirability of allowing end users to issue natural-language requests to image- and video-generation systems. Applying Chen’s NLU front-end to Tancik’s rendering pipeline directly produces the recited “receiving a query from a user relating to a location” followed by extracting or inferring the values for the conditions from that query . 07-21-aia AIA Claim 12 is rejected under 35 USC 103 as being unpatentable over Tancik in view of Martin-Brualla and further in view of Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions to Haque et al. (hereinafter Haque) . Per claim 12, Tancik combined with Martin-Brualla discloses claim 1. Tancik combined with Martin-Brualla does not expressly disclose but with Haque does teach the query comprises a text query that specifies one or more objects to be included in the scene and wherein the video depicts the one or more object included in the scene (Haque: Section 1 and Fig. 1…Haque takes a user text query (e.g., “Give him a cowboy hat”, “Turn the bear into a grizzly bear”) that specifies an object to be included in the NeRF scene, edits the scene accordingly, and renders the resulting NeRF such that the output video depicts the specified object included in the scene; under BRI this constitutes a text query specifying one or more objects to be included in the scene with the video depicting those objects, “we propose Instruct-NeRF2NeRF, a method for editing 3D NeRF scenes that requires as input only a text instruction… we can enable a wide variety of edits using flexible and expressive textual instruction such as “Give him a cowboy hat” or “Turn him into Albert Einstein”; Abstract…”Result videos can be found on the project website: https://instruct-nerf2nerf.github.io”). Tancik, Martin-Brualla and Haque are analogous art because they are all within the same field of endeavor, namely neural-radiance-field rendering of real-world scenes; indeed, Haque is co-authored by Tancik. They address the same problem solving area of producing photorealistic renderings of a captured scene, with Haque specifically targeting interactive user-driven adjustment of that scene via natural-language instructions. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to apply Haque’s text-instruction-based NeRF editing pipeline to Tancik’s per-location, conditioned Block-NeRFs so that an end user could issue an initial text query specifying objects to be rendered into the scene at the queried location. The combination is a straightforward stacking of Haque’s NeRF-editing front-end onto Tancik’s NeRF-rendering back-end, producing predictable results. The suggestion/motivation for doing so would have been Haque’s explicit statement that text-instruction-driven NeRF editing makes 3D scene editing accessible and intuitive for everyday users (Haque: Section 1), combined with Tancik’s already- large NeRF scenes that lend themselves to user-driven content adjustment. The two together directly produce the recited initial query plus objects . Allowable Subject Matter 12-151-08 AIA 07-43 12-51-08 Claim 14 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is the statement of reasons for the indication of allowable subject matter: The prior art disclosed by the applicant and cited by the Examiner fail to teach or suggest, alone or in combination, all the limitations of the independent claim 1, further including the particular notable limitations of receiving a further query from the user relating to the video; in response to receiving the further query, generating further conditioning parameters based at least in part on the further query, wherein the further conditioning parameters provide values for one or more further conditions associated with the scene to be rendered at the location; generating, using the generative machine-learned model, an adjusted video, wherein the adjusted video depicts the scene at the location and with the values for the one or more further conditions; and providing the adjusted video for presentation to the user. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN CHEN whose telephone number is (571)272-4143. The examiner can normally be reached M-F 10-7. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALAN CHEN/Primary Examiner, Art Unit 2125 Application/Control Number: 18/532,392 Page 2 Art Unit: 2125 Application/Control Number: 18/532,392 Page 3 Art Unit: 2125 Application/Control Number: 18/532,392 Page 4 Art Unit: 2125 Application/Control Number: 18/532,392 Page 5 Art Unit: 2125 Application/Control Number: 18/532,392 Page 6 Art Unit: 2125 Application/Control Number: 18/532,392 Page 7 Art Unit: 2125 Application/Control Number: 18/532,392 Page 8 Art Unit: 2125 Application/Control Number: 18/532,392 Page 9 Art Unit: 2125 Application/Control Number: 18/532,392 Page 10 Art Unit: 2125 Application/Control Number: 18/532,392 Page 11 Art Unit: 2125 Application/Control Number: 18/532,392 Page 12 Art Unit: 2125 Application/Control Number: 18/532,392 Page 13 Art Unit: 2125 Application/Control Number: 18/532,392 Page 14 Art Unit: 2125 Application/Control Number: 18/532,392 Page 15 Art Unit: 2125 Application/Control Number: 18/532,392 Page 16 Art Unit: 2125 Application/Control Number: 18/532,392 Page 17 Art Unit: 2125 Application/Control Number: 18/532,392 Page 18 Art Unit: 2125 Application/Control Number: 18/532,392 Page 19 Art Unit: 2125 Application/Control Number: 18/532,392 Page 20 Art Unit: 2125 Application/Control Number: 18/532,392 Page 21 Art Unit: 2125 Application/Control Number: 18/532,392 Page 22 Art Unit: 2125
Read full office action

Prosecution Timeline

Dec 07, 2023
Application Filed
May 21, 2026
Non-Final Rejection mailed — §103, §112
Aug 12, 2026
Examiner Interview Summary
Aug 12, 2026
Applicant Interview (Telephonic)
Aug 21, 2026
Response Filed
Sep 30, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737694
Architecture for Classification of a Decision Tree Ensemble and Method
3y 9m to grant Granted Sep 15, 2026
Patent 12737645
GRAPH MACHINE LEARNING MODEL BASED TECHNIQUES FOR EVALUATING KNOWLEDGE GRAPH DATASETS
3y 6m to grant Granted Sep 15, 2026
Patent 12737181
MANAGING OPERATIONAL RESILIENCE OF SYSTEM ASSETS USING AN ARTIFICIAL INTELLIGENCE MODEL
9m to grant Granted Sep 15, 2026
Patent 12731082
SMART COPY OPTIMIZATION IN CUSTOMER ACQUISITION AND CUSTOMER MANAGEMENT PLATFORMS
4y 8m to grant Granted Sep 08, 2026
Patent 12725070
GRADIENT-BASED QUANTUM ASSISTED HAMILTONIAN LEARNING
4y 0m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
91%
Grant Probability
98%
With Interview (+6.7%)
2y 9m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month