Prosecution Insights
Last updated: October 01, 2026
Application No. 19/015,191

VISUAL LOCALIZATION OF IMAGE VIEWPOINTS IN 3D SCENES USING NEURAL REPRESENTATIONS

Non-Final OA §102§103§112
Filed
Jan 09, 2025
Priority
Jan 12, 2024 — provisional 63/620,683
Examiner
KUDO, KEN
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
43 currently pending
Career history
40
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: reference characters 522 and 524 in FIG. 5A; and reference character 711, appearing three times in FIG. 11. With respect to FIG. 11, the corresponding description discusses inference and/or training logic 715. Applicant must correct 711 to 715, if intended, or otherwise make the drawing and description consistent. Applicant must either amend the specification to describe reference characters 522 and 524, provided such an amendment is supported by the original disclosure, or delete or correct the unintended reference characters. No new matter may be introduced. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. The drawings are objected to because: reference character 102 has been used to designate two different parts. Specifically, numeral 102 identifies a hill in FIG. 1A and a second building in FIG. 1B. In FIG. 11, lead lines associated with external graphics processor 1112 and graphics processors 1108 refer to reference numeral 711, whereas the corresponding hardware structure described in paragraphs [0069–0070] and shown across FIGS. 7A, 7B, 8, 9, 12, and 14 is designated with reference character 715. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification The disclosure is objected to because of the following informalities: [0025] uses “neural rendering field (NERF),” whereas the remainder of the disclosure generally uses “neural radiance field (NeRF).” Applicant must use the intended term and acronym consistently. [0025], “such an inference microservice” apparently should read “such as an inference microservice.” [0025], “model deployment an execution software” apparently should read “model deployment and execution software.” [0027] recites “between 3,600 matches between coarse image patches and 3D points.” One occurrence of “between” should be deleted or the sentence otherwise corrected. [0029], “digital of virtual representation” apparently should read “digital or virtual representation.” [0033] states that “training image 130 is illustrated in FIG. 1C.” Training image 130 is illustrated in FIG. 1B; FIG. 1C illustrates query image 160. The figure reference must be corrected. [0038], “reference posed” apparently should read “reference poses.” [0040], “ -- fine matching one a set” apparently should read “-- fine matching once a set.” [0041], “course matching process” apparently should read “coarse matching process.” [0045–0046] identify viewing direction d as belonging to a two-dimensional space, but subsequently use d with a three-dimensional origin o in the ray equation r(t)=o+td. Applicant must correct or clarify the dimensional notation consistently. [0049], “their associating NeRF features” apparently should read “their associated NeRF features.” [0064], “a shader 558” is inconsistent with FIG. 5B, which identifies 558 as light-sample generation and 562 as shader(s). Applicant must correct the reference numeral to 562, if intended. [0065], “client device that include” apparently should read “client device that includes.” [0066] refers to “content application 604,” whereas FIG. 6 identifies 604 as “Control Application.” Applicant must make the terminology consistent. [0066] refers to “visual localization module 6320,” whereas FIG. 6 identifies the visual-localization module by numeral 632. Applicant must correct the reference numeral. [0080] refers to “dynamic read-only memory.” This terminology is internally inconsistent. Applicant must identify and consistently recite the intended memory type, such as dynamic random-access memory if that is what was intended. [0115], “discreet external graphics processor” apparently should read “discrete external graphics processor.” [0116] refers to “graphics processor 1500.” Numeral 1500 identifies the process of FIG. 15A and does not correspond to the graphics processors in FIG. 11, which are identified as 1108 and 1112. Applicant must identify the intended graphics processor and correct the reference numeral accordingly. [0126] refers to “graphics core(s) 1202A–1202N,” whereas FIG. 12 identifies those components as processor cores 1202A–1202N. Applicant must make the terminology consistent. [0126] refers to “graphics processor 1200,” whereas FIG. 12 identifies graphics processor 1208. Applicant must correct the reference numeral. [0129], “prcocessing” should read “processing.” [0166] identifies pre-trained models by reference numerals 1506 and 1306. FIG. 15A identifies the pre-trained models as 1406, the customer dataset as 1506, and the deployment system as 1306. Applicant must correct each occurrence that purports to identify the pre-trained models and make the accompanying grammar consistent, such as “pre-trained models 1406 are trained using,” if intended. Appropriate correction is required. Claim Objections Claims 9, 16, and 18 are objected to because of the following informalities: Claim 9 recites “starting from locations for the set of coarse matches.” The phrase is grammatically incomplete or nonidiomatic and does not clearly state the intended relationship between the locations and the coarse matches. Applicant must supply the intended wording, such as “-- locations corresponding to the set of coarse matches” or “-- locations selected based on the set of coarse matches,” as appropriate. Because those alternatives may differ in scope. Claim 16 recites “a neural field representation (NeRF).” This acronym expansion is inconsistent with claims 1 and 9 and the remainder of the disclosure, which use “neural radiance field (NeRF).” Applicant must make the terminology and acronym consistent. Claim 18 recites “in a vicinity of an determined image pixel.” Applicant must replace “an” with “a” or “the,” as intended. Appropriate correction is required. Claim Rejections - 35 USC § 112(a) The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim 18 is rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 18 recites limitation "increase a feature dimensionality of individual point features associated with the coarse matches". The Examiner acknowledges that this claim language is reproduced in paragraphs [0230-0231] of the specification, which recite, as part of the "clauses" describing various embodiments: "18. The system of clause 17, wherein the one or more processors are further to [0231] increase a feature dimensionality of individual point features associated with the coarse matches; and [0232] align the individual point features, with the increased feature dimensionality, with corresponding local map features...". However, paragraphs [0230-0232] merely restate the claim language verbatim and do not provide any independent technical disclosure, such as an explanation of the mechanism, a worked numerical example, or any description of how or why the recited dimensionality increase is achieved, beyond what is already recited in claim 18 itself. A paragraph that simply mirrors claim language, without more, does not by itself demonstrate that the inventor possessed the claimed subject matter as of the filing date. See MPEP 2163(I) (the written description requirement is not satisfied by mere reproduction of the claim language in the specification without additional descriptive support); Ariad Pharms., Inc. v. Eli Lilly & Co., 598 F.3d 1336 (Fed. Cir. 2010) (en banc) (written description requires the specification to reasonably convey to a person of ordinary skill in the art that the inventor had possession of the claimed invention). The only portions of the specification providing an actual technical disclosure of the dimension-adjustment step at issue, i.e., the step preceding alignment of an individual 3D point (coarse-match) feature with a corresponding fine-level local image feature map, are paragraphs [0040] and [0053]. These paragraphs disclose the opposite of what is claimed: Paragraph [0053] discloses that "the cross-attended feature obtained from the coarse matching can be processed through a linear layer to adjust its feature dimension from Dc to Df," where Dc is the dimension of the coarse-level feature map and Df is the dimension of the fine-level feature map (see paragraph [0050], defining the coarse-level feature map as Fcm ∈ RNm x Dc and the fine-level feature map as Ffm . Paragraph [0040] provides the only concrete numerical embodiment of this adjustment, disclosing that "point NeRF features in a feature map used for fine matching may have a resolution of 256 by default, but may need to be downsampled to a value such as 128, which is the dimension of a corresponding high resolution feature map." This passage establishes that, in the disclosed embodiment, the coarse-level dimension (Dc, exemplified as 256) is greater than the fine-level dimension (Df, exemplified as 128), such that the adjustment from Dc to Df described in paragraph [0053] is a decrease (downsampling), not an increase. Because the specification's only enabling technical disclosure of this dimension-adjustment step teaches a decrease in feature dimensionality, and no portion of the specification, including paragraphs [0230-0231], which add no independent technical content beyond restating the claim, describes, illustrates, or otherwise conveys possession of an embodiment in which this dimensionality is increased, claim 18 is not adequately supported by the written description. Applicant is required to either (i) point to specific portions of the specification, apart from the claims or the "clauses" of paragraphs [0174-0256] themselves, that describe an embodiment in which feature dimensionality is increased at this step, or (ii) amend claim 18 to remove the directional limitation (e.g., reciting "adjusting" or "modifying" the feature dimensionality, consistent with paragraph [0053] and with claims 4 and 11) or to otherwise conform the claim to the disclosed decrease/ downsampling embodiment. Claim Rejections - 35 USC § 112(b) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3-5, 11-12, and 17-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 3, 10, and 17 recite the limitation “an image patch feature from a fine feature map.” Independent claims 1, 9, and 16 (from which claims 3, 10, and 17 respectively depend) recite only "an image feature map" (claim 1) or "a feature map" (claim 16). There is insufficient antecedent basis for this limitation in the claim. It is unclear whether the newly-introduced "fine feature map" of claims 3, 10 and 17 is the same feature map recited in the independent claim, a different resolution/ version of that same feature map (consistent with the two-resolution embodiment of claims 7, 14, and the specification at paragraph [0040]), or an entirely separate and distinct feature map. The metes and bounds of claims 3, 10, and 17 cannot be determined with reasonable certainty as a result. Claims 4, 11, and 18, which depend (directly or indirectly) from claims 3, 10, and 17, respectively, inherit the above indefiniteness and fail to cure the deficiency. Claims 4, 11, and 18 recite the limitation "individual point features" add further unantecedented terms of their own. This term is introduced in claims 4, 11, and 18 without antecedent basis. Independent claims 1, 9, and 16 recite "reference feature vectors" corresponding to 3D points, but the claims never establish that "point features" refers to these same reference feature vectors (or, per the specification at paragraph [0053], to the cross-attended features obtained during coarse matching). It is unclear whether "individual point features" is the same as, or different from, the previously-claimed "reference feature vectors.", "corresponding local map features" (claims 4, 11) and "local map features" (claim 18); this term likewise lacks antecedent basis and is not tied to the previously-recited "fine feature map" or "image patch feature.". Claim 5 recites the limitation "analyzing a probability distribution corresponding to determined probabilities to obtain a set of refined feature matches" in claim. However, claim 3 (from which claim 5 depends) contains no recitation of "probabilities" of any kind. The term "probabilities" is first introduced in claim 4, a claim that is not in claim 5's chain of dependency (claim 5 depends from claim 3, not claim 4). There is accordingly no antecedent basis in claim 5's dependency chain for "determined probabilities," rendering the scope of claim 5 indefinite. It appears claim 5 was intended to depend from claim 4, which recites "aligning individual point features with corresponding local map features to determine probabilities..." and would supply the missing antecedent. Claim 12 is rejected for the same reason. Claim 12 depends directly from claim 9, which contains no recitation of "probabilities." The term is first introduced in claim 11 (which depends from claim 10, which depends from claim 9). Claim 12 does not depend from claim 11, and therefore lacks antecedent basis for "determined probabilities." It appears claim 12 was intended to depend from claim 11. Claim 19 is rejected for the same reason. Claim 19 depends directly from claim 16, which contains no recitation of "probabilities." The term is first introduced in claim 18 (which depends from claim 17, which depends from claim 16). Claim 19 does not depend from claim 18, and therefore lacks antecedent basis for "determined probabilities." It appears claim 19 was intended to depend from claim 18. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1–17, and 19–20 are rejected under 35 U.S.C. §102(a)(1) as being anticipated by Liu (Liu et al. (2023). NeRF-Loc: Visual Localization with Conditional Neural Radiance Field. arXiv.Org). Regarding claim 1, Liu teaches a computer-implemented method, comprising: generating an image feature map corresponding to a two-dimensional (2D) query image; ( [Fig. 1], [§ III]: Liu processes a query image through a 2D backbone to generate coarse and fine query-feature maps. For coarse matching, Liu reshapes the coarse target/ query feature map into image descriptors (D2d ∈ RHW ×C). The disclosed pipeline is computer-implemented using neural networks and a GPU. ) identifying, as a set of coarse matches, a number of reference feature vectors closest to image feature vectors of the image feature map in a latent feature space, the reference feature vectors corresponding to three-dimensional (3D) points in a 3D environment reconstructed using a neural radiance field (NeRF) trained on a set of reference images of the 3D environment; and ( [§§ III.A; Figs. 1–2; Eq. (1)], [§§ III.C]: Liu represents the known 3D scene using a conditional/ generalizable NeRF conditioned on posed support/ reference images and depth maps. Sampled 3D reference points (X) are supplied to the neural 3D model (Mc) to produce respective point descriptors (D3d = Mc(X) ∈RN ×C). Liu transforms the 3D descriptors and the query-image descriptors using self- and cross-attention, computes the learned descriptor-space correlation score (S = sigmoid(mlp(D′3d ⊙D′2d)) ∈RN ×HW), and selects coarse 3D–2D correspondences using a score threshold and mutual-nearest checking. The mutually nearest descriptors having the best learned correlation are the claimed closest feature vectors in the latent feature space. ) determining, using a fine matching network and starting from locations selected based on the set of coarse matches, a camera pose corresponding to the 2D query image with respect to the 3D environment. ( [§§ III.C], [Fig. 1]: "In fine level matching, for each matched 3D-2D pair from coarse matching, we construct the fine level 2D descriptors Dfine2d ∈ R M × 7 × 7 × C by taking a feature patch centered at the coarse 2D location, and query fine level 3D descriptors Dfine3d = M f X m a t c h e d ∈ R M × C ". For every coarse 3D–2D correspondence, Liu extracts a (7×7) fine-feature patch centered at the coarse 2D location and obtains the fine descriptor for the matched 3D point. A self/ cross-attention fine matcher calculates a fine correlation distribution and refines the 2D location by spatial expectation. Liu then supplies the resulting 3D–2D correspondences to a PnP solver with RANSAC to determine the query camera pose (R,t). ) Regarding claim 2, Liu teaches the computer-implemented method of claim 1, wherein the camera pose is determined for at least one query image of a set of 2D query images without retraining the NeRF. ( [Page 2 > § III], [Page 3 > Fig. 1], [Page 5 > §§ IV.A – IV.C]: Liu separates multi-scene pretraining and per-scene optimization from test-time inference. Liu evaluates sets containing hundreds or thousands of frames using disclosed train/ test splits. After training, support images are selected at test time and each query frame is localized through the inference pipeline in approximately 250 ms. Accordingly, test-query images are localized using the already-trained conditional NeRF without repeating the disclosed training stages for each query image. ) Regarding claim 3, Liu teaches the computer-implemented method of claim 1, further comprising: obtaining, for individual coarse matches of the set of coarse matches, an image patch feature from a fine feature map associated with a corresponding location of the individual coarse match. ( [§§ III.C], [Fig. 1]: For each matched 3D–2D pair produced by coarse matching, Liu constructs fine-level 2D descriptors (Dfine2d ∈ R M × 7 × 7 × C ) by taking a (7×7) feature patch from the fine query-feature map, centered at the corresponding coarse 2D location. ) Regarding claim 4, Liu teaches the computer-implemented method of claim 3, further comprising: aligning individual point features with corresponding local map features to determine probabilities of the individual point features corresponding to pixels in a vicinity of a determined image pixel. ( [§§ III.C], [Fig. 1]: For each coarse match, Liu obtains the descriptor (Dfine3d = M f X m a t c h e d ) of the corresponding 3D point and a local (7×7) image-feature patch centered at the coarse image pixel. Liu aligns/ transforms the point descriptor and local image descriptors using self- and cross-attention and calculates correlation matrix (Sfine = softmax mlp D 3 d f i n e ' ⊙ D 2 d f i n e ' ). The softmax values represent the relative probabilities that the individual 3D point corresponds to respective pixels within the local (7×7) vicinity. ) Regarding claim 5, Liu teaches the computer-implemented method of claim 3, further comprising: analyzing a probability distribution corresponding to determined probabilities to obtain a set of refined feature matches; and ( [§§ III.C], [Fig. 1 and the (Sfine) formula]: Liu analyzes the softmax correlation distribution (Sfine) by calculating its spatial expectation (E^). The resulting sub-pixel 2D coordinates refine the respective coarse 3D–2D correspondences, collectively producing a set of refined feature matches. ) determining the camera pose based in part upon the set of refined feature matches. ( [§ III], [Fig. 1]: After the fine 3D–2D correspondences have been established, Liu processes those correspondences using PnP with RANSAC. Figure 1 routes the fine correlation and 2D-location-expectation results to the PnP block, which outputs the camera pose (R,t). ) Regarding claim 6, Liu teaches the computer-implemented method of claim 1, wherein the camera pose is determined by processing, using a perspective-n-point (PnP) solver, a set of fine matches output by the fine matching network. ( [§ III], [Fig. 1], and [§§ III.C]: Liu’s coarse-to-fine transformer matcher refines the matched 2D coordinates using a fine softmax correlation distribution and spatial expectation. Liu then processes the resulting fine 3D–2D correspondences using “a basic PnP solver with Ransac” to compute the camera pose (R,t). ) Regarding claim 7, Liu teaches the computer-implemented method of claim 1, further comprising: generating image feature maps for the 2D query image at two different resolutions, wherein a lower resolution 2D feature map of the image feature maps is used to determine the set of coarse matches, and wherein a higher resolution 2D feature map of the image feature maps is used by the fine matching network to determine the camera pose. ( [Fig. 1], [§§ III.B – III.C]: Liu’s 2D backbone generates separately depicted coarse and fine query-feature maps, with Figure 1 depicting the fine-feature grid as spatially denser than the coarse-feature grid. Liu reshapes the coarse target/ query map to obtain descriptors used for coarse matching, while the fine map supplies (7×7) local patches used for fine matching and sub-pixel refinement before PnP pose estimation. Liu additionally describes these features as levels of an image-feature pyramid. "Image feature pyramid is defined as Fi, where F0 refers to the original RGB image". "These two level 3D models are denoted as Mc and Mf, where Mc is based on F0 and F3, Mf uses F0 and F2": F3 corresponds to a coarser (lower-resolution, deeper pyramid level) feature map used to derive D2d for coarse matching, while F2 corresponds to a finer (higher-resolution, shallower pyramid level) feature map used to derive Dfine2d for fine-level matching to determine the camera pose. ) Regarding claim 8, Liu teaches the computer-implemented method of claim 7, wherein the fine matching network receives input based upon feature sampling of the set of coarse matches and the higher resolution 2D image feature map. ( [§§ III.C], [Fig. 1]: For every 3D–2D pair selected by coarse matching, Liu samples a (7×7) patch centered at the coarse 2D location from the fine query-feature map and obtains the fine 3D descriptor at ( X m a t c h e d ). Liu supplies these sampled descriptors to the fine self/ cross-attention and correlation operations; the fine level 2D descriptor input to the fine matching module (self/cross-attention transformation and correlation, id.) is obtained by sampling (taking a feature patch) the higher-resolution feature map (F2-based) at locations determined by (centered at 7×7) the set of coarse matches. ) Regarding claims 9-10, 12, 16-17 and 19, the rationale provided in the rejection of claims 1, 3 and 5 is incorporated herein. Further, the computer-implemented method of claims 1, 3 and 5 corresponds to the processor of claims 9-10 and 12, as well as the system of claims 16-17 and 19, ( [§§ IV.C, “Implementation Details”]: Liu's system performs the disclosed localization process on an NVIDIA V100 GPU. ), and performs the steps disclosed herein. Therefore, the claims are all rejected. Regarding claims 11 and 13-14, the rationale provided in the rejection of claims 4 and 6-7 is incorporated herein. Further, the computer-implemented method of claims 4 and 6-7 corresponds to the processor of claims 11 and 13-14 ( [§§ IV.C, “Implementation Details”]: Liu's system performs the disclosed localization process on an NVIDIA V100 GPU. ), and performs the steps disclosed herein. Therefore, the claims are all rejected. Regarding claim 15, Liu teaches the at least one processor of claim 9, wherein the processor is comprised in at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing light transport simulation; a system for rendering graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; a system for generating or presenting mixed reality (MR) content; a system incorporating one or more Virtual Machines (VMs); a system implemented at least partially in a data center; a system for performing hardware testing using simulation; a system for synthetic data generation; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a vision language model (VLM); a system implemented using one or more multi-modal language models; a system using or deploying one or more inference microservices; a system that incorporates one or more machine learning models deployed in a service or microservice along with an OS-level virtualization package (e.g., a container); a collaborative content creation platform for 3D assets; or a system implemented at least partially using cloud computing resources. ( [Abstract], [Fig. 1], [§§ III.C], and [IV.C, “Implementation Details”]: Liu’s system performs deep-learning operations using a conditional NeRF, a 2D neural-network backbone, self- and cross-attention transformer layers, an MLP, and coarse- and fine-level learned matching. Liu trains the network through multi-scene pretraining and per-scene optimization and performs the resulting localization operations on an NVIDIA V100 GPU. Accordingly, Liu’s processor is comprised in a system for performing deep-learning operations. Additionally, Liu’s conditional NeRF includes an auxiliary rendering head that performs neural rendering and is trained using a rendering loss, thereby also teaching a system for rendering graphical output. ) Regarding claim 20, the rationale provided in the rejection of claim 15 is incorporated herein. Further, the processor of claim 15 corresponds to the system of claim 20 ( [§§ IV.C, “Implementation Details”]: Liu's system performs the disclosed localization process on an NVIDIA V100 GPU. ), and performs the steps disclosed herein. Therefore, the claim is rejected. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 18 is rejected under 35 U.S.C. §103 as being unpatentable over Liu in view of Yuan (Yuan et al, US 2024/0177353 A1, 2024). Regarding claim 18, Liu teaches the system of claim 17, wherein the one or more processors are further to: While Liu teaches coarse-to-fine 3D-to-2D feature matching and local probability refinement via correlation tensors, Liu does not explicitly disclose increasing a feature dimensionality of coarse point features prior to local map alignment, where Yuan teaches: increase a feature dimensionality of individual point features associated with the coarse matches; and ( [0048–0052]], [Fig. 5]: Yuan begins with an individual point feature having (C) feature dimensions, duplicates the feature (r) times, and attaches an (m)-dimensional vector to each duplicated feature in the feature dimension. Yuan identifies (m=2) as an example and expressly states that the resulting feature dimension of each point is (C+2). Therefore, Yuan increases the feature dimensionality of an individual point feature from (C) to (C+m), including specifically from (C) to (C+2) associated with the "first feature" obtained during the "coarse feature expansion" stage of Yuan's point cloud refinement pipeline. ) It would have been obvious to one of ordinary skill in the art before the effective filing date to apply Yuan’s (m)-dimensional point-feature expansion to Liu’s point descriptors associated with the coarse matches before Liu’s attention-based fine matching. Both references process learned 3D point features using MLP and attention mechanisms. Liu seeks improved 3D descriptors and more accurate 3D–2D matching, while Yuan teaches that attaching the additional vector differentiates point features, enables the MLP to extract more feature details, and enhances feature interaction. Yuan also places the feature expansion immediately before MLP and attention processing, corresponding to Liu’s placement of the matched point descriptors immediately before fine self/cross-attention and correlation. The modification therefore would have applied a known point-feature augmentation technique at its ordinary insertion point, predictably producing richer and more distinguishishable point descriptors for Liu’s local matching operation, with a reasonable expectation of success and without changing Liu’s principle of operation. Accordingly, Liu [as modified by Yuan] continue to teach: align the individual point features, with the increased feature dimensionality, with corresponding local map features to determine probabilities of the individual point features corresponding to pixels in a vicinity of an determined image pixel. ( [Yuan > 0048–0052 and Fig. 5], [Liu > §§ III.C]: Yuan teaches increasing each individual point feature from (C) dimensions to (C+m) dimensions before MLP and attention processing. Liu teaches that, for each coarse match, the corresponding point descriptor (Dfine3d = M f X m a t c h e d ) is processed with a (7×7) local image-feature patch using self- and cross-attention. Liu then correlates the transformed point feature and local-map features and calculates (Sfine = softmax mlp D 3 d f i n e ' ⊙ D 2 d f i n e ' ). Accordingly, Yuan’s dimensionally expanded (C+m) point feature is supplied as the point-feature input to Liu’s fine-level attention and local-map correlation operation, with the attention projections configured to a common correlation width. Liu’s softmax operation then determines the probabilities that the dimensionally expanded point feature corresponds to respective pixels within the local vicinity. Therefore, Liu [as modified by Yuan] increases feature dimensionality, then calculates the alignment with corresponding local-map features and the determination of nearby-pixel probabilities. ) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. KEN KUDO Examiner Art Unit 2671 /KEN KUDO/Examiner, Art Unit 2671 /VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Jan 09, 2025
Application Filed
Aug 28, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month