Prosecution Insights
Last updated: October 02, 2026
Application No. 19/010,943

PERFORMING A THREE-DIMENSIONAL COMPUTER VISION TASK USING A NEURAL RADIANCE FIELD GRID REPRESENTATION OF A SCENE PRODUCED FROM TWO-DIMENSIONAL IMAGES OF AT LEAST A PORTION OF THE SCENE

Non-Final OA §103
Filed
Jan 06, 2025
Priority
Jan 04, 2024 — provisional 63/617,446 +1 more
Examiner
SALVUCCI, MATTHEW D
Art Unit
Tech Center
Assignee
GEORGIA TECH RESEARCH Corporation
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
357 granted / 494 resolved
+12.3% vs TC avg
Strong +27% interview lift
Without
With
+27.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
27 currently pending
Career history
512
Total Applications
across all art units

Statute-Specific Performance

§101
4.6%
-35.4% vs TC avg
§103
62.9%
+22.9% vs TC avg
§102
16.1%
-23.9% vs TC avg
§112
14.0%
-26.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 494 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4, 13, and 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over Barron et al. (NPL: Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields), hereinafter Barron, in view of Lorraine et al. (US Pub. 2025/0124640), hereinafter Lorraine. Regarding claim 1, Barron discloses a system, comprising: a processor (Section G: use a larger model: We have a view-dependent bottleneck size of 256, and we process those bottleneck vectors with a 3-layer MLP with 256 hidden units and a skip connection from the bottleneck to the second layer); and a memory (Section G: ablations in the paper caused training to run out of memory, so those experiments use half of the batch size used in other experiments and are trained for twice as many iterations) storing: a neural radiance field grid network module including instructions that, when executed by the processor, cause the processor to produce, from two-dimensional images of at least a portion of a scene, three-dimensional patches of a neural radiance field grid representation of the scene (Section 2: Mip-NeRF uses features that approximate the integral over the positional encoding of the coordinates within a sub volume, which is a conical frustum. This results in Fourier features whose amplitudes are small when the feature sinusoid’s period is larger than the standard deviation of the Gaussian — the features express the spatial location of the sub-volume only at the wavelengths that are larger than the size of the sub-volume. Because this feature encodes both position and scale, the MLP that consumes it is able to learn a multi-scale representation of the 3D scene that renders anti-aliased images. Grid-based representations like iNGP do not natively allow for sub-volumes to be queried, and in stead use trilinear interpolation at a single point to construct features for use in an MLP, which results in learned models that cannot reason about scale or aliasing); a three-dimensional shifted window visual transformer module including instructions that, when executed by the processor, cause the processor to produce, from the three-dimensional patches, a feature map (Section C: Along with features fℓ we also average and concatenate a featurized version of {ωj,ℓ }for use as input to our MLP…This feature takes ωj (shifted and scaled to [−1,1]) and scales it by the standard deviation of the values in Vℓ). Barron does not explicitly disclose at least one of: a first decoder module including instructions that, when executed by the processor, cause the processor to produce, from the feature map, the neural radiance field grid representation of the scene to train the system, or a second decoder module including instructions that, when executed by the processor, cause the processor to produce, from the feature map, the neural radiance field grid representation of the scene to perform a three-dimensional computer vision task for a cyber-physical system. However, Lorraine teaches 3D representations using NeRFs (Paragraph [0057]), further comprising: at least one of: a first decoder module including instructions that, when executed by the processor, cause the processor to produce, from the feature map, the neural radiance field grid representation of the scene to train the system (Fig. 3; Paragraphs [0067]-[0068]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder… FIG. 3 illustrates an example of a stratified sampling process 300 for sampling one or more data values from a sampling range, according to at least one embodiment. In at least one embodiment, sampling process 300 may be used to sample any stochastic parameters used in training one or more neural networks, for example as shown in FIG. 1, one or more camera orientation parameters, a diffusion timestep, a noise sample, an image type, a data augmentation parameter (e.g., crop, rotation, etc.), a text prompt, and/or the like), or a second decoder module including instructions that, when executed by the processor, cause the processor to produce, from the feature map, the neural radiance field grid representation of the scene to perform a three-dimensional computer vision task for a cyber-physical system (Fig. 9A; Paragraph [0067]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder; Paragraph [0132]: FIG. 9A illustrates an example of an autonomous vehicle 900, according to at least one embodiment. In at least one embodiment, autonomous vehicle 900 (alternatively referred to herein as “vehicle 900”) may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and/or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 900 may be a semi-tractor-trailer truck used for hauling cargo. In at least one embodiment, vehicle 900 may be an airplane, robotic vehicle, or other kind of vehicle). Lorraine teaches that this will add realism to two-dimensional (2D) and/or three-dimensional (3D) models (Paragraph [0537]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barron with the features of above as taught by Lorraine so as to add realism to models as presented by Lorraine. Regarding claim 4, Barron, in view of Lorraine teaches the system of claim 1, Barron discloses wherein: the neural radiance field grid network module includes a neural radiance field network, and the neural radiance field network lacks a three-dimensional shifted window visual transformer (Section 2: 3D coordinates are rotated into world coordinates by multiplying them by an orthonormal basis whose third vector is the ray direction (and whose first two vectors are an arbitrary frame that is perpendicular to the ray) and then shifted by the ray origin. By construction, the sample means and variances (along the ray and perpendicular to the ray) of these multisamples exactly match those of the conical frustum, analogously to mip-NeRF Gaussians; Section C: Along with features fℓ we also average and concatenate a featurized version of {ωj,ℓ }for use as input to our MLP…This feature takes ωj (shifted and scaled to [−1,1]) and scales it by the standard deviation of the values in Vℓ). Regarding claim 13, Barron, in view of Lorraine teaches the system of claim 1, wherein: a set of three-dimensional shifted windows visual transformer modules comprises the three-dimensional shifted windows visual transformer module and at least one other three-dimensional shifted windows visual transformer module (Barron: Section 3: interlevel loss used in mip-NeRF 360 is invariant to mono tonic transformations of distance, so it is unaffected by the choice of g(·). However, the prefiltering in our anti-aliased loss removes this invariance, and using mip-NeRF 360’s g(·) in our model results in catastrophic failure, so we must construct a new normalization. To do this, we construct a novel power transformation; Section 4: Individually disabling (E) multi sampling and (F) downweighting lets us measure the impact of each component; they both contribute evenly. (G) Not using ω as a feature hurts performance slightly. Alter native multisampling approaches like (H) sampling 6 random points from each mip-NeRF Gaussian or (G) using 7 unscented transform control points from each Gaussian [14] as multisamples are competitive alternatives to our hexagonal pattern, but have slightly worse speeds and/or accuracies. (J) Disabling our anti-aliased interlevel loss has little effect on our single-image metrics, but causes z-aliasing artifacts in our video results), at least one of: a set of first decoder modules comprises the first decoder module and at least one other first decoder module (Lorraine: Fig. 3; Paragraphs [0067]-[0068]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder… FIG. 3 illustrates an example of a stratified sampling process 300 for sampling one or more data values from a sampling range, according to at least one embodiment. In at least one embodiment, sampling process 300 may be used to sample any stochastic parameters used in training one or more neural networks, for example as shown in FIG. 1, one or more camera orientation parameters, a diffusion timestep, a noise sample, an image type, a data augmentation parameter (e.g., crop, rotation, etc.), a text prompt, and/or the like), or a set of second decoder modules comprises the second decoder module and at least one other second decoder module (Lorraine: Fig. 9A; Paragraph [0067]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder; Paragraph [0132]: FIG. 9A illustrates an example of an autonomous vehicle 900, according to at least one embodiment. In at least one embodiment, autonomous vehicle 900 (alternatively referred to herein as “vehicle 900”) may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and/or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 900 may be a semi-tractor-trailer truck used for hauling cargo. In at least one embodiment, vehicle 900 may be an airplane, robotic vehicle, or other kind of vehicle), a set of feature maps comprises the feature map and at least one other feature map (Lorraine: Fig. 3; Paragraphs [0067]-[0068]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder… FIG. 3 illustrates an example of a stratified sampling process 300 for sampling one or more data values from a sampling range, according to at least one embodiment. In at least one embodiment, sampling process 300 may be used to sample any stochastic parameters used in training one or more neural networks, for example as shown in FIG. 1, one or more camera orientation parameters, a diffusion timestep, a noise sample, an image type, a data augmentation parameter (e.g., crop, rotation, etc.), a text prompt, and/or the like), the three-dimensional shifted windows visual transformer module is: configured to produce the feature map, and connected to the first decoder module or the second decoder module (Lorraine: Fig. 3; Paragraphs [0067]-[0068]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder… FIG. 3 illustrates an example of a stratified sampling process 300 for sampling one or more data values from a sampling range, according to at least one embodiment. In at least one embodiment, sampling process 300 may be used to sample any stochastic parameters used in training one or more neural networks, for example as shown in FIG. 1, one or more camera orientation parameters, a diffusion timestep, a noise sample, an image type, a data augmentation parameter (e.g., crop, rotation, etc.), a text prompt, and/or the like), and the at least one other three-dimensional shifted windows visual transformer module is: configured to produce the at least one other feature map, and connected to the at least one other first decoder module or the at least one other second decoder module (Lorraine: Fig. 3; Paragraphs [0067]-[0068]: denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder… FIG. 3 illustrates an example of a stratified sampling process 300 for sampling one or more data values from a sampling range, according to at least one embodiment. In at least one embodiment, sampling process 300 may be used to sample any stochastic parameters used in training one or more neural networks, for example as shown in FIG. 1, one or more camera orientation parameters, a diffusion timestep, a noise sample, an image type, a data augmentation parameter (e.g., crop, rotation, etc.), a text prompt, and/or the like). Regarding claim 15, Barron, in view of Lorraine teaches the system of claim 13, Lorraine discloses wherein: the feature map is associated with a first degree of resolution, the at least one other feature map is associated with at least one other degree of resolution, and the first degree of resolution is larger than the at least one other degree of resolution (Paragraph [0067]: In at least one embodiment, denoising model 212 may be implemented through a structure referred to as “U-Net.” In at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps. In at least one embodiment, each up-sampling layer and an associated one or more convolutional layers may be considered a decoder). Regarding claim 16, Barron, in view of Lorraine teaches the system of claim 1, Lorraine discloses wherein the cyber-physical system comprises at least one of a robot or an automated vehicle (Fig. 9A; Paragraph [0132]: FIG. 9A illustrates an example of an autonomous vehicle 900, according to at least one embodiment. In at least one embodiment, autonomous vehicle 900 (alternatively referred to herein as “vehicle 900”) may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and/or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 900 may be a semi-tractor-trailer truck used for hauling cargo. In at least one embodiment, vehicle 900 may be an airplane, robotic vehicle, or other kind of vehicle). Regarding claim 17, the limitations of this claim substantially correspond to the limitations of claim 1; thus they are rejected on similar grounds. Regarding claim 18, Barron, in view of Lorraine teaches the method of claim 17, Lorraine discloses wherein the three-dimensional computer vision task for the cyber-physical system comprises a three-dimensional computer vision task for a control of a motion of the cyber-physical system (Paragraphs [0132]-[0135]: FIG. 9A illustrates an example of an autonomous vehicle 900, according to at least one embodiment. In at least one embodiment, autonomous vehicle 900 (alternatively referred to herein as “vehicle 900”) may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and/or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 900 may be a semi-tractor-trailer truck used for hauling cargo. In at least one embodiment, vehicle 900 may be an airplane, robotic vehicle, or other kind of vehicle… a steering system 954, which may include, without limitation, a steering wheel, is used to steer vehicle 900 (e.g., along a desired path or route) when propulsion system 950 is operating (e.g., when vehicle 900 is in motion). In at least one embodiment, steering system 954 may receive signals from steering actuator(s) 956. In at least one embodiment, a steering wheel may be optional for full automation (Level 5) functionality. In at least one embodiment, a brake sensor system 946 may be used to operate vehicle brakes in response to receiving signals from brake actuator(s) 948 and/or brake sensors). Regarding claim 19, Barron, in view of Lorraine teaches the method of claim 17, Lorraine discloses wherein the three-dimensional computer vision task comprises at least one of: an object detection operation performed on the neural radiance field grid representation of the scene, a semantic labeling operation performed on the neural radiance field grid representation of the scene (Paragraph [0191]: described herein allow for multiple neural networks to be performed simultaneously and/or sequentially, and for results to be combined together to enable Level 3-5 autonomous driving functionality. For example, in at least one embodiment, a CNN executing on a DLA or a discrete GPU (e.g., GPU(s) 920) may include text and word recognition, allowing reading and understanding of traffic signs, including signs for which a neural network has not been specifically trained. In at least one embodiment, a DLA may further include a neural network that is able to identify, interpret, and provide semantic understanding of a sign, and to pass that semantic understanding to path planning modules running on a CPU Complex), or a super-resolution imaging operation performed on the neural radiance field grid representation of the scene (Paragraph [0067]: at least one embodiment, U-Net structure may include a series of convolutional layers and pooling layers which generate progressively lower resolution multi-channel feature maps. In at least one embodiment, each pooling layer and an associated one or more convolutional layers may be considered an encoder. In at least one embodiment, the convolutional and pooling layers (i.e., encoders) may be followed by a series of up-sampling layers and convolutional layers which generate progressively higher resolution multi-channel feature maps; Paragraph [0649]: one or more components of systems and/or processors disclosed above can communicate with one or more CPUs, ASICs, GPUs, FPGAs, or other hardware, circuitry, or integrated circuit components that include, e.g., an upscaler or upsampler to upscale an image, an image blender or image blender component to blend, mix, or add images together, a sampler to sample an image (e.g., as part of a DSP), a neural network circuit that is configured to perform an upscaler to upscale an image (e.g., from a low resolution image to a high resolution image), or other hardware to modify or generate an image, frame, or video to adjust its resolution, size, or pixels; one or more components of systems and/or processors disclosed above can use components described in this disclosure to perform methods, operations, or instructions that generate or modify an image). Regarding claim 20, the limitations of this claim substantially correspond to the limitations of claim 1; thus they are rejected on similar grounds. Allowable Subject Matter Claims 2, 3, 5-12, and 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 2 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a system comprising operate a first multi-head self-attention module in which self-attention is computed for each window of the first set of windows, and produce a downstream internal version of the feature map, the downstream stage module is configured to partition, into a second set of windows, the downstream version of the feature map, a window, of the second set of windows, includes at least one three-dimensional patch of the three-dimensional patches, as presented in the environment of the remaining limitations of claim 2. It is noted that the closest prior art, Barron, shows the cited portions of the system of claim 1, and the three-dimensional shifted window visual transformer module comprises an upstream stage module and a downstream stage module, the upstream stage module is configured to partition, into a first set of windows, at least one of the neural radiance field grid representation or an upstream internal version of the feature map, a window, of the first set of windows, includes at least one three-dimensional patch of the three-dimensional patches, no window, of the first set of windows, overlaps any other window of the first set of windows, produce a three-dimensional shifted window visual transformer output version of the feature map. However, Baron fails to disclose or suggest operate a first multi-head self-attention module in which self-attention is computed for each window of the first set of windows, and produce a downstream internal version of the feature map, the downstream stage module is configured to partition, into a second set of windows, the downstream version of the feature map, a window, of the second set of windows, includes at least one three-dimensional patch of the three-dimensional patches, no window, of the second set of windows, overlaps any other window of the second set of windows, a boundary of at least one window, of the second set of windows, overlaps a boundary of at least one window of the first set of windows, and the downstream stage module is configured to: operate a second multi-head self-attention module in which self-attention is computed for each window of the second set of windows. Claim 5 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a system comprising embed, for each unmasked three-dimensional patch of the three-dimensional patches, three-dimensional position information about a three-dimensional position of a corresponding unmasked three-dimensional patch, of the three-dimensional patches, within the volume, as presented in the environment of the remaining limitations of claim 5. It is noted that the closest prior art, Barron, shows the cited portions of the system of claim 1, and wherein: the two-dimensional images comprise two-dimensional images produced by cameras, each camera, of the cameras, at a time of a production of a corresponding two-dimensional image, of the two-dimensional images: is at a specific position with respect to the scene, and has a specific viewing direction, and the instructions to produce the three-dimensional patches of the neural radiance field grid representation of the scene include instructions to: produce an initial neural radiance field representation of the scene, determine values at discrete positions, within a volume that defines the initial neural radiance field representation, to produce the neural radiance field grid representation. However, Baron fails to disclose or suggest partition the neural radiance field grid representation to produce the three-dimensional patches, and embed, for each unmasked three-dimensional patch of the three-dimensional patches, three-dimensional position information about a three-dimensional position of a corresponding unmasked three-dimensional patch, of the three-dimensional patches, within the volume. Claim 14 is allowable over the prior art of record since the cited references taken individually or in combination fails to particularly disclose or suggest a system comprising merge a set of the three-dimensional patches, in the feature map, into a new single three-dimensional patch to produce a modified feature map, as presented in the environment of the remaining limitations of claim 14. It is noted that the closest prior art, Barron, shows the cited portions of the system of claim 1, and communicate the modified feature map to one of the at least one other three-dimensional shifted windows visual transformer modules, and discrete positions of the three-dimensional patches, of the set of the three-dimensional patches and within a volume that defines the neural radiance field grid representation of the scene. However, Baron fails to disclose or suggest merge a set of the three-dimensional patches, in the feature map, into a new single three-dimensional patch to produce a modified feature map; and form a shape of a cube. Remaining claims 3 and 6-12 each depends from one of the above claims and would accordingly be allowable. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Arksey et al. (US Pub. 2022/0398806) teaches use of NeRF to construct voxel grids. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW D SALVUCCI whose telephone number is (571)270-5748. The examiner can normally be reached M-F: 7:30-4:00PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, XIAO WU can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW SALVUCCI/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Jan 06, 2025
Application Filed
Aug 31, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737976
METHOD FOR CONSTRUCTING STRUCTURAL SEMANTIC MAP UNDER UNDERGROUND WEAK-LIGHT AND LOW-TEXTURE ENVIRONMENT
2y 2m to grant Granted Sep 15, 2026
Patent 12731340
METHOD FOR INTRAOPERATIVE DISPLAY FOR SURGICAL SYSTEMS
4y 6m to grant Granted Sep 08, 2026
Patent 12729975
DISPLAY CONTROL APPARATUS, DISPLAY DEVICE, AND DISPLAY CONTROL METHOD
1y 11m to grant Granted Sep 08, 2026
Patent 12731222
IMAGE PROCESSING DEVICE AND OPERATING METHOD THEREOF
1y 10m to grant Granted Sep 08, 2026
Patent 12711643
TARGET DIGITAL TWIN MODEL GENERATION SYSTEM, CONTROL SYSTEM FOR ROBOT, VIRTUAL SHOP GENERATION SYSTEM, TARGET DIGITAL TWIN MODEL GENERATION METHOD, CONTROL METHOD FOR ROBOT, AND VIRTUAL SHOP GENERATION METHOD
2y 4m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+27.4%)
2y 11m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 494 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month