Prosecution Insights
Last updated: August 17, 2026
Application No. 18/664,737

DEPTH ESTIMATION BASED ON RELATIONSHIPS IN TWO-DIMENSIONAL AND THREE-DIMENSIONAL SPACE FOR AUTONOMOUS SYSTEMS AND APPLICATIONS

Non-Final OA §102§103§112
Filed
May 15, 2024
Examiner
PARK, SOO JIN
Art Unit
2675
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
602 granted / 735 resolved
+19.9% vs TC avg
Strong +17% interview lift
Without
With
+17.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
11 currently pending
Career history
744
Total Applications
across all art units

Statute-Specific Performance

§101
9.6%
-30.4% vs TC avg
§103
40.0%
+0.0% vs TC avg
§102
23.8%
-16.2% vs TC avg
§112
20.1%
-19.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 735 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Election/Restrictions Applicant’s election without traverse of Species I (pertaining to fig 7) in the reply filed on 06/29/2026 is acknowledged. Claims 9-17 have been canceled and claims 21-29 have been newly added. The applicant argues that all claims 1-8 and 18-29 belong to Species I. The examiner respectfully disagrees. The applicant is reminded that claims themselves are never species according to MPEP 806.04(e). The following are claims that are non-elected as belonging to Species II (pertaining to fig 8) and are withdrawn from consideration: Claims 4 and 22 each recites applying a modified tensor to convolutional layers, which belongs in Species II (B806 in fig 8), as described by the applicant’s specification [67]-[68]. Claims 5, 6, and 19 each recites obtaining second data of points having values corresponding to 3D locations, which belongs in Species II (B804 in fig 8), as described by the applicant’s specification [66]. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 20, 25, and 29 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 20, following are limitations recited in the claim that render the claim indefinite, because it is unclear and confusing what they refer to: “an autonomous or semi-autonomous machine”, “simulation operations”, “digital twin operations”, “light transport simulation”, “collaborative content creation for 3D assets”, “one or more deep learning operations”, “an edge device”, “a robot”, “one or more generative AI operations”, “one or more large language models (LLMs)”, “one or more vision language models (VLMs)”, “conversational AI operations”, “generating synthetic data”, “virtual reality content, augmented reality content, or mixed reality content”, “one or more virtual machines (VMs)”, “implemented partially in a data center”, and “partially using cloud computing resources”. Each of these terms are broad and does not provide significant details enough to have been obvious for one of ordinary skill in the art, as the metes and bounds of the claimed invention is not clearly set forth. Similar reasons apply to claim 29. Regarding claim 25, the limitation “absolute depth scale reference” renders the claim indefinite, because it is unclear and confusing what such reference refers to. The phrase appears once in applicant’s specification [29], and describes that the “absolute depth scale reference” may be “derived from entire scene data” somehow, however, does not specify details of such derivation. The phrase “absolute depth scale reference” does not appear to be a well-known technological term with a standard definition and the applicant does not provide any explicit definition either. It is noted that the applicant also uses the term “ground truth”, thereby indicating that the “absolute depth scale reference” is not the same as “ground truth”. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 7, 8, 18, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Bhanushali et al. (US 2024/0135667). Regarding claim 1, Bhanushali discloses: obtaining, using one or more sensors associated with a machine, sensor data representative of one or more images of an environment (see para [56], obtaining, using cameras of a vehicle, images of an environment); obtaining data representative of one or more depth distribution maps indicative of one or more distances, relative to the one or more sensors, associated with one or more locations in the environment that correspond to one or more pixels of the one or more images (see [56], obtaining depth fields indicative of distances, of all locations corresponding to pixels of the images, from the vehicle); determining, based at least on one or more machine learning models processing the sensor data and the data representative of the one or more depth distribution maps, one or more depth values associated with one or more objects in the environment (see [53] and [57], determining, based on a first machine-learning model detecting objects in the images and the depth fields, an average depth value associated with each detected object); and performing one or more operations associated with the machine based at least on the one or more depth values (see [57], performing weight assignment to each of the objects according to their average depth values). Regarding claim 2, Bhanushali further discloses: wherein the one or more distances comprise one or more average distances, relative to the one or more sensors, associated with the one or more locations in the environment (see [57] and rejection of claim 1, a particular depth field is an average of itself and is associated with a corresponding image capturing the environment), the one or more average distances determined based at least on one or more second images obtained using one or more second sensors associated with one or more second machines (see [57] and rejection of claim 1, the particular depth field is obtained via capturing a depth image with a depth sensor). Regarding claim 3, Bhanushali further discloses: generating, based at least on ground truth data obtained from second sensor data captured using one or more second sensors, the data representative of the one or more depth distribution maps, the data including one or more points having one or more values corresponding to the one or more distances (see [57], generating, based on data obtained from a depth sensor, the depth fields indicative of distances, of all locations corresponding to pixels of the images, from the vehicle). Regarding claim 7, Bhanushali further discloses: determining, based at least on the one or more machine learning models processing the sensor data and the data representative of the one or more depth distribution maps, one or more predicted locations of the one or more objects in the environment, the one or more objects depicted in the one or more images of the environment (see [57], determining, based on the first machine-learning model detecting objects in the images and the depth field, objects in the images). Regarding claim 8, Bhanushali further discloses: wherein a first depth distribution map of the one or more depth distribution maps corresponds to a first sensor of the one or more sensors and a second depth distribution map of the one or more depth distribution maps corresponds to a second sensor of the one or more sensors, the first sensor having a different point of view associated with the environment than the second sensor (see [56], each of the images is captured by a one of multiple cameras facing varying directions, wherein a corresponding depth field is then captured by a depth sensor; hence, there is disclosed multiple inherent depth sensors each facing the same direction as the corresponding cameras). Regarding claim 18, Bhanushali discloses: at least one processor comprising: one or more circuits to perform one or more operations associated with a machine based at least on one or more depth values computed using one or more machine learning models (see [69]-[70], a computer; and see [57], performing weight assignment to objects in images based on average depth values computed for each object using a first machine-learning model), wherein the one or more depth values are computed by modifying one or more inputs applied to the one or more machine learning models to include data representative of one or more average distances associated with one or more locations in an environment that correspond to one or more pixels included in one or more images applied to the one or more machine learning models (see [56]-[57] and fig 10, the average depth values of each object are computed by merging the images to be inputted into a second machine-learning model with depth fields (1008 in fig 10) to include data representative of average depths associated with pixels of each detected object in the images to be inputted into the second machine-learning model). Regarding claim 20, Bhanushali further discloses: wherein the processor is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations (see fig 5, deep learning); a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative Al operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing one or more conversational Al operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs);a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21, 23, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Bhanushali in view of Ashley (US 2019/0355171). Regarding claim 21, Bhanushali discloses everything claimed as applied above (see rejection of claim 1), however, does not disclose: cause the machine to perform one or more planning or control operations to navigate within the environment based at least on the one or more depth values (i.e., Bhanushali discloses, in [57], performing weight assignment to each of the objects, to aid in driving a vehicle through the environment, however, does not disclose that such vehicle performs planning or control operations because automated driving is not disclosed). In a similar field of endeavor of transforming images captured by surrounding cameras of a vehicle to aid in driving, Ashley discloses: cause the machine to perform one or more planning or control operations to navigate within the environment based at least on the one or more depth values (see [19]-[20] and [50], transforming images captured by surrounding cameras of a vehicle and recognizing depth around the vehicle; and see [91], automated driving of the vehicle inherently discloses planning). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Bhanushali with Ashley, and transform images captured by surrounding cameras of a vehicle to aid in driving, as disclosed by Bhanushali, wherein the vehicle drives automatically, as disclosed by Ashley, for the purpose of providing convenience to the drive (see Ashley [91]). Regarding claim 23, Bhanushali further discloses: wherein the depth distribution map is sensor-specific such that each pixel of the depth distribution map has a respective depth value representing an average depth for a corresponding pixel of images captured using a particular sensor of the one or more sensors (see [55]-[57], every image taken by a specific camera facing a specific direction has a corresponding depth field obtained, such that each pixel of a depth field of an object has a respective average depth value for the corresponding pixel in the image). Regarding claim 28, Bhanushali and Ashley further disclose: wherein causing the machine to perform the one or more planning or control operations comprises at least one of: determining a trajectory for the machine to follow based on the one or more depth values; providing the one or more depth values to a planning component of the machine; or sending a notification to another machine based at least on determining, using the one or more depth values, that a route is invalid or a map feature is incorrect (see rejection of claim 21, Bhanushali discloses transforming images and determining average depths of objects in the environment of a vehicle to aid in driving, while Ashley discloses automatic driving based on transformed images; therefore, Bhanushali in view of Ashley discloses transforming images and determining average depths of objects and inherently determining a trajectory for automatic driving). Regarding claim 29, Bhanushali further discloses: wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations (see fig 5, deep learning); a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative Al operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing one or more conversational Al operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs);a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Allowable Subject Matter Claims 24, 26, and 27 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art of record does not disclose the subject matter recited in claim 25, however, claim 25 is rejected under 112(b). Claim 25 would be allowable if amended to overcome the 112(b) rejection and rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SJ PARK whose telephone number is (571)270-3569. The examiner can normally be reached M-F 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, EMILY TERRELL can be reached at 571-270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SJ Park/Primary Examiner, Art Unit 2675
Read full office action

Prosecution Timeline

May 15, 2024
Application Filed
Jul 17, 2026
Non-Final Rejection mailed — §102, §103, §112
Aug 02, 2026
Interview Requested
Aug 13, 2026
Applicant Interview (Telephonic)
Aug 13, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12685501
Maskless 2D/3D Artificial Subtraction Angiography
3y 7m to grant Granted Jul 21, 2026
Patent 12676014
AUTOMATIC IMAGE QUALITY AND FEATURE CLASSIFICATION METHOD
3y 2m to grant Granted Jul 07, 2026
Patent 12676015
HIGH DIMENSIONAL SPATIAL ANALYSIS
2y 8m to grant Granted Jul 07, 2026
Patent 12639929
METHOD AND APPARATUS FOR EXTRACTING IMAGE FEATURE BASED ON VISION TRANSFORMER
3y 0m to grant Granted May 26, 2026
Patent 12639963
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
2y 5m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+17.3%)
2y 7m (~4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 735 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month