DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant’s election without traverse of Species I (pertaining to fig 7) in the reply filed on 06/29/2026 is acknowledged. Claims 9-17 have been canceled and claims 21-29 have been newly added. The applicant argues that all claims 1-8 and 18-29 belong to Species I.
The examiner respectfully disagrees. The applicant is reminded that claims themselves are never species according to MPEP 806.04(e). The following are claims that are non-elected as belonging to Species II (pertaining to fig 8) and are withdrawn from consideration:
Claims 4 and 22 each recites applying a modified tensor to convolutional layers, which belongs in Species II (B806 in fig 8), as described by the applicant’s specification [67]-[68].
Claims 5, 6, and 19 each recites obtaining second data of points having values corresponding to 3D locations, which belongs in Species II (B804 in fig 8), as described by the applicant’s specification [66].
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 20, 25, and 29 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 20, following are limitations recited in the claim that render the claim indefinite, because it is unclear and confusing what they refer to: “an autonomous or semi-autonomous machine”, “simulation operations”, “digital twin operations”, “light transport simulation”, “collaborative content creation for 3D assets”, “one or more deep learning operations”, “an edge device”, “a robot”, “one or more generative AI operations”, “one or more large language models (LLMs)”, “one or more vision language models (VLMs)”, “conversational AI operations”, “generating synthetic data”, “virtual reality content, augmented reality content, or mixed reality content”, “one or more virtual machines (VMs)”, “implemented partially in a data center”, and “partially using cloud computing resources”. Each of these terms are broad and does not provide significant details enough to have been obvious for one of ordinary skill in the art, as the metes and bounds of the claimed invention is not clearly set forth. Similar reasons apply to claim 29.
Regarding claim 25, the limitation “absolute depth scale reference” renders the claim indefinite, because it is unclear and confusing what such reference refers to. The phrase appears once in applicant’s specification [29], and describes that the “absolute depth scale reference” may be “derived from entire scene data” somehow, however, does not specify details of such derivation. The phrase “absolute depth scale reference” does not appear to be a well-known technological term with a standard definition and the applicant does not provide any explicit definition either. It is noted that the applicant also uses the term “ground truth”, thereby indicating that the “absolute depth scale reference” is not the same as “ground truth”.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-3, 7, 8, 18, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Bhanushali et al. (US 2024/0135667).
Regarding claim 1, Bhanushali discloses:
obtaining, using one or more sensors associated with a machine, sensor data representative of one or more images of an environment (see para [56], obtaining, using cameras of a vehicle, images of an environment);
obtaining data representative of one or more depth distribution maps indicative of one or more distances, relative to the one or more sensors, associated with one or more locations in the environment that correspond to one or more pixels of the one or more images (see [56], obtaining depth fields indicative of distances, of all locations corresponding to pixels of the images, from the vehicle);
determining, based at least on one or more machine learning models processing the sensor data and the data representative of the one or more depth distribution maps, one or more depth values associated with one or more objects in the environment (see [53] and [57], determining, based on a first machine-learning model detecting objects in the images and the depth fields, an average depth value associated with each detected object); and
performing one or more operations associated with the machine based at least on the one or more depth values (see [57], performing weight assignment to each of the objects according to their average depth values).
Regarding claim 2, Bhanushali further discloses:
wherein the one or more distances comprise one or more average distances, relative to the one or more sensors, associated with the one or more locations in the environment (see [57] and rejection of claim 1, a particular depth field is an average of itself and is associated with a corresponding image capturing the environment),
the one or more average distances determined based at least on one or more second images obtained using one or more second sensors associated with one or more second machines (see [57] and rejection of claim 1, the particular depth field is obtained via capturing a depth image with a depth sensor).
Regarding claim 3, Bhanushali further discloses: generating, based at least on ground truth data obtained from second sensor data captured using one or more second sensors, the data representative of the one or more depth distribution maps, the data including one or more points having one or more values corresponding to the one or more distances (see [57], generating, based on data obtained from a depth sensor, the depth fields indicative of distances, of all locations corresponding to pixels of the images, from the vehicle).
Regarding claim 7, Bhanushali further discloses: determining, based at least on the one or more machine learning models processing the sensor data and the data representative of the one or more depth distribution maps, one or more predicted locations of the one or more objects in the environment, the one or more objects depicted in the one or more images of the environment (see [57], determining, based on the first machine-learning model detecting objects in the images and the depth field, objects in the images).
Regarding claim 8, Bhanushali further discloses: wherein a first depth distribution map of the one or more depth distribution maps corresponds to a first sensor of the one or more sensors and a second depth distribution map of the one or more depth distribution maps corresponds to a second sensor of the one or more sensors, the first sensor having a different point of view associated with the environment than the second sensor (see [56], each of the images is captured by a one of multiple cameras facing varying directions, wherein a corresponding depth field is then captured by a depth sensor; hence, there is disclosed multiple inherent depth sensors each facing the same direction as the corresponding cameras).
Regarding claim 18, Bhanushali discloses:
at least one processor comprising: one or more circuits to perform one or more operations associated with a machine based at least on one or more depth values computed using one or more machine learning models (see [69]-[70], a computer; and see [57], performing weight assignment to objects in images based on average depth values computed for each object using a first machine-learning model), wherein
the one or more depth values are computed by modifying one or more inputs applied to the one or more machine learning models to include data representative of one or more average distances associated with one or more locations in an environment that correspond to one or more pixels included in one or more images applied to the one or more machine learning models (see [56]-[57] and fig 10, the average depth values of each object are computed by merging the images to be inputted into a second machine-learning model with depth fields (1008 in fig 10) to include data representative of average depths associated with pixels of each detected object in the images to be inputted into the second machine-learning model).
Regarding claim 20, Bhanushali further discloses: wherein the processor is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations (see fig 5, deep learning); a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative Al operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing one or more conversational Al operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs);a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 21, 23, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Bhanushali in view of Ashley (US 2019/0355171).
Regarding claim 21, Bhanushali discloses everything claimed as applied above (see rejection of claim 1), however, does not disclose: cause the machine to perform one or more planning or control operations to navigate within the environment based at least on the one or more depth values (i.e., Bhanushali discloses, in [57], performing weight assignment to each of the objects, to aid in driving a vehicle through the environment, however, does not disclose that such vehicle performs planning or control operations because automated driving is not disclosed).
In a similar field of endeavor of transforming images captured by surrounding cameras of a vehicle to aid in driving, Ashley discloses: cause the machine to perform one or more planning or control operations to navigate within the environment based at least on the one or more depth values (see [19]-[20] and [50], transforming images captured by surrounding cameras of a vehicle and recognizing depth around the vehicle; and see [91], automated driving of the vehicle inherently discloses planning).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Bhanushali with Ashley, and transform images captured by surrounding cameras of a vehicle to aid in driving, as disclosed by Bhanushali, wherein the vehicle drives automatically, as disclosed by Ashley, for the purpose of providing convenience to the drive (see Ashley [91]).
Regarding claim 23, Bhanushali further discloses: wherein the depth distribution map is sensor-specific such that each pixel of the depth distribution map has a respective depth value representing an average depth for a corresponding pixel of images captured using a particular sensor of the one or more sensors (see [55]-[57], every image taken by a specific camera facing a specific direction has a corresponding depth field obtained, such that each pixel of a depth field of an object has a respective average depth value for the corresponding pixel in the image).
Regarding claim 28, Bhanushali and Ashley further disclose: wherein causing the machine to perform the one or more planning or control operations comprises at least one of: determining a trajectory for the machine to follow based on the one or more depth values; providing the one or more depth values to a planning component of the machine; or sending a notification to another machine based at least on determining, using the one or more depth values, that a route is invalid or a map feature is incorrect (see rejection of claim 21, Bhanushali discloses transforming images and determining average depths of objects in the environment of a vehicle to aid in driving, while Ashley discloses automatic driving based on transformed images; therefore, Bhanushali in view of Ashley discloses transforming images and determining average depths of objects and inherently determining a trajectory for automatic driving).
Regarding claim 29, Bhanushali further discloses: wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations (see fig 5, deep learning); a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative Al operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing one or more conversational Al operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system incorporating one or more virtual machines (VMs);a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
Allowable Subject Matter
Claims 24, 26, and 27 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The prior art of record does not disclose the subject matter recited in claim 25, however, claim 25 is rejected under 112(b). Claim 25 would be allowable if amended to overcome the 112(b) rejection and rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SJ PARK whose telephone number is (571)270-3569. The examiner can normally be reached M-F 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, EMILY TERRELL can be reached at 571-270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SJ Park/Primary Examiner, Art Unit 2675