Prosecution Insights
Last updated: August 17, 2026
Application No. 18/305,185

AUTOMATIC PROPAGATION OF LABELS BETWEEN SENSOR REPRESENTATIONS FOR AUTONOMOUS SYSTEMS AND APPLICATIONS

Final Rejection §101§102§103
Filed
Apr 21, 2023
Examiner
HINTON, HENRY R
Art Unit
3665
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
NVIDIA Corporation
OA Round
2 (Final)
72%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
39 granted / 54 resolved
+20.2% vs TC avg
Strong +40% interview lift
Without
With
+40.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
12 currently pending
Career history
75
Total Applications
across all art units

Statute-Specific Performance

§101
12.2%
-27.8% vs TC avg
§103
56.6%
+16.6% vs TC avg
§102
17.4%
-22.6% vs TC avg
§112
12.2%
-27.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 54 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 4 and 17 are objected to because of the following informalities: In claim 4, the words “Project” and “Wherein” are incorrectly capitalized. In claim 17, the term “the determining the at least one annotation” appears to be a typo that was intended to be written as “the generating the at least one annotation.” Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-3, 5-6, 8, 10-16, and 18-20 are rejected under 35 U.S.C. 101 because they are directed to an abstract idea without significantly more. The Examiner will now proceed through the two-prong test laid out in MPEP § 2106 on claim 1 (the present claim) to illustrate how the broadest reasonable interpretation of the claims is directed toward the judicial exception. However, the other independent and dependent claims are also directed to a judicial exception unless otherwise specified. Firstly, the broadest reasonable interpretation (BRI) of the present claim is that of an image processing system that annotates objects in map data. Regarding Step 1, the present claim is directed to a [] because it describes []. The analysis proceeds to Step 2A. Regarding Step 2A, the present claim recites a judicial exception because it is (1) directed to an abstract idea; and (2) it does not recite additional elements that integrate the judicial exception into a practical application. The present claim is (1) directed to an abstract idea, particularly a mental process. A mental process is any concept that could be interpreted as being performed by the human mind or by a human mind with a physical aid. MPEP § 2106.04(a)(2)(III). In the present claim, the following claim limitations, when broadly interpreted, do not preclude a human from performing them in their mind or with a pen and piece of paper: “determine that an image, represented by first image data generated using one or more first sensors of a machine, is associated with a time; determine, based at least on the time, a portion of second image data that is generated using one or more second sensors of the machine; generate, based at least on the portion of the second image data, map data representative of one or more locations associated with one or more objects, and generate, based at least on the map data, at least one annotation associated with an object, of the one or more objects, depicted by the image.” These limitations are directed to a mental process because of the high level of generality with which the limitations are recited, and analysis proceeds to step (2). The present claim also (2) fails to integrate the judicial exception into a practical application. In a computing environment, a mental process may be integrated into a practical application where the claim goes “beyond generally linking the use of the judicial exception to a particular technological environment . . . .” MPEP § 2106.04(d)(1). Here, the generic recitation of “one or more processing units” does not appear to the Examiner as more than generally linking the mental process defined in step 2A to being generally performed by a computer. Therefore, the present claim does not integrate the mental process into a practical application. The present claim reciting to a mental process generically applied on computer hardware, the analysis proceeds to Step 2B. Regarding Step 2B, the claim does not recite additional elements that amount to significantly more than the judicial exception. Additional elements of computer components to an abstract idea do not amount to significantly more than the judicial exception when, considered as a whole, the claim appears to be “[s]imply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception.” MPEP § 2106.05(I)(A). The additional elements of generically-recited “images,” “sensors,” and “map data” in the claim appear to be appending the well-understood, routine, conventional activity of performing processes on a computer, at a high level of generality, to the mental process of the present claim. Therefore, the present claim does not recite significantly more than the judicial exception. The Examiner notes that while the above analysis was applied to claim 1 in particular, further steps recited in the other independent and dependent claims all feature similar issues that bar them from being considered eligible subject matter unless specified below. The Examiner notes that claims 4, 7, 9, and 17 appear to overcome the §101 rejection if incorporated into Claim 1. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-5, 7, 10-11, 13-15, and 19-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 20230046289 A1 to Thorsen, Justin et al. (“Thorsen”). Regarding claim 1, Thorsen teaches a system comprising: one or more processing units to: determine that an image, represented by first image data generated using one or more first sensors of a machine, is associated with a time (Thorsen [0061]: “At block 1320, second sensor data for a vehicle is identified. The second sensor data was captured by a second sensor of the vehicle at a second location at a second point in time outside of the first point in time . . . ” First image data of the present invention taken as the second sensor data of Thorsen. The second sensor data taught to comprise camera or lidar data.); determine, based at least on the time, a portion of second image data that is generated using one or more second sensors of the machine (Thorsen [0061]: “For instance, at block 1310, first sensor data for a vehicle is identified. The first sensor data was captured by a first sensor of the vehicle at a first location at a first point in time and the first sensor data is associated with a first label for an object.” The second image data of the present invention is taken as the first sensor data of Thorsen. It is determined based on time because it is temporally related to the first sensor data of Thorsen.); generate, based at least on the portion of the second image data, map data representative of one or more locations associated with one or more objects (Thorsen [0065]: “FIG. 7 is an example representation of LIDAR sensor data 700, for example, first sensor data for a first point in time corresponding to the point in time and location of the vehicle 100 as presented in FIG. 6.”); and generate, based at least on the map data, at least one annotation associated with an object, of the one or more objects, depicted by the image (Thorsen [0066]: “As noted above, the first sensor data may include bounding boxes, as well as one or more associated labels for objects detected by the vehicles perception system or another system which may have processed the sensor data in order to generate the bounding boxes and/or associated labels.”). Regarding claim 2, Thorsen teaches the system of claim 1, wherein: the second sensor data corresponds to a point cloud (Thorsen [0062]: “The first . . . sensor data may include data points generated by . . . LIDAR data points . . . ”); the one or more second sensors comprise one or more LiDAR sensors (Thorsen [0064]: “For example, the first sensor data may include data generated during a single spin of laser-based (e.g. LIDAR) sensor which rotates 360 degrees such as the LIDAR sensor of housing 312.”); and the determination of the portion of the point cloud comprises: determining a first portion of the point cloud that is associated with a spin of the one or more LiDAR sensors that occurred proximate to the time (Thorsen [0064]: “In one example, the first sensor data may represent a first point in time within a finite period of time or timeframe, such as 100 milliseconds or more or less, during which the first sensor data would have been captured or generated by the vehicle's perception system.”); and determining a second portion of the point cloud that is associated with one or more spins that occurred at least one of before the spin or after the spin, the portion of the point cloud including the first portion of the point cloud and the second portion of the point cloud (Thorsen [0068]: “For example, if the frequency of the LIDAR sensor is 10 Hz (or 10 revolutions per second) the minimum difference between the first and second points in time may be 0.1 second. Of course, the first and second points in time may be more or less than 0.1 second apart from one another.” Understood that the point cloud representing the environment is updated each spin. The data received by one such successive spin is broadly interpreted as determining a second portion of the point cloud.). Regarding claim 3, Thorsen teaches the system of claim 1, wherein the one or more processing units are further to: determine, based at least on one or more parameters associated with the one or more first sensors, a portion of the map data that is associated with the image (Thorsen [0073]: “Labels from the LIDAR data points of the first sensor data may be “transferred” by the one or more server computing devices 410 to the one or more camera images of the second sensor data. In other words, a first label from the first sensor data may be used to automatically generate a second label for the second sensor data, by simply associating the first label with the second sensor data. This may be especially useful as the visible range in a camera image may go well beyond the effective perceptive range of the LIDAR data.” Understood that labels from LIDAR data are only transferred to camera images where the labeled object is within visible range (taken as the one or more parameters). The labeled lidar data is taken as the portion of the map data associated with the image because its label is transferred to the image.), wherein the generation of the at least one annotation is further based at least on the portion of the map data (Thorsen [0073]: “Labels from the LIDAR data points of the first sensor data may be “transferred” by the one or more server computing devices 410 to the one or more camera images of the second sensor data.”). Regarding claim 4, Thorsen teaches the system of claim 1, wherein the one or more processing units are further to: Project, based at least on the map data, one or more points associated with the object onto the image, Wherein the generation of the at least one annotation is based at least on the projection of the one or more points (Thorsen FIGS. 8, 11: Understood that in order to pass label data between LIDAR and image datasets, the location of the object (taken as the one or more points) receiving the label in one dataset must be projected onto the location of the same object on the other dataset.). Regarding claim 5, Thorsen teaches the system of claim 1, wherein the one or more processing units are further to determine, based at least on the map data, one or more depth values associated with the object (Thorsen [0065]: Understood that LIDAR, Light Detection and Ranging, collects point data that contains a range for every point. These ranges are broadly interpreted as the depth values. Because the LIDAR detects the entire environment, it necessarily determines depth values associated with the object.). Regarding claim 10, Thorsen teaches the system of claim 1, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine (Thorsen [0036]: “In one aspect the computing devices 110 may be part of an autonomous control system capable of communicating with various components of the vehicle in order to control the vehicle in an autonomous driving mode.”); a perception system for an autonomous or semi-autonomous machine (The term “at least one of” requires consideration of only one of the listed options.); a system for performing simulation operations (The term “at least one of” requires consideration of only one of the listed options.); a system for performing digital twin operations (The term “at least one of” requires consideration of only one of the listed options.); a system for performing light transport simulation (The term “at least one of” requires consideration of only one of the listed options.); a system for performing collaborative content creation for 3D assets (The term “at least one of” requires consideration of only one of the listed options.); a system for performing deep learning operations (The term “at least one of” requires consideration of only one of the listed options.); a system implemented using an edge device (The term “at least one of” requires consideration of only one of the listed options.); a system implemented using a robot (The term “at least one of” requires consideration of only one of the listed options.); a system for performing conversational AI operations(The term “at least one of” requires consideration of only one of the listed options.); a system implementing one or more large language models (LLMs) (The term “at least one of” requires consideration of only one of the listed options.); a system for generating synthetic data (The term “at least one of” requires consideration of only one of the listed options.); a system incorporating one or more virtual machines (VMs) (The term “at least one of” requires consideration of only one of the listed options.); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (The term “at least one of” requires consideration of only one of the listed options.). Regarding claim 11, Thorsen teaches a method comprising: generating, based at least on first sensor data generated using one or more first sensors of a machine, map data representing one or more first locations of one or more static objects and one or more second locations of one or more dynamic objects (Thorsen [0070]: “In this regard, labels may be “transferred” forward or backwards in time. In this regard, although the example of FIG. 10 provides an image captured at a point in time that is after the first point in time, other images captured before the first point in time may also be automatically labeled in a similar way.” Thorsen’s first sensor data, taught to be LIDAR data, taken as the first sensor data of the present invention. It is associated with Thorsen’s second sensor data, taught to be image data depicting the location of various objects because its labels can be transferred between the two datasets.); determining, based at least on second sensor data generated using one or more second sensors of the machine, that an image represented by the second sensor data is associated with a portion of the map data (Thorsen [0046]: “As an example, a perception system software module of the perception system 172 may use sensor data generated by one or more sensors of an autonomous vehicle, such as cameras, LIDAR sensors, . . . to detect and identify objects and their features. These features may include location, type. heading, orientation, speed, acceleration, change in acceleration, size, shape, etc.” Image data taken as the second sensor data.); and generating, based at least on the portion of the map data, at least one annotation associated with a dynamic object, of the one or more dynamic objects, depicted by the image (Thorsen [0046]: “As an example, a perception system software module of the perception system 172 may use sensor data generated by one or more sensors of an autonomous vehicle, such as cameras, . . . to detect and identify objects and their features. These features may include location, type. heading, orientation, speed, acceleration, change in acceleration, size, shape, etc.” Understood that the labels generated for the map data include labels for dynamic objects because some of the label features include movement parameters. The lidar in particular annotates a dynamic object in [0066]. ). Regarding claim 13, Thorsen teaches the method of claim 11, wherein: the one or more second locations of the one or more dynamic objects are associated with a first time (Thorsen [0074]: “An object may be determined to be static if the localized position of that object over the entire time that the object is perceived by a LIDAR sensor of the perception system does not change or rather only changes to a predetermined amount such as a very slight degree (e.g. within an error of the perception system's localization of the object).” A period of time inherently has at least a first time and second time (beginning and end of period). The initial location of the dynamic object when detected taken as first time and second location.); the map data further represents one or more third locations of the one or more dynamic objects, the one or more third locations associated with a second time (Thorsen [0074]: A period of time inherently has at least a first time and second time. The last location of the dynamic object at the end of the period taken as the third location at a second time.); and the method further comprises: determining a third time that the second sensor data was generated by the one or more second sensors (Thorsen [0071]: “In some instances, the second sensor data may include a plurality of camera images captured by one or more cameras of the vehicle 100 at the second point in time.” Point in time of the capture of the image taken as the third time.); and determining to annotate the image using the one or more second locations of the one or more dynamic objects based at least on the first time, the second time, and the third time (Thorsen [0074]: “An object may be determined to be static if the localized position of that object over the entire time that the object is perceived by a LIDAR sensor of the perception system does not change or rather only changes to a predetermined amount such as a very slight degree (e.g. within an error of the perception system's localization of the object).”; [0075]: “Returning to FIG. 13, at block 1340, based on the determination that the object is a static object, the first label is used to automatically generate a second label for the second sensor data.” The choice between annotating and not annotating the image taken as the determining to annotate the image.). Regarding claim 14, Thorsen teaches the method of claim 11, further comprising: determining a location associated with the machine when the second sensor data was generated using the one or more second sensors (Thorsen [0069]: “The second location may be defined in both GPS coordinates (e.g. latitude, longitude, altitude) as well as in a smooth coordinate system or a local frame.”), wherein the determining that the image is associated with the portion of the map data is further based at least on the location associated with the machine (Thorsen [0070]: “In this regard, although the example of FIG. 10 provides an image captured at a point in time that is after the first point in time, other images captured before the first point in time may also be automatically labeled in a similar way.” Thorsen teaches that image data is not labeled unless taken within some distance from the object. Thus, annotation (taken as association with map data portions) occurs further based on the location of the vehicle when the image is taken.). Regarding claim 15, Thorsen teaches the method of claim 11, further comprising determining, based at least on the portion of the map data, one or more depth values associated with the dynamic object (Thorsen [0065]: Understood that LIDAR, Light Detection and Ranging, collects point data that contains a range for every point. These ranges are broadly interpreted as the depth values.). Regarding claim 19, Thorsen teaches a processor comprising: one or more processing units to generate, based at least on a portion of map data, an annotation associated with a dynamic object depicted by an image represented by first sensor data generated using one or more first sensors of a machine (Thorsen [0046]: “As an example, a perception system software module of the perception system 172 may use sensor data generated by one or more sensors of an autonomous vehicle, such as cameras, . . . to detect and identify objects and their features. These features may include location, type. heading, orientation, speed, acceleration, change in acceleration, size, shape, etc.” Understood that the labels generated for the map data include labels for dynamic objects because some of the label features include movement parameters.), wherein the map data is associated with second sensor data generated using one or more second sensors of the machine and represents at least a location of the dynamic object at a point in time (Thorsen [0070]: “In this regard, labels may be “transferred” forward or backwards in time. In this regard, although the example of FIG. 10 provides an image captured at a point in time that is after the first point in time, other images captured before the first point in time may also be automatically labeled in a similar way.” Thorsen’s first sensor data, taught to be LIDAR data, taken as the second sensor data of the present invention. It is associated with Thorsen’s second sensor data, taught to be image data depicting the location of various objects because its labels can be transferred between the two datasets, see for example FIG. 13.). Claim 20 is rejected over similar reasons to claim 10, applied to the system of claim 19. Claims 11-12 are further rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 20170220876 A1 to Gao, Peter et al. (“Gao”) in the alternative because Thorsen does not read on the Broadest Reasonable Interpretation of claim 11 in light of some of the dependent claims. Regarding claim 11, Gao teaches a method comprising: generating, based at least on first sensor data generated using one or more first sensors of a machine (Gao [0065]: “ Each of the lidar devices 546c receive lidar data and process the lidar data (e.g., packets of lidar return information) to generate a lidar point cloud 546. Each point cloud 546 is a three-dimensional set of points in a three-hundred and sixty (360) degree zone around the vehicle.”; [0038]: “The radar devices 240a output raw point clouds.” Radar and LIDAR taken as the one or more first sensors.), map data representing one or more first locations of one or more static objects and one or more second locations of one or more dynamic objects (Gao [0068]; [0038]: Gao teaches the 3D pointset generated by LIDAR data to classify objects, some of which have a speed. APOSITA would have understood that the objects of Gao read on static objects (speed of zero) and dynamic objects (speed other than zero). Gao teaches at [0038] creating point cloud information comprising stationary and dynamic objects.); determining, based at least on second sensor data generated using one or more second sensors of the machine, that an image represented by the second sensor data is associated with a portion of the map data (Gao [0066]: “The system can the select whatever image was sampled/obtained closest to the point in time during which the lidar data was acquired such that only images that were captured near a certain target time (i.e., when the lidar device is looking at the same region that a camera is pointing) will be processed.”); and generating, based at least on the portion of the map data, at least one annotation associated with a dynamic object, of the one or more dynamic object, depicted by the image (Gao FIG. 5, [0063]: “The object tracking module 590 processes the objects 574, the radar tracking information 556, and the object classification data 582 to generate object tracking information 592.”). Regarding claim 12, Gao teaches the method of claim 11, further comprising: determining a time that the second sensor data was generated by the one or more second sensors; and determining that the map is associated with the time (Gao [0066]: “The system can the select whatever image was sampled/obtained closest to the point in time during which the lidar data was acquired such that only images that were captured near a certain target time (i.e., when the lidar device is looking at the same region that a camera is pointing) will be processed.”), wherein the generating the at least one annotation is further based at least on the map data being associated with the time (Gao FIG. 5, [0068]: Understood the image data processed along with the LIDAR data to generate object tracking information is the data associated with the LIDAR scan with time as discussed in [0066].). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Thorsen, further in view of US 20230026278 A1 to Vianello Giacomo et al. (“Vianello”). Regarding claim 6, Thorsen teaches the system of claim 5, wherein the one or more processing units are further to: determine, based at least on the map data, one or more second depth values associated with a second object depicted by the image (Thorsen [0065]-[0066]: Understood that the LIDAR sensor detects multiple objects. LIDAR operates by measuring ranges to objects in the environment (Light Detection and Ranging). Another of the multiple objects detected by LIDAR, therefore, taken as a second object with one or more second depth values.). While teaching obtaining depth data for a given object in a map, Thorsen does not appear to expressly teach the processing units determine, based at least on the one or more depth values associated with the object and the one or more second depth values associated with the second object, that one of: the object is at least partially occluded by the second object; or the second object is at least partially occluded by the object. However, Vianello teaches processing units that determine, based at least on the one or more depth values associated with the object and the one or more second depth values associated with the second object (Vianello FIG. 1, [0032]: “The measurements can function as the basis for . . . occlusion detection . . . Examples of measurements can include: . . . point clouds (e.g., generated from LIDAR . . . ) . . . ” Given FIG. 1, one of ordinary skill in the art would have understood Vianello teaches using the measurements comprising point clouds to determine if an object is occluding an object of interest.), that one of: the object is at least partially occluded by the second object (Vianello FIG. 1: S200 depicts determining whether the object of interest is occluded by an occluding object. The Examiner notes that without further description in the present claims, the object of interest can read on the first object and the occluding object can read on the second object or vice versa.); or the second object is at least partially occluded by the object (Vianello FIG. 1: S200 depicts determining whether the object of interest is occluded by an occluding object. The Examiner notes that without further description in the present claims, the object of interest can read on the first object and the occluding object can read on the second object or vice versa.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have combined the system that detects the depths of multiple objects taught by Thorsen with the system that detects based on the depths of objects whether one occludes the other taught by Vianello. Doing so would have “produce[d] a more accurate depiction of the object of interest, thereby resulting in a more accurate OOI segment and/or attribute extraction” as discussed in [0022] of Vianello. Regarding claim 16, Thorsen teaches the method of claim 15, further comprising: determining, based at least on the portion of the map data, one or more second depth values associated with an object depicted by the image, the object including one of the one or more static objects or the one or more dynamic objects (Thorsen [0065]-[0066]: Understood that the LIDAR sensor detects multiple objects. LIDAR operates by measuring ranges to objects in the environment (Light Detection and Ranging). Another of the multiple objects detected by LIDAR, therefore, taken as a second object with one or more second depth values.). While teaching obtaining depth data for a given object in a map, Thorsen does not appear to expressly teach determining, based at least on the one or more depth values associated with the dynamic object and the one or more second depth values associated with the object, that one of: the dynamic object is at least partially occluded by the object; or the object is at least partially occluded by the dynamic object. However, Vianello teaches determining, based at least on the one or more depth values associated with the dynamic object and the one or more second depth values associated with the object (Vianello FIG. 1, [0032]: “The measurements can function as the basis for . . . occlusion detection . . . Examples of measurements can include: . . . point clouds (e.g., generated from LIDAR . . . ) . . . ” Given FIG. 1, one of ordinary skill in the art would have understood Vianello teaches using the measurements comprising point clouds to determine if an object is occluding an object of interest. The Examiner notes that Vianello teaches an occluding object may be dynamic at [0031].), that one of: the dynamic object is at least partially occluded by the object; or the object is at least partially occluded by the dynamic object (Vianello FIG. 1: S200 depicts determining whether the object of interest is occluded by an occluding object.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have combined the system that detects the depths of multiple objects taught by Thorsen with the system that detects based on the depths of objects whether one occludes the other taught by Vianello. Doing so would have “produce[d] a more accurate depiction of the object of interest, thereby resulting in a more accurate OOI segment and/or attribute extraction” as discussed in [0022] of Vianello. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Thorsen, further in view of US 20050190274 A1 to Yoshikawa, Seiji et al. (“Yoshikawa”). Regarding claim 7, Thorsen teaches the system of claim 1, wherein: the one or more first sensors include a first rolling shutter that is associated with a first time period (Thorsen [0064]: One spin of the LIDAR taken as the first rolling shutter.). Thorsen does not appear to expressly teach the one or more second sensors include a second rolling shutter that is associated with a second time period; and the determination of the at least one annotation associated with the dynamic object is further based at least on the first time period and the second time period. However, Yoshikawa teaches the one or more second sensors include a second rolling shutter that is associated with a second time period (Yoshikawa [0072]: “The image sensor 11 has shutter function including the global shutter function and the rolling shutter function.” APOSITA would have understood that take in combination with Thorsen above, the second sensor of Thorsen (an image sensor) would have comprised a rolling shutter type image sensor, which in term comprises a second time period of image capture.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have combined the system comprising an image sensor taught by Thorsen with the image sensor comprising a rolling shutter taught by Yoshikawa. Doing so would have given the image sensor a higher capture speed as taught by Yoshikawa [0118]. This combination further teaches the determining the at least one annotation associated with the dynamic object is further based at least on the first time period and the second time period (Thorsen [0068]: “The minimum difference between the first and second points in time may be determined based on any number of different metrics, including for example, a difference corresponding to the frequency of the LIDAR sensor used to capture the first sensor data. For example, if the frequency of the LIDAR sensor is 10 Hz (or 10 revolutions per second) the minimum difference between the first and second points in time may be 0.1 second.” APOSITA would have understood that Thorsen teaches the first point in time being a rolling shutter time period and Yoshikawa teaches the second point in time being a rolling shutter time period. At [], Thorsen teaches the dynamic object is determined based on its movement during the LIDAR scan (first rolling shutter). If the object is dynamic, the LIDAR label is not passed to the second sensor data ). Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Gao, further in view of US 20050190274 A1to Yoshikawa, Seiji et al. (“Yoshikawa”). Regarding claim 17, Gao teaches the system of claim 11, wherein: the one or more first sensors include a first rolling shutter that is associated with a first time period (Gao [0064]: LIDAR scanning taken as the first rolling shutter.). While Gao teaches detection of the objects using a camera, it does not appear to expressly teach the one or more second sensors include a second rolling shutter that is associated with a second time period; and the determining the at least one annotation associated with the dynamic object is further based at least on the first time period and the second time period. However, Yoshikawa teaches the one or more second sensors include a second rolling shutter that is associated with a second time period (Yoshikawa [0072]: “The image sensor 11 has shutter function including the global shutter function and the rolling shutter function.” APOSITA would have understood that take in combination with Thorsen above, the second sensor of Thorsen (an image sensor) would have comprised a rolling shutter type image sensor, which in term comprises a second time period of image capture.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have combined the system comprising an image sensor taught by Gao with the image sensor comprising a rolling shutter taught by Yoshikawa. Doing so would have given the image sensor a higher capture speed as taught by Yoshikawa [0118]. This combination further teaches the determining the at least one annotation associated with the dynamic object is further based at least on the first time period and the second time period (Gao FIG. 5: APOSITA would have understood that the object tracking information obtained based on LIDAR and image data would have used the data from the rolling shutters of the camera of the combination of Gao and Yoshikawa and LIDAR of Gao.). Claims 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over Thorsen, further in view of US 20190228537 A1 to Sekiguchi, Hiroyoshi et al. (“Sekiguchi”). Regarding claim 8, Thorsen teaches the system of claim 1, wherein the one or more processing units are further to: determine, based at least on the map data, one or more first depth values associated with the object (Thorsen [0065]: Understood that LIDAR, Light Detection and Ranging, collects point data that contains a range for every point. These ranges are broadly interpreted as the depth values.); determine, based at least on the first sensor data, that a portion of the image depicts the object (Thorsen [0066]: “As noted above, the first sensor data may include bounding boxes, as well as one or more associated labels for objects detected by the vehicles perception system or another system which may have processed the sensor data in order to generate the bounding boxes and/or associated labels.”, [0073]: “Labels from the LIDAR data points of the first sensor data may be “transferred” by the one or more server computing devices 410 to the one or more camera images of the second sensor data.” The system of Thorsen at least determines that part of the image data depicts the object by allowing the label data to be passed from the LIDAR dataset to the image dataset.). Thorsen does not appear to expressly teach the units determine, based at least on the one or more first depth values and the portion of the image, one or more second depth values associated with the object. However, Sekiguchi teaches the units determine, based at least on the one or more first depth values and the portion of the image, one or more second depth values associated with the object (Sekiguchi FIG. 5A: Sekiguchi teaches fusion of depth image data (portion of the image) with lidar distance information (map data first depth values) to obtain a fused range image.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have combined the system that creates an environmental map including depth information that associates the map with an image taught by Thorsen with the system that combines LIDAR depth information with image data including depth information to generate a fused range image taught by Sekiguchi. Doing so would have improved the accuracy of the depth information by creating a high-quality and high-resolution range images as discussed in [0088] of Sekiguchi. Regarding claim 9, the above combination of Thorsen and Sekiguchi further teaches the system of claim 8, wherein: the one or more first sensors are associated with a first resolution (Sekiguchi [0064]: “ . . . the distance resolution of a stereocamera becomes sharply larger in accordance with increase of the distance Z.”); the one or more second sensors are associated with a second resolution (Sekiguchi [0064]: “As illustrated in FIG. 1A, the distance resolution of LIDAR is almost constant regardless of a value of the distance Z . . . ”); and the determination of the one or more second depth values associated with the object occurs based at least on first resolution being greater than the second resolution (Sekiguchi FIG. 1A: Sekiguchi teaches that camera distance resolution is greater with respect to distance than LIDAR distance resolution. Sekiguchi, as discussed in Claim 8 above, teaches that fusing the two types of data improves the range estimation. Thus, Sekiguchi teaches that determining the second depth values (fusion of the data) is based on the greater resolution of cameras than LIDAR.). Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Gao, further in view of US 20190228537 A1 to Sekiguchi, Hiroyoshi et al. (“Sekiguchi”). Regarding claim 18, Gao teaches the method of claim 11, further comprising: determining, based at least on the portion of the map data, one or more first depth values associated with the dynamic object (Gao [0065]: Understood that LIDAR, Light Detection and Ranging, collects point data that contains a range for every point. These ranges are broadly interpreted as the depth values.); determining, based at least on the second sensor data, that a portion of the image depicts the dynamic object (Gao FIG. 5, [0063]: “ The image classification module 566 processes the rectified camera images 562 and the three-dimensional locations of objects 591 from the object tracking module 590 to generate the image classification data 568, and provides the image classification data 568 to the object classification module 580.”). Gao does not appear to expressly teach determining, based at least on the one or more first depth values and the portion of the image, one or more second depth values associated with the dynamic object. However, Sekiguchi teaches determining, based at least on the one or more first depth values and the portion of the image, one or more second depth values associated with the dynamic object (Sekiguchi FIG. 5A: Sekiguchi teaches fusion of depth image data (portion of the image) with lidar distance information (map data first depth values) to obtain a fused range image.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have combined the system that creates an environmental map including the dynamic object and captures an image of the object taught by Gao with the system that combines LIDAR depth information with image data including depth information to generate a fused range image taught by Sekiguchi. Doing so would have improved the accuracy of the depth information by creating a high-quality and high-resolution range images as discussed in [0088] of Sekiguchi. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Sheu, Kevin et al.. US 20220172495 A1. Instance Segmentation Using Sensor Data Having Different Dimensionalities. Wang, Peng et al.. US 20200364554 A1. Systems and Methods for Deep Localization and Segmentation with a 3D Semantic Map. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HENRY RICHARD HINTON whose telephone number is (703)756-1051. The examiner can normally be reached Monday-Friday 7:30-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hunter Lonsberry can be reached at (571) 272-7298. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HENRY R HINTON/ Examiner, Art Unit 3665 /HUNTER B LONSBERRY/ Supervisory Patent Examiner, Art Unit 3665
Read full office action

Prosecution Timeline

Apr 21, 2023
Application Filed
Dec 01, 2025
Non-Final Rejection mailed — §101, §102, §103
Feb 03, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12703396
UNKNOWN DRIVING HAZARD DETECTION AND RESPONSE SYSTEM
2y 7m to grant Granted Aug 11, 2026
Patent 12680830
METHOD AND SYSTEM FOR ESTIMATING THE POSITION OF A RAILWAY VEHICLE TRAVELLING ALONG A RAILWAY LINE, AND RAILWAY VEHICLE COMPRISING SUCH SYSTEM
2y 0m to grant Granted Jul 14, 2026
Patent 12662165
TRAJECTORY PREDICTION BY SAMPLING SEQUENCES OF DISCRETE MOTION TOKENS
2y 3m to grant Granted Jun 23, 2026
Patent 12644711
ON-PREMISES POSITIONING DETERMINATION AND ANALYTICS SYSTEM
3y 8m to grant Granted Jun 02, 2026
Patent 12637105
IMPLEMENTING MANOEUVRES IN AUTONOMOUS VEHICLES
3y 9m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+40.0%)
2y 10m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 54 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month