Prosecution Insights
Last updated: October 02, 2026
Application No. 18/917,905

SHARED VISION SYSTEM BACKBONE

Non-Final OA §103§DOUBLEPATENT
Filed
Oct 16, 2024
Priority
Apr 28, 2022 — continuation of 12/148,223
Examiner
SILVA-AVINA, EMMANUEL
Art Unit
Tech Center
Assignee
Toyota Motor Corporation
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
65 granted / 81 resolved
+20.2% vs TC avg
Moderate +10% lift
Without
With
+9.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
11 currently pending
Career history
92
Total Applications
across all art units

Statute-Specific Performance

§101
10.5%
-29.5% vs TC avg
§103
57.8%
+17.8% vs TC avg
§102
16.8%
-23.2% vs TC avg
§112
13.8%
-26.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 81 resolved cases

Office Action

§103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This communication is in response to Application No. 18/917,905 filed 10/16/2024. Claims 1-20 are pending. Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 01/24/2025 have been entered and considered. Initialed copies of the PTO-1449 by the examiner are attached. Specification The disclosure is objected to because of the following informalities: At paragraph [0068] lines 6-7 should recite, in part, “. Appropriate correction is required. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claim 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-15, 17-18 and 20 of U.S. Patent No. US12148223B2. Although the claims at issue are not identical, they are not patentably distinct from each other because the instant application and the conflicting Patent are claiming common subject matter, as follows: This Application No. 18/917,905 US Patent No. US12148223B2 Claim 1: A method for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: generating, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle; generating, via a sparse depth network, one or more sparse depth estimates of the environment, each sparse depth estimate associated with a respective sparse representation of one or more sparse representations; generating the dense LiDAR representation based on a dense depth estimate that is generated based on the depth estimate and the one or more sparse depth estimates; and controlling an action of the vehicle based on the dense LiDAR representation. Claim 1: A method for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: receiving, at a sparse depth network, one or more sparse representations of an environment within a vicinity of the vehicle; generating, at a depth estimation network, a depth estimate of the environment depicted in an image captured by an image capturing sensor integrated with the vehicle based on receiving the one or more sparse representation; generating, via the sparse depth network, one or more sparse depth estimates based on receiving the one or more sparse representations of the environment, each sparse depth estimate associated with a respective sparse representation of the one or more sparse representations; fusing, at a depth fusion network, the depth estimate and the one or more sparse depth estimates to generate a dense depth estimate; generating the dense LiDAR representation based on the dense depth estimate; and controlling an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation. Claim 9: An apparatus for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: at least one processor; and at least one memory coupled with the at least one processor and storing instructions operable, when executed by the at least one processor, to cause the apparatus to: generate, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle; generate, via a sparse depth network, one or more sparse depth estimates of the environment, each sparse depth estimate associated with a respective sparse representation of one or more sparse representations; generate the dense LiDAR representation based on a dense depth estimate that is generated based on the depth estimate and the one or more sparse depth estimates; and control an action of the vehicle based on the dense LiDAR representation. Claim 8: An apparatus for generating a dense light detection and ranging (LiDAR) representation at a vision system of a vehicle, the apparatus comprising: at least one processor; and at least one memory coupled with the at least one processor and storing instructions operable, when executed by the at least one processor, to cause the apparatus: receive, at a sparse depth network, one or more sparse representations of an environment within a vicinity of the vehicle; generate, at a depth estimation network, a depth estimate of the environment depicted in an image captured by an image capturing sensor integrated with the vehicle based on receiving the one or more sparse representation; generate, via the sparse depth network, one or more sparse depth estimates based on receiving the one or more sparse representations of the environment, each sparse depth estimate associated with a respective sparse representation of the one or more sparse representations; fuse, at a depth fusion network, the depth estimate and the one or more sparse depth estimates to generate a dense depth estimate; generate the dense LiDAR representation based on the dense depth estimate; and control an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation. Claim 17: A non-transitory computer-readable medium having program code recorded thereon for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, the program code executed by at least one processor and comprising: program code to generate, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle; program code to generate, via a sparse depth network, one or more sparse depth estimates of the environment, each sparse depth estimate associated with a respective sparse representation of one or more sparse representations; program code to generate the dense LiDAR representation based on a dense depth estimate that is generated based on the depth estimate and the one or more sparse depth estimates; and program code to control an action of the vehicle based on the dense LiDAR representation. Claim 15: A non-transitory computer-readable medium having program code recorded thereon for generating a dense light detection and ranging (LiDAR) representation at a vision system of a vehicle, the program code executed by a processor and comprising: program code to receive, at a sparse depth network, one or more sparse representations of an environment within a vicinity of the vehicle; program code to generate, at a depth estimation network, a depth estimate of the environment depicted in an image captured by an image capturing sensor integrated with the vehicle based on receiving the one or more sparse representation; program code to generate, via the sparse depth network, one or more sparse depth estimates based on receiving the one or more sparse representations of the environment, each sparse depth estimate associated with a respective sparse representation of the one or more sparse representations; program code to fuse, at a depth fusion network, the depth estimate and the one or more sparse depth estimates to generate a dense depth estimate; program code to generate the dense LiDAR representation based on the dense depth estimate; and program code to control an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation. The further limitations of the dependent claims are similar as indicated below: This Application No. 18/917,905 US Patent No. US12148223B2 Claim 2: The method of claim 1, further comprising performing one or more vision based tasks based on a combination of features associated with the image and the one or more sparse depth estimates. Claim 3: The method of claim 2, wherein the one or more vision based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment, or generating a semantic segmentation map of the environment. Claim 4: The method of claim 1, wherein: generating the dense LiDAR representation comprises: decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate; and the dense LiDAR representation is based on the 3D space. Claim 5: The method of claim 1, further comprising: receiving, at the sparse depth network, a semantic segmentation map; generating, via the sparse depth network, a sparse depth estimate of the semantic segmentation map based on receiving the semantic segmentation map; generating, at a segmentation fusion block, a fused segmentation representation by fusing the depth estimate and the sparse semantic segmentation map; and generating, via a lane segmentation network, a lane segmentation map of the environment based on a combination features associated with the image and the one or more sparse depth estimates. Claim 6: The method of claim 1, further comprising generating each sparse representation by a respective sparse representation sensor of one or more sparse representation sensors integrated with the vehicle. Claim 7: The method of claim 6, wherein: the one or more sparse representations include one or more of a sparse LiDAR representation or a radar representation; and the one or more sparse representation sensors include one or more of a sparse LiDAR sensor or a radar sensor. Claim 8: The method of claim 1, wherein the action of the vehicle is based on identifying a three-dimensional object in the dense LiDAR representation. Claim 10: The apparatus of claim 9, wherein execution of the instructions further causes the apparatus to perform one or more vision-based tasks based on a combination of features associated with the image and the one or more sparse depth estimates. Claim 11: The apparatus of claim 10, wherein the one or more vision-based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment, or generating a semantic segmentation map of the environment. Claim 12: The apparatus of claim 9, wherein: generating the dense LiDAR representation comprises: decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate; and the dense LiDAR representation is based on the 3D space. Claim 13: The apparatus of claim 9, wherein execution of the instructions further cause the apparatus to: receive, at the sparse depth network, a semantic segmentation map; generate a sparse depth estimate of the semantic segmentation map based on receiving the semantic segmentation map; generate, at a segmentation fusion block, a fused segmentation representation by fusing the depth estimate and the sparse semantic segmentation map; and generate, via a lane segmentation network, a lane segmentation map of the environment based on a combination of features associated with the image and the one or more sparse depth estimates. Claim 14: The apparatus of claim 9, further comprising instructions operable to generate each sparse representation by a respective sparse representation sensor of one or more sparse representation sensors integrated with the vehicle. Claim 15: The apparatus of claim 14, wherein: the one or more sparse representations include one or more of a sparse LiDAR representation or a radar representation; and the one or more sparse representation sensors include one or more of a sparse LiDAR sensor or a radar sensor. Claim 16: The apparatus of claim 9, wherein execution of the instructions causes the apparatus to control the vehicle's action based on identifying a three-dimensional object in the dense LiDAR representation. Claim 18: The non-transitory computer-readable medium of claim 17, further comprising program code to perform one or more vision-based tasks based on a combination of features associated with the image and the one or more sparse depth estimates. Claim 19: The non-transitory computer-readable medium of claim 18, wherein the one or more vision-based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment, or generating a semantic segmentation map of the environment. Claim 20: The non-transitory computer-readable medium of claim 17, wherein the program code further comprises: program code to receive, at the sparse depth network, a semantic segmentation map; program code to generate a sparse depth estimate of the semantic segmentation map based on receiving the semantic segmentation map; program code to generate, at a segmentation fusion block, a fused segmentation representation by fusing the depth estimate and the sparse semantic segmentation map; and program code to generate, via a lane segmentation network, a lane segmentation map of the environment based on a combination of features associated with the image and the one or more sparse depth estimates. Claim 2: The method of claim 1, further comprising: generating, via a feature extraction network, features associated with the image; and performing one or more vision based tasks based on a combination of the features and the one or more sparse depth estimates. Claim 3: The method of claim 2, wherein the one or more vision based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment, or generating a semantic segmentation map of the environment. Claim 4: The method of claim 1, wherein: generating the dense LiDAR representation comprises: decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate; and the dense LiDAR representation is based on the 3D space. Claim 5: The method of claim 1, further comprising: receiving, at the sparse depth network, a semantic segmentation map; generating, via the sparse depth network, a sparse depth estimate of the semantic segmentation map based on receiving the semantic segmentation map; generating, at a segmentation fusion block, a fused segmentation representation by fusing the depth estimate and the sparse semantic segmentation map; and generating, via a lane segmentation network, a lane segmentation map of the environment based on a combination features associated with the image and the one or more sparse depth estimates, wherein the features are generated via a feature extraction network. Claim 6: The method of claim 1, further comprising generating each sparse representation by a respective sparse representation sensor of one or more sparse representation sensors integrated with the vehicle. Claim 7: The method of claim 6, wherein: the one or more sparse representations include one or more of a sparse LiDAR representation or a radar representation; and the one or more sparse representation sensors include one or more of a sparse LiDAR sensor or a radar sensor. Claim 1: controlling an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation. Claim 9: The apparatus of claim 8, wherein execution of the instructions further cause the apparatus to: generate, via a feature extraction network, features associated with the image; and perform one or more vision based tasks based on a combination of the features and the one or more sparse depth estimates. Claim 10: The apparatus of claim 9, wherein the one or more vision based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment, or generating a semantic segmentation map of the environment. Claim 11: The apparatus of claim 8, wherein: execution of the instructions that cause the apparatus to generate the dense LiDAR representation further cause the apparatus to: decode the depth estimate via a depth decoder; and convert a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate; and the dense LiDAR representation is based on the 3D space. Claim 12: The apparatus of claim 8, wherein execution of the instructions further cause the apparatus to: receive, at the sparse depth network, a semantic segmentation map; generate, via the sparse depth network, a sparse depth estimate of the semantic segmentation map based on receiving the semantic segmentation map; generate, at a segmentation fusion block, a fused segmentation representation by fusing the depth estimate and the sparse semantic segmentation map; and generate, via a lane segmentation network, a lane segmentation map of the environment based on a combination features associated with the image and the one or more sparse depth estimates, wherein the features are generated via a feature extraction network. Claim 13: The apparatus of claim 8, wherein execution of the instructions further cause the apparatus to generate each sparse representation by a respective sparse representation sensor of one or more sparse representation sensors integrated with the vehicle. Claim 14: The apparatus of claim 13, wherein: the one or more sparse representations include one or more of a sparse LiDAR representation or a radar representation; and the one or more sparse representation sensors include one or more of a sparse LiDAR sensor or a radar sensor. Claim 8: control an action of the vehicle based on identifying a three-dimensional object in the dense LiDAR representation. Claim 17: The non-transitory computer-readable medium of claim 15, wherein the program code further comprises: program code to generate, via a feature extraction network, features associated with the image; and program code to perform one or more vision based tasks based on a combination of the features and the one or more sparse depth estimates. Claim 18: The non-transitory computer-readable medium of claim 17, wherein the one or more vision based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment, or generating a semantic segmentation map of the environment. Claim 20: The non-transitory computer-readable medium of claim 15, wherein the program code further comprises: program code to receive, at the sparse depth network, a semantic segmentation map; program code to generate, via the sparse depth network, a sparse depth estimate of the semantic segmentation map based on receiving the semantic segmentation map; program code to generate, at a segmentation fusion block, a fused segmentation representation by fusing the depth estimate and the sparse semantic segmentation map; and program code to generate, via a lane segmentation network, a lane segmentation map of the environment based on a combination features associated with the image and the one or more sparse depth estimates, wherein the features are generating via a feature extraction network. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6-12, and 14-19 are rejected under 35 U.S.C. 103 as being unpatentable over Urtasun et al. (US 20200160559A1, hereinafter referred to as “Urtasun”) in view of Guizilini et al. (“Sparse auxiliary networks for unified monocular depth prediction and completion”, 2021, hereinafter referred to as “Guizilini”). Regarding claim 1, Urtasun teaches a method for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: generating, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle (“the model architecture can further include a machine-learned depth completion model. The depth completion model can be configured to receive the image feature map generated by the machine-learned image backbone model and to produce a depth completion map. For example, the depth completion map can describe a depth value for each pixel of the image” Urtasun, [0029]; wherein the sensor is that of “at least one camera configured to capture an image of an environment surrounding the autonomous vehicle” Urtasun, [0006]); each sparse depth estimate associated with a respective sparse representation of one or more sparse representations (“a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]); generating the dense LiDAR representation based on a dense depth estimate that is generated based on the depth estimate and the one or more sparse depth estimates (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image. Specifically, given the projected (into the image plane) sparse depth 210 from the LIDAR point cloud 206 and a camera image 208, the models 204 and 214 cooperate to output dense depth 222 at the same resolution as the input image 208” Urtasun, [0086]-[0087]; “Depth completion is exploited to learn better cross-modality feature representation and achieve dense feature map fusion by transforming predicted dense depth image into dense pseudo LIDAR points” Urtasun, [0066]; see additionally, [0091]; wherein the pseudo-LIDAR points are in 3D space, see Urtasun [0088]); and controlling an action of the vehicle based on the dense LiDAR representation (“or example, output(s) can be provided to one or more of the perception system 124, prediction system 126, motion planning system 128, and vehicle control system 138 to implement additional autonomy processing functionality based on the output(s). For example, motion planning system 128 of FIG. 1 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) based at least in part on the output(s) of the system illustrated in FIG. 2. Stated differently, given information about the current locations of objects detected via the output(s) and/or predicted future locations and/or moving paths of proximate objects detected via the output(s), the motion planning system 128 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) that best navigates the autonomous vehicle (e.g., vehicle 102) along a determined travel route relative to the objects at such locations. The motion planning system 128 then can provide the selected motion plan to a vehicle control system 138 that controls one or more vehicle controls (e.g., actuators or other devices that control gas flow, steering, braking, etc.) to execute the selected motion plan” Urtasun, [0065]; wherein Fig. 2 illustrates the generation of a dense depth image of lidar representation). Urtasun fails to explicitly teach generating, via a sparse depth network, one or more sparse depth estimates of the environment. However, Guizilini explicitly teaches generating, via a sparse depth network, one or more sparse depth estimates of the environment (“the use of sparse convolutions to process input depth maps, while RGB images are still processed using standard convolutions. More specifically, we use Minkowski convolutions [4], a highly efficient generalized sparse convolution recently introduced to address high-dimensional problems. In this work we focus on the 2D application of Minkowski convolutions (image processing)” Guizilini, pg. 3 Col 1; wherein the sparse depth network is that of a Sparse Auxiliary Network (SANS) described in Fig. 2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Urtasun of having a method for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: generating, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle, with the teachings of Guizilini of having generating, via a sparse depth network, one or more sparse depth estimates of the environment. Wherein having Urtasun’s multi-sensor fusion system having generating, via a sparse depth network, one or more sparse depth estimates of the environment. The motivation behind the modification would have been to obtain multi-sensor fusion for three-dimensional object detection, since both Urtasun and Guizilini are systems for predicting dense depth using camera and LIDAR. Wherein Urtasun multi-sensor fusion system provides a multi-sensor detector that reasons about a target task of 2D and/or 3D object detection in addition to one or more auxiliary tasks such as ground estimation, depth completion, and/or other tasks such as other vision tasks, while Guizilini enables a monocular depth prediction network to also perform depth completion in the presence of optional sparse 3D measurements at inference time. Please see Urtasun et al. (US 20200160559 A1), Paragraph [0021] and Guizilini et al. (“Sparse auxiliary networks for unified monocular depth prediction and completion”, 2021), pg. 1 Col 2). Regarding claim 2, Urtasun in view of Guizilini teach the method of claim 1, Urtasun further teaches further comprising performing one or more vision based tasks based on a combination of features associated with the image and the one or more sparse depth estimates (“fusing image features from image data (e.g., image data obtained at 704) with LIDAR features from the BEV representation of the LIDAR data. In some implementations, fusing at 712 can include executing one or more continuous convolutions to fuse image features from a first data stream with LIDAR features from a second data stream... generating a feature map comprising the fused image features and LIDAR features determined at 712.... detecting three-dimensional objects of interest based on the fused ROI crops generated at 714. In some implementations, detecting objects of interest at 716 can include providing the feature map generated at 714 as input to a machine-learned refinement model” Urtasun, [0109]-[0111]; wherein the “first data stream descriptive of image data and a second data stream descriptive of LIDAR point cloud data” Urtasun, [0149] in which the “parse depth image 210 can be generated from the LIDAR point cloud 206” Urtasun, [0088]). Regarding claim 3, Urtasun in view of Guizilini teach the method of claim 2, Urtasun further teaches wherein the one or more vision based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment (“detecting objects of interest at 716 can include providing the feature map generated at 714 as input to a machine-learned refinement model. In response to receiving the feature map, the machine-learned refinement model can be trained to generate as output a plurality of detections corresponding to identified objects of interest within the feature map. In some implementations, detecting objects of interest at 716 can include determining a plurality of object classifications and/or bounding shapes corresponding to the detected objects of interest. For example, in one implementation, the plurality of objects detected at 716 can include a plurality of bounding shapes at locations within the feature map(s) having a confidence score associated with an object likelihood that is above a threshold value. In some implementations, detecting objects of interest at 716 can include determining one or more of a classification indicative of a likelihood that each of the one or more objects of interest comprises a class of object from a predefined group of object classes (e.g., vehicle, bicycle, pedestrian, etc.) and a bounding shape representative of a size, a location, and an orientation of each the one or more objects of interest” Urtasun, [0111]), or generating a semantic segmentation map of the environment. Regarding claim 4, Urtasun in view of Guizilini teach the method of claim 1, Urtasun further teaches wherein: generating the dense LiDAR representation comprises: the dense LiDAR representation is based on the 3D space (“example implementations of the present disclosure use the machine-learned depth completion model to predict dense depth. The predicted depth can be used as pseudo-LIDAR points to find dense correspondences between multi-sensor feature maps” Urtasun, [0030]; “Depth completion is exploited to learn better cross-modality feature representation and achieve dense feature map fusion by transforming predicted dense depth image into dense pseudo LIDAR points” Urtasun, [0066]; see additionally, [0091]; wherein the pseudo-LIDAR points are in 3D space, see Urtasun [0088]). Urtasun fails to explicitly teach decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate. However, Guizilini teaches decoding the depth estimate via a depth decoder (“The dense RGB module can be any encoder-decoder depth prediction network that uses skip connections. In our work we consider two baseline state-of-the-art network architectures: Pack Net and BTS” Guizilini, pg. 4 Col 1); and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate (“The depth maps predicted by PackNet-SAN were projected into 3D as pseudo-LiDAR pointclouds” Guizilini, pg. 7 Col 1). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Urtasun of having a method for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: generating, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle, with the teachings of Guizilini of having decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate. Wherein having Urtasun’s multi-sensor fusion system of having decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate. The motivation behind the modification would have been to obtain multi-sensor fusion for three-dimensional object detection, since both Urtasun and Guizilini are systems for predicting dense depth using camera and LIDAR. Wherein Urtasun multi-sensor fusion system provides a multi-sensor detector that reasons about a target task of 2D and/or 3D object detection in addition to one or more auxiliary tasks such as ground estimation, depth completion, and/or other tasks such as other vision tasks, while Guizilini enables a monocular depth prediction network to also perform depth completion in the presence of optional sparse 3D measurements at inference time. Please see Urtasun et al. (US 20200160559 A1), Paragraph [0021] and Guizilini et al. (“Sparse auxiliary networks for unified monocular depth prediction and completion”, 2021), pg. 1 Col 2). Regarding claim 6, Urtasun in view of Guizilini teach the method of claim 1, Urtasun further teaches further comprising generating each sparse representation by a respective sparse representation sensor of one or more sparse representation sensors integrated with the vehicle (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image” Urtasun, [0086]; “a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]). Regarding claim 7, Urtasun in view of Guizilini teach the method of claim 6, Urtasun further teaches wherein: the one or more sparse representations include one or more of a sparse LiDAR representation or a radar representation (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image” Urtasun, [0086]; “a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]); and the one or more sparse representation sensors include one or more of a sparse LiDAR sensor or a radar sensor (“As one example, an autonomous vehicle can include one or more cameras and a LIDAR system... Other sensors can include radio detection and ranging (RADAR) sensors, ultrasound sensors, and/or the like” Urtasun, [0022]). Regarding claim 8, Urtasun in view of Guizilini teach the method of claim 1, Urtasun further teaches wherein the action of the vehicle is based on identifying a three-dimensional object in the dense LiDAR representation (“For example, motion planning system 128 of FIG. 1 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) based at least in part on the output(s) of the system illustrated in FIG. 2. Stated differently, given information about the current locations of objects detected via the output(s) and/or predicted future locations and/or moving paths of proximate objects detected via the output(s), the motion planning system 128 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) that best navigates the autonomous vehicle (e.g., vehicle 102) along a determined travel route relative to the objects at such locations” Urtasun, [0065]; wherein “FIG. 2 illustrates one example architecture of a multi-task multi-sensor fusion model for 2D and 3D object detection” Urtasun, [0066]). Regarding claim 9, Urtasun teaches an apparatus for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, comprising: at least one processor; and at least one memory coupled with the at least one processor and storing instructions operable, when executed by the at least one processor, to cause the apparatus to (“The one or more computing devices of the operations computing system 104 can include one or more processors and one or more memory devices. The one or more memory devices of the operations computing system 104 can store instructions that when executed by the one or more processors cause the one or more processors to perform operations and functions associated with operation of one or more vehicles” Urtasun, [0044]): generate, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle (“the model architecture can further include a machine-learned depth completion model. The depth completion model can be configured to receive the image feature map generated by the machine-learned image backbone model and to produce a depth completion map. For example, the depth completion map can describe a depth value for each pixel of the image” Urtasun, [0029]; wherein the sensor is that of “at least one camera configured to capture an image of an environment surrounding the autonomous vehicle” Urtasun, [0006]); each sparse depth estimate associated with a respective sparse representation of one or more sparse representations (“a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]); generate the dense LiDAR representation based on a dense depth estimate that is generated based on the depth estimate and the one or more sparse depth estimates (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image. Specifically, given the projected (into the image plane) sparse depth 210 from the LIDAR point cloud 206 and a camera image 208, the models 204 and 214 cooperate to output dense depth 222 at the same resolution as the input image 208” Urtasun, [0086]-[0087]; “Depth completion is exploited to learn better cross-modality feature representation and achieve dense feature map fusion by transforming predicted dense depth image into dense pseudo LIDAR points” Urtasun, [0066]; see additionally, [0091]; wherein the pseudo-LIDAR points are in 3D space, see Urtasun [0088]); and control an action of the vehicle based on the dense LiDAR representation (“output(s) can be provided to one or more of the perception system 124, prediction system 126, motion planning system 128, and vehicle control system 138 to implement additional autonomy processing functionality based on the output(s). For example, motion planning system 128 of FIG. 1 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) based at least in part on the output(s) of the system illustrated in FIG. 2. Stated differently, given information about the current locations of objects detected via the output(s) and/or predicted future locations and/or moving paths of proximate objects detected via the output(s), the motion planning system 128 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) that best navigates the autonomous vehicle (e.g., vehicle 102) along a determined travel route relative to the objects at such locations. The motion planning system 128 then can provide the selected motion plan to a vehicle control system 138 that controls one or more vehicle controls (e.g., actuators or other devices that control gas flow, steering, braking, etc.) to execute the selected motion plan” Urtasun, [0065]; wherein Fig. 2 illustrates the generation of a dense depth image of lidar representation). Urtasun fails to explicitly teach generate, via a sparse depth network, one or more sparse depth estimates of the environment. However, Guizilini teaches generate, via a sparse depth network, one or more sparse depth estimates of the environment (“the use of sparse convolutions to process input depth maps, while RGB images are still processed using standard convolutions. More specifically, we use Minkowski convolutions [4], a highly efficient generalized sparse convolution recently introduced to address high-dimensional problems. In this work we focus on the 2D application of Minkowski convolutions (image processing)” Guizilini, pg. 3 Col 1; wherein the sparse depth network is that of a Sparse Auxiliary Network (SANS) described in Fig. 2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Urtasun of having an apparatus for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, with the teachings of Guizilini of having generate, via a sparse depth network, one or more sparse depth estimates of the environment. Wherein having Urtasun’s multi-sensor fusion system of having generate, via a sparse depth network, one or more sparse depth estimates of the environment. The motivation behind the modification would have been to obtain multi-sensor fusion for three-dimensional object detection, since both Urtasun and Guizilini are systems for predicting dense depth using camera and LIDAR. Wherein Urtasun multi-sensor fusion system provides a multi-sensor detector that reasons about a target task of 2D and/or 3D object detection in addition to one or more auxiliary tasks such as ground estimation, depth completion, and/or other tasks such as other vision tasks, while Guizilini enables a monocular depth prediction network to also perform depth completion in the presence of optional sparse 3D measurements at inference time. Please see Urtasun et al. (US 20200160559 A1), Paragraph [0021] and Guizilini et al. (“Sparse auxiliary networks for unified monocular depth prediction and completion”, 2021), pg. 1 Col 2). Regarding claim 10, Urtasun in view of Guizilini teach the apparatus of claim 9, Urtasun further teaches wherein execution of the instructions further causes the apparatus to perform one or more vision-based tasks based on a combination of features associated with the image and the one or more sparse depth estimates (“fusing image features from image data (e.g., image data obtained at 704) with LIDAR features from the BEV representation of the LIDAR data. In some implementations, fusing at 712 can include executing one or more continuous convolutions to fuse image features from a first data stream with LIDAR features from a second data stream... generating a feature map comprising the fused image features and LIDAR features determined at 712.... detecting three-dimensional objects of interest based on the fused ROI crops generated at 714. In some implementations, detecting objects of interest at 716 can include providing the feature map generated at 714 as input to a machine-learned refinement model” Urtasun, [0109]-[0111]; wherein the “first data stream descriptive of image data and a second data stream descriptive of LIDAR point cloud data” Urtasun, [0149] in which the “parse depth image 210 can be generated from the LIDAR point cloud 206” Urtasun, [0088]). Regarding claim 11, Urtasun in view of Guizilini teach the apparatus of claim 10, Urtasun further teaches wherein the one or more vision-based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment (“detecting objects of interest at 716 can include providing the feature map generated at 714 as input to a machine-learned refinement model. In response to receiving the feature map, the machine-learned refinement model can be trained to generate as output a plurality of detections corresponding to identified objects of interest within the feature map. In some implementations, detecting objects of interest at 716 can include determining a plurality of object classifications and/or bounding shapes corresponding to the detected objects of interest. For example, in one implementation, the plurality of objects detected at 716 can include a plurality of bounding shapes at locations within the feature map(s) having a confidence score associated with an object likelihood that is above a threshold value. In some implementations, detecting objects of interest at 716 can include determining one or more of a classification indicative of a likelihood that each of the one or more objects of interest comprises a class of object from a predefined group of object classes (e.g., vehicle, bicycle, pedestrian, etc.) and a bounding shape representative of a size, a location, and an orientation of each the one or more objects of interest” Urtasun, [0111]), or generating a semantic segmentation map of the environment. Regarding claim 12, Urtasun in view of Guizilini teach the apparatus of claim 9, Urtasun further teaches wherein: generating the dense LiDAR representation comprises: and the dense LiDAR representation is based on the 3D space (“example implementations of the present disclosure use the machine-learned depth completion model to predict dense depth. The predicted depth can be used as pseudo-LIDAR points to find dense correspondences between multi-sensor feature maps” Urtasun, [0030]; “Depth completion is exploited to learn better cross-modality feature representation and achieve dense feature map fusion by transforming predicted dense depth image into dense pseudo LIDAR points” Urtasun, [0066]; see additionally, [0091]; wherein the pseudo-LIDAR points are in 3D space, see Urtasun [0088]). Urtasun fails to explicitly teach decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate. However, Guizilini teaches decoding the depth estimate via a depth decoder (“The dense RGB module can be any encoder-decoder depth prediction network that uses skip connections. In our work we consider two baseline state-of-the-art network architectures: Pack Net and BTS” Guizilini, pg. 4 Col 1); and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate (“The depth maps predicted by PackNet-SAN were projected into 3D as pseudo-LiDAR pointclouds” Guizilini, pg. 7 Col 1). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Urtasun of having An apparatus for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, with the teachings of Guizilini of having decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate. Wherein having Urtasun’s multi-sensor fusion system of having decoding the depth estimate via a depth decoder; and converting a two-dimensional representation of the environment to a 3D space based on the decoded depth estimate. The motivation behind the modification would have been to obtain multi-sensor fusion for three-dimensional object detection, since both Urtasun and Guizilini are systems for predicting dense depth using camera and LIDAR. Wherein Urtasun multi-sensor fusion system provides a multi-sensor detector that reasons about a target task of 2D and/or 3D object detection in addition to one or more auxiliary tasks such as ground estimation, depth completion, and/or other tasks such as other vision tasks, while Guizilini enables a monocular depth prediction network to also perform depth completion in the presence of optional sparse 3D measurements at inference time. Please see Urtasun et al. (US 20200160559 A1), Paragraph [0021] and Guizilini et al. (“Sparse auxiliary networks for unified monocular depth prediction and completion”, 2021), pg. 1 Col 2). Regarding claim 14, Urtasun in view of Guizilini teach the apparatus of claim 9, Urtasun further teaches further comprising instructions operable to generate each sparse representation by a respective sparse representation sensor of one or more sparse representation sensors integrated with the vehicle (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image” Urtasun, [0086]; “a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]). Regarding claim 15, Urtasun in view of Guizilini teach the apparatus of claim 14, Urtasun further teaches wherein: the one or more sparse representations include one or more of a sparse LiDAR representation or a radar representation (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image” Urtasun, [0086]; “a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]); and the one or more sparse representation sensors include one or more of a sparse LiDAR sensor or a radar sensor (“As one example, an autonomous vehicle can include one or more cameras and a LIDAR system... Other sensors can include radio detection and ranging (RADAR) sensors, ultrasound sensors, and/or the like” Urtasun, [0022]). Regarding claim 16, Urtasun in view of Guizilini teach the apparatus of claim 9, Urtasun further teaches wherein execution of the instructions causes the apparatus to control the vehicle's action based on identifying a three-dimensional object in the dense LiDAR representation (“For example, motion planning system 128 of FIG. 1 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) based at least in part on the output(s) of the system illustrated in FIG. 2. Stated differently, given information about the current locations of objects detected via the output(s) and/or predicted future locations and/or moving paths of proximate objects detected via the output(s), the motion planning system 128 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) that best navigates the autonomous vehicle (e.g., vehicle 102) along a determined travel route relative to the objects at such locations” Urtasun, [0065]; wherein “FIG. 2 illustrates one example architecture of a multi-task multi-sensor fusion model for 2D and 3D object detection” Urtasun, [0066]). Regarding claim 17, Urtasun teaches a non-transitory computer-readable medium having program code recorded thereon for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, the program code executed by at least one processor and comprising (“one or more tangible, non-transitory, computer readable media can store instructions that when executed by the one or more processors” Urtasun, [0051]): program code to generate, at a depth estimation network, a depth estimate of an environment depicted in an image captured by an image capturing sensor integrated with the vehicle (“the model architecture can further include a machine-learned depth completion model. The depth completion model can be configured to receive the image feature map generated by the machine-learned image backbone model and to produce a depth completion map. For example, the depth completion map can describe a depth value for each pixel of the image” Urtasun, [0029]; wherein the sensor is that of “at least one camera configured to capture an image of an environment surrounding the autonomous vehicle” Urtasun, [0006]); each sparse depth estimate associated with a respective sparse representation of one or more sparse representations (“a three-channel sparse depth image 210 can be generated from the LIDAR point cloud 206, representing the sub-pixel offsets and the depth value” Urtasun, [0088]); program code to generate the dense LiDAR representation based on a dense depth estimate that is generated based on the depth estimate and the one or more sparse depth estimates (“LIDAR provides long range 3D information for accurate 3D object detection. However, the observation is sparse especially at long range. The present disclosure proposes to densify LIDAR points by depth completion from both sparse LIDAR observation and RGB image. Specifically, given the projected (into the image plane) sparse depth 210 from the LIDAR point cloud 206 and a camera image 208, the models 204 and 214 cooperate to output dense depth 222 at the same resolution as the input image 208” Urtasun, [0086]-[0087]; “Depth completion is exploited to learn better cross-modality feature representation and achieve dense feature map fusion by transforming predicted dense depth image into dense pseudo LIDAR points” Urtasun, [0066]; see additionally, [0091]; wherein the pseudo-LIDAR points are in 3D space, see Urtasun [0088]); and program code to control an action of the vehicle based on the dense LiDAR representation (“output(s) can be provided to one or more of the perception system 124, prediction system 126, motion planning system 128, and vehicle control system 138 to implement additional autonomy processing functionality based on the output(s). For example, motion planning system 128 of FIG. 1 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) based at least in part on the output(s) of the system illustrated in FIG. 2. Stated differently, given information about the current locations of objects detected via the output(s) and/or predicted future locations and/or moving paths of proximate objects detected via the output(s), the motion planning system 128 can determine a motion plan for the autonomous vehicle (e.g., vehicle 102) that best navigates the autonomous vehicle (e.g., vehicle 102) along a determined travel route relative to the objects at such locations. The motion planning system 128 then can provide the selected motion plan to a vehicle control system 138 that controls one or more vehicle controls (e.g., actuators or other devices that control gas flow, steering, braking, etc.) to execute the selected motion plan” Urtasun, [0065]; wherein Fig. 2 illustrates the generation of a dense depth image of lidar representation). Urtasun fails to explicitly teach program code to generate, via a sparse depth network, one or more sparse depth estimates of the environment. However, Guizilini teaches program code to generate, via a sparse depth network, one or more sparse depth estimates of the environment (“the use of sparse convolutions to process input depth maps, while RGB images are still processed using standard convolutions. More specifically, we use Minkowski convolutions [4], a highly efficient generalized sparse convolution recently introduced to address high-dimensional problems. In this work we focus on the 2D application of Minkowski convolutions (image processing)” Guizilini, pg. 3 Col 1; wherein the sparse depth network is that of a Sparse Auxiliary Network (SANS) described in Fig. 2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Urtasun of having A non-transitory computer-readable medium having program code recorded thereon for generating a dense light detection and ranging (LiDAR) representation by a vision system of a vehicle, with the teachings of Guizilini of having program code to generate, via a sparse depth network, one or more sparse depth estimates of the environment. Wherein having Urtasun’s multi-sensor fusion system of having program code to generate, via a sparse depth network, one or more sparse depth estimates of the environment. The motivation behind the modification would have been to obtain multi-sensor fusion for three-dimensional object detection, since both Urtasun and Guizilini are systems for predicting dense depth using camera and LIDAR. Wherein Urtasun multi-sensor fusion system provides a multi-sensor detector that reasons about a target task of 2D and/or 3D object detection in addition to one or more auxiliary tasks such as ground estimation, depth completion, and/or other tasks such as other vision tasks, while Guizilini enables a monocular depth prediction network to also perform depth completion in the presence of optional sparse 3D measurements at inference time. Please see Urtasun et al. (US 20200160559 A1), Paragraph [0021] and Guizilini et al. (“Sparse auxiliary networks for unified monocular depth prediction and completion”, 2021), pg. 1 Col 2). Regarding claim 18, Urtasun in view of Guizilini teach the non-transitory computer-readable medium of claim 17, Urtasun further teaches further comprising program code to perform one or more vision-based tasks based on a combination of features associated with the image and the one or more sparse depth estimates (“fusing image features from image data (e.g., image data obtained at 704) with LIDAR features from the BEV representation of the LIDAR data. In some implementations, fusing at 712 can include executing one or more continuous convolutions to fuse image features from a first data stream with LIDAR features from a second data stream... generating a feature map comprising the fused image features and LIDAR features determined at 712.... detecting three-dimensional objects of interest based on the fused ROI crops generated at 714. In some implementations, detecting objects of interest at 716 can include providing the feature map generated at 714 as input to a machine-learned refinement model” Urtasun, [0109]-[0111]; wherein the “first data stream descriptive of image data and a second data stream descriptive of LIDAR point cloud data” Urtasun, [0149] in which the “parse depth image 210 can be generated from the LIDAR point cloud 206” Urtasun, [0088]). Regarding claim 19, Urtasun in view of Guizilini teach the non-transitory computer-readable medium of claim 18, Urtasun further teaches wherein the one or more vision-based tasks include one or more of generating an instance segmentation map of the environment, identifying a two-dimensional object in the environment (“detecting objects of interest at 716 can include providing the feature map generated at 714 as input to a machine-learned refinement model. In response to receiving the feature map, the machine-learned refinement model can be trained to generate as output a plurality of detections corresponding to identified objects of interest within the feature map. In some implementations, detecting objects of interest at 716 can include determining a plurality of object classifications and/or bounding shapes corresponding to the detected objects of interest. For example, in one implementation, the plurality of objects detected at 716 can include a plurality of bounding shapes at locations within the feature map(s) having a confidence score associated with an object likelihood that is above a threshold value. In some implementations, detecting objects of interest at 716 can include determining one or more of a classification indicative of a likelihood that each of the one or more objects of interest comprises a class of object from a predefined group of object classes (e.g., vehicle, bicycle, pedestrian, etc.) and a bounding shape representative of a size, a location, and an orientation of each the one or more objects of interest” Urtasun, [0111]), or generating a semantic segmentation map of the environment. Allowable Subject Matter Claims 5, 13 and 20 would be allowable if rewritten to overcome the non-statutory double patenting rejection set forth in this Office action and to include all the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. YOO et al. (US 20230230269 A1) – receiving an RGB image and a sparse image through a camera and LiDAR, generating a dense first depth map by processing color information of the RGB image through a first branch based on an encoder-decoder, generating a dense second depth map by up-sampling the sparse image through a second branch based on an encoder-decoder, generating a third depth map by fusing the first depth map and the second depth map. Ma et al. (“Self-supervised sparse-to-dense: Self-supervised depth completion from lidar and monocular camera”, 2019) – a deep regression model to learn direct mapping from sparse depth and color images to dense depth. Inquiries Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMMANUEL SILVA-AVINA whose telephone number is (571)270-0729. The examiner can normally be reached Monday - Friday 11 AM - 8 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EMMANUEL SILVA-AVINA/Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Oct 16, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743865
Apparatus For Recognizing Object And Method Thereof
2y 5m to grant Granted Sep 22, 2026
Patent 12694644
RE-IDENTIFICATION SYSTEM
2y 4m to grant Granted Jul 28, 2026
Patent 12682446
Non-Destructive Wire Bonding Inspection Method
2y 10m to grant Granted Jul 14, 2026
Patent 12659443
VIEWPOINT SYNTHESIS WITH ENHANCED 3D PERCEPTION
2y 11m to grant Granted Jun 16, 2026
Patent 12639791
IMAGE ENHANCEMENT USING TEXTURE MATCHING GENERATIVE ADVERSARIAL NETWORKS
2y 8m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
90%
With Interview (+9.5%)
2y 11m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 81 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month