Prosecution Insights
Last updated: October 02, 2026
Application No. 18/627,159

OBJECT DETECTION USING DENSE DEPTH AND LEARNED FUSION OF DATA OF CAMERA AND LIGHT DETECTION AND RANGING SENSORS

Non-Final OA §103
Filed
Apr 04, 2024
Examiner
PATEL, PINALBEN V
Art Unit
Tech Center
Assignee
TORC Robotics Inc.
OA Round
1 (Non-Final)
89%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
504 granted / 565 resolved
+29.2% vs TC avg
Moderate +10% lift
Without
With
+9.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
27 currently pending
Career history
579
Total Applications
across all art units

Statute-Specific Performance

§101
8.2%
-31.8% vs TC avg
§103
59.3%
+19.3% vs TC avg
§102
5.2%
-34.8% vs TC avg
§112
18.6%
-21.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 565 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Foreign priority is not claimed. Information Disclosure Statement The information disclosure statement (IDS) submitted on 04/04/2024, 06/12/2025 and 11/10/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Urtasun et al. (US Pub No. 20200160559 A1) in view of Unger et al. (“Multi-camera Bird’s Eye View Perception for Autonomous Driving”, Sep. 2023, as provided). Regarding Claim 1, Urtasun discloses A perception system, comprising: at least one memory configured to store machine executable instructions; and at least one processor configured to execute the stored executable instructions to: extract camera features from stereo images; (Urtasun, [0005], discloses a computing system configured to perform multi-task multi-sensor fusion for detecting objects. The computing system includes one or more processors and one or more non-transitory computer-readable media that collectively store: a machine-learned light detection and ranging (LIDAR) backbone model configured to receive a bird's eye view (BEV) representation of a LIDAR point cloud generated for an environment surrounding an autonomous vehicle and to process the BEV representation of the LIDAR point cloud to generate a LIDAR feature map; a machine-learned image backbone model configured to receive an image of the environment surrounding the autonomous vehicle and to process the image to generate an image feature map; a machine-learned refinement model configured to receive respective region of interest (ROI) feature crops from each of the LIDAR feature map and the image feature map; image feature map is generated from the captured camera image) extract LiDAR features from a LiDAR point cloud; (Urtasun, [0005], discloses a computing system configured to perform multi-task multi-sensor fusion for detecting objects. The computing system includes one or more processors and one or more non-transitory computer-readable media that collectively store: a machine-learned light detection and ranging (LIDAR) backbone model configured to receive a bird's eye view (BEV) representation of a LIDAR point cloud generated for an environment surrounding an autonomous vehicle and to process the BEV representation of the LIDAR point cloud to generate a LIDAR feature map; a machine-learned image backbone model configured to receive an image of the environment surrounding the autonomous vehicle and to process the image to generate an image feature map; a machine-learned refinement model configured to receive respective region of interest (ROI) feature crops from each of the LIDAR feature map and the image feature map, to perform ROI-wise fusion to fuse respective pairs of ROI feature crops to generate fused ROI feature crops, and to generate one or more object detections based on the fused ROI feature crops, wherein each of the one or more object detections indicates a location of a detected object within the environment surrounding the autonomous vehicle; LiDAR data is obtained) transform the LiDAR features in the BEV space; (Urtasun, [0005], discloses a computing system configured to perform multi-task multi-sensor fusion for detecting objects. The computing system includes one or more processors and one or more non-transitory computer-readable media that collectively store: a machine-learned light detection and ranging (LIDAR) backbone model configured to receive a bird's eye view (BEV) representation of a LIDAR point cloud generated for an environment surrounding an autonomous vehicle and to process the BEV representation of the LIDAR point cloud to generate a LIDAR feature map; a machine-learned image backbone model configured to receive an image of the environment surrounding the autonomous vehicle and to process the image to generate an image feature map; a machine-learned refinement model configured to receive respective region of interest (ROI) feature crops from each of the LIDAR feature map and the image feature map, to perform ROI-wise fusion to fuse respective pairs of ROI feature crops to generate fused ROI feature crops, and to generate one or more object detections based on the fused ROI feature crops, wherein each of the one or more object detections indicates a location of a detected object within the environment surrounding the autonomous vehicle; LiDAR data is obtained and converted to BEV space) and fuse the transformed camera features and LiDAR features in the BEV space using a learned fusion with attention technique to generate the fused camera features and LiDAR features in the BEV space. (Urtasun, [0005], discloses a machine-learned refinement model configured to receive respective region of interest (ROI) feature crops from each of the LIDAR feature map and the image feature map, to perform ROI-wise fusion to fuse respective pairs of ROI feature crops to generate fused ROI feature crops, and to generate one or more object detections based on the fused ROI feature crops, wherein each of the one or more object detections indicates a location of a detected object within the environment surrounding the autonomous vehicle; and a machine-learned depth completion model configured to receive the image feature map generated by the machine-learned image backbone model and to produce a depth completion map; ROI crops are built from image feature maps (BEV space of image data) and LiDAR BEV space and fused together) Urtasun does not explicitly disclose transform the camera features in a bird’s-eye-view (BEV) space; Unger discloses transform the camera features in a bird’s-eye-view (BEV) space; (Unger, pg. 2-5, Introduction, Fig. 1-2, discloses more recent approaches use deep neural networks to output directly in BEV space. These methods transform camera images into BEV space using geometric constraints implicitly or explicitly in the network; Camera-based systems depended on two types of image representations: Perspective View (PS) representation and Bird;s eye view (BEV) representation. BEV representation have gained much attention due to their efficacy in different parts of the automated driving pipeline; PNG media_image1.png 194 462 media_image1.png Greyscale ; PNG media_image2.png 288 467 media_image2.png Greyscale PNG media_image3.png 322 510 media_image3.png Greyscale ; camera is converted to BEV space representation and fused with BEV space representation of LiDAR image) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Urtasun in view of Unger having a method of fusing camera image and BEV space representation of LiDAR image, with the teachings of Unger having system, converting camera image to BEV space representation and further fusing with LiDAR data in order to accurately depict the scene before in applications including autonomous vehicle navigation. Regarding Claim 2, The combination of Urtasun and Unger further discloses wherein to transform the camera features in the BEV space, the at least one processor is further configured to use dense depth or per pixel depth (Urtasun, [0030-0032], discloses Exploiting the task of depth completion can assist in learning better cross-modality feature representation and more importantly, as is discussed further elsewhere herein, in achieving dense point-wise feature fusion with pseudo LIDAR points from dense depth. In particular, existing approaches which attempt to perform sensor fusion through projection of the LIDAR data suffer when the LIDAR points are sparse (e.g., the points within a certain region of the environment that is relatively distant from the LIDAR system). To address this issue, example implementations of the present disclosure use the machine-learned depth completion model to predict dense depth. The predicted depth can be used as pseudo-LIDAR points to find dense correspondences between multi-sensor feature maps; example implementations of the proposed detection architecture perform both point-wise and ROI-wise feature fusion. In particular, in some implementations, the machine-learned LIDAR backbone model can include one or more point-wise fusion layers. The point-wise fusion layer(s) can be configured to receive image feature data from one or more intermediate layers of the machine-learned image backbone model and perform point-wise fusion to fuse the image feature data with one or more intermediate LIDAR feature maps generated by one or more intermediate layers of the machine-learned LIDAR backbone model. Thus, point-wise feature fusion can be applied to fuse multi-scale image features from the image stream to the BEV stream. In some implementations, as indicated above, the depth completion map generated by the machine-learned depth completion model can be used to assist in performing dense fusion; aspects of the present disclosure are directed to a multi-task multi-sensor detection model that jointly reasons about 2D and 3D object detection, ground estimation, and/or depth completion. Point-wise and ROI-wise feature fusion can both be applied to achieve full multi-sensor fusion, while multi-task learning provides additional map prior and geometric clues enabling better representation learning and denser feature fusion; depth is obtained from LiDAR feature points to be fused with ROI of camera image and represented in 2D or 3D form space representation) in the stereo images to project the stereo images into a three-dimensional (3D) representation in the BEV space. (Unger, pg. 4, pg. 15, discloses where the main aim is to safely interact with the static and dynamic elements of the scene. Thus, both PV and BEV representations form an integral part of automated driving and are crucial for an automated vehicle's safe and efficient operation; PV and BEV representations are typically employed at distinct stages of the perception pipeline, with PV mainly used in the earlier stages for tasks such as segmentation and tracking. In contrast, BEV is employed in subsequent sensor fusion and path planning tasks. Thus, there is a disconnect within the pipeline since the generation of BEV representation requires a depth estimate that cannot be directly obtained from the PV representation. Multiple approaches address this disconnect by either explicitly predicting the scene's depth using depth estimation networks or by extracting depth information from range-based sensors such as LiDARs. These depth es timates are then combined with the intermediate outputs from the PV to generate the required BEV representations. These multi-stage approaches, however, generate sub-optimal BEV representation due to scale inconsistencies in depth estimation networks and the sparsity of range-based sensors. Many recent works have proposed various deep learning-based approaches to generate BEV representations directly from PV images, following an end-to-end learning strategy to alleviate this limitation. These approaches learn the complex characteristics of PV-BEV mapping using neural networks, thus generating highly accurate representations in the BEV; Argoverse provides both 3D bounding boxes and semantic maps and provides even the benefits ofhaving, in addition to the 360° field of view, a pair offront-facing stereo cameras. Argoverse contains fewer scenes than nuScenes or the Waymo Open dataset, which is why it is less used. The KITTI-360 dataset provides 3D bounding boxes and semantic map information for a pair of front-viewing stereo cameras and two fisheye cameras; depth of image is converted to 2D or 3D BEV space representation). Additionally, the rational and motivation to combine the references Urtasun and Unger as applied in rejection of claim 1 apply to this claim. Regarding Claim 3, The combination of Urtasun and Unger further discloses wherein the dense depth or per pixel depth in the stereo images is determined based upon at least a focal length of the stereo cameras, a baseline corresponding to a distance between two lenses of the stereo cameras, and a disparity corresponding to a horizontal displacement between a pair of corresponding pixels on the stereo images. (Unger, pg. 11, geometry-based transformation approaches use the camera parameters like posi tion, orientation, or focal length, which are given by the extrinsic and intrinsic camera parameters and are known for cameras used in the autonomous driving context. Based on those parameters, one can calculate the 3D position of the image pixel up to the ambiguous depth of the pixel; depth is obtained by focal length of camera). Additionally, the rational and motivation to combine the references Urtasun and Unger as applied in rejection of claim 1 apply to this claim. Regarding Claim 4, The combination of Urtasun and Unger further discloses wherein to transform the LiDAR features in the BEV space, the at least one processor is further configured to flatten the LiDAR features along an axis in which the LiDAR features have higher granularity in comparison to LiDAR features along other axes. (Unger, pg. 8, discloses grid size defines the number of grid cells and determines the spatial size of the BEV segmentation output. The larger the grid size, the higher the information that can be represented in the segmentation output. However, a larger grid size increases the computational complexity, which subsequently increases the computational requirements as well as the runtime of the model. In contrast, grid resolution determines the real-world size of a single grid cell. Given a fixed grid size, a finer resolution allows the BEV grid to capture finer details, such as elements belonging to lane markings and pedestrians, at the expense of a smaller real-world area. In comparison, a coarser grid resolution increases the real-world area covered by the BEV grid at the cost of a decreased segmentation granularity; features are concatenated into smaller representation by converting them to flattening features by their granularity along specific axis). Additionally, the rational and motivation to combine the references Urtasun and Unger as applied in rejection of claim 1 apply to this claim. Regarding Claim 5, The combination of Urtasun and Unger further discloses wherein the camera features are extracted from the stereo images using a camera encoder stack including a series of convolutional layers configured to extract different levels of features from the stereo images. (Unger, pg. 9-10, fig. 1.6, discloses 1.3.1 Image encoder The image encoder, also known as a feature extractor or backbone, forms the first component of a deep neural network. Its main task is to extract features from the input image for use in the downstream modules. This module often consumes a substantial portion of computational resources compared to the whole network. The structure of the network backbone is often non-trivial. It involves engineering several hyper parameters, such as the number of layers, the size of each layer, and connections between layers, among many others. Consequently, several works have explored various strategies to develop optimal network backbones. Several works reuse these backbone designs in BEV perception since the feature extraction step is similar in both PV and BEV perception pipelines; which applies cross-attention to decide for each BEV grid cell which PV features are relevant. The EfficientDet image encoder extends EfficientNet with a Bi-directional Feature Pyramid Network (BiFPN) layerfor multi-scale feature fusion. It is used by Panop ticBEV to generate the image features for the task of BEV panoptic segmentation. Another popular class of image encoders are attention-based vision transformers such as Swin. Through their success in image classification, attention based backbones are also employed for BEV perception. It is important to note that switching from one image encoder to another is relatively easy. However, the model runtime and the training convergence rate heavily depend on the choice of the image encoder. This decision becomes even more relevant when multiple cameras are used since the inference time of the model can decide whether a model is deployed in the real world or not. Table 1.2 and Table 1.1 present an overview of the image encoder used by multiple popular BEV perception approaches; different types of encoders are layered to extract different features from image). Additionally, the rational and motivation to combine the references Urtasun and Unger as applied in rejection of claim 1 apply to this claim. Regarding Claim 6, The combination of Urtasun and Unger further discloses wherein the LiDAR features are extracted from the LiDAR point cloud using a LiDAR encoder stack including a series of convolutional layers configured to extract semantic information of the LiDAR point cloud. (Unger, pg. 9-11, 1.3.1 Image encoder, Fig. 1.6, discloses The image encoder, also known as a feature extractor or backbone, forms the first component of a deep neural network. Its main task is to extract features from the input image for use in the downstream modules. This module often consumes a substantial portion of computational resources compared to the whole network. The structure of the network backbone is often non-trivial. It involves engineering several hyper parameters, such as the number of layers, the size of each layer, and connections between layers, among many others. Consequently, several works have explored various strategies to develop optimal network backbones. Several works reuse these backbone designs in BEV perception since the feature extraction step is similar in both PV and BEV perception pipelines; Geometry-based BEV transformation approaches show great performance since they model the real world using a 3D model that is condensed in a subsequent step. Recent architectures use a four-step transformation approach: First, a primarily CNN-based image encoder generates image features from the input images. Second, the view transformation module transforms the image features from the 2D perspective space into the BEV space. Then a BEV encoder processes the features using a convolutional network in the BEV space. Finally, the encoded BEV features are fed into one or several task-specific heads, e.g., for semantic segmentation, 3D object detection, or both; encoders are stacked to extract semantic features from LiDAR point or image). Additionally, the rational and motivation to combine the references Urtasun and Unger as applied in rejection of claim 1 apply to this claim. Regarding Claim 7, The combination of Urtasun and Unger further discloses wherein the at least one processor is further configured to decode the fused camera features and LiDAR features in the BEV space for lane line segmentation, lane marking detection, or three-dimensional object detection.(Unger, pg.3-4, discloses On the other hand, the BEV representation depicts the scene from the viewpoint of a downward-facing virtual orthographic camera placed above the automated vehicle. Historically, surround vision systems commonly used the IPM principle to generate a BEV image and to show it on display to the driver, as depicted in Figure 1.1. It captures the scene's depth proportional to the metric scale, which allows it to be directly used for distance-sensitive tasks such as collision avoidance. It also can explicitly capture occlusions in the scene, allowing subsequent tasks, such as path planning and control, to handle the ambiguity associated with such regions gracefully. Lastly, being an orthographic projection, BEV representation does not suffer from perspective distortion inherent to PV, simplifying the representation and processing of lane geometry and road markings. These characteristics make BEV representation apt 4 Multi-camera Bird's Eye View Perception for Autonomous Driving for decision-based tasks such as trajectory estimation and control, where the main aim is to safely interact with the static and dynamic elements of the scene. Thus, both PV and BEV representations form an integral part of automated driving and are crucial for an automated vehicle's safe and efficient operation. PV and BEV representations are typically employed at distinct stages of the perception pipeline, with PV mainly used in the earlier stages for tasks such as segmentation and tracking. In contrast, BEV is employed in subsequent sensor fusion and path planning tasks. Thus, there is a disconnect within the pipeline since the generation of BEV representation requires a depth estimate that cannot be directly obtained from the PV representation. Multiple approaches address this disconnect by either explicitly predicting the scene's depth using depth estimation networks or by extracting depth information from range-based sensors such as LiDARs. These depth estimates are then combined with the intermediate outputs from the PV to generate the required BEV representations; BEV space representation is obtained to accurately estimate lane and road markings at varying stages of sensor fusion). Additionally, the rational and motivation to combine the references Urtasun and Unger as applied in rejection of claim 1 apply to this claim. Claims 8-14 recite vehicle system with elements corresponding to the system elements recited in Claims 1-7 respectively. Therefore, the recited elements of the vehicle system claims 8-14 are mapped to the proposed combination in the same manner as the corresponding elements of Claim 8-14 respectively. Additionally, the rationale and motivation to combine the Urtasun and Unger references presented in rejection of Claim 1, apply to these claims. Furthermore, the combination of Urtasun and Unger further discloses A vehicle, comprising: a stereo camera configured to capture stereo images; a light detection and ranging (LiDAR) sensor configured to generate data of a LiDAR point cloud; at least one memory configured to store machine executable instructions; and at least one processor configured to execute the stored executable instructions (Urtasun, [0043-0044], discloses operations computing system 104 can be associated with a service provider that can provide one or more vehicle services to a plurality of users via a fleet of vehicles that includes, for example, the vehicle 102. The vehicle services can include transportation services (e.g., rideshare services), courier services, delivery services, and/or other types of services; operations computing system 104 can include multiple components for performing various operations and functions. For example, the operations computing system 104 can include and/or otherwise be associated with the one or more computing devices that are remote from the vehicle 102. The one or more computing devices of the operations computing system 104 can include one or more processors and one or more memory devices. The one or more memory devices of the operations computing system 104 can store instructions that when executed by the one or more processors cause the one or more processors to perform operations and functions associated with operation of one or more vehicles (e.g., a fleet of vehicles), with the provision of vehicle services, and/or other operations as discussed herein). Claims 15-20 recite method with steps corresponding to the system elements recited in Claims 1-4, (5 and 6), and 7 respectively. Therefore, the recited steps of the method claims 15-20 are mapped to the proposed combination in the same manner as the corresponding elements of Claim 1-4, (5 and 6), and 7 respectively. Additionally, the rationale and motivation to combine the Urtasun and Unger references presented in rejection of Claim 1, apply to these claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: US-20180232947-A1 (Nehmadi et al., A system and method for generating a high-density three-dimensional (3D) map are disclosed. The system comprises acquiring at least one high density image of a scene using at least one passive sensor; acquiring at least one new set of distance measurements of the scene using at least one active sensor; acquiring a previously generated 3D map of the scene comprising a previous set of distance measurements; merging the at least one new set of distance measurements with the previous set of upsampled distance measurements, wherein merging the at least one new set of distance measurements further includes accounting for a motion transformation between a previous high-density image frame and the acquired high density image and the acquired distance measurements; and overlaying the new set of distance measurements on the high-density image via an upsampling interpolation, creating an output 3D map, Abstract) Any inquiry concerning this communication or earlier communications from the examiner should be directed to PINALBEN V PATEL whose telephone number is (571)270-5872. The examiner can normally be reached M-F: 10am - 8pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at 571-272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Pinalben Patel/Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Apr 04, 2024
Application Filed
Aug 31, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749174
OPTIMAL DETERMINATION OF AN OVERLAY TARGET USING MACHINE LEARNING
2y 7m to grant Granted Sep 29, 2026
Patent 12743934
SYSTEMS AND METHODS FOR DYNAMICALLY DETECTING DISABILITIES
3y 0m to grant Granted Sep 22, 2026
Patent 12737897
OBJECT TRACKING USING PREDICTED POSITIONS
2y 10m to grant Granted Sep 15, 2026
Patent 12725236
LOW-LIGHT IMAGE ENHANCEMENT METHOD AND DEVICE BASED ON WAVELET TRANSFORM AND RETINEX-NET
2y 3m to grant Granted Sep 01, 2026
Patent 12718605
MACHINE-LEARNING MODELS FOR IMAGE PROCESSING
1y 0m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
89%
Grant Probability
99%
With Interview (+9.5%)
2y 3m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 565 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month