Prosecution Insights
Last updated: October 01, 2026
Application No. 19/077,562

SCALABLE MULTI-MODAL PERCEPTION FRAMEWORK FOR AUTONOMOUS SYSTEMS AND APPLICATIONS

Non-Final OA §102§103
Filed
Mar 12, 2025
Priority
Apr 30, 2024 — provisional 63/640,750
Examiner
HAKALA, ALAN GREGORY
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
24 currently pending
Career history
20
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 6, 8-12, 14-18, 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Liu (BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation). Regarding claim 16, 1, Liu teaches: At least one processor comprising one or more circuits (Liu 3.2 “Though simple, BEV pooling is surprisingly inefficient and slow, taking more than 500ms on an RTX 3090 GPU” Note: Liu teaches the specific processor used is an Nvidia RTX 3090, teaching at least one processor comprised of one or more circuits.) to assemble components of a multi-modal perception pipeline according to configuration data that identifies the components, (Liu Abstract “Multi-sensor fusion is essential for an accurate and reliable autonomous driving system. Recent approaches are based on point-level fusion: augmenting the LiDARpoint cloud with camera features. However, the camera-to-LiDAR projection throws away the semantic density of camera features, hindering the effectiveness of such methods, especially for semantic-oriented tasks (such as 3D scene segmentation). In this paper, we break this deeply-rooted convention with BEVFusion, an efficient and generic multi-task multi-sensor fusion framework. It unifies multi modal features in the shared bird’s-eye view (BEV) representation space, which nicely preserves both geometric and semantic information … BEVFusion is fundamentally task-agnostic and seamlessly supports different 3D perception tasks with almost no architectural changes. It establishes the new state of the art on nuScenes, achieving 1.3% higher mAP and NDS on 3D object detection and 13.6% higher mIoU on BEV map segmentation, with 1.9× lower computation cost.” 3 “Given different sensory inputs, we first apply modality-specific encoders to extract their features. We transform multi-modal features into a unified BEV representation that preserves both geometric and semantic information.” Note: A pipeline in this machine learning context refers to the ordered sequence of accepting inputs, transforming them into usable means, and generating an output based on the inputs. Liu teaches a model that capable of producing outputs like object detection and segmentation from a multi-modal perception input. Specifically, the multi-model perception comes from LiDAR captures and camera image captures that are input to the model. It is known that these perception components are identified specifically as Liu 3 teaches modality-specific encoders are used, which requires that the type of data input is known. Thus, Liu teaches a multi-modal preception pipeline according to configuration data that identifies components. This pipeline data taught by Liu meets the definition of “configuration data” which defines structured set up of a model.) the multi-modal perception pipeline to: synchronize, using a first component of the components, first sensor data corresponding to a first sensor modality and second sensor data corresponding to a second sensor modality into one or more synchronized frames; (Liu 1 “Data from different sensors are expressed in fundamentally different modalities: e.g., cameras capture data in perspective view and LiDAR in 3D view. To resolve this view discrepancy, we have to find a unified representation that is suitable for multi-task multi-modal feature fusion … In this paper, we propose BEVFusion to unify multi-modal features in a shared bird’s-eye view (BEV) representation space for task-agnostic learning” 4.2 “For each frame, we only perform the evaluation in the [-50m, 50m]×[-50m, 50m] region around the ego car following [39, 70, 63, 25]. In BEVFusion, we use a single model that jointly performs binary segmentation for all classes instead of following the conventional approach to train a separate model for each class. This results in 6× faster inference and training. We reproduced the results of all open-source competing methods” 3.1 “3.1 Unified Representation Different features can exist in different views. For instance, camera features are in the perspective view, while LiDAR/radar features are typically in the 3D/bird’s-eye view. Even for camera features, each one of them has a distinct viewing angle (i.e., front, back, left, right). This view discrepancy” PNG media_image1.png 496 1226 media_image1.png Greyscale Note: Liu 1 teaches that a first sensor data, in this case camera captured images, and a second sensor data, LiDAR point cloud data, are unified, or synchronized, to form a single bird’s eye view, the view in 4.2 is also referred to as a ‘frame’. PNG media_image2.png 526 820 media_image2.png Greyscale As seen from Fig. 2 of Liu and Fig. 7D of the specifications, both teach accepting the first and second sensor inputs in different modalities, which are in both cases is camera and lidar data. Both teach specific encoders for the different modalities of the inputs and a camera to BEV transformer for the camera capture data before a final BEV encoder produces the fused/synchronized frame output.) compute, using one or more multi-modal inference models of a second component of the components processing the one or more synchronized frames, inference data indicating 3D information associated with the first sensor data and the second sensor data; ( Liu 3.4 “We apply multiple task-specific heads to the fused BEV feature map. Our method is applicable to most 3D perception tasks. We showcase two examples: 3D object detection and BEV map segmentation. Detection. We use a class-specific center heatmap head to predict the center location of all objects and a few regression heads to estimate the object size, rotation, and velocity. We refer the readers to previous 3D detection papers [1, 67, 68] for more details. Segmentation. Different map categories may overlap (e.g., crosswalk is a subset of drivable space). Therefore, we formulate this problem as multiple binary semantic segmentation, one for each class. We follow CVT [70] to train the segmentation head with the standard focal loss [29].”3 “Method BEVFusion focuses on multi-sensor fusion (i.e., multi-view cameras and LiDAR) for multi-task 3D perception (i.e., detection and segmentation). We provide an overview of our framework in Figure 2. Given different sensory inputs, we first apply modality-specific encoders to extract their features. We transform multi-modal features into a unified BEV representation that preserves both geometric and semantic information. We identify the efficiency bottleneck of the view transformation and accelerate BEVpooling with precomputation and interval reduction. We then apply the convolution-based BEV encoder to the unified BEV features to alleviate the local misalignment between different features. Finally, we append a few task-specific heads to support different 3D tasks. “ Note: Liu teaches that once a synchronized frame, which in Liu is a unified bird’s eye view, a task specific head can be added to the end of the pipeline, naming 3D information tasks specifically. Examples that are given are 3D object detection and segmentation. It is known these fit the claims definition of 3D tasks as the specifications state ¶118 “For example, the frame 780 may correspond to a rendered frame 612 of FIG. 6. In this example, the frame 780 displays six cameras in six different views (top row and bottom row of the rendered result). The top row shows different front views of a vehicle (e.g., autonomous vehicle 1000 of FIG. 10A), such as front-left, front, and front-right views. The bottom row shows the rear views of the vehicle, such as rear-left, rear, and rear-right views. In this example, one single LiDAR data source 704 is rendered in two different views (middle row of the rendered result). The two different views include a top view and front view. As shown in FIG. 7E, 3D bounding boxes or shapes (e.g., a bounding shape 790) from the 3D inference result are displayed in all the views, with other intrinsic/extrinsic parameters.” As seen from the provided citation the specifications state object detection resulting in bounding boxes is an example of a 3D inference.) and generate, using a third component of the components and the 3D information, a rendering including multiple views associated with the first sensor modality and the second sensor modality. (.( PNG media_image3.png 718 1328 media_image3.png Greyscale Note: The specifications provide examples of the first and second sensor views being used with the 3D information, like object detection and segmentation, to produce a rendering as seen in Fig. 7E, cited below and described previously in ¶188 cited above. From Liu Fig. 4 it can be seen that Liu produces the same type of renderings where multiple views of both sensors, LiDAR and camera capture, are combined with 3D inference outputs like bounding boxes from object detection to produce the renderings. PNG media_image4.png 498 760 media_image4.png Greyscale ) Regarding claims 17, 2, Liue teaches: The at least one processor of claim 16, wherein the multi-modal perception pipeline is further to match, using a fourth component of the components and calibration data associated with the first sensor data and the second sensor data, first data points corresponding to the first sensor data with second data points corresponding to the second sensor data to align the first sensor data with the second sensor data in the one or more synchronized frames, (Liu 3 “We transform multi-modal features into a unified BEV representation that preserves both geometric and semantic information.” 3.1 “On the one hand, the LiDAR-to-BEV projection flattens the sparse LiDAR features along the height dimension, thus does not create geometric distortion” 3.2 “Camera-to-BEV transformation is non-trivial because the depth associated with each camera feature pixel is inherently ambiguous. Following LSS [39] and BEVDet [20, 19], we explicitly predict the discrete depth distribution of each pixel. We then scatter each feature pixel into D discrete points along the camera ray and rescale the associated features by their corresponding depth probabilities (Figure 3a). This generates a camera feature point cloud of size NHWD, where N is the number of cameras and (H,W) is the camera feature map size. Such 3D feature point cloud is quantized along the x,y axes with a step size of r (e.g., 0.4m). We use the BEV pooling operation to aggregate all features within each r × r BEV grid and flatten the features along the z-axis.” 3.3 “With all sensory features converted to the shared BEV representation, we can easily fuse them together with an elementwise operator (such as concatenation). Though in the same space, LiDAR BEV features and camera BEV features can still be spatially misaligned to some extent due to the inaccurate depth in the view transformer. To this end, we apply a convolution-based BEV encoder (with a few residual blocks) to compensate for such local misalignments. Our method could potentially benefit from more accurate depth estimation (e.g., supervising the view transformer with ground-truth depth [43, 38]), which we leave for future work” Note: Liu teaches that the first and second sensor inputs are converted to BEVs so that they share a common format/space. Before the sensor data can be fused Liu teaches that there may be misalignments between the data that should be corrected. In this case the ‘calibration data’ is the BEV representation made by making pixels of the camera capture into points to project and flattening the z axis of the LiDAR data. As the fused BEV representation has already been shown to be analogous to the claims synchronized frame in claim 1, Liu teaches producing a synchronized frame from a first and second sensor data with calibration data that involves the aligning of the two sensor’s data.) and the computing of the inference data is based at least on the matching. ( PNG media_image1.png 496 1226 media_image1.png Greyscale Note: As seen from the pipeline depicted in Liu Fig.2, the BEV encoder which performs the fusion task described in Liu 3.3 produces the fused BEV features which is then directly used by task-specific heads to infer the final 3D inference data.) Regarding claims 18, 3, Liu teaches: The at least one processor of claim 16, wherein the first component is derived, at least in part, from a first interface element of a multi-modal sensor fusion framework, the first interface element providing a set of predefined synchronization methods and data structures used to perform the synchronizing. ( PNG media_image1.png 496 1226 media_image1.png Greyscale Note: Claim 1 has already established that the first component is the means through which the synchronized frame, which in Liu is the fused BEV, is taught by Liu. The claims “first interface element” which produces data structures and synchronization methods has similarly already been shown to be taught by Liu. As seen from Fig. 2 above, the portion which encompasses the first component that produces the frame accepts the two-sensor data inputs and encodes them. The process of encoding implicitly places the raw data into a data structure, and the transformation to a BEV and subsequent BEV fusion teaches the method to perform the synchronizing.) Regarding claim 6, Liu teaches: The computer-implemented method of claim 1, further comprising converting, using a fourth component of the components, the first sensor data into a unified data structure format that is shared with the second sensor data, wherein the synchronizing is performed on the first sensor data and the second sensor data in the unified data structure format. (Liu 3 “We transform multi-modal features into a unified BEV representation that preserves both geometric and semantic information.” 3.1 “On the one hand, the LiDAR-to-BEV projection flattens the sparse LiDAR features along the height dimension, thus does not create geometric distortion” 3.2 “Camera-to-BEV transformation is non-trivial because the depth associated with each camera feature pixel is inherently ambiguous. Following LSS [39] and BEVDet [20, 19], we explicitly predict the discrete depth distribution of each pixel. We then scatter each feature pixel into D discrete points along the camera ray and rescale the associated features by their corresponding depth probabilities (Figure 3a). This generates a camera feature point cloud of size NHWD, where N is the number of cameras and (H,W) is the camera feature map size. Such 3D feature point cloud is quantized along the x,y axes with a step size of r (e.g., 0.4m). We use the BEV pooling operation to aggregate all features within each r × r BEV grid and flatten the features along the z-axis.” 3.3 “With all sensory features converted to the shared BEV representation, we can easily fuse them together with an elementwise operator (such as concatenation). Though in the same space, LiDAR BEV features and camera BEV features can still be spatially misaligned to some extent due to the inaccurate depth in the view transformer. To this end, we apply a convolution-based BEV encoder (with a few residual blocks) to compensate for such local misalignments. Our method could potentially benefit from more accurate depth estimation (e.g., supervising the view transformer with ground-truth depth [43, 38]), which we leave for future work” Note: Liu teaches that the first and second sensor inputs are converted to BEVs so that they share a common format/space. Before the sensor data can be fused Liu teaches that there may be misalignments between the data that should be corrected. In this case the ‘calibration data’ is the BEV representation made by making pixels of the camera capture into points to project and flattening the z axis of the LiDAR data. As the fused BEV representation has already been shown to be analogous to the claims synchronized frame in claim 1, Liu teaches producing a synchronized frame from a first and second sensor data with calibration data that involves the aligning of the two sensor’s data. PNG media_image1.png 496 1226 media_image1.png Greyscale Note: The fourth component referenced here is also referenced in claim 2 where the content of this claim is listed in a different manner. The unified data structure format of the claims is the BEV representation which both the camera data and the LiDAR point cloud data is put into. Liu similarly teaches that once the sensor data is in the unified data structure format, aka BEV representation, the synchronizing is performed. In the case of Liu, the synchronizing is the fusing of the BEV representations for LiDAR and Camera data to produce the fused BEV features, as shown in Fig. 2 above.) Regarding claim 8, Liu teaches: The computer-implemented method of claim 1, wherein the one or more synchronized frames include a HashMap storing key-value pairs representing the first sensor data and the second sensor data. (Liu 3.2 “Camera-to-BEV transformation is non-trivial because the depth associated with each camera feature pixel is inherently ambiguous. Following LSS [39] and BEVDet [20, 19], we explicitly predict the discrete depth distribution of each pixel. We then scatter each feature pixel into D discrete points along the camera ray and rescale the associated features by their corresponding depth probabilities (Figure 3a). This generates a camera feature point cloud of size NHWD, where N is the number of cameras and (H,W) is the camera feature map size. Such 3D feature point cloud is quantized along the x,y axes with a step size of r (e.g., 0.4m). We use the BEV pooling operation to aggregate all features within each r × r BEV grid and flatten the features along the z-axis … The first step of BEV pooling is to associate each point in the camera feature point cloud with a BEV grid. Different from LiDAR point clouds, the coordinates of the camera feature point cloud are fixed (as long as the camera intrinsics and extrinsics stay the same, which is usually the case after proper calibration). Motivated by this, we precompute the 3D coordinate and the BEV grid index of each point. We also sort all points according to grid indices and record the rank of each point. During inference, we only need to reorder all feature points based on the precomputed ranks. This caching mechanism can reduce the latency of grid association from 17ms to 4ms. Interval Reduction. After grid association, all points within the same BEV grid will be consecutive in the tensor representation. The next step of BEV pooling is then to aggregate the features within each BEV grid by some symmetric function” 3.3 “Though in the same space, LiDAR BEV features and camera BEV features can still be spatially misaligned to some extent due to the inaccurate depth in the view transformer. To this end, we apply a convolution-based BEV encoder (with a few residual blocks) to compensate for such local misalignments.” 3.4 “Multi-Task Heads We apply multiple task-specific heads to the fused BEV feature map. Our method is applicable to most 3D perception tasks. We showcase two examples: 3D object detection and BEV map segmentation.” PNG media_image5.png 416 922 media_image5.png Greyscale Note: As shown by Liu 3.4 and 3.2, the aforementioned BEV representations that are made are BEV feature maps, where the index/key in the map denotes the location of the respective sensor information at that index value. As Liu teaches the two sensor inputs are converted to a map where both maps are fused to create a single fused map where each index/position value maps to both the first and second sensor data Liu teaches the claims synchronized frame/Fused BEV can be a hash map.) Regarding claim 9, Liu teaches: The computer-implemented method of claim 1, wherein the rendering includes first representations of 3D bounding shapes overlaid on one or more first frames corresponding to the first sensor data and second representations of the 3D bounding shapes overlaid on one or more second frames corresponding to the second sensor data. ( PNG media_image3.png 718 1328 media_image3.png Greyscale Note: As seen in the renderings produced by Liu in Fig. 4, renderings for both sensor modalities camera and LiDAR point cloud are shown with bounding shapes placed over detected objects. As both LiDAR and camera captured data have 3D bounding shapes overlaid denoting the results of 3D object detection Liu teaches a first and second representation made from the first and second sensor modalities respectively.) Regarding claim 10, Liu teaches: A system comprising: one or more processors to perform operations including: assembling components of a multi-modal perception pipeline according to configuration data that identifies the components;(Liu Abstract and 3 have been shown to teach a multi-modal perception pipeline according to a configuration data that identifies the components in claim 1.) synchronizing, using a first component of the components, first inference data corresponding to a first sensor modality and second inference data(Liu 3.2 “Camera-to-BEV transformation is non-trivial because the depth associated with each camera feature pixel is inherently ambiguous. Following LSS [39] and BEVDet [20, 19], we explicitly predict the discrete depth distribution of each pixel. We then scatter each feature pixel into D discrete points along the camera ray and rescale the associated features by their corresponding depth probabilities (Figure 3a). This generates a camera feature point cloud of size NHWD, where N is the number of cameras and (H,W) is the camera feature map size. Such 3D feature point cloud is quantized along the x,y axes with a step size of r (e.g., 0.4m). We use the BEV pooling operation to aggregate all features within each r × r BEV grid and flatten the features along the z-axis … The first step of BEV pooling is to associate each point in the camera feature point cloud with a BEV grid. Different from LiDAR point clouds, the coordinates of the camera feature point cloud are fixed (as long as the camera intrinsics and extrinsics stay the same, which is usually the case after proper calibration). Motivated by this, we precompute the 3D coordinate and the BEV grid index of each point. We also sort all points according to grid indices and record the rank of each point. During inference, we only need to reorder all feature points based on the precomputed ranks. This caching mechanism can reduce the latency of grid association from 17ms to 4ms. Interval Reduction. After grid association, all points within the same BEV grid will be consecutive in the tensor representation. The next step of BEV pooling is then to aggregate the features within each BEV grid by some symmetric function” 3.3 “Though in the same space, LiDAR BEV features and camera BEV features can still be spatially misaligned to some extent due to the inaccurate depth in the view transformer. To this end, we apply a convolution-based BEV encoder (with a few residual blocks) to compensate for such local misalignments.” 3.4 “Multi-Task Heads We apply multiple task-specific heads to the fused BEV feature map. Our method is applicable to most 3D perception tasks. We showcase two examples: 3D object detection and BEV map segmentation.” PNG media_image5.png 416 922 media_image5.png Greyscale Note: Liu establishes that its BEV representations are more formally defined as BEV feature maps where an input index/key that also denotes a location is input and maps to the sensor data for that location. The specifications define inference data as ¶126 “For example, the pipeline manager 102 may use the mixer(s) 108 to synchronize first inference data (e.g., generated using the inference environment(s) 106A) corresponding to a first sensor modality (e.g., corresponding to the data source(s) 132) and second inference data (e.g., generated using the inference environment(s) 106B) corresponding to a second sensor modality (e.g., corresponding to the data source(s) 134) into one or more synchronized frames, such as a synchronized frame(s) 340 of FIG. 3” PNG media_image6.png 388 618 media_image6.png Greyscale The means through which the first and second inference data is produced is said to be 106A and B, which as seen from Fig. 1C can output a hash map as a type of inference. As claim 8 has already established that the Liu citations above defining a BEV as a feature map teach the claims hash map, and a hash map is defined as an acceptable definition for a first and second inference data, Liu teaches making a first and second inference data which is used to create the synchronized frame.) corresponding to a second sensor modality into one or more synchronized frames; (Liu 1, 4.2 , 3.1, and- Fig. 2 cited in claim 1, teach synchronized frames made from synchronizing the first and second sensor modalities.) matching, using a second component of the components and the one or more synchronized fames, first data points corresponding to the first inference data with second data points corresponding to the second inference data to generate 3D information corresponding to the first data points fused with the second data points;(Liu 3.1, 3.2, and 3.2 cited claim 1 teach that the first and second sensor data are both made into BEVs. The LiDAR data is flattened along its Z axis to obtain the resulting points and the camera data’s pixels are taught to be made into points of a BEV representation. Once both sensor inputs are in a BEV representation of the same shared space they have first and second data points that can be fused together to produce a fused BEV. As established previously, these BEV representations are feature maps, aka the claims first and second inference data.) and generating, using a third component of the components and the 3D information, a rendering including multiple views associated with the first sensor modality and the second sensor modality.(Liu Fig. 4, cited in claim 1, provides multiple examples of renderings from multiple views from both the camera capture data, the first sensor modality, and the LiDAR data, the second modality.) Regarding claim 11, Liu teaches: The system of claim 10, wherein the first component is derived, at least in part, from a first interface element of a multi-modal sensor fusion framework, the first interface element providing a set of predefined synchronization methods and data structures used to perform the synchronizing. As claim 11 is identical to claims 3/18, other than the preamble referencing claim 10 which has been shown to be rejected, it is rejected under the same rationale. Regarding claim 12, Liu teaches: The system of claim 10, wherein the operations further include computing the first inference data using one or more first inference models of one or more fourth components of the components processing first sensor data, and the second inference data using one or more second inference models of the one or more fourth components processing second sensor data. (Liu 3.2, 3.3, 3.4, cited in the rejection of claim 10 and claim 8, establish that the claims first and second inference data, the same referenced here, can be defined as a hash map as taught by the specifications in Fig. 1C. It was shown in claim 8 that the first and second inference data, the hash maps, are fused into a single hash map. The Liu citations formally define the BEV as a BEV feature map which sensor data to a key/index denoting a location, making it analogous to a hash map. As this inference data is taught to be made for each input modality separately before the inference data is fused Liu teaches a first and second means to produce the inference data, the claims first and second model.) Regarding claim 14, Liu teaches: The system of claim 10, wherein the operations further include converting, using a fourth component of the components, the first inference data into a unified data structure format that is shared with the second inference data, wherein the synchronizing is performed on the first inference data and the second inference data in the unified data structure format. As claim 14 is identical to claim 6, other than the preamble referencing claim 10 which has been shown to be rejected, it is rejected under the same rationale. Regarding claim 15, Liu teaches: The system of claim 10, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; (Liu Abstract “Multi-sensor fusion is essential for an accurate and reliable autonomous driving system. Recent approaches are based on point-level fusion: augmenting the Li DARpoint cloud with camera features … BEVFusion is fundamentally task-agnostic and seamlessly supports different 3D perception tasks with almost no architectural changes. It establishes the new state of the art on nuScenes, achieving 1.3% higher mAP and NDS on 3D object detection and 13.6% higher mIoU on BEVmapsegmentation, with 1.9× lower computation cost.” Note: Liu teaches that its model is intended to replace existing models that perform perception tasks from multi-sensor input like 3D object detection for systems such as autonomous driving systems. Liu specifically lists results from testing on the nuScenes dataset, a known and popular dataset for training autonomous driving vehicle models.)… The remainder of the claim is not listed here as it contains a long list of possible things the system can comprise of where only at least one is needed, which has already been taught. Regarding claim 20, Liu teaches: The at least one processor of claim 16, wherein the at least one processor is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; (Liu Abstract “Multi-sensor fusion is essential for an accurate and reliable autonomous driving system. Recent approaches are based on point-level fusion: augmenting the Li DARpoint cloud with camera features … BEVFusion is fundamentally task-agnostic and seamlessly supports different 3D perception tasks with almost no architectural changes. It establishes the new state of the art on nuScenes, achieving 1.3% higher mAP and NDS on 3D object detection and 13.6% higher mIoU on BEVmapsegmentation, with 1.9× lower computation cost.” Note: Liu teaches that its model is intended to replace existing models that perform perception tasks from multi-sensor input like 3D object detection for systems such as autonomous driving systems. Liu specifically lists results from testing on the nuScenes dataset, a known and popular dataset for training autonomous driving vehicle models.)… The remainder of the claim is not listed here as it contains a long list of possible things the system can comprise of where only at least one is needed, which has already been taught. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 4, 19, are rejected under 35 U.S.C. 103 as being unpatentable over Liu (BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation) in view of Sun (CN 116502230 A). Regarding claims 19, 4, Liu teaches: The at least one processor of claim 16, wherein the second component includes While Liu has been shown to teach a second component Liu does not detail the use of an API to host the model as a means to interact with it. This is taught by Sun which teaches an Application Programming Interface (API) client of an API server that hosts the one or more multi-modal inference models. (Sun ¶53 “The deep learning model of the present invention is a multi-modal input multi-output deep learning model, which can balance the influence of various vulnerability attribute fields on the vulnerability authority, and through this model, information of various vulnerability attribute fields can be input at the same time” ¶92 “Use Keras to design multi-modal input models. Keras can handle multiple inputs and multiple outputs through its function API. First, three inputs of the Keras neural network are defined. The vulnerability description is (128*768), and the cwe data is 3 Dimensions, other inputs are 49 dimensions. The visualization model architecture is shown in Figure 4.” ¶113 “Based on this understanding, the essence of the technical solution of the present invention or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including Several instructions are used to make a computer device (which may be a personal computer, a server, or a network device, etc.)” Note: Sun teaches a multi-modal model that will generate an output, aka inference, that can be interacted with through an API. Sun further teaches that its described software product, the multi-modal model and API, can be stored on a server, thus teaching the ability to interact with a multi-modal inference model hosted on a server with an API.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Liu with Sun where a model with the described second is hosted on a server and can be interacted with via an API. There are several reasons that would motivate one to do so, a common aim when developing machine learning models is to monetize them by hosting the already trained model on fast, high-powered servers to provide the model as a service to users in return for payment. This could be easily implemented via an API which would allow end users to efficiently interface with the model. Claims 5, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Liu (BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation) in view of the BEVfusion github repository (https://github.com/mit-han-lab/bevfusion) hereinafter BEVfusion. Regarding claim 5, Liu teaches: The computer-implemented method of claim 1, wherein the configuration data While Liu has been shown to teach configuration data which defines the structures and steps of the model Liu is a research paper and aims to describe solutions found for current problems in the art their results, not implementation details such as file types, code, etc… The code for Liu is linked to in a github repository in the paper, referred to as BEV fusion, which teaches that the configuration data is a graph-based schema configuration file that identifies the components and interconnection specifications corresponding to two or more components of the components ( PNG media_image7.png 954 812 media_image7.png Greyscale PNG media_image8.png 122 348 media_image8.png Greyscale PNG media_image9.png 964 850 media_image9.png Greyscale Note: The above github repo originates form Liu’s BEVfusion research paper and details the code for the model in the paper. Though linked to in Liu the repo is on a separate site with its own publication info and is thus treated as a separate source while acknowledging the content it discusses is the same model taught by Liu. Two screenshots of the same default.yaml file, a known type of configuration file, are included above. The claims first component as defined in claim 1 accepts the two sensor inputs and makes a fused BEV of them both. The claims second and third components involve inferring 3D information from the fused BEV/synchronized frame and making a rendering from the inferred 3D info. The first and second screenshots show the two sensor modalities being accepted, the camera capture and LiDAR point cloud data along with a BEVsegmentation function teaching that the BEV has been found. The third screenshot shows the loading of the sensor inputs and the defining of an object detection tasks labelling the different classes it can detect such as car, truck bus, etc… The object detection task is the second component, the inferring of 3D data, and the aforementioned accepting of sensor data and creation of a synchronized frame/fused BEV is the first component. Thus, BEVfusion has been shown to teach a configuration file where two or more components are identified and connected. It is known that this .yaml is a graph-based schema configuration file as the specifications state ¶32 “the configuration data may comprise a single graph-based schema configuration file (e.g., a JavaScript Object Notation (JSON) file, a Tom's Obvious, Minimal Language (TOML) file, an initialization (INI) file, YAML, etc.”, clarifying that it is an acceptable file type.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Liu with BEVFusion where a model’s configuration data is stored on a graph based schema configuration file. There are several reasons that would motivate one to do so, configuration files like .yaml files referenced by both BEVFusion and the specifications provide a human readable and straightforward way to define large scale steps of the model that make up the overall pipeline. If one wished for their model’s pipeline to be defined efficiently, succinctly, and in a human readable format in a single source a configuration file provides a means to do so. Regarding claim 13, Liu teaches: The system of claim 10, wherein the configuration data is a graph-based schema configuration file that identifies the components and interconnection specifications corresponding to two or more components of the components. As claim 13 is identical to claim 5, other than the preamble referencing claim 10 which has been shown to be rejected, it is rejected under the same rationale. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Liu (BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation) in view of Ozbilgin (US 11435752 B2). Regarding claim 7, Liu teaches: The computer-implemented method of claim 1, wherein the first component While Liu teaches the generation of a synchronized frame it does not detail policies for dropping or interpolating frames of the first sensor data as part of the synchronization process. This is taught by Ozbilgin which details using one or more policies and a target framerate to generate the one or more synchronized frames, (Ozbilgin Abstract “A sensor data fusion system for a vehicle with multiple sensors includes a first-sensor, a second-sensor, and a controller-circuit. The first-sensor is configured to output a first-frame of data and a subsequent-frame of data indicative of objects present in a first-field-of-view. The first-frame is characterized by a first-time-stamp, the subsequent-frame of data characterized by a subsequent-time-stamp different from the first-time-stamp. The second-sensor is configured to output a second-frame of data indicative of objects present in a second-field-of-view that overlaps the first-field-of-view. The second-frame is characterized by a second-time-stamp temporally located between the first-time-stamp and the subsequent-time-stamp. The controller-circuit is configured to synthesize an interpolated-frame from the first-frame and the subsequent-frame … the controller-circuit fuses the interpolated-frame with the second-frame to provide a fused-frame of data characterized by the interpolated-time-stamp, and operates the host-vehicle in accordance with the fused-frame.” Note: Ozbilgin teaches that frames from multiple sensor inputs are accepted with their timestamps. As the specific time of each frame is known, Ozbilgin implicitly teaches that framerate is known. Ozbilgin uses this info and the frames to fuse them and generate a single frame from multiple frames, thus teaching policies and a target frame rate to generate one or more frames. As the generated frame is the result of data from two different sensor inputs that are fused, it is analogous to the claim synchronized frame.) performs one or more of dropping or interpolating one or more frames corresponding to the first sensor data. (Ozbilgin Abstract, cited above, teaches the process in which a fused frame is produced from a second frame from a second sensor data and a first frame from a first sensor data. Ozbilgin specifically states the frames of the fist sensor data are interpolated to make a single interpolated frame from multiple, this single interpolated frame is then fused with one frame from the second sensor data to make the final synchronized frame.) It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify Liu with Ozbilgin where a second component which generates a synchronized frame from two sensor data inputs, one of which provides frame data, uses frame rate and interpolation or frame dropping methods for frames of the first sensor data. There are several reasons that would motivate one to do so, when working with sensor data from different modalities it is likely that the different sensors have different rates of capturing data. As LiDAR sensors collect millions of data points in a 3D space as opposed to converting visible light into a grid of pixels as a digital camera does LiDAR sensors will often run at a lower rate than modern cameras. To account for this common temporal mismatch to align the sensor data across modalities it would be beneficial to limit the amount of camera frame data via interpolation or frame dropping. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN GREGORY HAKALA whose telephone number is (571)272-7863. The examiner can normally be reached 8:00am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571) 270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALAN GREGORY HAKALA/ Examiner, Art Unit 2617 /KING Y POON/Supervisory Patent Examiner, Art Unit 2617
Read full office action

Prosecution Timeline

Mar 12, 2025
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month