Prosecution Insights
Last updated: October 02, 2026
Application No. 18/608,250

EFFICIENT CLOUD-BASED DYNAMIC MULTI-VEHICLE BEV FEATURE FUSION FOR EXTENDED ROBUST COOPERATIVE PERCEPTION

Final Rejection §102§103§112
Filed
Mar 18, 2024
Examiner
ALFONSO, DENISE G
Art Unit
2662
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
5m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
99 granted / 130 resolved
+14.2% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
10 currently pending
Career history
144
Total Applications
across all art units

Statute-Specific Performance

§101
7.5%
-32.5% vs TC avg
§103
59.8%
+19.8% vs TC avg
§102
19.2%
-20.8% vs TC avg
§112
9.4%
-30.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 130 resolved cases

Office Action

§102 §103 §112
DETAILED ACTIONS Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (“IDS”) filed on 05/22/2026 was reviewed and the listed references were noted. Status of Claims Claims 1-30 remain pending in the application. Response to Amendment The amendment filed 05/22/2026 has been entered in full. Claims 1-30 remain pending in the application. Applicant’s amendment to the Claims have overcome each and every 112(b) rejections previously set forth in the Non-Final Office Action mailed February 25th, 2026. Response to Arguments Applicant’s arguments with respect to claims 1, 27, and 30 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 3, 5-6, 12, 16, and 24 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Xu et al., " V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception", (2023), hereinafter referred to as Xu. Claim 1 Xu discloses a system for processing data from a plurality of vehicles (Xu, Fig. 7), the system comprising: one or more memories (Xu, Section 5.1, “GPU (RTX3090)”) for storing grid-free vehicle data (Xu, Section 3.1, “We collect the V2V4Real via two experimental connected automated vehicles including a Tesla vehicle (Fig. 2a) and a Ford Fusion vehicle (Fig. 2b) retrofitted by Transportation Research Center(TRC) company and AutonomouStuff (AStuff) Company respectively. Both vehicles are equipped with a Velodyne VLP-32 LiDAR sensor”, “We sample the frames at 10Hz, resulting in a total of 20K frames of LiDAR point cloud and 40K frames of RGB images”, LiDAR point clouds from both vehicles are considered grid-free vehicle data) from each of the plurality of vehicles (Xu, Fig. 2, Tesla vehicle and Ford Fusion vehicle), the grid-free vehicle data being grid-free not including corresponding two-dimension grid data associated with each of the plurality of vehicles (Xu, Section 3.1, “We sample the frames at 10Hz, resulting in a total of 20K frames of Li DAR point cloud and 40K frames of RGB images”, point cloud data does not include two-dimension grid data); and one or more processors in communication with the one or more memories (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”), the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”) configured to: determine one or more first features from first grid-free vehicle data of the grid-free vehicle data (Xu, Section 3.1, “We employ SusTech Point , a powerful opensource labeling tool, to annotate 3D bounding boxes for the collected LiDAR data. We hire two groups of professional annotators. One group is responsible for the initial labeling, and the other further re fines the annotations. There are five object classes in total, including cars, vans, pickup trucks, semi-truck, and buses. For each object, we annotate its 7-degree-of-freedom 3D bounding box containing x,y,z for the centroid position and l,w,h,yaw for the bounding box extent and yaw angles.”, “Since the bounding boxes are annotated separately for the two collection vehicles, an object in the Tesla’s frame could have the same id as a different object in Ford Fusion’s frame. To avoid such issues, all the object ids in Tesla are labeled between 0 − 1000, while ids in Ford Fusion range from 1001 − 2000), the first grid-free vehicle data being from a first vehicle of the plurality of vehicles (Xu, Fig. 2, Tesla Vehicle); determine one or more second features from second vehicle data of the vehicle data (Xu, Section 3.1, “We employ SusTech Point , a powerful opensource labeling tool, to annotate 3D bounding boxes for the collected LiDAR data. We hire two groups of professional annotators. One group is responsible for the initial labeling, and the other further re fines the annotations. There are five object classes in total, including cars, vans, pickup trucks, semi-truck, and buses. For each object, we annotate its 7-degree-of-freedom 3D bounding box containing x,y,z for the centroid position and l,w,h,yaw for the bounding box extent and yaw angles.”, “Since the bounding boxes are annotated separately for the two collection vehicles, an object in the Tesla’s frame could have the same id as a different object in Ford Fusion’s frame. To avoid such issues, all the object ids in Tesla are labeled between 0 − 1000, while ids in Ford Fusion range from 1001 − 2000), the second vehicle data being from a second vehicle of the plurality of vehicles (Xu, Fig. 2, Ford Fusion vehicle); fuse the one or more first features and the one or more second features to generate fused features (Xu, Section 3.1, “The HD map generation pipeline refers to generating a global point cloud map and vector map. To generate the point cloud map, we fuse a sequence of point cloud frames together. More specifically, we first preprocess each LiDAR frame by removing the dynamic objects while keeping the static elements. Then, a Normal Transformation Distribution scan matching algorithm is ap plied to compute the relative transformation between two consecutive LiDAR frames. The LiDAR odometry can then be constructed by taking the transformation. However, the noise imbued in the LiDAR data can lead to accumulated errors in the estimated transformation matrix as the frame index increases.”, Section 4.1, “Intermediate Fusion: The collaborators will first project their LiDAR to the ego vehicle’s coordinate system and then extract intermediate features using a neural feature extractor. Afterward, the encoded features are compressed and broadcasted to the ego vehicle for cooperative feature fusion.”); and generate a bird’s-eye-view (BEV) representation based on the fused features (Xu, Fig. 1, the aggregated LiDAR data, HD map, and Sattelite map are bird’s eye view representation based on the fused annotation or extracted features as shown in Fig. 1). PNG media_image1.png 320 340 media_image1.png Greyscale Claim 3 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein at least one processor of the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”) that is configured to generate the BEV representation (Xu, Fig. 1, the aggregated LiDAR data, HD map, and Sattelite map are bird’s eye view representation based on the fused annotation or extracted features as shown in Fig. 1) is located in the first vehicle or the second vehicle (Xu, Section 4.1, “The collaborators will first project their LiDAR to the ego vehicle’s coordinate system and then extract intermediate features using a neural feature extractor. Afterward, the encoded features are compressed and broadcasted to the ego vehicle for cooperative feature fusion.”). Claim 5 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”) are further configured to at least one of receive a first indication of the one or more first features or receive a second indication of the one or more second features (Xu, Section 4.1, “Each vehicle detects 3D objects utilizing its own sensor observations and delivers the pre dictions to others. Then the receiver applies Non maximum suppression to produce the final outputs.”, “The collaborators will first project their LiDAR to the ego vehicle’s coordinate system and then extract intermediate features using a neural feature extractor. Afterward, the encoded features are compressed and broadcasted to the ego vehicle for cooperative feature fusion.”). Claim 6 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”) are further configured to at least one of send a first indication of the one or more first features or receive a second indication of the one or more second features (Xu, Section 4.1, “Each vehicle detects 3D objects utilizing its own sensor observations and delivers the pre dictions to others. Then the receiver applies Non maximum suppression to produce the final outputs.”, “The collaborators will first project their LiDAR to the ego vehicle’s coordinate system and then extract intermediate features using a neural feature extractor. Afterward, the encoded features are compressed and broadcasted to the ego vehicle for cooperative feature fusion.”). Claim 12 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein the BEV representation is a unified BEV representation for the plurality of vehicles (Xu, Fig. 1, aggregated LiDAR data). Claim 16 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein a first feature of the one or more first features (Xu, Section 3.1, “Since the bounding boxes are annotated separately for the two collection vehicles, an object in the Tesla’s frame could have the same id as a different object in Ford Fusion’s frame. To avoid such issues, all the object ids in Tesla are labeled between 0 − 1000, while ids in Ford Fusion range from 1001 − 2000. Moreover, identical objects could have different ids in the annotation files of the two collection vehicles. To solve this issue, we transform the objects from different coordinates to a unified coordinate system and calculate the BEV IoU between all objects. For the objects that have IoU larger than a certain threshold, we assign them the same object id and unify their bounding box sizes.”) corresponds to a second feature of the one or more second features (Xu, Section 3.1, “Since the bounding boxes are annotated separately for the two collection vehicles, an object in the Tesla’s frame could have the same id as a different object in Ford Fusion’s frame. To avoid such issues, all the object ids in Tesla are labeled between 0 − 1000, while ids in Ford Fusion range from 1001 − 2000. Moreover, identical objects could have different ids in the annotation files of the two collection vehicles. To solve this issue, we transform the objects from different coordinates to a unified coordinate system and calculate the BEV IoU between all objects. For the objects that have IoU larger than a certain threshold, we assign them the same object id and unify their bounding box sizes.”), and wherein the first feature is represented with a different level of distortion than the second feature (Xu, Section 3.3, “The number of Vans and Semi-Trucks are similar, while Bus has the least quantities. Fig. 6 shows the LiDAR points density distribution inside different objects bounding boxes and the bounding boxes’ size distribution. As we may see in the left figure, when there is only one vehicle (Tesla) scanning the environment, the number of LiDAR points within bounding boxes drops dramatically as the radial distance increases. Enhanced by the shared visual information from the other vehicle (Ford Fusion), the LiDAR point density of each object increases significantly and still retains at a high level even when the distance reaches 100 m. This validates the great benefits that cooperative perception can bring to the system.”). Claim 24 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein the system comprises an intermediate collaboration system (Xu, Fig. 7, intermediate fusion), wherein the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”) are further configured to: receive, from the first vehicle (Xu, Fig. 2, Tesla vehicle), an indication of the one or more first features (Xu, Section 4.1, “The collaborators will first project their LiDAR to the ego vehicle’s coordinate system and then extract intermediate features using a neural feature extractor. Afterward, the encoded features are compressed and broadcasted to the ego vehicle for cooperative feature fusion.”); and receive, from the second vehicle (Xu, Fig. 2, Ford Fusion vehicle), an indication of the one or more second features (Xu, Section 4.1, “The collaborators will first project their LiDAR to the ego vehicle’s coordinate system and then extract intermediate features using a neural feature extractor. Afterward, the encoded features are compressed and broadcasted to the ego vehicle for cooperative feature fusion.”). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2, 4, 9, 15, 17-20, 22, 26-27, and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Chang et al., "BEV-V2X: Cooperative Birds-Eye-View Fusion and Grid Occupancy Prediction via V2X-Based Data Sharing", (2023), hereinafter referred to as Chang. Claim 2 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”). Xu does not explicitly disclose wherein the one or more processors are further configured to send the BEV representation to at least one vehicle of the plurality of vehicles. However, Chang teaches wherein the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) are further configured to send the BEV representation to at least one vehicle of the plurality of vehicles (Chang, Section I, “By extracting the single vehicle BEV data in the historical time horizons, we can integrate the perception information of different CAVs and predict the global BEV occupancy grid map in the future time horizons. This article focuses on BEV fusion and prediction. The fusion and prediction results can help achieve accurate environment perception and strengthen the understanding of the global scenario. Based on the results, the system can provide real-time driving risk warning, formulate the corresponding planning scheme, and send the messages to vehicles in the control area.”, Section 1, “Each connected and automated vehicle (CAV) regularly reports its own information to other vehicles or roadside units. By aggregating and fusing the data information from different CAVs, we can get a more accurate understanding of the global scenario”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein the one or more processors are further configured to send the BEV representation to at least one vehicle of the plurality of vehicles. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to can real-time driving risk warning (Chang, Section I). Claim 4 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein at least one processor of the one or more processors that is configured to generate the BEV representation is located outside of the first vehicle and the second vehicle. However, Chang teaches wherein at least one processor of the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) that is configured to generate the BEV representation is located outside of the first vehicle and the second vehicle (Chang, Fig. 1, roadside unit, Section 1, “The roadside unit collects the local BEV of all CAVs in the control area, and periodically extracts the historical data.”, “In fact, there are basically two modes of vehicle-to-vehicle (V2V) communication and vehicle-to-infrastructure (V2I) communication to achieve BEV fusion via V2X technique. We recommend that the corresponding models be deployed at the roadside or cloud center.”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein the one or more processors are further configured to send the BEV representation to at least one vehicle of the plurality of vehicles. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to achieve higher accuracy (Chang, Abstract). Claim 9 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein the vehicle data comprises at least one of vehicle pose, vehicle location, or vehicle trajectory, and wherein the one or more processors are further configured to determine to not fuse a third one or more features from third vehicle data, the third vehicle data being from a third vehicle of the plurality of vehicles as part of generating the BEV representation based on the at least one of vehicle pose, vehicle location, or vehicle trajectory. However, Chang teaches wherein the vehicle data comprises at least one of vehicle pose, vehicle location, or vehicle trajectory (Chang, Fig. 2, the BEV data shows the location and pose of the vehicle), and wherein the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) are further configured to determine to not fuse a third one or more features from third vehicle data (Chang, Fig. 1, the roadside unit only fuses the data from the vehicles, ABD, and not CE), the third vehicle data being from a third vehicle of the plurality of vehicles (Chang, Fig.1, Vehicles C or E, Section I, “After raw perception data are aggregated to the BEV, the corresponding grid position and associated confidence of vehicle C in the local BEV of A and B are also different. It is a critical issue to collect and fuse the local BEV of different vehicles to obtain global BEV with higher reliability and more comprehensive scenario understanding” ) as part of generating the BEV representation based on the at least one of vehicle pose, vehicle location, or vehicle trajectory (Chang, Section A, “The roadside unit collects the local BEV of all CAVs in the control area, and periodically extracts the historical data. Using the data, the system calls the deep learning model to fuse the information from different vehicles, and predict the future cooperative BEV occupancy grid map.”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein the vehicle data comprises at least one of vehicle pose, vehicle location, or vehicle trajectory, and wherein the one or more processors are further configured to determine to not fuse a third one or more features from third vehicle data, the third vehicle data being from a third vehicle of the plurality of vehicles as part of generating the BEV representation based on the at least one of vehicle pose, vehicle location, or vehicle trajectory. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to achieve higher accuracy (Chang, Abstract). Claim 15 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein the one or more are configured to obtain vehicle data generated over time, wherein the vehicle data comprises information indicative of at least one of a change in pose of the first vehicle or a change in pose of the second vehicle, and wherein the change in pose of the first vehicle or the change in pose of the second vehicle comprise a change in at least one of rotation or translation. However, Chang teaches wherein the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) are configured to obtain vehicle data generated over time (Chang, Section III.A, “CAV sends its real-time SBEV data package to the roadside unit in a timer-trigger style”), wherein the vehicle data comprises information indicative of at least one of a change in pose of the first vehicle or a change in pose of the second vehicle (Chang, Section III.a, “We apply the symbol Pc(x,y) to denote the probability that the BEV position (x,y) is occupied by category c.[Pc(x,y)]C× H× W is the occupancy probability matrix of C elements in the grid network with the size of H × W.”), and wherein the change in pose of the first vehicle or the change in pose of the second vehicle comprise a change in at least one of rotation or translation (Chang, Section IV, “The simulation experiment of BEV fusion and prediction requires naturalistic driving scenario data, which contain the movement information of traffic participants in a certain spatial area and a continuous time range, as well as the environment in formation”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein the one or more are configured to obtain vehicle data generated over time, wherein the vehicle data comprises information indicative of at least one of a change in pose of the first vehicle or a change in pose of the second vehicle, and wherein the change in pose of the first vehicle or the change in pose of the second vehicle comprise a change in at least one of rotation or translation. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to achieve higher accuracy (Chang, Abstract). Claim 17 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein a first feature of the one or more first features is representative of a dynamic object. However, Chang teaches wherein a first feature of the one or more first features is representative of a dynamic object (Chang, Section III.A, “we divide the scenario elements into different categories, i.e., dynamic traffic participants such as vehicles and pedestrians, and static road environment information such as drivable areas, lanes, traffic infrastructures, channelization”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein a first feature of the one or more first features is representative of a dynamic object. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to achieve higher accuracy (Chang, Abstract). Claim 18 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein at least a portion of the system resides in a cloud-computing environment However, Chang teaches wherein at least a portion of the system resides in a cloud-computing environment (Chang, Section V, “The third is the method adopted by our model, which applies roadside unit or cloud center to collect all CAVs’ information in the control area in a unified and centralized manner for global fusion and prediction.”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein at least a portion of the system resides in a cloud-computing environment. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to make the system more robust. Claim 19 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein the one or more processors are further configured to transmit BEV feature configuration information to the first vehicle and the second vehicle for configuring BEV processing of the first vehicle and the second vehicle, wherein the BEV feature configuration information comprises at least one of BEV feature vector size, a model index for a model to transform two-dimensional camera images to BEV feature vectors, or the model. However, Chang teaches wherein the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) are further configured to transmit BEV feature configuration information (Chang, Section III.A, “The model outputs the BEV occupancy grid estimate for the whole control area in future time horizon F. The output [Pc(x,y)]C× HO× WO is the occupancy probability of C elements in the global grid network with the size of HO×WO. Further, the occupancy state representation and visual image of the global CBEV are generated by the transformation rules in Formula (1) and (2), which constitute the final fusion and prediction information.”) to the first vehicle and the second vehicle for configuring BEV processing of the first vehicle and the second vehicle (Chang, Section I, “By extracting the single vehicle BEV data in the historical time horizons, we can integrate the perception information of different CAVs and predict the global BEV occupancy grid map in the future time horizons. This article focuses on BEV fusion and prediction. The fusion and prediction results can help achieve accurate environment perception and strengthen the understanding of the global scenario. Based on the results, the system can provide real-time driving risk warning, formulate the corresponding planning scheme, and send the messages to vehicles in the control area.”, Section 1, “Each connected and automated vehicle (CAV) regularly reports its own information to other vehicles or roadside units. By aggregating and fusing the data information from different CAVs, we can get a more accurate understanding of the global scenario”), wherein the BEV feature configuration information comprises at least one of BEV feature vector size, a model index for a model to transform two-dimensional camera images to BEV feature vectors, or the model (Chang, Section III.A, “The model outputs the BEV occupancy grid estimate for the whole control area in future time horizon F. The output [Pc(x,y)]C× HO× WO is the occupancy probability of C elements in the global grid network with the size of HO×WO. Further, the occupancy state representation and visual image of the global CBEV are generated by the transformation rules in Formula (1) and (2), which constitute the final fusion and prediction information.”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein the one or more processors are further configured to transmit BEV feature configuration information to the first vehicle and the second vehicle for configuring BEV processing of the first vehicle and the second vehicle, wherein the BEV feature configuration information comprises at least one of BEV feature vector size, a model index for a model to transform two-dimensional camera images to BEV feature vectors, or the model.. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to achieve higher accuracy (Chang, Abstract). Claim 20 The combination of Xu in view of Chang discloses the system of claim 19 (Xu, Fig. 7), wherein the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) are further configured to receive, from the first vehicle (Chang, Section III.A, “CAV sends its real-time SBEV data package to the roadside unit in a timer-trigger style”), a BEV feature vector in accordance with the BEV feature vector size, wherein the BEV feature vector comprises a raw BEV feature vector or a compressed BEV feature vector (Chang, Section III.A, “With the help of various sensors, such as cameras and Lidar, the single vehicle perceives the surrounding environment. Then, the vehicle system converts the raw sensory data, such as images and point clouds into BEV space, and generates the local BEV centered on its own coordinates. BEV is a semantically composite data structure, which uses matrices to represent the occupancy of scenario elements within a certain spatial area. Each matrix element corresponds to the occupancy probability or state of each grid in the driving environment, which can be further summarized and displayed as RGB image.”). The proposed combination as well as the motivation for combining the Xu and Chang references presented in the rejection of Claim 19, apply to Claim 20 and are incorporated herein by reference. Thus, the system recited in Claim 20 is met by Xu and Chang. Claim 22 The combination of Xu in view of Chang discloses the system of claim 20 (Xu, Fig. 7), wherein the compressed BEV feature vector comprises soft BEV features or hard BEV features (Chang, Section III.A, “With the help of various sensors, such as cameras and Lidar, the single vehicle perceives the surrounding environment. Then, the vehicle system converts the raw sensory data, such as images and point clouds into BEV space, and generates the local BEV centered on its own coordinates. BEV is a semantically composite data structure, which uses matrices to represent the occupancy of scenario elements within a certain spatial area. Each matrix element corresponds to the occupancy probability or state of each grid in the driving environment, which can be further summarized and displayed as RGB image.”), wherein the soft BEV features comprise probabilities or likelihoods of object presence or object attributes, and wherein hard BEV features comprise binary or categorical representations of object presence or attributes (Chang, Section III.A, “At each grid location of the BEV, the occupying objects may include both vehicles and road elements, and they are not in conflict with each other. Therefore, we divide the scenario elements into different categories, i.e., dynamic traffic participants such as vehicles and pedestrians, and static road environment information such as drivable areas, lanes, traffic infrastructures, channelization, etc. We apply the symbol Pc(x,y) to denote the probability that the BEV position(x,y) is occupied by category c.[Pc(x,y)]C× H× W is the occupancy probability matrix of C elements in the grid network with the size of H × W.”). The proposed combination as well as the motivation for combining the Xu and Chang references presented in the rejection of Claim 19, apply to Claim 20 and are incorporated herein by reference. Thus, the system recited in Claim 20 is met by Xu and Chang. Claim 26 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein the one or more processors are further configured to, prior to or as part of generating the BEV representation, align grids associated with the first vehicle and the second vehicle. However, Chang teaches wherein the one or more processors (Chang, Section 4.A, “the CPU of the machine is Intel 10900X, and the GPU is RTX 3090. Our operating system is Ubuntu 18.04LTS with 128GB RAM”) are further configured to, prior to or as part of generating the BEV representation, align grids associated with the first vehicle and the second vehicle (Chang, Section III.B, “Since the map information is pre-stored on the roadside, the system only needs to take the predicted grid occupancy state of vehicles and pedestrians, and then concatenate the tensors with the pre-stored standard map occupancy state as the final results.”). Xu and Chang are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Chang wherein the one or more processors are further configured to, prior to or as part of generating the BEV representation, align grids associated with the first vehicle and the second vehicle. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to achieve higher accuracy (Chang, Abstract). Claim 27 is rejected for similar reasons as those described in claims 1-2. The additional elements in Claim 27 (the combination of Xu in view of Chang) discloses includes: a method for processing data from a plurality of vehicles (Xu, Fig. 7). The proposed combination as well as the motivation for combining the Xu and Chang references presented in the rejection of Claim 2, apply to Claim 27 and are incorporated herein by reference. Thus, the method recited in Claim 27 is met by Xu and Chang. Claim 30 is rejected for similar reasons as those described in claims 1-2. The additional elements in Claim 30 (the combination of Xu in view of Chang) discloses includes: (the combination of Xu in view of Chang) discloses includes: a method for processing data from a plurality of vehicles (Xu, Fig. 7). The proposed combination as well as the motivation for combining the Xu and Chang references presented in the rejection of Claim 2, apply to Claim 30 and are incorporated herein by reference. Thus, the method recited in Claim 30 is met by Xu and Chang. Claims 7, 25, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Li et al., " Learning for Vehicle-to-Vehicle Cooperative Perception under Lossy Communication", (2023), hereinafter referred to as Li. Claim 7 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein the first vehicle data comprises first grid-free kernels, the first grid-free kernels not including grid information associated with the first vehicle, wherein the second vehicle data comprises second grid-free kernels, the second grid-free kernels not including grid information associated with the second, and wherein as part of fusing the one or more first features and the one or more second features, the one or more processors are configured to fuse the first grid-free kernels and the second grid-free kernels. However, Li teaches wherein the first vehicle data (Li, Fig. 2, CAV 1) comprises first grid-free kernels, the first grid-free kernels not including grid information associated with the first vehicle (Li, Fig. 2, projected LiDAR, point cloud is a grid-free data, Section III.B, “The framework of the LC-aware repair network is shown in Fig. 3, which is an encoder-decoder architecture with skip connections. This network generates a specific per-tensor filter kernel to jointly align and recover the input damaged feature to produce a recovered version of the output feature. The input feature for LC-aware repair network is S ∈ Rc×h×w, then a tensor-wise kernel K is generated and applied to S to produce the recovered output feature ˆ S ∈Rc×h×w.”), wherein the second vehicle data (Li, Fig. 2, CAV 2) comprises second grid-free kernels, the second grid-free kernels not including grid information associated with the second vehicle (Li, Fig. 2, projected LiDAR, point cloud is a grid-free data, Section III.B, “The framework of the LC-aware repair network is shown in Fig. 3, which is an encoder-decoder architecture with skip connections. This network generates a specific per-tensor filter kernel to jointly align and recover the input damaged feature to produce a recovered version of the output feature. The input feature for LC-aware repair network is S ∈ Rc×h×w, then a tensor-wise kernel K is generated and applied to S to produce the recovered output feature ˆ S ∈Rc×h×w.”), and wherein as part of fusing the one or more first features and the one or more second features (Li, Fig. 2, Section III.A, “The intermediate features aggregated from other surrounding CAVs are fed into the major component of our framework i.e., LC-Aware Repair Network for recovering the intermediate feature map in lossy communication by using tensor-wise filtering, and V2V Attention module for iterative inter-vehicle as well as intra-vehicle feature fusion utilizing attention mechanisms.”), the one or more processors are configured to fuse the first grid-free kernels and the second grid-free kernels (Li, Fig. 2, Fig. 3, “The LC-aware Repair architecture for feature recovery is based on the encoder-decoder structure, which outputs per-tensor feature kernels. These kernels then are applied to the input lossy features.”);. Xu and Li are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Li wherein the first vehicle data comprises first grid-free kernels, the first grid-free kernels not including grid information associated with the first vehicle, wherein the second vehicle data comprises second grid-free kernels, the second grid-free kernels not including grid information associated with the second, and wherein as part of fusing the one or more first features and the one or more second features, the one or more processors are configured to fuse the first grid-free kernels and the second grid-free kernels. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to mitigate the negative impact of LC and enhance the interaction between the ego vehicle and other vehicle (Li, Abstract). Claim 25 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein as part of generating the BEV representation the one or more processors are configured to discretize grid-free kernels for at least one vehicle of the plurality of vehicles. However, Li teaches wherein as part of generating the BEV representation (Li, Fig. 2, output, Section I, “to generate bird’s-eye-view representations for 3D object detection”), the one or more processors (Li, Fig. 2) are configured to discretize grid-free kernels for at least one vehicle of the plurality of vehicles (Li, Section III.B, “This network generates a specific per-tensor filter kernel to jointly align and recover the input damaged feature to produce a recovered version of the output feature. The input feature for LC-aware repair network is S ∈ Rc×h×w, then a tensor-wise kernel K is generated and applied to S to produce the recovered output feature ˆ S ∈Rc×h×w”, “To acquire the repaired output feature ˆS, the tensor-wise filtering of the input damaged feature could largely preserve the feature detailed without corruption. Therefore, a large kernel size k is desired to leverage the rich neighborhood information of each tensor fully. In our experiment, the kernel size k is set to 5 due to memory limitations”). Xu and Li are both considered to be analogous to the claimed invention because they are in the same field of vehicle-to-vehicle collaboration perception. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Li wherein as part of generating the BEV representation the one or more processors are configured to discretize grid-free kernels for at least one vehicle of the plurality of vehicles. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to mitigate the negative impact of LC and enhance the interaction between the ego vehicle and other vehicle (Li, Abstract). Claim 28 is rejected for similar reasons as those described in claim 7. The additional elements in Claim 28 (the combination of Xu in view of Li) discloses includes: a method for processing data from a plurality of vehicles (Xu, Fig. 7). The proposed combination as well as the motivation for combining the Xu and Li references presented in the rejection of Claim 7, apply to Claim 28 and are incorporated herein by reference. Thus, the method recited in Claim 28 is met by Xu and Li. Claims 8 and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Li in further view of Tang et al., " Prototypical Variational Autoencoder for Few-shot 3D Point Cloud Object Detection", (2023), hereinafter referred to as Tang. Claim 8 The combination of Xu in view of Li discloses the system of claim 7 (Xu, Fig. 7). The combination of Xu in view of Li does not explicitly disclose wherein the first grid-free kernels and the second grid-free kernels comprise variational autoencoder-Gaussian mixture model (VAE-GMM) kernels. However, Tang teaches wherein the first grid-free kernels and the second grid-free kernels (Tang teaches point cloud as the input and using kernels for the neural network, Li teaches the first and second grid-free kernels) comprise variational autoencoder-Gaussian mixture model (VAE-GMM) kernels (Tang, page 2, “we propose a variational autoencoder approach particularly designed for prototype learning, named Prototypical VAE (abbr. P-VAE). As Fig. 1(a) illustrates, instead of directly learning the features, our P-VAE learns the distribution parameters, by which we can construct a Gaussian Mixture Model(GMM)-based posterior for sampling features from the probabilistic latent space.”, Fig. 2, input point cloud is processed using the VAE-GMM). Xu, Li, and Tang are all considered to be analogous to the claimed invention because they are in the same field of object detection. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Tang wherein the first grid-free kernels and the second grid-free kernels comprise variational autoencoder-Gaussian mixture model (VAE-GMM) kernels. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to gather richer information to generate good-quality representative prototypes for the novel classes (Tang, page 2). Claim 29 is rejected for similar reasons as those described in claim 8. The additional elements in Claim 29 (the combination of Xu in view of Li in view of Tang) discloses includes: a method for processing data from a plurality of vehicles (Xu, Fig. 7). The proposed combination as well as the motivation for combining the Xu, Li, and Tang references presented in the rejection of Claim 8, apply to Claim 29 and are incorporated herein by reference. Thus, the method recited in Claim 29 is met by Xu, Li, and Tang. Claims 11 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Ren et al., "Collaborative Perception for Autonomous Driving: Current Status and Future Trend", (2022), hereinafter referred to as Ren. Claim 11 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein at least one processor of the one or more processors (Xu, Section 5.1, “All the detection models employ PointPillar [18] as the backbone to extract 2D features from the point cloud. we train all models with 60 epochs, a batch size of 4 per GPU (RTX3090), a learning rate of 0.001 , and we decay the learning rate with a cosine annealing”). Xu does not explicitly disclose to generate a mask based on overlapping fields of view of at least one sensor system of the first vehicle and at least one sensor system of the second vehicle; and apply the mask to a plurality of first features as part of determining the one or more first features and apply the mask to a plurality of second feature as part of determining the one or more second features. However, Ren teaches to generate a mask based on overlapping fields of view of at least one sensor system of the first vehicle and at least one sensor system of the second vehicle; and apply the mask to a plurality of first features as part of determining the one or more first features and apply the mask to a plurality of second feature as part of determining the one or more second features (Ren, Fig. 4, collaborative perception mask). Xu and Ren are both considered to be analogous to the claimed invention because they are in the same field of multiple vehicle data sharing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Ren to generate a mask based on overlapping fields of view of at least one sensor system of the first vehicle and at least one sensor system of the second vehicle; and apply the mask to a plurality of first features as part of determining the one or more first features and apply the mask to a plurality of second feature as part of determining the one or more second features. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to improve the accuracy of environmental perception as well as robustness and safety of transportation systems (Ren, Section 1). Claim 14 Xu discloses the system of claim 1 (Xu, Fig. 7), wherein the first vehicle data is based on first sensor data from a first plurality of sensor systems of the first vehicle data (Xu, Section 3.1, “We employ SusTech Point , a powerful opensource labeling tool, to annotate 3D bounding boxes for the collected LiDAR data. We hire two groups of professional annotators. One group is responsible for the initial labeling, and the other further re fines the annotations. There are five object classes in total, including cars, vans, pickup trucks, semi-truck, and buses. For each object, we annotate its 7-degree-of-freedom 3D bounding box containing x,y,z for the centroid position and l,w,h,yaw for the bounding box extent and yaw angles.”, “Since the bounding boxes are annotated separately for the two collection vehicles, an object in the Tesla’s frame could have the same id as a different object in Ford Fusion’s frame. To avoid such issues, all the object ids in Tesla are labeled between 0 − 1000, while ids in Ford Fusion range from 1001 − 2000) and the second vehicle data is based on second sensor data from a second plurality of sensor systems of the second vehicle data (Xu, Section 3.1, “We employ SusTech Point , a powerful opensource labeling tool, to annotate 3D bounding boxes for the collected LiDAR data. We hire two groups of professional annotators. One group is responsible for the initial labeling, and the other further re fines the annotations. There are five object classes in total, including cars, vans, pickup trucks, semi-truck, and buses. For each object, we annotate its 7-degree-of-freedom 3D bounding box containing x,y,z for the centroid position and l,w,h,yaw for the bounding box extent and yaw angles.”, “Since the bounding boxes are annotated separately for the two collection vehicles, an object in the Tesla’s frame could have the same id as a different object in Ford Fusion’s frame. To avoid such issues, all the object ids in Tesla are labeled between 0 − 1000, while ids in Ford Fusion range from 1001 − 2000) Xu does not explicitly disclose wherein at least one sensor system of the first plurality of sensor systems is of a different type than each sensor system of the second plurality of sensor systems. However, Ren teaches wherein at least one sensor system of the first plurality of sensor systems is of a different type than each sensor system of the second plurality of sensor systems (Ren, Section 4.2, “Collaborative semantic segmentation of 3D scenes targets to produce semantic segmentation masks for each agent given observations (images, LIDAR point clouds, etc.) of 3D scenes from several agents”). Xu and Ren are both considered to be analogous to the claimed invention because they are in the same field of multiple vehicle data sharing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Ren wherein at least one sensor system of the first plurality of sensor systems is of a different type than each sensor system of the second plurality of sensor systems. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to improve the accuracy of environmental perception as well as robustness and safety of transportation systems (Ren, Section 1). Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Cheng et al., (US 2022/0036098 A1), hereinafter referred to as Cheng. Claim 13 Xu discloses the system of claim 1 (Xu, Fig. 7), Xu does not explicitly disclose wherein the first vehicle data is based on first sensor data from a first plurality of sensor systems of the first vehicle and the second vehicle data is based on second sensor data from a second plurality of sensor systems of the second vehicle, wherein the first sensor data and the second sensor data have a different resolution. However, Cheng teaches wherein the first vehicle data (Cheng, Fig. 6, vehicle 602A) is based on first sensor data from a first plurality of sensor systems of the first vehicle (Cheng, [0073], “ The vehicle 602A can use one or more sensors 240 (such as one or more cameras 246) to acquire first environment data of at least a portion 608A of the external environment of the vehicle 602A”) and the second vehicle data (Cheng, Fig. 6, vehicle 602B) is based on second sensor data from a second plurality of sensor systems of the second vehicle (Cheng, [0073], “ The vehicle 602B can use one or more sensors 240 (such as one or more cameras 246) to acquire second environment data of at least a portion 608B of the external environment of the vehicle 602B.”), wherein the first sensor data and the second sensor data have a different resolution (Cheng, [0076], The vehicle 602A can identify the first environment data that is located in the common region 610, such the first environment data that includes the person 606. The vehicle 602A can reduce the resolution of the first portion of the first environment data that is located in the common region 610. As an example, the vehicle 602A can downsample the first portion using a Canny edge detector. In some instances, the vehicle 602A can also compress the first portion using a neural network encoder/decoder and/or any suitable compression method.”). Xu and Cheng are both considered to be analogous to the claimed invention because they are in the same field of multiple vehicle data sharing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Chang to incorporate the teachings of Cheng wherein the first vehicle data is based on first sensor data from a first plurality of sensor systems of the first vehicle and the second vehicle data is based on second sensor data from a second plurality of sensor systems of the second vehicle, wherein the first sensor data and the second sensor data have a different resolution. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to reduce redundant sensor data being transferred from the ego vehicle to the other vehicle (Cheng, [0016]). Claims 10, 21, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Chang in further view of Tu et al., "CoBEVT: Cooperative Bird’s Eye View Semantic Segmentation with Sparse Transformers", hereinafter referred to as Tu. Claim 10 The combination of Xu in view of Chang discloses the system of claim 1 (Xu, Fig. 7). The combination of Xu in view of Chang does not explicitly disclose wherein as part of determining to not fuse the third one or more features, the one or more processors are configured to determine that the third one or more features based on a determination that the third vehicle will be outside a neighborhood in less than or less than or equal to a predetermined threshold amount of time, the neighborhood comprising a geographical area including the plurality of vehicles. However, Tu teaches wherein as part of determining to not fuse the third one or more features, the one or more processors are configured to determine that the third one or more features based on a determination that the third vehicle will be outside a neighborhood in less than or less than or equal to a predetermined threshold amount of time, the neighborhood comprising a geographical area including the plurality of vehicles (Tu, Section 4.2, “We assume all the AVs have a 70m communication range following , and all the vehicles out of this broadcasting radius of ego vehicle will not have any collaboration. For the OPV2V camera-track, we choose ResNet34 [52] as the image feature extractor in SinBEVT. The transmitted BEV intermediate representation has a resolution of 32 × 32 × 128. For the multi agent fusion, our FuseBEVT component has 3 encoded layers and a window size of 8 for both local and global attention”). Xu, Chang, and Tu are all considered to be analogous to the claimed invention because they are in the same field of multiple vehicle data sharing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu and Chang to incorporate the teachings of Tu wherein as part of determining to not fuse the third one or more features, the one or more processors are configured to determine that the third one or more features based on a determination that the third vehicle will be outside a neighborhood in less than or less than or equal to a predetermined threshold amount of time, the neighborhood comprising a geographical area including the plurality of vehicles. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to retrieve better accuracy (Tu, Section 2.2). Claim 21 The combination of Xu in view of Chang discloses the system of claim 19 (Xu, Fig. 7). The combination of Xu in view of Chang does not explicitly disclose wherein the compressed BEV feature vector is compressed using at least one of quantization, pruning, hashing, or transformation. However, Tu teaches wherein the compressed BEV feature vector is compressed (Tu, Fig. 1, Compressed BEV features) using at least one of quantization, pruning, hashing, or transformation (Tu, Section 1, “Each AV computes its own BEV representation from its camera rigs with the SinBEVTTransformer and then transmits it to others after compression. The receiver (i.e. other AVs) transforms the received BEV features onto its coordinate system, and employs the proposed FuseBEVT for BEV-level aggregation. The core ingredient of these two transformers is a novel fused axial attention(FAX) module, which can search over the whole BEV or camera image space across all agents or camera views via local and global spatial sparsity.”). Xu, Chang and Tu are both considered to be analogous to the claimed invention because they are in the same field of multiple vehicle data sharing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu and Chang to incorporate the teachings of Tu wherein the compressed BEV feature vector is compressed using at least one of quantization, pruning, hashing, or transformation. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to retrieve better accuracy (Tu, Section 2.2). Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable over Xu in view of Tu. Claim 23 Xu discloses the system of claim 1 (Xu, Fig. 7). Xu does not explicitly disclose wherein prior to fusing the one or more first features and the one or more second features, compress the one or more first features and the one of more second features using at least one of quantization, pruning, hashing, or transformation. However, Tu teaches wherein prior to fusing the one or more first features and the one or more second features (Tu, Fig. 2, aggregated BEV Features), compress the one or more first features and the one of more second features (Tu, Fig. 1, Compressed BEV features) using at least one of quantization, pruning, hashing, or transformation (Tu, Section 1, “Each AV computes its own BEV representation from its camera rigs with the SinBEVTTransformer and then transmits it to others after compression. The receiver (i.e. other AVs) transforms the received BEV features onto its coordinate system, and employs the proposed FuseBEVT for BEV-level aggregation. The core ingredient of these two transformers is a novel fused axial attention(FAX) module, which can search over the whole BEV or camera image space across all agents or camera views via local and global spatial sparsity.”). Xu and Tu are both considered to be analogous to the claimed invention because they are in the same field of multiple vehicle data sharing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system as taught by Xu to incorporate the teachings of Tu wherein prior to fusing the one or more first features and the one or more second features, compress the one or more first features and the one of more second features using at least one of quantization, pruning, hashing, or transformation. Such a modification is the result of combining prior art elements according to known methods to yield predictable results. The motivation for the proposed modification would have been to retrieve better accuracy (Tu, Section 2.2). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DENISE G ALFONSO whose telephone number is (571)272-1360. The examiner can normally be reached Monday - Friday 7:30 - 5:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DENISE G ALFONSO/Examiner, Art Unit 2662 /AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Mar 18, 2024
Application Filed
Feb 25, 2026
Non-Final Rejection mailed — §102, §103, §112
Apr 30, 2026
Interview Requested
May 15, 2026
Examiner Interview Summary
May 15, 2026
Applicant Interview (Telephonic)
May 22, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718568
OBSTACLE RECOGNITION METHOD AND APPARATUS, AND DEVICE, MEDIUM AND WEEDING ROBOT
3y 3m to grant Granted Aug 25, 2026
Patent 12711175
VIDEO RETRIEVAL METHOD AND APPARATUS USING VECTORIZING SEGMENTED VIDEOS
4y 0m to grant Granted Aug 18, 2026
Patent 12711767
SYSTEM AND METHOD FOR IDENTIFYING EVENTS IN A VIDEO STREAM
2y 10m to grant Granted Aug 18, 2026
Patent 12705895
VIDEO COMPARISON METHOD AND APPARATUS, COMPUTER DEVICE, AND STORAGE MEDIUM
4y 3m to grant Granted Aug 11, 2026
Patent 12688698
OBSTACLE RECONGNITION METHOD APPLIED TO AUTOMATIC TRAVELING DEVICE AND AUTOMATIC TRAVELING DEVICE
3y 2m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
91%
With Interview (+14.5%)
3y 0m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 130 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month