DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Claims 9-20 are withdrawn from further consideration pursuant to 37 CFR 1.142(b) as being drawn to a nonelected subcombination, there being no allowable generic or linking claim. Election was made without traverse in the reply filed on 07/24/2026.
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. 18898120, filed on 09/26/2024.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 01/08/2025, 06/23/2025, and 11/25/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 4-8, 21, 24-30, and 32 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li (BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers).
Regarding claim 1, Li teaches “A method comprising: determining, based at least on processing image data obtained using a plurality of cameras located within an environment, one or more first features associated with a plurality of images represented by the image data;” (Li, Figure 2 left panel; Section 3.1 Paragraph 2 (see below).)
PNG
media_image1.png
336
1065
media_image1.png
Greyscale
“determining, based at least on one or more spatial encoders processing the one or more first features and calibration data that relates one or more three-dimensional (3D) coordinates associated with the environment to one or more two-dimensional (2D) coordinates associated with the plurality of images, one or more second features associated with the environment; (Li, Figure 2 and Section 3.3 (see below). Note that the second features are mapped to current BEV Bt, which is determined based on the above spatial encoder processing.)
PNG
media_image2.png
817
891
media_image2.png
Greyscale
“determining, based at least on one or more decoders processing the one or more second features, one or more 3D locations associated with one or more objects located within the environment; and performing one or more operations based at least on the one or more 3D locations.” (Li, Section 3.5 and end of Page 3 (see below).)
PNG
media_image3.png
397
875
media_image3.png
Greyscale
PNG
media_image4.png
422
876
media_image4.png
Greyscale
Regarding claim 4, Li teaches “The method of claim 1,”
“wherein: a first portion of the one or more first features is associated with a first image of the plurality of images and a second portion of the one or more first features is associated with a second image of the plurality of images;” (Li, Figure 2 left panel; Section 3.1 Paragraph 2 (see below).)
PNG
media_image1.png
336
1065
media_image1.png
Greyscale
“and the determining the one or more second features is by at least aggregating the first portion of the one or more first features with the second portion of the one or more first features.” (Li, as described in the excerpt above, Ft is expressed as a vector of the portions of the first features, which amounts to an aggregation of these portions. Under another interpretation, the claim is also taught by the first Paragraph of section 3.1 (see below, last sentence). Note that the aggregation of Li is part of the pipeline to determine the second features Bt.)
PNG
media_image5.png
173
878
media_image5.png
Greyscale
Regarding claim 5, Li teaches “The method of claim 1,”
“wherein the calibration data represents at least matrices for projecting the 3D coordinates within the environment to the 2D coordinates associated with the plurality of images.” (Li, Figure 2 and Section 3.3 (see below).)
PNG
media_image2.png
817
891
media_image2.png
Greyscale
Regarding claim 6, Li teaches “The method of claim 1,”
“wherein the determining the one or more second features comprises: determining, based at least on the one or more spatial encoders processing the calibration data, that 3D points within the environment are associated with 2D points within the plurality of images; determining, based at least on the one or more first features, that the 2D points are associated with the one or more second features; and associating, based at least on the 2D points being associated with the one or more second features, the one or more second features with the 3D points. (Li, Figure 2 and Section 3.3 (see below). Figure 2 shows the association of 3D points in the environment with 2D points within the plurality of images (2b), which is subsequently output and aggregated to generate the second features (Bt, a 2D feature map) and ultimately identify the 3D points associated with the second features.)
PNG
media_image2.png
817
891
media_image2.png
Greyscale
Regarding claim 7, Li teaches “The method of claim 1,”
“wherein: the plurality of cameras is located within the environment and oriented such that the plurality of cameras includes fields-of-view representing at least a portion of an interior of the environment;” (Li, see the camera locations on the vehicle in Figure 2. When the vehicle is within an environment (i.e. city, street block, parking garage, etc.), the fields of view represent an interior of the environment (i.e. in the city, in the street block, in the parking garage, etc.).
“and the one or more objects include one or more dynamic objects located within the interior of the environment.” (Li, Figures 4 and 8-9 show detected vehicles (dynamic objects) as the one or more objects.)
Regarding claim 8, Li teaches “The method of claim 1,”
“wherein the one or more operations include at least one of:
determining one or more tracks associated with the one or more objects within environment;
determining one or more classifications associated with the one or more objects;
determining one or more 2D locations associated with the one or more objects within the plurality of images;
or causing a presentation of information associated with the one or more 3D locations.” (Li, Section 3.5 teaches various types of object the classification (determine semantic categories, determine bounding boxes, determine velocities, etc.). Figure 4 also shows presentation of information associated with the locations.)
PNG
media_image3.png
397
875
media_image3.png
Greyscale
Regarding claims 21 and 24-28, these claims recite a system with elements corresponding to the steps recited in Claims 1 and 4-8. Therefore, the recited elements of these claims are mapped to the analogous steps in the corresponding method claims. Additionally, Li teaches a processor (The architecture of Figure 2 inherently requires a processor. Li also extensively discusses computational costs throughout the publication).
Regarding claim 29, Li teaches “The method of claim 21,”
“wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing one or more simulation operations;
a system for performing one or more digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing one or more deep learning operations;
a system implemented using an edge device;
a system implemented using a robot;
a system for performing one or more generative AI operations;
a system for performing operations using one or more small language models;
a system for performing operations using one or more large language models;
a system for performing operations using one or more vision language models (VLMs);
a system for performing one or more conversational AI operations;
a system for generating synthetic data;
a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.” (Li, Introduction, Paragraph 5, “To this end, we present a transformer-based bird’s-eye-view (BEV) encoder, termed BEVFormer, which can effectively aggregate spatiotemporal features from multi-view cameras and history BEV features. The BEV features generated from the BEVFormer can simultaneously support multiple 3D perception tasks such as 3D object detection and map segmentation, which is valuable for the autonomous driving system. As shown in Fig. 1, our BEVFormer contains three key designs, which are (1) grid-shaped BEV queries to fuse spatial and temporal features via attention mechanisms flexibly, (2) spatial cross-attention module to aggregate the spatial features from multi-camera images, and (3) temporal self-attention module to extract temporal information from history BEV features, which benefits the velocity estimation of moving objects and the detection of heavily occluded objects, while bringing negligible computational overhead. With the unified features generated by BEVFormer, the model can collaborate with different task-specific heads such as Deformable DETR [56] and mask decoder [22], for end-to-end 3D object detection and map segmentation.”)
Regarding claims 30 and 32, these claims recite a processor with elements corresponding to the elements of the system recited in Claims 21 and 29. Therefore, the recited elements of these claims are mapped to the analogous elements in the corresponding system claims. Additionally, Li recites a processor with circuitry (see above rejection of claims 21 and 24-28).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2-3, 22-23, and 31 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li.
Regarding claim 2, Li teaches “The method of claim 1,”
Li does not expressly disclose “generating, based at least on one or more temporal encoders processing the one or more second features and one or more previous features associated with the environment, one or more fused features,”
However, in a second architecture region, Li does disclose generating fused features based on a temporal encoder processing a feature and a previous feature associated with the environment. (Li, Section 3 Paragraph 1 and Figure 2 describe the feature fusion based on temporal features (spatiotemporal feature aggregation); Section 3.4 (see below) and Figure 2 describe the temporal encoder processing a feature (BEV Queries Q) and a previous feature associated with the environment (History BEV Bt-1).)
PNG
media_image6.png
387
605
media_image6.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to perform the generating fused features (aggregated spatiotemporal features) based on a temporal encoder processing a feature and a previous feature associated with the environment (History BEV Bt-1), taught by the second architectural region of Li, wherein the second features (Bt) of Li is used as the first input feature (instead of BEV Queries Q).
The motivation for doing so would have been to further improve the perception tasks. Paragraph 4 of the Introduction of Li explains:
PNG
media_image7.png
252
809
media_image7.png
Greyscale
Therefore, incorporating a temporal encoder step as outlined in the combination above, to generate fused features for input to the detection and segmentation heads, would predictably result in better perception tasks. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Li with the above teaching of the second architectural region of Li to fully disclose, “generating, based at least on one or more temporal encoders processing the one or more second features and one or more previous features associated with the environment, one or more fused features,”
Li in view of the second architectural region of Li further disclose, “wherein the determining the one or more 3D locations is based at least on the one or more decoders processing the one or more fused features.” (As the different teachings of the Li reference are combined above, the fused features are input to the decoders that determine the 3D locations.)
Regarding claim 3, Li in view of the second architectural region of Li teaches “The method of claim 2,”
“wherein: the one or more second features are associated with the image data obtained using the plurality of cameras during a first time period;” (As outlined in the rejection of claim 1, the second feature is mapped to current BEV Bt, which is associated with the image data obtained using the plurality of cameras during a first time period (see figure 2).)
“and the one or more previous features are associated with second image data obtained using the plurality of cameras during a second time period that precedes the first time period.” (As the different teachings of the Li reference are combined above in the rejection of claim 2, the previous features are mapped to History BEV Bt-1, which is associated with other image data obtained by the cameras at a preceding time period. See Figure 2 and sections 3.4.)
Regarding claims 22, 23, and 31: claims 22 and 23 recite a system with elements corresponding to the steps recited in claims 2-3. Claim 31 recites a processor with elements corresponding to the steps recited in claim 2. Therefore, the recited elements of these claims are mapped to the analogous steps in the corresponding method claims. The rationale and motivation to combine the references apply here. Additionally, Li teaches a processor (The architecture of Figure 2 inherently requires a processor with circuitry. Li also extensively discusses computational costs throughout the publication).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wang (US 20260120445 A1) teaches driving environment perception system that fuses multi-modal data to track object (such as pedestrian and car) trajectory, including temporal encoding and spatial encoding.
Wang (US 20260087803 A1) teaches an environmental image processing and object detection strategy that includes extracting feature maps comprising 2D pixel coordinates and depth values, then translating these 2D pixel coordinates into a 3D scene. Wang also processes spatial and temporal features.
Zhou (WO 2024186974 A1) teaches a perception system for tracking objects in a vehicle scene based, in part, on features extracted from images and spatial-temporal information.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AARON JOSEPH SORRIN whose telephone number is (703)756-1565. The examiner can normally be reached Monday - Friday 9am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AARON JOSEPH SORRIN/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672