DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 12/1/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 2, 10-12, and 15 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation," arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026).
Regarding claim 1, Xie discloses a 3D target detection method based on multi-view fusion, comprising:
acquiring at least one image captured from a multi-camera view (section 3.1, Our framework takes N RGB images from multi-view cameras as input and corresponding extrinsic/intrinsic parameters);
performing feature extraction on the at least one image to obtain respective feature data corresponding to the at least one image in a multi-camera view space (fig. 3; section 3.1, Part1 states we run the forward-pass with a shared CNN backbone for all images, e.g. ResNet, and a feature pyramid network (FPN) to create 4-level features), wherein the feature data comprises a feature of a target object (section 3.1, 4-level features);
mapping, based on internal parameters and vehicle parameters of a multi-camera system, the respective feature data corresponding to the at least one image in the multi-camera view space to a same bird's-eye view space so as to obtain respective feature data corresponding to the at least one image in the bird's-eye view space (fig. 3, Given N images at timestamp T and corresponding intrinsic and extrinsic camera parameters as input, the encoder first extracts 2D features from the multi-view images, then the 2D features are unprojected to the 3D ego-car coordinate frame to generate a Bird’s-Eye View (BEV) feature representation; section 3.2);
performing feature fusion on the respective feature data corresponding to the at least one image in the bird's-eye view space to obtain a bird's-eye view fusion feature (section 3.1, The multi-view features are combined and projected to 3D space to obtain a voxel. The voxel feature contains image features with all the views thus it is a unified feature representation. In the next step, the voxel feature is fed to 3D BEV encoder to obtain the BEV feature); and
performing target prediction on the target object of the bird's-eye view fusion feature to obtain 3D spatial information of the target object (Overview of section 3.1, The outputs are 3D bounding boxes of objects; Part4 of section 3.1, Specifically, we directly adopt the detection head from PointPillars, which generates dense 3D anchors in BEV and then predicts the category, box size, and direction of each object).
Regarding claim 2, Xie discloses a method wherein the mapping, based on internal parameters and vehicle parameters of a multi-camera system, the respective feature data corresponding to the at least one image in the multi-camera view space to a same bird's-eye view space so as to obtain respective feature data corresponding to the at least one image in the bird's-eye view space comprises:
determining, based on the internal parameters and vehicle parameters of the multi-camera system, a transformation matrix of multi-camera of the multi-camera system from a camera coordinate system to a bird's-eye view coordinate system (fig. 3, Given N images at timestamp T and corresponding intrinsic and extrinsic camera parameters as input, the encoder first extracts 2D features from the multi-view images, then the 2D features are unprojected to the 3D ego-car coordinate frame to generate a Bird’s-Eye View (BEV) feature representation; Preliminary in section 3.2); and
transforming, based on the transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system, the respective feature data corresponding to the at least one image in the multi-camera view space from the multi-camera view space to the bird's-eye view space to obtain the respective feature data corresponding to the at least one image in the bird's-eye view space (section 3.2).
Regarding claims 10-11, the claim recites similar subject matter as claim 1 and is rejected for the same reasons as stated above.
Regarding claims 12 and 15, the claim recites similar subject matter as claim 2 and is rejected for the same reasons as stated above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 3, 13, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation", arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026) in view of Zeng et al (US20150341629).
Regarding claim 3, Xie fails to teach a method wherein the determining, based on the internal parameters and vehicle parameters of the multi-camera system, a transformation matrix of multi-camera of the multi-camera system from a camera coordinate system to a bird's-eye view coordinate system comprises:
acquiring internal parameters of a camera and external parameters of a camera of the multi-camera in the multi-camera system, respectively, and acquiring a transformation matrix from a vehicle coordinate system to the bird's-eye view coordinate system; and
determining, based on the external parameters of the camera and the internal parameters of the camera of the multi-camera, and the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system.
However Zeng teaches acquiring internal parameters of a camera and external parameters of a camera (para. [0003], The technique described herein provides an approach for autonomously estimating extrinsic camera parameters and for refining intrinsic parameters of the cameras by merging multiple views of pre-defined patterns) of the multi-camera in the multi-camera system (para. [0006]), respectively, and acquiring a transformation matrix from a vehicle coordinate system to the bird's-eye view coordinate system (para. [0015]; para. [0026], In combining the estimated camera parameters in real world coordinates and the transformation from the vehicle coordinates to the world coordinates, camera parameters in vehicle coordinates may be calculated which is the extrinsic calibration objective); and
determining, based on the external parameters of the camera and the internal parameters of the camera of the multi-camera, and the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system (para. [0003]-[0004]).
Therefore taking the combined teachings of Xie and Zeng as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Zeng into the method of Xie. The motivation to combine Zeng and Xie would be to provide a less complex approach for camera alignment and synchronization which would increase accuracy as to the location of the detected objects (para. [0014] of Zeng).
Regarding claims 13 and 16, the claim recites similar subject matter as claim 3 and is rejected for the same reasons as stated above.
Claim(s) 4-5, 14, and 17-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation", arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026) in view of Yin et al ("Center-based 3d object detection and tracking", Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, Pages 11784-11793, retrieved from the Internet on 7/17/2026).
Regarding claim 4, Xie fails to teach a method wherein the performing target prediction on the target object of the bird's-eye view fusion feature to obtain 3D spatial information of the target object comprises:
acquiring, using a predictive network, a heat map for determining a first preset coordinate value of the target object in a bird's-eye view coordinate system, and acquiring other attribute maps for determining a second preset coordinate value, a size and an orientation angle of the target object in the bird's-eye view coordinate system from the bird's-eye view fusion feature;
determining the first preset coordinate value of the target object in the bird's-eye view coordinate system based on peak information in the heat map, and determining the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate from the other attribute maps based on the first preset coordinate value of the target object in the bird's-eye view coordinate system; and
determining 3D spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system.
However Yin teaches acquiring, using a predictive network, a heat map for determining a first preset coordinate value of the target object (section 4, Center heatmap head) in a bird's-eye view coordinate system (section 3, bird eye view), and acquiring other attribute maps for determining a second preset coordinate value (section 4, At inference time, we extract all properties by indexing into dense regression head outputs at each object’s peak location), a size and an orientation angle of the target object in the bird's-eye view coordinate system from the bird's-eye view fusion feature (section 4, We store several object properties at center-features of objects: a sub-voxel location refinement, height-above-ground, the 3D size, and a yaw rotation angle);
determining the first preset coordinate value of the target object in the bird's-eye view coordinate system based on peak information in the heat map (section 3, Each
local maximum (i.e., pixels whose value is greater than its eight neighbors) in the output heatmap corresponds to the center of a detected object) and determining the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate from the other attribute maps based on the first preset coordinate value of the target object in the bird's-eye view coordinate system (section 4, At inference time, we extract all properties by indexing into dense regression head outputs at each object’s peak location); and
determining 3D spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system (section 3, Each bounding box b = (u, v, d, w, l, h, α) consists of a center location (u, v, d), relative to the objects ground plane, and 3D size (w, l, h), and rotation expressed by yaw α).
Therefore taking the combined teachings of Xie and Yin as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Yin into the method of Xie. The motivation to combine Yin and Xie would be to provide a detection and tracking algorithm which is simple, efficient, and effective (abstract of Yin).
Regarding claim 5, the modified method of Xie teaches a method further comprising:
constructing, in a training stage of the predictive network, a first loss function between the heat map predicted by the predictive network and a true value heat map (section 4 of Yin, focal loss), and a second loss function between the other attribute maps predicted by the predictive network and other true value attribute maps (section 4 of Yin, We train all outputs using an L1 loss at the ground truth center location); and
determining a total loss function of the predictive network during the training stage based on the first loss function and the second loss function to supervise a training process of the predictive network (section 4 of Yin, CenterPoint combines all heatmap and regression losses in one common objective and jointly optimizes them).
Regarding claims 14 and 17, the claim recites similar subject matter as claim 4 and is rejected for the same reasons as stated above.
Regarding claim 18, the claim recites similar subject matter as claim 5 and is rejected for the same reasons as stated above.
Claim(s) 7-8 and 20-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation", arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026) in view of Huang et al ("Bevdet: High-performance multi-camera 3d object detection in bird-eye-view", arXiv preprint arXiv:2112.11790 (2021), pages 1-11, retrieved from the Internet on 7/17/2026).
Regarding claim 7, Xie fails to teach a method wherein the performing target prediction on the target object of the bird's-eye view fusion feature to obtain 3D spatial information of the target object comprises:
performing, using a neural network, feature extraction on the bird's-eye view fusion feature to obtain bird's-eye view fusion feature data comprising the feature of the target object; and
performing, using a predictive network, target prediction on the target object of the bird's-eye view fusion feature data comprising the feature of the target object to obtain 3D spatial information of the target object.
However Huang teaches performing, using a neural network, feature extraction on the bird's-eye view fusion feature to obtain bird's-eye view fusion feature data comprising the feature of the target object (fig. 2; BEV Encoder in section 3, Though the structure of this module is similar to that of the image-view encoder with a backbone and a neck, it perceives some pivotal cues like depth, scale, orientation, and speed, which is difficult for predicting in image view. We follow to utilize ResNet with classical residual block); and
performing, using a predictive network, target prediction on the target object of the bird's-eye view fusion feature data comprising the feature of the target object to obtain 3D spatial information of the target object (fig. 2; Head in section 3, The task-specific head is constructed upon the BEV feature. In common sense, 3D object detection in automatic pilot aims at the position, scale, orientation, and speed of movable objects like pedestrians, vehicles, barriers, and so on).
Therefore taking the combined teachings of Xie and Huang as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Huang into the method of Xie. The motivation to combine Yin and Huang would be to achieve multi-tasks learning with both high performance and high efficiency (section 2.2 of Huang).
Regarding claim 8, Xia fails to teach a method wherein the performing feature extraction on the at least one image to obtain respective feature data corresponding to the at least one image in a multi-camera view space, wherein the feature data comprises a feature of a target object comprises:
performing, using a depth neural network, convolution calculation on images corresponding to respective views, to obtain feature data of a plurality of different resolutions respectively corresponding to the images of respective views in the multi-camera view space, wherein the feature data comprises a feature of a target object.
However Huang teaches performing, using a depth neural network (Image-View Encoder in section 3), convolution calculation (Image-View Encoder in section 3, ResNet and FPN) on images corresponding to respective views (fig. 2), to obtain feature data of a plurality of different resolutions respectively corresponding to the images of respective views in the multi-camera view space (fig. 2; Image-View Encoder in section 3, FPN-LSS simply upsamples the feature with 1/32 input resolution to 1/16 input resolution and concatenates it with the one generated
by the backbone with 1/16 input resolution), wherein the feature data comprises a feature of a target object (section 5, BEVDet successfully pushes the performance boundary and is particularly good at predicting the target’s translation, orientation, and velocity).
Therefore taking the combined teachings of Xie and Huang as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Huang into the method of Xie. The motivation to combine Yin and Huang would be to achieve multi-tasks learning with both high performance and high efficiency (section 2.2 of Huang).
Regarding claim 20, the claim recites similar subject matter as claim 7 and is rejected for the same reasons as stated above.
Regarding claim 21, the claim recites similar subject matter as claim 8 and is rejected for the same reasons as stated above.
Allowable Subject Matter
Claims 6 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Related Art
Theverapperuma et al (US20220024485) – see figs. 7-8, para. [0096], [0108]
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEON VIET Q NGUYEN whose telephone number is (571)270-1185. The examiner can normally be reached Mon-Fri 11AM-7PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gregory Morse can be reached at 571-272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LEON VIET Q NGUYEN/ Primary Examiner, Art Unit 2663