Prosecution Insights
Last updated: October 04, 2026
Application No. 18/866,789

3D TARGET DETECTION METHOD AND APPARATUS BASED ON MULTI-VIEW FUSION

Non-Final OA §102§103
Filed
Nov 18, 2024
Priority
May 18, 2022 — CN 202210544237.0 +1 more
Examiner
NGUYEN, LEON VIET Q
Art Unit
Tech Center
Assignee
Beijing Horizon Robotics Technology Research And Development Co. Ltd.
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
973 granted / 1141 resolved
+25.3% vs TC avg
Moderate +10% lift
Without
With
+9.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
39 currently pending
Career history
1160
Total Applications
across all art units

Statute-Specific Performance

§101
5.1%
-34.9% vs TC avg
§103
66.5%
+26.5% vs TC avg
§102
16.8%
-23.2% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1141 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/1/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1, 2, 10-12, and 15 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation," arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026). Regarding claim 1, Xie discloses a 3D target detection method based on multi-view fusion, comprising: acquiring at least one image captured from a multi-camera view (section 3.1, Our framework takes N RGB images from multi-view cameras as input and corresponding extrinsic/intrinsic parameters); performing feature extraction on the at least one image to obtain respective feature data corresponding to the at least one image in a multi-camera view space (fig. 3; section 3.1, Part1 states we run the forward-pass with a shared CNN backbone for all images, e.g. ResNet, and a feature pyramid network (FPN) to create 4-level features), wherein the feature data comprises a feature of a target object (section 3.1, 4-level features); mapping, based on internal parameters and vehicle parameters of a multi-camera system, the respective feature data corresponding to the at least one image in the multi-camera view space to a same bird's-eye view space so as to obtain respective feature data corresponding to the at least one image in the bird's-eye view space (fig. 3, Given N images at timestamp T and corresponding intrinsic and extrinsic camera parameters as input, the encoder first extracts 2D features from the multi-view images, then the 2D features are unprojected to the 3D ego-car coordinate frame to generate a Bird’s-Eye View (BEV) feature representation; section 3.2); performing feature fusion on the respective feature data corresponding to the at least one image in the bird's-eye view space to obtain a bird's-eye view fusion feature (section 3.1, The multi-view features are combined and projected to 3D space to obtain a voxel. The voxel feature contains image features with all the views thus it is a unified feature representation. In the next step, the voxel feature is fed to 3D BEV encoder to obtain the BEV feature); and performing target prediction on the target object of the bird's-eye view fusion feature to obtain 3D spatial information of the target object (Overview of section 3.1, The outputs are 3D bounding boxes of objects; Part4 of section 3.1, Specifically, we directly adopt the detection head from PointPillars, which generates dense 3D anchors in BEV and then predicts the category, box size, and direction of each object). Regarding claim 2, Xie discloses a method wherein the mapping, based on internal parameters and vehicle parameters of a multi-camera system, the respective feature data corresponding to the at least one image in the multi-camera view space to a same bird's-eye view space so as to obtain respective feature data corresponding to the at least one image in the bird's-eye view space comprises: determining, based on the internal parameters and vehicle parameters of the multi-camera system, a transformation matrix of multi-camera of the multi-camera system from a camera coordinate system to a bird's-eye view coordinate system (fig. 3, Given N images at timestamp T and corresponding intrinsic and extrinsic camera parameters as input, the encoder first extracts 2D features from the multi-view images, then the 2D features are unprojected to the 3D ego-car coordinate frame to generate a Bird’s-Eye View (BEV) feature representation; Preliminary in section 3.2); and transforming, based on the transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system, the respective feature data corresponding to the at least one image in the multi-camera view space from the multi-camera view space to the bird's-eye view space to obtain the respective feature data corresponding to the at least one image in the bird's-eye view space (section 3.2). Regarding claims 10-11, the claim recites similar subject matter as claim 1 and is rejected for the same reasons as stated above. Regarding claims 12 and 15, the claim recites similar subject matter as claim 2 and is rejected for the same reasons as stated above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 3, 13, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation", arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026) in view of Zeng et al (US20150341629). Regarding claim 3, Xie fails to teach a method wherein the determining, based on the internal parameters and vehicle parameters of the multi-camera system, a transformation matrix of multi-camera of the multi-camera system from a camera coordinate system to a bird's-eye view coordinate system comprises: acquiring internal parameters of a camera and external parameters of a camera of the multi-camera in the multi-camera system, respectively, and acquiring a transformation matrix from a vehicle coordinate system to the bird's-eye view coordinate system; and determining, based on the external parameters of the camera and the internal parameters of the camera of the multi-camera, and the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system. However Zeng teaches acquiring internal parameters of a camera and external parameters of a camera (para. [0003], The technique described herein provides an approach for autonomously estimating extrinsic camera parameters and for refining intrinsic parameters of the cameras by merging multiple views of pre-defined patterns) of the multi-camera in the multi-camera system (para. [0006]), respectively, and acquiring a transformation matrix from a vehicle coordinate system to the bird's-eye view coordinate system (para. [0015]; para. [0026], In combining the estimated camera parameters in real world coordinates and the transformation from the vehicle coordinates to the world coordinates, camera parameters in vehicle coordinates may be calculated which is the extrinsic calibration objective); and determining, based on the external parameters of the camera and the internal parameters of the camera of the multi-camera, and the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system (para. [0003]-[0004]). Therefore taking the combined teachings of Xie and Zeng as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Zeng into the method of Xie. The motivation to combine Zeng and Xie would be to provide a less complex approach for camera alignment and synchronization which would increase accuracy as to the location of the detected objects (para. [0014] of Zeng). Regarding claims 13 and 16, the claim recites similar subject matter as claim 3 and is rejected for the same reasons as stated above. Claim(s) 4-5, 14, and 17-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation", arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026) in view of Yin et al ("Center-based 3d object detection and tracking", Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, Pages 11784-11793, retrieved from the Internet on 7/17/2026). Regarding claim 4, Xie fails to teach a method wherein the performing target prediction on the target object of the bird's-eye view fusion feature to obtain 3D spatial information of the target object comprises: acquiring, using a predictive network, a heat map for determining a first preset coordinate value of the target object in a bird's-eye view coordinate system, and acquiring other attribute maps for determining a second preset coordinate value, a size and an orientation angle of the target object in the bird's-eye view coordinate system from the bird's-eye view fusion feature; determining the first preset coordinate value of the target object in the bird's-eye view coordinate system based on peak information in the heat map, and determining the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate from the other attribute maps based on the first preset coordinate value of the target object in the bird's-eye view coordinate system; and determining 3D spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system. However Yin teaches acquiring, using a predictive network, a heat map for determining a first preset coordinate value of the target object (section 4, Center heatmap head) in a bird's-eye view coordinate system (section 3, bird eye view), and acquiring other attribute maps for determining a second preset coordinate value (section 4, At inference time, we extract all properties by indexing into dense regression head outputs at each object’s peak location), a size and an orientation angle of the target object in the bird's-eye view coordinate system from the bird's-eye view fusion feature (section 4, We store several object properties at center-features of objects: a sub-voxel location refinement, height-above-ground, the 3D size, and a yaw rotation angle); determining the first preset coordinate value of the target object in the bird's-eye view coordinate system based on peak information in the heat map (section 3, Each local maximum (i.e., pixels whose value is greater than its eight neighbors) in the output heatmap corresponds to the center of a detected object) and determining the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate from the other attribute maps based on the first preset coordinate value of the target object in the bird's-eye view coordinate system (section 4, At inference time, we extract all properties by indexing into dense regression head outputs at each object’s peak location); and determining 3D spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system (section 3, Each bounding box b = (u, v, d, w, l, h, α) consists of a center location (u, v, d), relative to the objects ground plane, and 3D size (w, l, h), and rotation expressed by yaw α). Therefore taking the combined teachings of Xie and Yin as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Yin into the method of Xie. The motivation to combine Yin and Xie would be to provide a detection and tracking algorithm which is simple, efficient, and effective (abstract of Yin). Regarding claim 5, the modified method of Xie teaches a method further comprising: constructing, in a training stage of the predictive network, a first loss function between the heat map predicted by the predictive network and a true value heat map (section 4 of Yin, focal loss), and a second loss function between the other attribute maps predicted by the predictive network and other true value attribute maps (section 4 of Yin, We train all outputs using an L1 loss at the ground truth center location); and determining a total loss function of the predictive network during the training stage based on the first loss function and the second loss function to supervise a training process of the predictive network (section 4 of Yin, CenterPoint combines all heatmap and regression losses in one common objective and jointly optimizes them). Regarding claims 14 and 17, the claim recites similar subject matter as claim 4 and is rejected for the same reasons as stated above. Regarding claim 18, the claim recites similar subject matter as claim 5 and is rejected for the same reasons as stated above. Claim(s) 7-8 and 20-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xie et al ("M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation", arXiv preprint arXiv:2204.05088, 4/19/2022, pages 1-21, retrieved from the Internet on 7/15/2026) in view of Huang et al ("Bevdet: High-performance multi-camera 3d object detection in bird-eye-view", arXiv preprint arXiv:2112.11790 (2021), pages 1-11, retrieved from the Internet on 7/17/2026). Regarding claim 7, Xie fails to teach a method wherein the performing target prediction on the target object of the bird's-eye view fusion feature to obtain 3D spatial information of the target object comprises: performing, using a neural network, feature extraction on the bird's-eye view fusion feature to obtain bird's-eye view fusion feature data comprising the feature of the target object; and performing, using a predictive network, target prediction on the target object of the bird's-eye view fusion feature data comprising the feature of the target object to obtain 3D spatial information of the target object. However Huang teaches performing, using a neural network, feature extraction on the bird's-eye view fusion feature to obtain bird's-eye view fusion feature data comprising the feature of the target object (fig. 2; BEV Encoder in section 3, Though the structure of this module is similar to that of the image-view encoder with a backbone and a neck, it perceives some pivotal cues like depth, scale, orientation, and speed, which is difficult for predicting in image view. We follow to utilize ResNet with classical residual block); and performing, using a predictive network, target prediction on the target object of the bird's-eye view fusion feature data comprising the feature of the target object to obtain 3D spatial information of the target object (fig. 2; Head in section 3, The task-specific head is constructed upon the BEV feature. In common sense, 3D object detection in automatic pilot aims at the position, scale, orientation, and speed of movable objects like pedestrians, vehicles, barriers, and so on). Therefore taking the combined teachings of Xie and Huang as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Huang into the method of Xie. The motivation to combine Yin and Huang would be to achieve multi-tasks learning with both high performance and high efficiency (section 2.2 of Huang). Regarding claim 8, Xia fails to teach a method wherein the performing feature extraction on the at least one image to obtain respective feature data corresponding to the at least one image in a multi-camera view space, wherein the feature data comprises a feature of a target object comprises: performing, using a depth neural network, convolution calculation on images corresponding to respective views, to obtain feature data of a plurality of different resolutions respectively corresponding to the images of respective views in the multi-camera view space, wherein the feature data comprises a feature of a target object. However Huang teaches performing, using a depth neural network (Image-View Encoder in section 3), convolution calculation (Image-View Encoder in section 3, ResNet and FPN) on images corresponding to respective views (fig. 2), to obtain feature data of a plurality of different resolutions respectively corresponding to the images of respective views in the multi-camera view space (fig. 2; Image-View Encoder in section 3, FPN-LSS simply upsamples the feature with 1/32 input resolution to 1/16 input resolution and concatenates it with the one generated by the backbone with 1/16 input resolution), wherein the feature data comprises a feature of a target object (section 5, BEVDet successfully pushes the performance boundary and is particularly good at predicting the target’s translation, orientation, and velocity). Therefore taking the combined teachings of Xie and Huang as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Huang into the method of Xie. The motivation to combine Yin and Huang would be to achieve multi-tasks learning with both high performance and high efficiency (section 2.2 of Huang). Regarding claim 20, the claim recites similar subject matter as claim 7 and is rejected for the same reasons as stated above. Regarding claim 21, the claim recites similar subject matter as claim 8 and is rejected for the same reasons as stated above. Allowable Subject Matter Claims 6 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Related Art Theverapperuma et al (US20220024485) – see figs. 7-8, para. [0096], [0108] Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEON VIET Q NGUYEN whose telephone number is (571)270-1185. The examiner can normally be reached Mon-Fri 11AM-7PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gregory Morse can be reached at 571-272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LEON VIET Q NGUYEN/ Primary Examiner, Art Unit 2663
Read full office action

Prosecution Timeline

Nov 18, 2024
Application Filed
Sep 22, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738028
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, IMAGING DEVICE, VEHICLE DEVICE, AND MEDICAL ROBOT DEVICE
3y 3m to grant Granted Sep 15, 2026
Patent 12731275
IMAGE PROCESSING DEVICE, COMPONENT GRIPPING SYSTEM, IMAGE PROCESSING METHOD AND COMPONENT GRIPPING METHOD
2y 6m to grant Granted Sep 08, 2026
Patent 12725276
MOTION FEEDBACK METHOD AND SYSTEM USING NORMALIZED DATA
2y 4m to grant Granted Sep 01, 2026
Patent 12718594
VISUAL PRESENTATION OF VEHICLE POSITIONING RELATIVE TO SURROUNDING OBJECTS
3y 1m to grant Granted Aug 25, 2026
Patent 12718388
NON-LINE-OF-SIGHT IMAGING VIA NEURAL TRANSIENT FIELD
2y 8m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
95%
With Interview (+9.9%)
2y 6m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1141 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month