Prosecution Insights
Last updated: October 01, 2026
Application No. 19/025,091

Vision-Based Perception System

Non-Final OA §103
Filed
Jan 16, 2025
Priority
Aug 29, 2022 — continuation of PCTCN2022115508
Examiner
HWANG, JINSU
Art Unit
Tech Center
Assignee
Shenzhen Yinwang Intelligent Technology Co., Ltd.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
41 granted / 51 resolved
+20.4% vs TC avg
Minimal -2% lift
Without
With
+-2.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
15 currently pending
Career history
62
Total Applications
across all art units

Statute-Specific Performance

§101
6.7%
-33.3% vs TC avg
§103
53.1%
+13.1% vs TC avg
§102
31.8%
-8.2% vs TC avg
§112
6.2%
-33.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 51 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 0/2/13/2025, 05/20/2025, 12/12/2025, 04/29/2025, and 05/21/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Allowable Subject Matter Claims 4-5, 7-8, 13-14 and 16-17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 9-12, is/are rejected under 35 U.S.C. 103 as being unpatentable over Bogdoll et al. (Bogdoll, Daniel et al. “Multimodal Detection of Unknown Objects on Roads for Autonomous Driving.” 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (2022): 325-332, hereinafter “Bogdoll”) in view of Armstrong et al. (US Patent Number 2021/0158308-A1, hereinafter “Armstrong”). Regarding claim 1, Bogdoll teaches: A vision-based perception system for monitoring a physical environment of the system comprising: (Fig. 2) a Light Detection and Ranging (LIDAR) device configured to obtain a temporal sequence of point cloud data sets representing the environment; (III. Method C. 3D Lidar Detection, "We first utilized CenterPoint++ [49] to label objects as known in the 3D lidar space. CenterPoint++ ranked 2nd in the Waymo Real-time 3D detection challenge [10], as it achieved an mAPH of 72.8 and an inference speed of 57.1ms. Besides the outstanding performance, we chose CenterPoint++ as the model is only based on lidar data and its implementation is open-sourced. To be more precise, we used the two staged model architecture with the VoxelNet [50] backbone. Moreover, the model’s input is the multi-sweep aggregation of the current and the last two-point cloud frames.") a camera device configured to capture a temporal sequence of images of the environment; (III. Method C. 3D Lidar Detection, "We first utilized CenterPoint++ [49] to label objects as known in the 3D lidar space. CenterPoint++ ranked 2nd in the Waymo Real-time 3D detection challenge [10], as it achieved an mAPH of 72.8 and an inference speed of 57.1ms. Besides the outstanding performance, we chose CenterPoint++ as the model is only based on lidar data and its implementation is open-sourced. To be more precise, we used the two staged model architecture with the VoxelNet [50] backbone. Moreover, the model’s input is the multi-sweep aggregation of the current and the last two-point cloud frames.") and a processing unit comprising a first neural network and a second neural network different from the first neural network and configured to: (III. Method A. Road Segmentation, "In order to focus primarily on objects on the road, we applied a semantic segmentation model in the image domain as a first step. "; D. 2D Camera Detection, "Clustered objects that were not classified by the 3D detector are further processed in the 2D image space. Therefore, we mapped the detected 3D bounding boxes onto the corresponding position in the image space. … For the 2D classification, we used CLIP [51], a zero-shot model that can be used for visual classification tasks by providing the names of the visual categories to be recognized.") recognize an object present in the temporal sequence of images by means of the first neural network; (III. Method A. Road Segmentation, "In order to focus primarily on objects on the road, we applied a semantic segmentation model in the image domain as a first step. By using camera data, we made use of the higher information density through given shape and texture properties compared to the sparse lidar point cloud.") determine an unrecognized object present in the temporal sequence of images; (III. Method A. Road Segmentation, Fig. 2(f); Method D. 2D camera detection, "Clustered objects that were not classified by the 3D detector are further processed in the 2D image space.") determine a bounding box for the determined unrecognized object based on at least one of the point cloud data sets; (III. Method D. 2D camera detection, "Therefore, we mapped the detected 3D bounding boxes onto the corresponding position in the image space. ") obtain a neural network representation of the unrecognized object based on the temporal sequence of images and the determined bounding box by means of the second neural network (III. Method D. 2D camera detection, "Next, we constructed 2D candidate bounding boxes from the area of each 3D box and passed them to the 2D classifier for recognition."); Bogdoll does not teach: And train the first neural network for recognition of the unrecognized object based on the obtained neural network representation of the unrecognized object. However, Armstrong does teach: And train the first neural network for recognition of the unrecognized object based on the obtained neural network representation of the unrecognized object. (Armstrong, [0045], "The method can optionally include classifying the image (or set thereof) as “unknown.” …The labeled images can subsequently be used as training data to update the contamination detection system (e.g., train or update the neural network).") At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify lidar segmentation through machine learning (as taught by Bogdoll) to include training based on unrecognized objects (as taught by Armstrong) because such a modification is the result of combining prior art elements according to known methods to yield predictable results. More specifically, lidar segmentation through machine learning as modified by training based on unrecognized objects can yield a predictable result of allowing for better future performance by teaching it based on data the model had previous issue with. Thus, a person of ordinary skill would have appreciated including in lidar segmentation through machine learning the ability to do training based on unrecognized objects since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable. Regarding claim 2, Bogdoll in view of Armstrong teaches: The vision-based perception system according to claim 1, wherein the processing is further configured to: cluster points of each of the point cloud data sets to obtain point clusters for each of the point cloud data sets; and determine the unrecognized object by determining that for at least a some of the point cloud data sets one of the point clusters does not correspond to any object recognized by means of the first neural network. (Bogdoll, III. Method B. Corner case proposal generation, "After we identified the point cloud on the road, we generated corner case proposals by clustering the points via DBSCAN. We chose an = 1 as a distance to neighbors in a cluster and required that a cluster consists of at least 30 points. The clustering into single objects makes our pipeline an open-set detection, as clusters are initially considered unknown. However, we used state-of-the-art closed-set detection architectures to rule out clusters as known objects.") Regarding claim 3, Bogdoll in view of Armstrong teaches: The vision-based perception system according to claim 2, wherein the processing unit is configured to perform: ground segmentation based on the point cloud data sets to determine ground; and to determine the unrecognized object by determining that that the unrecognized object is located on the determined ground. (Bogdoll, III. Method B. Corner case proposal generation, "Therefore, we flattened all points and reduced the point cloud to two dimensions, i.e., setting the height to zero. A point lies on the road if the number of intersections with the scene of the estimated alpha shape is even, and thus the flattened point lies within the surface.") Regarding claim 9, Bogdoll in view of Armstrong teaches: The vision-based perception system according to claim 1, wherein the vision- based perception system is configured to be installed in a vehicle and the temporal sequence of point cloud data sets and the temporal sequence of images represent a driving scene of the vehicle. (Bogdoll, III. Method A. Road Segmentation, Fig. 2) Regarding claim 10, claim 10 has been analyzed with regard to claim 1 and is rejected for the same reasons of obviousness as used above. Regarding claim 11, claim 11 has been analyzed with regard to claim 2 and is rejected for the same reasons of obviousness as used above. Regarding claim 12, claim 12 has been analyzed with regard to claim 3 and is rejected for the same reasons of obviousness as used above. Claim(s) 6 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bogdoll et al. (Bogdoll, Daniel et al. “Multimodal Detection of Unknown Objects on Roads for Autonomous Driving.” 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (2022): 325-332, hereinafter “Bogdoll”) in view of Armstrong et al. (US Patent Number 2021/0158308-A1, hereinafter “Armstrong”) and Xie et al. (Xie, Christopher & Park, Keunhong & Martin Brualla, Ricardo & Brown, Matthew. (2021). FiG-NeRF: Figure-Ground Neural Radiance Fields for 3D Object Category Modelling. 10.48550/arXiv.2104.08418, hereinafter “Xie”). Regarding claim 6, Bogdoll in view of Armstrong does not teach: The vision-based perception system according to claim 1, wherein the second neural network comprises a first Multilayer Perceptron (MLP) trained based on the Neural Radiance Field technique and a second MLP trained based on the Neural Radiance Field technique different from the first MLP; and the first MLP is configured to obtain the neural network representation of the unrecognized object and the second MLP is configured to obtain a neural network representation of a background of the unrecognized object. However, Xie does teach: The vision-based perception system according to claim 1, wherein the second neural network comprises a first Multilayer Perceptron (MLP) trained based on the Neural Radiance Field technique and a second MLP trained based on the Neural Radiance Field technique different from the first MLP; and the first MLP is configured to obtain the neural network representation of the unrecognized object and the second MLP is configured to obtain a neural network representation of a background of the unrecognized object. (Xie, 1. Introduction, "We propose Figure-Ground Neural Radiance Fields (FiG-NeRF), which uses two NeRF models to model the objects and background, respectively."; "The key to our approach is to decompose the neural radiance field into two components: a foreground component Ff and a background component Fb, each modeled by a conditional NeRF.") At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify lidar segmentation through machine learning (as taught by Bogdoll in view of Armstrong) to include neuro radiance field technique (as taught by Xie) because such a modification is the result of combining prior art elements according to known methods to yield predictable results. More specifically, lidar segmentation through machine learning as modified by neuro radiance field technique can yield a predictable result of allowing for 3d reconstructions from 2d data that allows for more accurate segmentation. Thus, a person of ordinary skill would have appreciated including in lidar segmentation through machine learning the ability to do neuro radiance field technique since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable. Regarding claim 15, claim 15 has been analyzed with regard to claim 6 and is rejected for the same reasons of obviousness as used above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jinsu Hwang whose telephone number is (703)756-1370. The examiner can normally be reached Mon -Thu 10am-8am EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571) 272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JINSU HWANG/Examiner, Art Unit 2667 /MATTHEW C BELLA/Supervisory Patent Examiner, Art Unit 2667
Read full office action

Prosecution Timeline

Jan 16, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725387
OBJECT DETECTION DEVICE, OBJECT DETECTION METHOD, AND PROGRAM
3y 0m to grant Granted Sep 01, 2026
Patent 12710376
DATA PROCESSING APPARATUS, DATA PROCESSING SYSTEM, DATA PROCESSING METHOD, AND DATA PROCESSING PROGRAM
3y 8m to grant Granted Aug 18, 2026
Patent 12705892
EXPLAINABILITY FOR EVENT ALERTS IN VIDEO DATA
3y 9m to grant Granted Aug 11, 2026
Patent 12688633
APPARATUS AND METHOD FOR DEEP-LEARNING-BASED SCATTER ESTIMATION AND CORRECTION
3y 8m to grant Granted Jul 21, 2026
Patent 12675898
METHOD AND APPARATUS FOR DETERMINING A POSE OF A VEHICLE, AND VEHICLE CONTAINING SAME
3y 6m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
78%
With Interview (-2.5%)
2y 11m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 51 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month