Prosecution Insights
Last updated: August 17, 2026
Application No. 18/954,611

RARE OBJECT DETECTION SYSTEM AND METHOD FOR IMAGE CORPUS BUILDING

Non-Final OA §103
Filed
Nov 21, 2024
Examiner
ESQUINO, CALEB LOGAN
Art Unit
2677
Tech Center
2600 — Communications
Assignee
GM Global Technology Operations LLC
OA Round
1 (Non-Final)
56%
Grant Probability
Moderate
1-2
OA Rounds
1y 2m
Est. Remaining
70%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
14 granted / 25 resolved
-6.0% vs TC avg
Moderate +14% lift
Without
With
+14.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
19 currently pending
Career history
49
Total Applications
across all art units

Statute-Specific Performance

§101
5.2%
-34.8% vs TC avg
§103
61.3%
+21.3% vs TC avg
§102
15.5%
-24.5% vs TC avg
§112
16.1%
-23.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 25 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to the application filed on November 21st, 2024. Claims 1-20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on November 3rd, 2025 is being considered by the examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3-10, 12-20 are rejected under 35 U.S.C. 103 as being unpatentable over “Towards Corner Case Detection for Autonomous Driving” (herein after referred to by its primary author, Bolte) in view of “Can gaze control steering?” (herein after referred to by its primary author, Tuhkanen) and “OW-Adapter: Human-Assisted Open-World Object Detection with a Few Examples” (herein after referred to by its primary author, Jamonnak). In regards to claim 1, Bolte teaches receiving, via an image interface, an image (Bolte Figure 3 “Input sequence”); generating a depth information in response to the image (Bolte Section IV C “To identify a relevant location, in a real system, typically one would use a perception approach based on a light detection and ranging (LIDAR) sensor to assign depth information to the image pixels [45], [46]. On the basis of that, a time to collision on a pixel basis can be estimated [47]. For the purpose of this work, however, instead we adopt a rather simple approach for the reason of conciseness of presentation and evaluation. We simply assume that objects being further above the bottom-line of the image are more distant to the ego vehicle.”); estimating an object closeness and an object vertical position in response to the depth information (Bolte Section IV C “The squared errors of the relevant classes e2t,rel(i) are then weighted depending on their distance from the bottom of the image and summed up resulting in an error score”); estimating a Bolte Figure 2; Section IV “we need a detection system that processes the information from both image prediction and semantic segmentation by information fusion, comprising a check, whether the nonpredictable (e.g., jumping) relevant class (e.g., pedestrian) is in a relevant location (will cross trajectory).” Examiner note: This section teaches that, when detecting anomalies, it is important to determine if the relevant object will cross the path of the vehicle. As will be discussed later with respect to Tuhkanen, a driver’s gaze can influence the path of travel, which will in turn affect whether the object is in a relevant location. For example, if an anomaly is detected in the right hand side of an image (such as a traffic cone on a right side street), it is relevant to determine if the drivers gaze is indicative of a right hand turn.); detecting a probability of a known object within the image in response to the image (Bolte Section IV A “The input image xt ε GHxWxC with image pixel xt(i) ε G, where G is the set of gray values, H and W are the image height and width in pixels and C = 3 is the number of color channels from set C = {1, 2, 3}, is fed into a fully-convolutional neural network. It maps the input to output scores Pt ε IHxWx|S|, where S denotes the set of classes with cardinality |S| = 19 and I = [0, 1]. For each pixel position i ε I, the third dimension of the output scores provides a posterior probability (score) Pt(i, s) for each class s ε S.” Examiner note: This section teaches that the input image is segmented into regions, and that each region is given a probability score for the known class that it is assigned.); calculating a joint probability of an occurrence of an unknown object within the image in response to the depth information, the object closeness and the object vertical position, the probability of the known object within the image (Bolte Abstract “The challenging task of corner case detection in video, which is also somehow related to unusual event or anomaly detection, aims at detecting these unusual situations, which could become critical, and to communicate this to the autonomous driving system (online use case).”; Section IV C “Finally, the corner case score is obtained by normalizing the error score εt to a value range from 0 to 1 using [equation (7)]”); Bolte does not teach a method of building an image corpus for training a neural network; estimating a driver gaze probability in response to the image; identifying the image for annotation to generate an annotated image indicative of the unknown object in response to the joint probability exceeding a threshold value; adding the annotated image to the image corpus; and training the neural network in response to the image corpus. However, Tuhkanen teaches estimating a driver gaze probability in response to the image (Tuhkanen Discussion Section “Overall, the models were able to produce successful steering across two different scenarios using gaze information alone, producing steering patterns that were similar to those of human drivers. This suggests that the human driver’s gaze placement is at least in principle able to facilitate steering (i.e., gaze contains sufficient information to guide steering).” Examiner note: This reference teaches that a driver’s gaze can influence the direction of travel. This reference, in combination with Bolte, suggests that an object is in a relevant location (will cross the path of the vehicle) if the driver’s gaze is fixated upon the object.). Tuhkanen is considered to be analogous to the claimed invention because they are both in the same field of unknown obstacle avoidance (Tuhkanen Experiment 1 “The focus of the current study was on modeling steering-to-intercept behavior”). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the system of Bolte to include the teachings of Tuhkanen, to provide the advantage of being able to determine a driver’s intended direction of travel from only their gaze direction (Tuhkanen Discussion Section “Each controller was made to steer-by-gaze without access to any visual information from the environment (“steering based on where you are looking,” not “steering based on what you see”) in order to gauge whether successful steering can be produced from gaze alone.”) Furthermore, Jamonnak teaches a method of building an image corpus for training a neural network (Jamonnak Figure 3) comprising: calculating a probability of an occurrence of an unknown object within the image (Jamonnak Figure 3(b)); identifying the image for annotation to generate an annotated image indicative of the unknown object in response to the joint probability exceeding a threshold value (Jamonnak Figure 3 (c) “We start with a pre-trained general detector (a) and use it to extract proposals of potential objects (b). Then, we group the proposals into clusters and rank the clusters (c) such that users can identify potential unknown classes efficiently.” Examiner note: As can be seen in the “top proposals” box of figure 3, images are sent for annotation based on their probability of containing a known/unknown object); adding the annotated image to the image corpus (Jamonnak Figure 3 (d) “For a newly discovered unknown class, we recommend relevant proposals for users to annotate(d).”); and training the neural network in response to the image corpus (Jamonnak Figure 3 (e) “With users ’pseudo-labeling annotations (users just need to specify unknown class names, NOT detailed bounding boxes), we train a lightweight classifier to detect the new classes (e).”). Jamonnak is considered to be analogous to the claimed invention because they are both in the same field of identifying unknown objects. Therefore it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the system of Bolte in view of Tuhkanen to include the teachings of Jamonnak, to provide the advantage of reduced computational cost (Jamonnak Section 4.4 “First, to minimize users’ annotation efforts, we only ask them to annotate a few examples, which is not sufficient to fine-tune the entire neural network. Second, following the trend of parameter-efficient fine-tuning, we train a lightweight adapter to reduce the computation cost and provide online feedback to users.”) In regards to claim 3, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of claim 1, further comprising controlling a vehicle along a motion path in response to a subsequent unknown object identified by the neural network in a subsequent image (Bolte Figure 1 “driving parameters of the vehicle”). In regards to claim 4, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of claim 3, wherein the neural network is a convolutional neural network (Jamonnak Section 4 “Given an image dataset and a pre-trained object detector, we first freeze the detector and generate proposals of potential objects (Section 4.1). Each proposal box contains an image patch that is likely to be an object with semantic meanings. Then, we group the proposal patches into clusters and rank the clusters to help users identify important unknown classes (Section 4.2). After identifying an unknown class, we recommend relevant proposal patches that are likely to belong to the target class for users to annotate (Section 4.3). The annotated examples are then used to train a lightweight classifier, which is used as a plug-in to the original detector to handle unknown classes (Section 4.4).”; Section 6.1.1 “Training and Testing We use a pre-trained Faster R-CNN [55] model from Detectron2 Model Zoo [66]. The model is trained on PASCAL VOC with 20 known object classes. Faster R-CNN is a popular object detection model that uses deep convolutional neural networks to detect and localize objects in images.” Examiner note: Section 4 teaches that a pre-trained model is used, and a classifier is plugged into the pre-trained model after annotation of the unknown objects. Section 6.1.1 teaches that the pre-trained model is a convolutional neural network.). In regards to claim 5, Bolte in view of Tuhkanen and Jamonnak teaches the method of claim 1, wherein the image corpus includes a plurality of images having a plurality of joint probabilities of the occurrence of the unknown object exceeding the threshold value (Jamonnak Figure 3(d) Examiner note: Each image in the corpus used to train the lightweight plugin classifier was determined from an unknown cluster, as only the orange clusters (which represent unknowns) are used). In regards to claim 6, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of claim 1, wherein the image includes a plurality of pixels and wherein a calculation of the joint probability generates the probability of the occurrence of the unknown object for each of the plurality of pixels (Bolte Section IV A “Here, I is the set of pixel indices in the image, and |I| = H x W the number of pixels”; Equation (6) Examiner note: The corner case score ε’t is calculated for each pixel and summation is performed for each pixel in the image in equation 6). In regards to claim 7, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of claim 1, wherein the object vertical position is indicative of an object height (Bolte Section IV C “The squared errors of the relevant classes e2t,rel(i) are then weighted depending on their distance from the bottom of the image and summed up resulting in an error score” Examiner note: The object’s vertical position (distance from the bottom) is indicative of the object’s height in the image). In regards to claim 8, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of claim 1, wherein a joint probability calculation is higher for an image pixel in response to the driver gaze probability for the image pixel overlapping a low depth area of the image (Bolte Figure 2; Section IV “we need a detection system that processes the information from both image prediction and semantic segmentation by information fusion, comprising a check, whether the nonpredictable (e.g., jumping) relevant class (e.g., pedestrian) is in a relevant location (will cross trajectory)” Section IV C “Thereby, our simple definition of a relevant location weights the bottom row squared errors by a one, and the top row squared errors of the relevant classes by a zero.” Examiner note: As previously discussed with respect to Tuhkanen, when a driver’s gaze is located towards an area, that area is likely to be in the intended direction of travel. Then, as Bolte figure 2 and section IV teaches, if an object is in the vehicles direction of travel), it is considered to be more likely to be relevant and therefore a corner case/anomaly. Furthermore, section IV C teaches that an object with a lower depth (which is therefore closer to the bottom of the image) has a higher weight than one with a higher depth). In regards to claim 9, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of Claim 1, wherein the image corpus is formed from a plurality of annotated images indicative of the unknown object in response to the joint probability exceeding the threshold value for each of the plurality of annotated images (Jamonnak Figure 3(c)-(e)). In regards to claim 10, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 1. In regards to claim 12, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 5. In regards to claim 13, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 8. In regards to claim 14, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 9. In regards to claim 15, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 7. In regards to claim 16, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 6. In regards to claim 17, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 9. In regards to claim 18, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 3. In regards to claim 19, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 1. In regards to claim 20, Bolte in view of Tuhkanen and Jamonnak renders obvious the claim limitations as in the consideration of claim 3. Claims 2 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Bolte in view of Tuhkanen and Jamonnak as applied to the claims above, and further in view of “Generalized Few-Shot 3D Object Detection of LiDAR Point Cloud for Autonomous Driving” (herein after referred to by its primary author, Liu). In regards to claim 2, Bolte in view of Tuhkanen and Jamonnak teaches the method of building the image corpus for training the neural network of claim 1, but fails to teach wherein the probability of the known object is detected in response to a long-class tail probability. However, Liu teaches wherein the probability of the known object is detected in response to a long-class tail probability (Liu Figure 1; Section 1 “We design a sample adaptive balance (SAB) loss function to deal with the severe long-tail distribution in point clouds.”; Figure 2 “Novel class”). Liu is considered to be analogous to the claimed invention because they are both in the same field of identifying unknown objects. Therefore it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the system of Bolte in view of Tuhkanen and Jamonnak to include the teachings of Liu, to provide the advantage of accurately detecting known classes while still detecting novel classes (Liu Section V C “From setting 4 and 5, the experiments show that by fixing the network parameters of θE, we do effectively preserve the detection effect of the base class. This shows that keeping the feature extraction part of the model unchanged can effectively retain the very good feature extraction capability of the original model.”) In regards to claim 11, Bolte in view of Tuhkanen, Jamonnak, and Liu renders obvious the claim limitations as in the consideration of claim 2. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: “Addressing Open-set Object Detection for Autonomous Driving perception: A focus on road objects” teaches a method of classifying road objects, including an unknown class. “Identifying Unknown Instances for Autonomous Driving” teaches a method of identifying novel objects that do not belong to a known class. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CALEB LOGAN ESQUINO whose telephone number is (703)756-1462. The examiner can normally be reached M-Fr 8:00AM-4:00PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at (571) 270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CALEB L ESQUINO/ Examiner, Art Unit 2677 /ANDREW W BEE/ Supervisory Patent Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Nov 21, 2024
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682427
Optical Method
3y 1m to grant Granted Jul 14, 2026
Patent 12675894
A METHOD AND A SYSTEM FOR ASSESSING AND OPTIONALLY MONITORING OR CONTROLLING THE TEXTURE OF A SURFACE
3y 1m to grant Granted Jul 07, 2026
Patent 12639781
REDUCING IMAGE SCALING ARTIFACTS VIA TILE SIZE SELECTION
3y 6m to grant Granted May 26, 2026
Patent 12602924
Method for Semantic Localization of an Unmanned Aerial Vehicle
4y 0m to grant Granted Apr 14, 2026
Patent 12602813
DEEP APERTURE
3y 6m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
56%
Grant Probability
70%
With Interview (+14.0%)
2y 10m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 25 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month