Prosecution Insights
Last updated: October 02, 2026
Application No. 18/602,504

METHOD FOR ASCERTAINING A DESCRIPTOR IMAGE FOR AN IMAGE OF AN OBJECT

Non-Final OA §103
Filed
Mar 12, 2024
Priority
Mar 31, 2023 — DE 10 2023 203 021.7
Examiner
WANG, YUEHAN
Art Unit
2617
Tech Center
2600 — Communications
Assignee
Robert Bosch GmbH
OA Round
3 (Non-Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
416 granted / 504 resolved
+20.5% vs TC avg
Moderate +14% lift
Without
With
+13.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
26 currently pending
Career history
544
Total Applications
across all art units

Statute-Specific Performance

§101
4.9%
-35.1% vs TC avg
§103
72.4%
+32.4% vs TC avg
§102
6.9%
-33.1% vs TC avg
§112
6.2%
-33.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 504 resolved cases

Office Action

§103
DETAILED ACTION Response to Amendment Applicant’s amendments filed on 21 April 2026 have been entered. Claims 1-9 have been amended. Claim 10 has been added. Claims 1-10 are still pending in this application, with claims 1 and 7-9 being independent. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 21 April 2026 has been entered. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6, 8-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dal Mutto et al. (US 20190108396 A1), referred herein as Mutto in view of ZHOU et al. (US 20220164595 A1), referred herein as ZHOU. Regarding Claim 1, Mutto in view of ZHOU teaches a method for ascertaining a descriptor image for an image of an object, comprising the following steps (Mutto [0002] systems and methods for the automatic identification of objects; [0006] compute one or more descriptors of the one or more 3-D models, each descriptor corresponding to a fixed-length feature vector; and retrieve metadata identifying the one or more objects based on the one or more descriptors): training a plurality of machine learning model instances, wherein each machine learning model instance of the plurality of machine learning model instances corresponds to an object class of a plurality of object classes ((Mutto FIG. 12.CNN1s; [0135] convolutional neural networks (CNNs) may be used for multi-view object classification; [0138] the neural network is trained based on training data, which may include a set of 3-D models of objects and their corresponding labels (e.g., the correct classifications of the objects)); [0142] FIGS. 12 and 13 are illustration of max-pooling according to one embodiment of the present invention. As shown in FIG. 13, each of the n views is supplied to the first stage CNN.sub.1 of the descriptor generator 314 to generate n feature vectors), the training including: training a first machine learning model instance to map images of objects of a first object class to descriptor images (Mutto [0132] Techniques for computing a descriptor of a 3-D model are based on a forward evaluation of a Multi-View Convolutional Neural Network (MV-CNN) or by a Volumetric Convolutional Neural Network (V-CNN). Such networks are usually trained for object classification; [0140] As shown in FIG. 11, the values computed by the first stage CNN.sub.1 (the convolutional stage) and supplied to the second stage CNN.sub.2 (the fully connected stage) are referred to herein as a descriptor (or feature vector) f), and Mutto disclosed regions of a reference 3-D model (see [0212]), but does not specifically teaches a first/second reference descriptors. However, ZHOU teaches storing first reference descriptors output by the first machine learning model instance corresponding to one or more objects of the first object class (ZHOU [0004] obtaining a set of reference descriptors; [0073] To obtain the set of reference descriptors 147 associated with the set of keypoints 143, the computing device 120 may first determine a reference descriptor map of the reference image 140 and then obtain, from the reference descriptor map, a plurality of reference descriptors (namely, the set of reference descriptors 147) corresponding to respective keypoints of the set of keypoints 143. For this purpose, the computing device 120 may generate the reference descriptor map from the reference image 140 in the same manner as that for generating the image descriptor map 160 from the captured image 130); and Mutto in view of ZHOU further teaches training a second machine learning model instance, that is different than the first machine learning model instance, to map images of objects of a second object class, that is different than the first object class, to descriptor images (Mutto [0141] The architecture of a classifier 270 described above with respect to FIG. 11 can be applied to classifying multi-view shape representations of 3-D objects based on n different 2-D views of the object; [0142] FIGS. 12 and 13 are illustration of max-pooling according to one embodiment of the present invention. As shown in FIG. 13, each of the n views is supplied to the first stage CNN.sub.1 of the descriptor generator 314 to generate n feature vectors; [0145] The extracted feature vector can then be supplied to a classifier to classify the object as being a member of one of a particular set of k different classes C, thereby resulting in classification of the query object 1), and storing second reference descriptors, that are different than the first reference descriptors, output by the second machine learning model instance corresponding to one or more objects of the second object class (ZHOU [0071] the computing device 120 or another entity (for example, another computing device) may have generated and stored a set of keypoints, a set of reference descriptors and a set of spatial coordinates in association for each reference image in the set of reference images of the external environment 105); and receiving an image of an object of an unknown object class (Mutto [0074] misclassification of one or more items in a checkout process generally leads to disputes and low customer satisfaction; [0079] the 3-D scanner 100 captures images of the object (operation 1200) and the object tracking system 400 tracks (operation 1400); [0208] the unknown input can be predicted as normal or abnormal by calculating a difference between the unknown sample and the transformed output; the analysis system 300 performs defect detection on the scanned objects to detect whether the scanned object has one or more defects or is non-defective); generating, by the first machine learning model instance, a first descriptor image for the object of the unknown object class (Mutto [0023] compute one or more first descriptors of the one or more first 3-D models; and identify the one or more goods in the shopping basket based on the one or more first descriptors of the one or more first 3-D models; [0133] in FIG. 7, the descriptor is computed from 2-D views 16 of the 3-D model 240, as rendered by the view generation module 312 in operation 1312); generating, by the second machine learning model instance, a second descriptor image for the object of the unknown object class that is different than the first descriptor image (Mutto [0023] receive the one or more second 3-D models from the second 3-D scanning system; compute one or more second descriptors of the one or more second 3-D models; [0133] In operation 1314, the synthesized 2-D views are supplied to a descriptor generator 314 to extract a descriptor or feature vector for each view); determining whether a first distance between the first descriptor image and the first reference descriptors is less than a second distance between the second descriptor image and the second reference descriptors (ZHOU [0088] Referring back to FIG. 2, at block 240, the computing device 120 may determine a plurality of similarities 170 between the plurality of sets of image descriptors 165 and the set of reference descriptors 147; [0091] in the case that the image descriptors and the reference descriptors are represented in the form of an n-dimensional vector, for each pair of corresponding “image descriptor-reference descriptor,” the computing device 120 may calculate the difference between the two descriptors as an L2 distance between the two paired descriptors; [0092] Subsequent to determining the plurality of differences associated with the plurality of descriptor pairs between the first set of image descriptors 165-1 and the set of reference descriptors 147, the computing device 120 may determine, based on the plurality of differences, a similarity between the first set of image descriptors 165-1 and the set of reference descriptors 147, namely the first similarity 170-1 of the plurality of similarities 170); and assigning, in response to determining the first distance is less than the second distance, the object of the unknown object class to the first object class, the assigning including selecting the first descriptor image as the descriptor image for the object (Mutto [0136] The output p of the second stage is a class-assignment probability distribution. For example, if the entire CNN is trained to assign input images to one of k different classes, then the output of the second stage CNN.sub.2 is a vector p that includes k different values, each value representing the probability (or “confidence”) that the input image should be assigned the corresponding class; [0147] A similarity metric is defined to measure the distance between any two given descriptors (vectors) F and F.sub.ds(m)… A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class (as measured from examples in the training data) and maximizes the average distance between vector pairs belonging to different classes; FIG. 12.CLASS). ZHOU discloses techniques for determining a plurality of similarities between the plurality of sets of image descriptors and the set of reference descriptors. ZHOU is analogous to the present patent application. It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mutto to incorporate the teachings of ZHOU, and applying the computing device 120 that may determine the first similarity 170-1 between the first set of image descriptors 165-1 and the set of reference descriptors 147 into the systems and methods for the automatic identification of objects. Doing so would simplify a scene interpretation and classification algorithm of an environment perception module, and allow the localization system of a vehicle to reach centimeter-level accuracy. Regarding Claim 2, Mutto in view of ZHOU teaches the method according to claim 1, and further teaches wherein the determining comprises: determining the first distance by: computing, for each reference descriptor of the first reference descriptors, a first plurality of distances to descriptors in the first descriptor image (ZHOU [0144] As shown in FIG. 16, the apparatus 1600 may include a first obtaining module 1610, a second obtaining module 1620, a first determining module 1630, a second determining module 1640, and an updating module 1650; [0039] the computing device may determine a plurality of similarities between the plurality of sets of image descriptors and the set of reference descriptors); selecting, from each first plurality of distances, a minimum distance to generate a first plurality of minimum distances (Mutto [0254] A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class); and aggregating the first plurality of minimum distances to generate the first distance (Mutto [0254] A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class (as measured from examples in the training data) and maximizes the average distance between vector pairs belonging to different classes); and determining the second distance by: computing, for each reference descriptor of the second reference descriptors, a second plurality of distances to descriptors in the second descriptor image (ZHOU [0149] a plurality of differences between a plurality of image descriptors of the first set of image descriptors and corresponding reference descriptors of the set of reference descripto); selecting, from each second plurality of distances, a minimum distance to generate a second plurality of minimum distances (Mutto [0254] A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class); and aggregating the second plurality of minimum distances to generate the second distance (Mutto [0254] A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class (as measured from examples in the training data) and maximizes the average distance between vector pairs belonging to different classes). Regarding Claim 3, Mutto in view of ZHOU teaches the method according to claim 1, and further teaches wherein the first machine learning model instance and the second machine learning model instance share a sub-model that is used as part of training the plurality of machine learning model instances (Mutto [0246] the first stage CNN.sub.1 can be applied independently to each of the n 2-D views used to represent the 3-D shape, thereby computing a set of n feature vectors f(1), f(2), . . . , f(n) (one for each of the 2-D views); FIG. 12.CNN1s; FIG. 13.CNN1s; ZHOU FIG. 16.1610, 1620). Regarding Claim 4, Mutto in view of ZHOU teaches the method according to claim 3, and further teaches further comprising training the sub-model using training data containing objects from plurality of the object classes (ZHOU [0088] For other sets of image descriptors among the plurality of sets of image descriptors 165, the computing device 120 may determine likewise the similarities between them and the set of reference descriptors 147 to finally obtain the plurality of similarities 170). Regarding Claim 6, Mutto in view of ZHOU teaches the method according to claim 1, and further teaches wherein at least one of the plurality of machine learning model instances is a neural network (Mutto FIG. 12.CNN1s; FIG. 13.CNN1s; ZHOU FIG. 16.1610, 1620; ZHOU [0060] the computing device 120 may input the captured image 130 into a trained machine learning model and then gain the image descriptor map 160 at the output of the machine learning model) . Regarding Claims 8 and 9, Mutto in view of ZHOU teaches a control unit configured to and a non-transitory computer-readable medium on which is stored instructions ascertaining a descriptor image for an image of an object, the instructions for controlling a robot to pick up or process an object, comprising the following steps (Mutto [0006] an analysis agent including a processor and memory, the memory storing instructions; ZHOU Abst: a method, an apparatus, an electronic device and a storage medium for vehicle localization, which relates to the technical fields of autonomous driving, electronic map, deep learning, image processing, and the like… a computing device obtains an image descriptor map corresponding to a captured image; Fig. 8). The metes and bounds of the claims substantially correspond to the limitations set forth in claim 1; thus they are rejected on similar grounds and rationale as their corresponding limitations. Regarding Claim 10, Mutto in view of ZHOU teaches the method according to claim 2, and further teaches wherein: the first plurality of distances and the second plurality of distances are computed according to a Euclidean distance (Mutto [0147] Some simple examples of similarity metrics are a Euclidean vector distance and a Mahalanobis vector distance); aggregating the first plurality of minimum distances to generate the first distance includes averaging the first plurality of minimum distances (Mutto [0254] A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class (as measured from examples in the training data) and maximizes the average distance between vector pairs belonging to different classes; FIG. 12.CNN1s; FIG. 13.CNN1s); and aggregating the second plurality of minimum distances to generate the second distance includes averaging the second plurality of minimum distances (Mutto [0254] A metric learning algorithm may learn a linear or non-linear transformation of feature vector space that minimizes the average distance between vector pairs belonging to the same class (as measured from examples in the training data) and maximizes the average distance between vector pairs belonging to different classes; FIG. 12.CNN1s; FIG. 13.CNN1s). Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dal Mutto et al. (US 20190108396 A1), referred herein as Mutto in view of ZHOU et al. (US 20220164595 A1), referred herein as ZHOU and WILLIAMS et al. (US 20220405363 A1), referred herein as WILLIAMS. Regarding Claim 5, Mutto in view of ZHOU teaches the method according to claim 1. However, in view of WILLIAMS, the prior art further teaches wherein: the first machine learning model instance is trained using a first training data set that contains images of objects of the first object class, and the objects of the first object class are overrepresented in the first training data set (Mutto FIG. 12.CNN1s; FIG. 13.CNN1s; WILLIAMS [0014] In the example of speaker authentication, using an overly large data set can lead to a neural network overfitting to the training data, especially if certain categories of speakers are overrepresented in a training data set); and the second machine learning model instance is trained using a second training data set that is different than the first training data set, and that contains images of objects of the second object class, and the objects of the second object class are overrepresented in the second training data set (Mutto FIG. 12.CNN1s; FIG. 13.CNN1s; WILLIAMS [0014] In the example of speaker authentication, using an overly large data set can lead to a neural network overfitting to the training data, especially if certain categories of speakers are overrepresented in a training data set). WILLIAMS discloses a method of generating a biometric signature of a user for use in authentication using a neural network. WILLIAMS is analogous to the present patent application. It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mutto to incorporate the teachings of WILLIAMS, and applying the overrepresented training data into the plurality of CNN1s for the automatic identification of objects. Doing so would improve the performance of neural networks. Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mutto et al. (US 20190108396 A1), referred herein as Mutto in view of ZHOU et al. (US 20220164595 A1), referred herein as ZHOU and Shrivastava et al. (US 20230186587 A1), referred herein as Shrivastava. Regarding Claim 7, Mutto in view of ZHOU teaches a method for controlling a robot to pick up or process an object, comprising the following steps (ZHOU Abst: a computing device obtains an image descriptor map corresponding to a captured image; Mutto [0165] replace the improper or anomalous component or to automatically replace the component (e.g., using a robot picking system)). The metes and bounds of the claims substantially correspond to the limitations set forth in claim 1; thus they are rejected on similar grounds and rationale as their corresponding limitations. Mutto in view of ZHOU further teaches ascertaining a position of a location or a pose for picking up or processing the object in a current control scenario from the ascertained descriptor image (ZHOU [0059] the image descriptor map 160 may include descriptors of respective image points in the captured image 130. For example, in the image descriptor map 160, a position corresponding to an image point in the captured image 130 records a descriptor of the image point); and controlling the robot to pick up or process the object according to the ascertained position or location or according to the ascertained pose (Shrivastava [0012] Robot guidance can include guiding a robot end effector, for example a gripper, to pick up a part and orient the part for assembly in an environment that includes a plurality of parts). Shrivastava discloses a method for vehicle localization, which relates to the technical fields of autonomous driving, electronic map, deep learning, image processing. Shrivastava is analogous to the present patent application. It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mutto to incorporate the teachings of Shrivastava, and applying the deep neural network trained based on object class prediction loss functions into the systems and methods for the automatic identification of objects. Doing so can improve localization accuracy and robustness of the vehicle visual localization algorithm. Response to Arguments Applicant's arguments filed on 21 April 2026, with respect to the 103 rejection have been fully considered but are moot in view of the new grounds of rejection. Examiner notes that independent claims 1, 9 and 13 have been amended to include new limitation. Examiner finds these limitations to be unpatentable as can be found in above detail action. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Samantha (Yuehan) Wang whose telephone number is (571)270-5011. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached on (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Samantha (YUEHAN) WANG/ Primary Examiner Art Unit 2617
Read full office action

Prosecution Timeline

Mar 12, 2024
Application Filed
Sep 30, 2025
Non-Final Rejection mailed — §103
Dec 04, 2025
Response Filed
Jan 23, 2026
Final Rejection mailed — §103
Apr 21, 2026
Request for Continued Examination
Apr 24, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731302
RAPID RENDERING AND/OR REALISTIC VISUALIZATION OF APPAREL DESIGN DRAFT FILES THROUGH APPLICATION OF ONE OR MORE GENERATIVE ARTIFICIAL NEURAL NETWORKS
2y 6m to grant Granted Sep 08, 2026
Patent 12718478
VIRTUAL ENVIRONMENT GUIDING METHOD AND SYSTEM
2y 3m to grant Granted Aug 25, 2026
Patent 12700146
SYSTEM FOR CONTEXTUAL DIMINISHED REALITY FOR METAVERSE IMMERSIONS
2y 3m to grant Granted Aug 04, 2026
Patent 12700192
AUGMENTED REALITY SYSTEM
2y 3m to grant Granted Aug 04, 2026
Patent 12694620
DEPTH RENDERING FROM NEURAL RADIANCE FIELDS FOR 3D MODELING
2y 2m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
96%
With Interview (+13.6%)
2y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 504 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month