Prosecution Insights
Last updated: October 02, 2026
Application No. 19/014,029

OPEN-VOCABULARY OBJECT DETECTION IN IMAGES

Non-Final OA §DP
Filed
Jan 08, 2025
Priority
May 06, 2022 — provisional 63/339,165 +2 more
Examiner
LIEW, ALEX KOK SOON
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
88%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
976 granted / 1114 resolved
+27.6% vs TC avg
Moderate +7% lift
Without
With
+7.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
30 currently pending
Career history
1129
Total Applications
across all art units

Statute-Specific Performance

§101
11.2%
-28.8% vs TC avg
§103
63.9%
+23.9% vs TC avg
§102
17.3%
-22.7% vs TC avg
§112
4.4%
-35.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1114 resolved cases

Office Action

§DP
DETAILED ACTION [1] Remarks I. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . II. Claims 1-20 are pending and have been examined, where claims 1-20 is/are rejected. Explanations will be provided below. III. Inventor and/or assignee search were performed and determined no double patenting rejection(s) is/are necessary. IV. Patent eligibility (updated in 2019) shown by the following: Claims 1-20 pass patent eligibility test because there is/are no limitation or a combination of limitations amounting to an abstract idea. Also, the following limitation or the combinations of the limitations: “the set of query embeddings, using a classification subnetwork of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings, wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the region of the input image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding” effects a transformation or a reduction of a particular article to a different state or thing / adds a specific limitation(s) other than what is well-understood, routine and conventional in the field, or adding unconventional steps that confine the claim to a particular useful application and providing improvements to the technical field of Deep learning, which recite additional elements that integrate the judicial exception into a practical application and amounting significant more. [2] Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. Use of the word “means” (or “step for”) in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function. Absence of the word “means” (or “step for”) in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function. Claim elements in this application that use the word “means” (or “step for”) are presumed to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Similarly, claim elements that do not use the word “means” (or “step for”) are presumed not to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Claim(s) 12-16 are not interpreted under 35 U.S.C. 112(f) or pre-AIA U.S.C. 112 6th paragraph because of the following reason(s): limitations are modified by sufficient structure or material for performing the claimed function. Claim(s) 2-11 and 17-21 do not require 35 U.S.C. 112(f) or pre-AIA U.S.C. 112 6th paragraph interpretation because they are method claims and / or they are CRM claims. Upon examination of the specification and claims, the examiner has determined, under the best understanding of the scope of the claim(s), rejection(s) under 35 U.S.C. 112(a)/(b) is not necessitated because of the following reasons: sufficient support are provided in the written description / drawings of the invention. [3] Grounds of Rejection Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the "right to exclude" granted by a patent and to prevent possible harassment by multiple assignees. See In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent is shown to be commonly owned with this application. See 37 CFR 1.130(b). Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b). Claim 2 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claim 1 of U.S. Patent No. 11,928,854. The conflicting claims are not identical because patent claim 1 requires the additional step of “processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (underlined portion)”, not required by claim 2. However, the conflicting claims are not patentably distinct from each other because: Claim 2 and patent claim 1 recite common subject matter; whereby claim 2, which recites the open-ended transitional phrase “comprising”, does not preclude the additional elements recited by patent claim 1, and whereby the elements of claim 2 are fully anticipated by patent claim 1. Claim 12 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claim 14 of U.S. Patent No. 11,928,854. The conflicting claims are not identical because patent claim 14 requires the additional step of “processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (underlined portion)”, not required by claim 12. However, the conflicting claims are not patentably distinct from each other because: Claim 12 and patent claim 14 recite common subject matter; whereby claim 12, which recites the open-ended transitional phrase “comprising”, does not preclude the additional elements recited by patent claim 14, and whereby the elements of claim 12 are fully anticipated by patent claim 14. Claim 17 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claim 15 of U.S. Patent No. 11,928,854. The conflicting claims are not identical because patent claim 15 requires the additional step of “processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (underlined portion)”, not required by claim 17. However, the conflicting claims are not patentably distinct from each other because: Claim 17 and patent claim 15 recite common subject matter; whereby claim 17, which recites the open-ended transitional phrase “comprising”, does not preclude the additional elements recited by patent claim 15, and whereby the elements of claim 17 are fully anticipated by patent claim 15. Note: All claims will be indicated rejected because all independent claims are rejected under Double Patenting. [4] Allowable Subject Matter Claims 2-21 are allowable / patentable if applicant(s) overcome double patenting rejections. The following is an examiner’s statement of reasons for allowance by comparing claims to closest references. The references are divided into primary and secondary, where primary would have been utilized in a USC 102 or main USC 103 reference and secondary would have been utilized a secondary USC 103 reference, but these references do not cover enough of the claim’s scope to warrant a rejection. Primary reference, Minderer “Simple Open-Vocabulary Object Detection with Vision Transformers” (publication date: 05/06/2022) discloses a method performed by one or more computers, the method comprising: obtaining: (i) an image, and (ii) a set of one or more query embeddings, wherein each query embedding represents a respective category of object (see figure 1, illustration below, image of giraffes and its text are input to the transformer); processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image (see figure 1, MLP is read as the neural network), comprising: processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (see figure 1, the image is divided into four equal regions, each sub regions is read as its own self attention, see “Vision Transformer Encoder” and “Text Transformer Encoder”): PNG media_image1.png 467 1472 media_image1.png Greyscale processing each object embedding using a localization subnetwork of the object detection neural network to generate localization data defining a corresponding region of the image (see figure 1, output of MLP is read as localization data); and processing: (i) the set of object embeddings, and (ii) the set of query embeddings, using a classification subnetwork of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings (see figure 1, illustration below the output of Linear Projections is read as classification scores), wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the text corresponding to the category represented by the query embedding (see figure 1 illustration below, the embedding corresponds to the likelihood of text input): PNG media_image2.png 623 1519 media_image2.png Greyscale . Minderer is silent in disclosing but not a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding. Also, Minderer does not qualify as a prior art because its publication date (May 12, 2022) is after the earliest priority of the current application (05/06/2022). Primary reference, Wang (US 20230019211) discloses a method performed by one or more computers, the method comprising: obtaining: (i) an image (see figure 2, 120), and (ii) a set of one or more query embeddings, wherein each query embedding represents a respective category of object (see figure 2, 140); processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image (see figure 2, 104 receives text embedding, also see illustrations below, see paragraph 63, trains one or more neural networks from data comprising a mix of paired and unpaired data), comprising: processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings (see paragraph 70, a pre-training framework obtains or otherwise receives as input an image 120, calculates an image embedding 122, which is input to a multi-scale image encoder 124 to calculate an image embedding 126, see paragraph 215, a CNN include a region-based or regional convolutional neural networks, RCNNs, and Fast RCNNs, e.g., as used for object detection or other type of CNN, 124 is read as the subnetwork, 122 is read as the image embedding); PNG media_image3.png 395 897 media_image3.png Greyscale processing each object embedding using a localization subnetwork of the object detection neural network to generate localization data defining a corresponding region of the image (see paragraph 245, a CNN for emergency vehicle detection and identification may use data from microphones 1496 to detect and identify emergency vehicle sirens); and processing: (i) the set of object embeddings (see figure 1, 108), and (ii) the set of query embeddings, using a classification subnetwork of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings (see paragraph 97, a softmax function is a generalization of a logistic function to multiple dimensions, a softmax function normalizes inputs into a probability distribution comprising probabilities), but is silent in disclosing wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding. Secondary reference, Lund (US 20200372610) discloses using a classification subnetwork of the object detection neural network to generate a respective classification score distribution over the set of query embeddings (see figure 3 and figure 5B), but not for each object embedding: PNG media_image4.png 417 1465 media_image4.png Greyscale ; wherein the respective classification score distribution for each of the PNG media_image5.png 264 844 media_image5.png Greyscale . Lund is silent in disclosing wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding. Secondary reference, Ermolov “Hyperbolic Vision Transformers: Combining Improvements in Metric Learning” discloses a method performed by one or more computers, the method comprising: obtaining (1) an image (see figure 1 illustrations below), and (ii) a set of one or more query embeddings, wherein each query embedding represents a respective category of object (see figure 1 illustrations below). Ermolov is silent in disclosing processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image. Secondary reference, AGARWAL (US 20190080225) discloses wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood paragraph 36, a view to force the network to learn better separation of the embeddings, query embeddings, the above loss may be increased slightly for all predictions, irrespective of whether the prediction is right or wrong, a square-root of all the probabilities in the prediction distribution Pi and then re-normalize to obtain the new probability distribution Qi, FIG. 4 is a graphical representation illustrating a predicted Probability Distribution (P), new probability distribution obtained after square-root and normalization of P, and T is the target distribution): PNG media_image6.png 308 712 media_image6.png Greyscale . AGARWAL is silent in disclosing a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding. AGARWAL’s network is not a CNN used for object detection. Wang, Lund, Ermolov and AGARWAL, taken alone or in combination with each other, are silent in disclosing all the limitations of claims. CONTACT INFORMATION Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEX LIEW (duty station is located in New York City) whose telephone number is (571)272-8623 (FAX 571-273-8623), cell (917)763-1192 or email alexa.liew@uspto.gov. Please note the examiner cannot reply through email unless an internet communication authorization is provided by the applicant. The examiner can be reached anytime. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MISTRY ONEAL R, can be reached on (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALEX KOK S LIEW/Primary Examiner, Art Unit 2674 Telephone: 571-272-8623 Date: 9/3/26
Read full office action

Prosecution Timeline

Jan 08, 2025
Application Filed
Sep 08, 2026
Non-Final Rejection mailed — §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743772
IDENTIFYING BLOOD VESSELS IN ULTRASOUND IMAGES
2y 8m to grant Granted Sep 22, 2026
Patent 12744872
EXPANDED FIELD OF VIEW USING MULTIPLE CAMERAS
2y 6m to grant Granted Sep 22, 2026
Patent 12727764
DENTAL CARIES DETECTION DEVICE
2y 6m to grant Granted Sep 08, 2026
Patent 12725391
Performance Recording System, Performance Recording Method, and Recording Medium
2y 5m to grant Granted Sep 01, 2026
Patent 12718408
CAMERA CALIBRATION METHOD AND APPARATUS
2y 6m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
88%
Grant Probability
95%
With Interview (+7.2%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1114 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month