Prosecution Insights
Last updated: August 17, 2026
Application No. 18/794,786

USING VISUAL LANGUAGE MODELS TO DETERMINE LOCATIONS OF IMAGE ELEMENTS WITHIN GRAPHICAL IMAGES

Non-Final OA §101§112
Filed
Aug 05, 2024
Examiner
PATEL, JAYESH A
Art Unit
2677
Tech Center
2600 — Communications
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
758 granted / 907 resolved
+21.6% vs TC avg
Minimal +5% lift
Without
With
+5.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
35 currently pending
Career history
932
Total Applications
across all art units

Statute-Specific Performance

§101
9.3%
-30.7% vs TC avg
§103
46.1%
+6.1% vs TC avg
§102
15.9%
-24.1% vs TC avg
§112
22.3%
-17.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 907 resolved cases

Office Action

§101 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 3 recites the limitation "the data structure" in line 2. There is insufficient antecedent basis for this limitation in the claim. Claims 6 and 9 recites the limitation "the one or more properties" in line 1 respectively. There is insufficient antecedent basis for this limitation in the claims. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 20 is rejected under 35 U.S.C 101. The claimed invention is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter because the broadest reasonable interpretation of “One or more computer storage media” encompasses a signal and signals are non-statutory. See MPEP 2106.03 I. Applicant is advised to amend the claim as “One or more non-transitory computer storage media-------” in-order to make the claim statutory. Allowable Subject Matter Claim 1 is allowed. Regarding independent claim 1, NPL1 (Flamingo: a Visual language Model for Few-Shot Learning, Jean-Baptiste Alayrac et al., arXiv, Nov 2021, Pages 1-54) hereafter NPL1 discloses “ A method performed by one or more computers and for training a visual language model to identify locations of image elements within a graphical image, the method comprising: generating a plurality of training data items, each training data item including (i) a graphical image rendered according to a corresponding set of instructions (fig 1 shows the graphical images in the input prompt section) (ii) a natural language query for identifying at least one image element of the graphical image (fig 1 shows the text description (natural language query) in the input prompt identifying at least one image element of the graphical image (i.e “This is a Chinchilla--- This is a Shiba. They are very popular in Japan etc meeting the claim limitations), and (iii) a target location for the at least one image element, the target location being determined from the set of instructions (fig 1 shows the third image on the top with the text query “This is” in the input prompt section and the arrow to Completion “a flamingo. They are found in the Caribbean and South America (i.e a target location for the at least one image element in the third image in the input prompt, examiner notes that the specifics of a target location/place are not required by the current claim) meeting the claim limitations); for each of the training data items, processing the corresponding graphical image and natural language query using a visual language model to generate a corresponding model output comprising a predicted location of an image element identified from the natural language query (fig 1 shows the bottom row 4th image on the top (i.e the Graphical image) based on the context “This is cityscape. It looks like Chicago” and with the natural language query “what makes you think this is Chicago?” as the natural language query and the answer “I think it is Chicago because of the Shedd Aquarium in the background (i.e a predicted location for the at least one image element identified from the natural language query), examiner notes that the specifics of a predicted location/place are not required by the current claim) meeting the claim limitations). NPL1 and the other cited arts alone or in combination however fail to disclose “adjusting parameters of the visual language model to optimize, for each of the training data items, an objective function that depends on a comparison between the predicted location of the model output corresponding to the training data item and the target location of the training data item.”, therefore claim 1 is allowed. Dependent claims 2, 4-5, 7-8 and 10-18 depending directly or indirectly on claim 1 are also allowed. Claim 19 is allowed. Regarding independent claim 19, NPL1 (Flamingo: a Visual language Model for Few-Shot Learning, Jean-Baptiste Alayrac et al., arXiv, Nov 2021, Pages 1-54) hereafter NPL1 discloses “ A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform operations for training a visual language model to identify locations of image elements within a graphical image, the operations comprising: generating a plurality of training data items, each training data item including (i) a graphical image rendered according to a corresponding set of instructions (fig 1 shows the graphical images in the input prompt section) (ii) a natural language query for identifying at least one image element of the graphical image (fig 1 shows the text description (natural language query) in the input prompt identifying at least one image element of the graphical image (i.e “This is a Chinchilla--- This is a Shiba. They are very popular in Japan etc meeting the claim limitations), and (iii) a target location for the at least one image element, the target location being determined from the set of instructions (fig 1 shows the third image on the top with the text query “This is” in the input prompt section and the arrow to Completion “a flamingo. They are found in the Caribbean and South America (i.e a target location for the at least one image element in the third image in the input prompt, examiner notes that the specifics of a target location/place are not required by the current claim) meeting the claim limitations); for each of the training data items, processing the corresponding graphical image and natural language query using a visual language model to generate a corresponding model output comprising a predicted location of an image element identified from the natural language query (fig 1 shows the bottom row 4th image on the top (i.e the Graphical image) based on the context “This is cityscape. It looks like Chicago” and with the natural language query “what makes you think this is Chicago?” as the natural language query and the answer “I think it is Chicago because of the Shedd Aquarium in the background (i.e a predicted location for the at least one image element identified from the natural language query), examiner notes that the specifics of a predicted location/place are not required by the current claim) meeting the claim limitations). NPL1 and the other cited arts alone or in combination however fail to disclose “adjusting parameters of the visual language model to optimize, for each of the training data items, an objective function that depends on a comparison between the predicted location of the model output corresponding to the training data item and the target location of the training data item.”, therefore claim 19 is allowed. NOTE: Claims 3, 6, 9 and 20 will be allowed after overcoming the 35 U.S.C 112 and 35 U.S.C 101 rejections. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAYESH PATEL whose telephone number is (571)270-1227. The examiner can normally be reached IFW Mon-FRI. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAYESH A PATEL/Primary Examiner, Art Unit 2677 /JAYESH PATEL/ Primary Examiner Art Unit 2677
Read full office action

Prosecution Timeline

Aug 05, 2024
Application Filed
Jul 20, 2026
Non-Final Rejection mailed — §101, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705857
EXTRACTING FEATURES FROM SENSOR DATA
3y 0m to grant Granted Aug 11, 2026
Patent 12700129
FEATURE DETECTION AND LOCALIZATION FOR AUTONOMOUS SYSTEMS AND APPLICATIONS
3y 0m to grant Granted Aug 04, 2026
Patent 12694545
ENCODING IRREGULAR SHAPES USING ANGLE-BASED CONTOUR DESCRIPTORS
3y 1m to grant Granted Jul 28, 2026
Patent 12688610
CCD CAMERA CALIBRATION SYSTEM, METHOD, COMPUTING DEVICE AND STORAGE MEDIUM
2y 6m to grant Granted Jul 21, 2026
Patent 12675847
COMPUTER-IMPLEMENTED METHOD FOR OBTAINING A COMBINED IMAGE
3y 0m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
89%
With Interview (+5.0%)
2y 11m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 907 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month