Prosecution Insights
Last updated: October 04, 2026
Application No. 18/439,322

TILE-BASED IMAGE UNDERSTANDING IN VISION AND LANGUAGE MODELS

Non-Final OA §101§103
Filed
Feb 12, 2024
Examiner
WELLS, HEATH E
Art Unit
2664
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Non-Final)
80%
Grant Probability
Favorable
2-3
OA Rounds
6m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
84 granted / 105 resolved
+18.0% vs TC avg
Moderate +6% lift
Without
With
+6.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
24 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
13.9%
-26.1% vs TC avg
§103
70.1%
+30.1% vs TC avg
§102
4.0%
-36.0% vs TC avg
§112
8.8%
-31.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 105 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments The reply filed on 5 June 2026 has been entered. Applicant’s arguments with respect to claims 1-22 have been considered but are moot in view of new ground(s) of rejection. As the previously applied art was not prior art, this rejection is non-final. Priority Receipt is acknowledged that application is a National Stage application of PCT PCT/US25/15393. Priority to PCT/US25/15393 with a priority date of 11 February 2025 is acknowledged under 35 USC 119(e) and 37 CFR 1.78. Information Disclosure Statement The IDSs dated 12 February 2024 and 21 July 2025 that have been previously considered remain placed in the application file. Specification - Abstract The abstract has been amended. The objection to the Abstract is withdrawn. Claim Rejections - 35 USC § 101 Claims 1 and 19 have been amended. The rejection of claims 1-22 under 35 USC 101 as being directed to a mental process is withdrawn. 1st Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-8, 10-17 and 19-22 (all claims except 9 and 18) are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2017 0371898 A1, (Sharma et al.) in view of US Patent Publication 2024 0160853 A1, (Li et al.). The references are listed in a PTO-892 from the Office Action in which they are first used. If a reference is not identifiable (e.g., due to a typo), it can be identified by searching for the quoted text. If a reference is not identifiable (e.g., due to a typo), it can be identified by searching for the quoted text. Claim 1 Regarding Claim 1, Sharma et al. teach a method implemented by one or more processors ("methods that include the actions of receiving (i) a query image, and (ii) a user tap location; processing the received query image based on the user tap location; identifying one or more entities associated with the processed query image; and in response to receiving (i) the query image, and (ii) the user tap location, providing information about the identified one or more of the entities," paragraph [0004]), the method comprising: receiving an input query associated with a client device, the input query comprising an input image and an input text query ("The image processing module 240 may use the OCR engines to process a received query image by running one or more of the engines on the query image to detect one or more areas of text in the received query image, e.g., one or more lines of text," paragraph [0051]); generating, from the input image and the input query, a plurality of image tiles, wherein each image tile is a sub-image of the input image ("the recognition engine 250 may process the photograph 100 using a shallow neural network to identify one or more entities including "buildings," "bridge," "city" or "sky scraper." In addition, the recognition engine may process a processed query image including a cropped version of photograph 100," paragraph [0055]); providing, to one or more image analysis models, the plurality of image tiles ("the recognition engine 250 may process the photograph 100 using a shallow neural network to identify one or more entities including "buildings," "bridge," "city" or "sky scraper." In addition, the recognition engine may process a processed query image including a cropped version of photograph 100," paragraph [0055]); receiving, from the one or more image analysis models, a plurality of image facts, each image fact corresponding to a respective one or more of the image tiles ("the recognition engine 250 may identify the entity "The Gherkin" and can identify additional terms associated with "The Gherkin" such as "Norman foster," (architect) "Standard Life," (tenant) or "City of London" (location)," paragraph [0059] ); and causing the response to the input query to be rendered at the client device ("the example search results page 110 may be provided by a system in response to receiving and processing example query image 100 and user tap location 106," paragraph [0035]). Sharma et al. is not relied upon to explicitly teach all of sequence to sequence vision and language model. However, Li et al. teach generating, using a sequence-to-sequence vision and language model ("unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework," paragraph [0115]), a response to the input query based on the input query, the plurality of image tiles, and the plurality of image facts ("the multi-modal vision-language model may generate a text response to a text question accompanying an input image," paragraph [0025]), wherein generating, using the vision and language model, the response to the input query comprises inputting, sequentially or in parallel, the plurality of image tiles and respective image facts into the vision and language model ("although FIGS. 1-5 show a single input image 105a or 115, multiple images may be used as an input," paragraph [0057]); Therefore, taking the teachings of Sharma et al. and Li et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify “Visual Recognition using User Tap locations” as taught by Sharma et al. to use “Vision-Language Pretraining Framework” as taught by Li et al., showing that Sharma et al. and Li et al. are analogous art because both are Visual recognition systems. The suggestion/motivation for combination is that, “there is a need for training efficiency and expanded capabilities of vision-language models.” as noted by the Li et al. disclosure in paragraph [0006], which also motivates combination because the combination would predictably have a higher efficiency as there is a reasonable expectation that images can contain multiple elements and queries may only apply to some of the elements; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Claim 2 Regarding claim 2, Sharma et al. teach the method of claim 1, as noted above. Sharma et al. is not relied upon to explicitly teach all of search requests based on the input query, the plurality of image tiles and the plurality of image facts. However, Li et al. teach further comprising: generating, using the vision and language model or a further vision and language model, one or more search requests based on the input query, the plurality of image tiles, and the plurality of image facts ("A multi-modal self-attention mask may be applied to the set of queries (e.g., 106 in FIG. 2) and the text (e.g., 105b in FIG. 2)," paragraph [0088]); providing, to a search engine, the one or more search requests ("At step 901, an input image (e.g., 115 in FIG. 5) and an input utterance (e.g., 116 in FIG. 5) relating to the input image may be received from a communication interface. For example, the input utterance indicates an expected output text to generate based on visual content of the image, such as but not limited to a question relating to visual content of the input image, a guided request on what to generate about the image and/or the like," paragraph [0097]); and receiving, from the search engine, one or more search responses to the one or more search requests ("At step 911, a response is presented via the communication interface based on the decoded output text in response to the input utterance," paragraph [0103]), wherein generating, using the vision and language model, the response to the input query is further based on the one or more search responses ("At step 911, a response is presented via the communication interface based on the decoded output text in response to the input utterance," paragraph [0103]). Sharma et al. and Li et al. are combined as per claim 1. Claim 3 Regarding claim 3, Sharma et al. teach the method of claim 1, wherein generating, from the input image and the input query, the plurality of image tiles comprises: inputting, into the vision and language model, the input image and the input query ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016]); processing the input image and the input query using the vision and language model to generate the plurality of image tiles ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016] where the entities are image tiles); and outputting from the vision and language model, the plurality of image tiles ("In further implementations providing information about the identified one or more entities comprises providing a representative search query for output in response to receiving (i) the query image," paragraph [0018]). Claim 4 Regarding claim 4, Sharma et al. teach the method of claim 1, wherein generating, from the input image and the input query, the plurality of image tiles comprises generating, from the input image and the input query, the plurality of image tiles using an object detection model ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016] where a descriptor matching engine is an object detection model). Claim 5 Regarding claim 5, Sharma et al. teach the method of claim 1, wherein providing, to the one or more image analysis models, the plurality of image tiles comprises: determining, using the vision and language model, a classification for one or more of the plurality of image tiles ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016] where a descriptor matching engine is an object detection model where identify one or more entities is classification); and providing each of the one or more image tiles to a respective one or more image analysis models in a plurality of image analysis models based at least in part on the respective image classification of the tile ("In further implementations providing information about the identified one or more entities comprises providing a representative search query for output in response to receiving (i) the query image," paragraph [0018]). Claim 6 Regarding claim 6, Sharma et al. teach the method of claim 1, wherein providing, to the one or more image analysis models, the plurality of image tiles comprises providing the plurality of image tiles to the one or more image analysis models in parallel ("In certain circumstances, multitasking and parallel processing may be advantageous," paragraph [0108]). Claim 7 Regarding claim 7, Sharma et al. teach the method of claim 1, wherein the plurality of image tiles comprises a plurality of bounding boxes, each bounding box corresponding to a respective one or more objects in the input image ("defining a bounding box around each identified one or more entities that are associated with the processed query image," paragraph [0021]). Claim 8 Regarding claim 8, Sharma et al. teach the method of claim 7, wherein one or more of the image facts relate to one or more objects in a bounding box of the bounding boxes ("defining a bounding box around each identified one or more entities that are associated with the processed query image," paragraph [0021]). Claim 10 Regarding claim 10, Sharma et al. teach the method of claim 1, wherein generating, using the vision and language model, the response to the input query comprises sequentially inputting the plurality of image tiles and respective image facts into the vision and language model ("while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order," paragraph [0108]). Claim 11 Regarding claim 11, Sharma et al. teach the method of claim 1, wherein generating, using the vision and language model, the response to the input query comprises inputting the plurality of image tiles and respective image facts into the vision and language model in parallel ("In certain circumstances, multitasking and parallel processing may be advantageous," paragraph [0108]). Claim 12 Regarding claim 12, Sharma et al. teach the method of claim 1, wherein the method further comprises: generating, based on one or more of the image facts, one or more sub-tiles of an image tile ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016]); providing, to the one or more image analysis models, the one or more sub-tiles ("Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable sub-combination," paragraph [0107]); and receiving, from the one or more image analysis models, one or more further image facts, each further image fact corresponding to a respective one or more of the sub-tiles ("although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination," paragraph [0107]), wherein generating, using the vision and language model, the response to the input query is further based on the one or more further image facts ("the example search results page 110 may be provided by a system in response to receiving and processing example query image 100 and user tap location 106," paragraph [0035]). Claim 13 Regarding claim 13, Sharma et al. teach the method of claim 1: wherein generating, using a vision and language model, a response to the input query comprises: generating a natural language response to the input query ("The system can provide information about the identified one or more entities associated with the processed query image as output to the user," paragraph [0043]); and associating an element of the natural language response with a respective one or more portions of the input image, the respective one or more portions of the input image corresponding to portions of the input image relevant to the element of the natural language response ("For example, the system may provide one or more knowledge cards relating to the identified one or more entities, a search results page relating to one or more of the identified entities, or one or more representative search queries relating to the identified one or more entities. In some implementations the system may provide information about the identified one or more entities based at least on the contextualized classified one or more entities in the processed query image, as described above with reference to step 306," paragraph [0086]); wherein causing the response to the input query to be rendered at the client device comprises causing the natural language response to the input query to be rendered at the client device ("For example, the system may use the contextualized classified one or more entities to generate a contextualized knowledge card, search results page or representative search query for identified one or more entities, e.g., a knowledge card or search results page relating to the NBA league as opposed to a knowledge card or search results page relating to shopping for basket balls," paragraph [0086]); and wherein the method further comprises: receiving an indication that the element of the natural language response has been selected at the client device ("FIG. 4 presents an example process 400 for providing a representative search query for output in response to receiving a query image and user tap location. For example, the process 400 can be performed by the system 200 in response to receiving a query image and user tap location by a user 204 at user device 202," paragraph [0087]); and causing an indication of the respective one or more portions of the input image to be rendered at the client device ("The system identifies, for one or more identified entities associated with a processed query image, one or more candidate search queries that are pre-associated with the one or more entities (step 402)," paragraph [0088]). Claim 14 Regarding claim 14 Sharma et al. teach the method of claim 13, wherein the respective one or more portions of the input image comprise one or more image tiles ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016] where the entities are image tiles). Claim 15 Regarding claim 15, Sharma et al. teach the method of claim 14, wherein the indication of the respective one or more portions of the input image comprises one or more bounding boxes, each bounding box corresponding to a respective image tile ("defining a bounding box around each identified one or more entities that are associated with the processed query image," paragraph [0021]). Claim 16 Regarding claim 16, Sharma et al. teach the method of claim 13, wherein the method further comprises: receiving an indication that one or more of the respective one or more portions of the input image has been selected at the client device ("methods that include the actions of receiving (i) a query image, and (ii) a user tap location; processing the received query image based on the user tap location; identifying one or more entities associated with the processed query image; and in response to receiving (i) the query image, and (ii) the user tap location, providing information about the identified one or more of the entities," paragraph [0004]); and causing an indication of the element of the natural language response to be rendered at the client device ("the example search results page 110 may be provided by a system in response to receiving and processing example query image 100 and user tap location 106," paragraph [0035]). Claim 17 Regarding claim 17, Sharma et al. teach the method of claim 1: wherein generating, using a vision and language model, a response to the input query comprises: generating a natural language response to the input query ("The system can provide information about the identified one or more entities associated with the processed query image as output to the user," paragraph [0043]); and associating an element of the natural language response with a respective one or more portions of the input image, the respective one or more portions of the input image corresponding to portions of the input image relevant to the element of the natural language response ("For example, the system may provide one or more knowledge cards relating to the identified one or more entities, a search results page relating to one or more of the identified entities, or one or more representative search queries relating to the identified one or more entities. In some implementations the system may provide information about the identified one or more entities based at least on the contextualized classified one or more entities in the processed query image, as described above with reference to step 306," paragraph [0086]); wherein causing the response to the input query to be rendered at the client device comprises causing the natural language response to the input query to be rendered at the client device ("For example, the system may use the contextualized classified one or more entities to generate a contextualized knowledge card, search results page or representative search query for identified one or more entities, e.g., a knowledge card or search results page relating to the NBA league as opposed to a knowledge card or search results page relating to shopping for basket balls," paragraph [0086]); and wherein the method further comprises: receiving an indication that one or more of the respective one or more portions of the input image has been selected at the client device ("FIG. 4 presents an example process 400 for providing a representative search query for output in response to receiving a query image and user tap location. For example, the process 400 can be performed by the system 200 in response to receiving a query image and user tap location by a user 204 at user device 202," paragraph [0087]); and causing an indication of element of the natural language response to be rendered at the client device ("The system identifies, for one or more identified entities associated with a processed query image, one or more candidate search queries that are pre-associated with the one or more entities (step 402)," paragraph [0088]). Claim 19 Regarding claim 19, Sharma et al. teach a method implemented by one or more processors ("methods that include the actions of receiving (i) a query image, and (ii) a user tap location; processing the received query image based on the user tap location; identifying one or more entities associated with the processed query image; and in response to receiving (i) the query image, and (ii) the user tap location, providing information about the identified one or more of the entities," paragraph [0004]), the method comprising: receiving an input query associated with a client device, the input query comprising an input image and an input text query ("The image processing module 240 may use the OCR engines to process a received query image by running one or more of the engines on the query image to detect one or more areas of text in the received query image, e.g., one or more lines of text," paragraph [0051]); generating, from the input image and the input query, a plurality of image tiles, wherein each image tile is a sub-image of the input image ("the recognition engine 250 may process the photograph 100 using a shallow neural network to identify one or more entities including "buildings," "bridge," "city" or "sky scraper." In addition, the recognition engine may process a processed query image including a cropped version of photograph 100," paragraph [0055]); generating a natural language response to the input query based on the input query and the plurality of image tiles ("The system can provide information about the identified one or more entities associated with the processed query image as output to the user," paragraph [0043]); and associating an element of the natural language response with a respective one or more portions of the input image, the respective one or more portions of the input image corresponding to portions of the input image relevant to the element of the natural language response ("For example, the system may provide one or more knowledge cards relating to the identified one or more entities, a search results page relating to one or more of the identified entities, or one or more representative search queries relating to the identified one or more entities. In some implementations the system may provide information about the identified one or more entities based at least on the contextualized classified one or more entities in the processed query image, as described above with reference to step 306," paragraph [0086]), and associating the element of the natural language response comprising generating linking data indicating a link between the element of the natural language response and the respective one or more portions of the input image ("For example, the system may provide one or more knowledge cards relating to the identified one or more entities, a search results page relating to one or more of the identified entities, or one or more representative search queries relating to the identified one or more entities. In some implementations the system may provide information about the identified one or more entities based at least on the contextualized classified one or more entities in the processed query image, as described above with reference to step 306," paragraph [0086]); causing the natural language response to the input query to be rendered at the client device ("The system identifies, for one or more identified entities associated with a processed query image, one or more candidate search queries that are pre-associated with the one or more entities (step 402)," paragraph [0088]); receiving an indication that the element of the natural language response has been selected at the client device ("FIG. 4 presents an example process 400 for providing a representative search query for output in response to receiving a query image and user tap location. For example, the process 400 can be performed by the system 200 in response to receiving a query image and user tap location by a user 204 at user device 202," paragraph [0087]); in response to receiving the indication that the element of the natural language has been selected at the client device, and based on the generated linking data; causing an indication of the respective one or more portions of the input image to be rendered at the client device ("the example search results page 110 may be provided by a system in response to receiving and processing example query image 100 and user tap location 106," paragraph [0035]). Sharma et al. is not relied upon to explicitly teach all of based on the input query and the plurality of image tiles. However, Li et al. teach generating, using a vision and language model, a response to the input query based on the input query and the plurality of image tiles ("the multi-modal vision-language model may generate a text response to a text question accompanying an input image," paragraph [0025]), comprising: Sharma et al. and Li et al. are combined as per claim 1. Claim 20 Regarding claim 20, Sharma et al. teach the method of claim 19, wherein the respective one or more portions of the input image comprise one or more image tiles ("In some implementations identifying one or more entities associated with the processed query image comprises processing the processed query image using a descriptor matching engine to identify one or more entities," paragraph [0016] where the entities are image tiles). Claim 21 Regarding claim 21, Sharma et al. teach the method of claim 20, wherein the indication of the respective one or more portions of the input image comprises one or more bounding boxes, each bounding box corresponding to a respective image tile ("defining a bounding box around each identified one or more entities that are associated with the processed query image," paragraph [0021]). Claim 22 Regarding claim 22, Sharma et al. teach the method of claim 19, wherein the method further comprises: receiving an indication that one or more of the respective one or more portions of the input image has been selected at the client device ("methods that include the actions of receiving (i) a query image, and (ii) a user tap location; processing the received query image based on the user tap location; identifying one or more entities associated with the processed query image; and in response to receiving (i) the query image, and (ii) the user tap location, providing information about the identified one or more of the entities," paragraph [0004]); and causing an indication of the element of the natural language response to be rendered at the client device ("the example search results page 110 may be provided by a system in response to receiving and processing example query image 100 and user tap location 106," paragraph [0035]). 2nd Claim Rejections - 35 USC § 103 Claim 9 is rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2017 0371898 A1, (Sharma et al.) and US Patent Publication 2024 0160853 A1, (Li et al.) in view of US Patent Publication 2021 0263962 A1, (Chang et al.). The references are listed in a PTO-892 from the Office Action in which they are first used. If a reference is not identifiable (e.g., due to a typo), it can be identified by searching for the quoted text. If a reference is not identifiable (e.g., due to a typo), it can be identified by searching for the quoted text. Claim 9 Regarding Claim 9, Sharma et al. and Li et al. teach the method of claim 7, as noted above. Sharma et al. and Li et al. are not relied upon to explicitly teach all of rotated bounding boxes. However, Chang et al. teach wherein one or more of the bounding boxes are rotated bounded boxes ("In alternative implementations, the relative object position model 1004 can utilize the same separate relative position threshold area and rotate the image before determining if the relative position is satisfied. For example, for the relative object position operator of RELATIVE_POSITION (Object_l, Object_2, Right), the relative object position model 1004 can rotate the image 90° clockwise and apply the relative position threshold area 1110 shown in FIG 11," paragraph [0192]). Therefore, taking the teachings of Sharma et al., Li et al. and Chang et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify “Visual Recognition using User Tap locations” as taught by Sharma et al. and “Vision-Language Pretraining Framework” as taught by Li et al. to use “Utilizing Natural Language Processing and Multiple Object Detection Models to Automatically Select Objects in Images” as taught by Chang et al., showing that Sharma et al., Li et al. and Chang et al. are analogous art because all are Visual recognition systems. The suggestion/motivation for combination is that, “These, along with additional problems and issues exist in image editing systems with respect to detecting and selecting objects in digital images..” as noted by the Chang et al. disclosure in paragraph [0007], which also motivates combination because the combination would predictably have a higher efficiency as there is a reasonable expectation that images can contain multiple elements and queries may only apply to some of the elements; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. 3rd Claim Rejections - 35 USC § 103 Claim 18 is rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2017 0371898 A1, (Sharma et al.) and US Patent Publication 2024 0160853 A1, (Li et al.) in view of US Patent Publication 2024 0354336 A1, (Gopalkrishna et al.). The references are listed in a PTO-892 from the Office Action in which they are first used. If a reference is not identifiable (e.g., due to a typo), it can be identified by searching for the quoted text. If a reference is not identifiable (e.g., due to a typo), it can be identified by searching for the quoted text. Claim 18 Regarding Claim 18, Sharma et al. and Li et al. teach the method of claim 1, as noted above. Sharma et al. and Li et al. are not relied upon to explicitly teach all of wherein the input image at a second resolution that is lower than the first resolution. However, Gopalkrishna et al. teach wherein receiving the input query associated with a client device comprises: receiving an initial input image that is at a first resolution ("Block 204 describes the reception of the input query. This is the juncture where the system interfaces with users, accepting image submissions for semantic analysis," paragraph [0042]); and generating the input image from the initial input image, wherein the input image is at a second resolution that is lower than the first resolution ("The preprocessing of these images involves normalization, resolution adjustment, and possibly feature enhancement to optimize them for encoding," paragraph [0042] where resolution adjustment includes lowering resolution). Therefore, taking the teachings of Sharma et al., Li et al. and Gopalkrishna et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify “Visual Recognition using User Tap locations” as taught by Sharma et al. and “Vision-Language Pretraining Framework” as taught by Li et al., to use “Multimodal Semantic Analysis and Image Retrieval” as taught by Gopalkrishna et al. showing that Sharma et al., Li et al. and Gopalkrishna et al. are analogous art because all are Visual recognition systems. The suggestion/motivation for combination is that, “The limitations of these conventional systems underscore the need for more advanced, intuitive, and context-aware semantic search technologies capable of bridging the gap between simple keyword or visual pattern matching and the rich, nuanced understanding required for today's diverse and dynamic data landscapes.” as noted by the Gopalkrishna et al. disclosure in paragraph [0004], which also motivates combination because the combination would predictably have a higher efficiency as there is a reasonable expectation that images can contain multiple elements and queries may only apply to some of the elements; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Reference Cited The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. US Patent Publication 2011 0125735 A1 to Petrou discloses processing a visual query by sending it to a plurality of parallel search systems, each implementing a distinct visual query search process. These parallel search systems may include but are not limited to optical character recognition (OCR), facial recognition, product recognition, bar code recognition, object-or-object-category recognition, named entity recognition, and color recognition. Then at least one search result is sent to the client system. Non Patent Publication “A tale of two interfaces: vitrivr at the lifelog search challenge” to Heller et al. discloses two systems for lifelog retrieval: vitrivr and vitrivr-VR, which share a common retrieval model and backend for multi-modal multimedia retrieval. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HEATH E WELLS whose telephone number is (703)756-4696. The examiner can normally be reached Monday-Friday 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-7882. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Heath E. Wells/Examiner, Art Unit 2664 Date: 8 September 2026
Read full office action

Prosecution Timeline

Feb 12, 2024
Application Filed
Feb 05, 2026
Non-Final Rejection mailed — §101, §103
May 05, 2026
Examiner Interview Summary
May 05, 2026
Applicant Interview (Telephonic)
Jun 05, 2026
Response Filed
Sep 11, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743867
IMPROVED LUMINANCE ADJUSTMENT BETWEEN DIFFERENT RESOLUTIONS
3y 3m to grant Granted Sep 22, 2026
Patent 12739416
TECHNIQUES FOR JOINTLY TRAINING A DOWNSCALER AND AN UPSCALER FOR VIDEO STREAMING
3y 10m to grant Granted Sep 15, 2026
Patent 12737876
CRANKSHAFT SHAPE INSPECTION METHOD, ARITHMETIC UNIT, PROGRAM, AND SHAPE INSPECTION APPARATUS
3y 1m to grant Granted Sep 15, 2026
Patent 12738033
SELF-SUPERVISED TRAINING AT SCALE WITH WEAKLY-SUPERVISED LATENT SPACE STRUCTURE
2y 11m to grant Granted Sep 15, 2026
Patent 12737869
SEMICONDUCTOR YIELD PREDICTION METHOD AND APPARATUS
2y 11m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
80%
Grant Probability
86%
With Interview (+6.2%)
3y 2m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 105 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month