Prosecution Insights
Last updated: August 18, 2026
Application No. 18/959,367

System For Extracting Relevant Passages As Context For Multimodal Queries

Final Rejection §103
Filed
Nov 25, 2024
Examiner
DAUD, ABDULLAH AHMED
Art Unit
2164
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
2 (Final)
55%
Grant Probability
Moderate
3-4
OA Rounds
2y 0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 55% of resolved cases
55%
Career Allowance Rate
96 granted / 174 resolved
At TC average
Strong +31% interview lift
Without
With
+31.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
23 currently pending
Career history
209
Total Applications
across all art units

Statute-Specific Performance

§101
13.6%
-26.4% vs TC avg
§103
73.5%
+33.5% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
7.5%
-32.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 174 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement IDS submitted on 3/23/2026 has been considered by the examiner. Response to Amendment This Office action is in response to Applicant's amendment filed on 2/19/2026. Claim 1-20 are pending. Claim 1, 19 and 20 are amended. Claim 1-20 are rejected. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 1-2, 4-5, 7-17 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu, Xinliang et al(US Patent No. 12210516), hereafter referred as to “Zhu”, in view of Popat, Dalsukhbhai et al (PGPUB Document No. 20230342411), hereafter, referred to as “Popat”, in further view of KI, Sehwan et al (PGPUB Document No. 20250157213), hereafter, referred to as “KI”. Regarding Claim 1 (Currently Amened), Zhu teaches A computer-implemented method for processing multimodal input queries, the method comprising: receiving, by a computing system, a multimodal input query, wherein the multimodal input query comprises image content(Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)); receiving, by the computing system, a plurality of search results from a search engine based on the multimodal input query(Zhu, Fig. 5 and col 9:40-22 disclose multimodal search system generating results (element 514 of Fig. 5) “The combined vector is then processed against the MIM index 112, which may be pre-generated offline, and leads to search result 514. As shown, these results include a number of stuffed bears, which the user may browse through for selection to purchase, among other options”); But Zhu does not explicitly teach processing, by the computing system, the multimodal input query and the plurality of search results with a passage-scoring model to generate a result score for each respective search result in the plurality of search results, wherein the passage-scoring model comprises a machine-learned multimodal model configured to simultaneously process both the image content from the multimodal input query and textual content from each respective search result to generate the result score for each respective search result; selecting, by the computing system, a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results; generating, by the computing system, a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing, by the computing system, the model input with a response generation model to generate a model output based on the model input, wherein the model output comprises a natural language response to the multimodal input query; and transmitting, by the computing system, the natural language response to the multimodal input query for display at a user computing device. However, in the same field of endeavor of paragraph/passage ranking Popat teaches processing, by the computing system, the multimodal input query and the plurality of search results (Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)) with a passage-scoring model to generate a result score for each respective search result in the plurality of search results(Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”), wherein the passage-scoring model comprises a machine-learned multimodal model configured to simultaneously process both the image content from the multimodal input query (Zhu, Fig. 5 discloses multimodal search system generating results (element 514 of Fig. 5) by simultaneously process text and image query))and textual content from each respective search result to generate the result score for each respective search result (Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”) selecting, by the computing system, a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results(Popat, para 0060 discloses top rated passages are getting selected for subsequent input (subset) ”in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1))”); generating, by the computing system, a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing, by the computing system, the model input with a response generation model to generate a model output based on the model input (Popat, Fig. 4 and para 0060-0061 disclose inputting query and selected result set to a machine learning model for output generation “input into the prediction engine to a form appropriate for input into a machine learning engine such as a neural network. For example, in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1)), a title of the top-rated passage”; where Popat further in para 0060 further discloses that inputted search result sub-set is being used as context to the query” the inputs may include additional titles (e.g., corresponding to the additional passages). The additional passages are referred to as context passages”); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of scoring passages of documents of Popat into the feature of content retrieval by multimodal queries of Zhu to produce an expected result of extracting relevant contents. The modification would be obvious because one of ordinary skill in the art would be motivated to improve the accuracy of the retrieved contents using a relevance score prediction engine(Popat, para 0003). But Zhu and Popat don’t explicitly teach wherein the model output comprises a natural language response to the multimodal input query; and transmitting, by the computing system, the natural language response to the multimodal input query for display at a user computing device. However, in the same field of endeavor of input processing KI teaches wherein the model output comprises a natural language response to the multimodal input query(KI, 0058 discloses output result in natural language “the IQA apparatus may output the IQA result based on an actual natural language rather than the IQA score 135……” ); and transmitting, by the computing system, the natural language response to the multimodal input query for display at a user computing device (KI, Fig. 10 and para 0118 disclose transmitting output to display device of the system “the output device 1070 may be an output interface or a display device. For example, when the output device 1070 is a display, the output device 1070 may display the assessment score calculated by the processor 1050 on a screen in response to the input image. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of incorporate the feature of providing output in natural language KI into the feature of content retrieval by multimodal queries of Zhu and Popat to produce an expected result of providing results in natural language. The modification would be obvious because one of ordinary skill in the art would be motivated to display output in user friendly way using natural language(KI, para 0058). Claim 2(Original), Zhu, Popat and KI teach all the limitations of 1 and Zhu further teaches wherein the multimodal input query includes textual content or speech content(Zhu, col 2:46-52 disclose multimodal input such as image, text or speech or audio “The inputs may include a textual input (e.g., a search string typed in by a user), an audio input (which may be converted to a textual input using one or more natural language processing systems or may be processed as an auditory input), an image input, a video input, or a combination thereof, such as text-based input with an accompanying audio or image sample”). Claim 4(Original), Zhu, Popat and KI teach all the limitations of 1 and Popat further teaches wherein the plurality of search results are ranked and wherein providing, by the computing system, the plurality of search results to a passage-scoring model to generate a result score for each respective search result in the plurality of search results further comprises (Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”): selecting, by the computing system, a predetermined number of search results to provide to the passage-scoring model, wherein the search results are selected, at least in part, based on their ranking(Popat, para 0046-0047 disclose only top-rated passages are fed into a model for accuracy scores “when there is more than one top-ranked passage, the accuracy score manager 240 selects the top-ranked passage with the highest accuracy score”). Regarding Claim 5(Original), Zhu, Popat and KI teach all the limitations of 4 and Zhu further teaches wherein the plurality of search systems comprise one or more of: an image search system, a multimodal search system, and a text-based search system (Zhu, Fig. 6C discloses a multimodal search system). Regarding Claim 7(Original), Zhu, Popat and KI teach all the limitations of 5 and Zhu further teaches wherein the multimodal search system is configured to: generate a query image embedding and a query text embedding based on the multimodal input query(Zhu, element 510 (“Vector Generator”) of Fig. 5 discloses embeddings/vectorization and in response to multimodal query); access a database of embedded multimodal documents; generate a similarity score for each embedded multimodal document in the database of embedded images based on a calculated similarity to the query embedding(Zhu, Fig. 2 and col 5:64-66 ~col 6: 1-3 disclose index database having vectorized/embedded product catalog/document for matching with query embeddings “the source catalog 202 may be processed to populate the MIM index 112 with different combined feature vectors for given items within the source catalog 202. ……. the MIM index 112 will be prepared for execution when a search request is received”); and select a plurality of embedded multimodal documents based on the similarity scores to return as search results(Zhu, col 7:64-67~col 8:1-4 further teaches displaying matched documents/catalog based on relevance in embedding/vector matching “The results can be used to provide relevant information (e.g., title and description) for products and services that are determined to be at least somewhat relevant to the search query. The relevant information can be pulled from a product catalog 414, for example, and combined with web content from a web content repository 416 or other such location, in order to provide a webpage with search results to return to the client device 402 for display”). Regarding Claim 8(Original), Zhu, Popat and KI teach all the limitations of 5 and Zhu don’t explicitly teach wherein the text-based search system is configured to: generate a textual representation of an image included in the multimodal input query(Zhu, Fig. 6C discloses textual representation of an image in response to multimodal query); generate a query text embedding based on the textual representation of the image and a textual portion of the multimodal input query(Zhu, element 510 (“Vector Generator”) of Fig. 5 discloses embeddings/vectorization and in response to multimodal query); access a database of embedded documents; generate a similarity score for each embedded document in the database of embedded images based on a calculated similarity to the query text embedding(Zhu, Fig. 2 and col 5:64-66 ~col 6: 1-3 disclose index database having vectorized/embedded product catalog/document for matching with query embeddings “the source catalog 202 may be processed to populate the MIM index 112 with different combined feature vectors for given items within the source catalog 202. ……. the MIM index 112 will be prepared for execution when a search request is received”); and select a plurality of embedded documents based on the similarity scores to return as search results(Zhu, col 7:64-67~col 8:1-4 further teaches displaying matched documents/catalog based on relevance in embedding/vector matching “The results can be used to provide relevant information (e.g., title and description) for products and services that are determined to be at least somewhat relevant to the search query. The relevant information can be pulled from a product catalog 414, for example, and combined with web content from a web content repository 416 or other such location, in order to provide a webpage with search results to return to the client device 402 for display”). Regarding Claim 9(Original), Zhu, Popat and KI teach all the limitations of 8 and Zhu further teaches wherein generating a textual representation of an image included in the multimodal input query comprises: providing, by the computing system, the image to a description generation for processing; and receiving, by the computing system, a model output from the description generation based on the image (Zhu, Fig. 6C discloses textual representation of an image(m) in response to multimodal query)). Regarding Claim 10(Original), Zhu, Popat and KI teach all the limitations of 1 and KI further teaches wherein the passage-scoring model is a large vision language model (KI, 0024 discloses vision language model can be used for scoring contents “calculate an assessment score of image quality corresponding to the input image by inputting the input image to the pre-trained target encoder, wherein the target encoder is pre-trained to simulate a visual-language model (VLM) based on data obtained by applying, to the VLM, a text prompt corresponding to representation related to image quality of an image” ). Regarding Claim 11(Original), Zhu, Popat and KI teach all the limitations of 1 and KI further teaches wherein the model output comprises a natural language response to the input query (KI, 0058 discloses output result in natural language “the IQA apparatus may output the IQA result based on an actual natural language rather than the IQA score 135……” ). Claim 12(Original), Zhu, Popat and KI teach all the limitations of 1 and Popat further teaches wherein the model input includes citation data for each search result in the subset of search results (Popat, element 180 of Fig. 1B and para 0024-0025 disclose citation(url) as input or data resource ” a resource 105 is data provided over the network 102 and that is associated with a resource address, e.g., a uniform resource locator (URL)”). Claim 13(Original), Zhu, Popat and KI teach all the limitations of 1 and Popat further teaches wherein the model output comprises citation data for each search result in the subset of search results provided to the response generation model (Popat, element 180 of Fig. 1B and para 0030 disclose citation(url) in the output results” An example search result 112 can include a web page title, a snippet of text or a portion of an image extracted from the web page, and the URL of the web page”). Claim 14(Original), Zhu, Popat and KI teach all the limitations of 1 and Popat further teaches wherein the model output is displayed on a page of search results (Zhu, Fig. 1B discloses model output on a webpage “The relevant information can be pulled from a product catalog 414, for example, and combined with web content from a web content repository 416 or other such location, in order to provide a webpage with search results to return to the client device 402 for display”). Claim 15(Original), Zhu, Popat and KI teach all the limitations of 1 and Zhu further teaches wherein the search results are multimodal (Zhu, Fig. 6c disclose multimodal search results). Regarding Claim 16(Original), Zhu, Popat and KI teach all the limitations of 1 and Popat further teaches wherein selecting, by the computing system, a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results further comprises: for a respective search result in the plurality of search results(Popat, para 0060 discloses top rated passages are getting selected for subsequent input (subset) ”in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1))”): segmenting, by the computing system, the respective search results into one or more passages; determining, by the computing system, a relevance score for each passage in the one or more passages; and adding, by the computing system, one or more passages to the subset of search results based on the relevance score for each passage(Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”; para 0029 discloses selection based on relevance scoring “In response to receiving a search query 109, the search system 120 accesses the search index 122 to identify resources 105 that are relevant to, e.g., have at least a minimum specified relevance score for, the search query 109”). Regarding claim 17(Original), Zhu, Popat and KI teach all the limitations of 16 and Popat further teaches wherein a number of search results in the subset of search results is determined, at least in part, based on a size limit for input to the response generation model(Popat, claim 12 discloses input size limit “wherein a number of search results in the subset of search results is determined, at least in part, based on a size limit for input to the response generation model.” ). Regarding Claim 18 (Original), Zhu, Popat and Ki teach all the limitations of 1 and Ki further teaches wherein the response generation model is a large vision language model. However, in the same field of endeavor of multimodal input processing KI teaches wherein the response generation model is a large vision language model (KI, 0024 discloses vision language model being used for output contents by scoring “calculate an assessment score of image quality corresponding to the input image by inputting the input image to the pre-trained target encoder, wherein the target encoder is pre-trained to simulate a visual-language model (VLM) based on data obtained by applying, to the VLM, a text prompt corresponding to representation related to image quality of an image” ). Regarding Claim 19 (Currently Amened), Zhu teaches A computing system, comprising: one or more processors; and one or more non-transitory computer-readable media that store instructions wherein, when executed by the one or more processors, the instructions cause the one or more processors to perform operations(Zhu, Fig. 11 discloses a computing system to perform multimodal queries having memory storages), the operations comprising: receiving a multimodal input query, wherein the multimodal input query comprises image content(Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)); receiving a plurality of search results from a search engine based on the multimodal input query(Zhu, Fig. 5 and col 9:40-22 disclose multimodal search system generating results (element 514 of Fig. 5) “The combined vector is then processed against the MIM index 112, which may be pre-generated offline, and leads to search result 514. As shown, these results include a number of stuffed bears, which the user may browse through for selection to purchase, among other options”); But Zhu does not explicitly teach processing the multimodal input query and the plurality of search results with a passage-scoring model to generate a result score for each respective search result in the plurality of search results, wherein the passage-scoring model comprises a machine-learned multimodal model configured to simultaneously process both the image content from the multimodal input query and textual content from each respective search result to generate the result score for each respective search result; selecting a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results; generating a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing the model input with a response generation model to generate a model output based on the model input, wherein the model output comprises a natural language response to the multimodal input query; and transmitting the natural language response to the multimodal input query. However, in the same field of endeavor of paragraph/passage ranking Popat teaches processing the multimodal input query and the plurality of search results(Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)) with a passage-scoring model to generate a result score for each respective search result in the plurality of search results (Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”), wherein the passage-scoring model comprises a machine-learned multimodal model configured to simultaneously process both the image content from the multimodal input query (Zhu, Fig. 5 discloses multimodal search system generating results (element 514 of Fig. 5) by simultaneously process text and image query))and textual content from each respective search result to generate the result score for each respective search result (Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”); selecting a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results (Popat, para 0060 discloses top rated passages are getting selected for subsequent input (subset) ”in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1))”); generating a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing the model input with a response generation model to generate a model output based on the model input (Popat, Fig. 4 and para 0060-0061 disclose inputting query and selected result set to a machine learning model for output generation “input into the prediction engine to a form appropriate for input into a machine learning engine such as a neural network. For example, in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1)), a title of the top-rated passage”; where Popat further in para 0060 further discloses that inputted search result sub-set is being used as context to the query” the inputs may include additional titles (e.g., corresponding to the additional passages). The additional passages are referred to as context passages”); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of scoring passages of documents of Popat into the feature of content retrieval by multimodal queries of Zhu to produce an expected result of extracting relevant contents. The modification would be obvious because one of ordinary skill in the art would be motivated to improve the accuracy of the retrieved contents using a relevance score prediction engine(Popat, para 0003). But Zhu and Popat don’t explicitly teach wherein the model output comprises a natural language response to the multimodal input query; and transmitting the natural language response to the multimodal input query. However, in the same field of endeavor of input processing KI teaches wherein the model output comprises a natural language response to the multimodal input query(KI, 0058 discloses output result in natural language “the IQA apparatus may output the IQA result based on an actual natural language rather than the IQA score 135……” ); and transmitting the natural language response to the multimodal input query (KI, Fig. 10 and para 0118 disclose transmitting output to display device of the system “the output device 1070 may be an output interface or a display device. For example, when the output device 1070 is a display, the output device 1070 may display the assessment score calculated by the processor 1050 on a screen in response to the input image”; where Zhu teaches multi-modal input query). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of incorporate the feature of providing output in natural language KI into the feature of content retrieval by multimodal queries of Zhu and Popat to produce an expected result of providing results in natural language. The modification would be obvious because one of ordinary skill in the art would be motivated to display output in user friendly way using natural language(KI, para 0058). Regarding Claim 20 (Currently Amended), Zhu teaches One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations(Zhu, Fig. 11 discloses a computing system to perform multimodal queries having memory storages), receiving a multimodal input query, wherein the multimodal input query comprises image content(Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)); receiving a plurality of search results from a search engine based on the multimodal input query(Zhu, Fig. 5 and col 9:40-22 disclose multimodal search system generating results (element 514 of Fig. 5) “The combined vector is then processed against the MIM index 112, which may be pre-generated offline, and leads to search result 514. As shown, these results include a number of stuffed bears, which the user may browse through for selection to purchase, among other options”); But Zhu does not explicitly teach processing the multimodal input query and the plurality of search results with a passage-scoring model to generate a result score for each respective search result in the plurality of search results, wherein the passage-scoring model comprises a machine-learned multimodal model configured to simultaneously process both the image content from the multimodal input query and textual content from each respective search result to generate the result score for each respective search result; selecting a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results; generating a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing the model input with a response generation model to generate a model output based on the model input, wherein the model output comprises a natural language response to the multimodal input query; and transmitting the natural language response to the multimodal input query. However, in the same field of endeavor of paragraph/passage ranking Popat teaches processing the multimodal input query and the plurality of search results(Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)) with a passage-scoring model to generate a result score for each respective search result in the plurality of search results (Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”), wherein the passage-scoring model comprises a machine-learned multimodal model configured to simultaneously process both the image content from the multimodal input query (Zhu, Fig. 5 discloses multimodal search system generating results (element 514 of Fig. 5) by simultaneously process text and image query))and textual content from each respective search result to generate the result score for each respective search result (Popat, para 0039-0040 disclose search result obtains passages (element 235(1)) its corresponding ranking/scoring (element 236(1) of Fig. 2)of passages “234(N) includes a link to a web site that includes respective passages 235(1), . . . , 235(N) and a respective ranking 236(1), . . . , 236(N)”); selecting a subset of search results from the plurality of search results based on based on the result score for each respective search result in the plurality of search results (Popat, para 0060 discloses top rated passages are getting selected for subsequent input (subset) ”in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1))”); generating a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing the model input with a response generation model to generate a model output based on the model input (Popat, Fig. 4 and para 0060-0061 disclose inputting query and selected result set to a machine learning model for output generation “input into the prediction engine to a form appropriate for input into a machine learning engine such as a neural network. For example, in some implementations the inputs into the accuracy score prediction engine include a query (e.g., query data 231), a top-rated passage (e.g., passage 235(1)), a title of the top-rated passage” where Popat further in para 0060 further discloses that inputted search result sub-set is being used as context to the query” the inputs may include additional titles (e.g., corresponding to the additional passages). The additional passages are referred to as context passages”); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of scoring passages of documents of Popat into the feature of content retrieval by multimodal queries of Zhu to produce an expected result of extracting relevant contents. The modification would be obvious because one of ordinary skill in the art would be motivated to improve the accuracy of the retrieved contents using a relevance score prediction engine(Popat, para 0003). But Zhu and Popat don’t explicitly teach wherein the model output comprises a natural language response to the multimodal input query; and transmitting the natural language response to the multimodal input query. However, in the same field of endeavor of input processing KI teaches wherein the model output comprises a natural language response to the multimodal input query (KI, 0058 discloses output result in natural language “the IQA apparatus may output the IQA result based on an actual natural language rather than the IQA score 135……” ); and transmitting the natural language response to the multimodal input query (KI, Fig. 10 and para 0118 disclose transmitting output to display device of the system “the output device 1070 may be an output interface or a display device. For example, when the output device 1070 is a display, the output device 1070 may display the assessment score calculated by the processor 1050 on a screen in response to the input image”; where Zhu teaches multi-modal input query). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of incorporate the feature of providing output in natural language KI into the feature of content retrieval by multimodal queries of Zhu and Popat to produce an expected result of providing results in natural language. The modification would be obvious because one of ordinary skill in the art would be motivated to display output in user friendly way using natural language(KI, para 0058). Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Zhu, Xinliang et al(US Patent No. 12210516), hereafter referred as to “Zhu”, in view of Popat, Dalsukhbhai et al (PGPUB Document No. 20230342411), hereafter, referred to as “Popat”, in view of KI, Sehwan et al (PGPUB Document No. 20250157213), hereafter, referred to as “KI”, in further view of Kang, Woo-shik et al (PGPUB Document No. 20190392009), hereafter, referred to as “Kang”. Regarding Claim 3 (Origtnal), Zhu, Popat and Ki teach all the limitations of 2 and Zhu further teaches wherein receiving, by the computing system, a plurality of search results from a search engine based on the multimodal input query further comprises (Zhu, Fig. 5 discloses a multimodal search system that takes multimodal query input such as image (element 502 for Fig. 5) and text (element 504 for Fig. 5)): But Zhu, Popat and Ki don’t explicitly teach providing, by the computing system, multimodal input query to a plurality of search systems; receiving, by the computing system, preliminary search results from each search system in the plurality of search systems; and combining, by the computing system, the preliminary search results into the plurality of search results. However, in the same field of endeavor of multimodal query processing Kang teaches providing, by the computing system, multimodal input query to a plurality of search systems; receiving, by the computing system, preliminary search results from each search system in the plurality of search systems(Kang, Fig. 66 and para 0242-0243 disclose a multimodal query systems that runs different type of query (element 6620) by different query processors and combining the results obtained from different query processors “The combination query 6610 may be processed by a plurality of the search methods 6620, thereby acquiring a search result” ); and combining, by the computing system, the preliminary search results into the plurality of search results(Kang, Fig. 66 and para 0245 disclose combining the results obtained from different query processors “A search method mergence module 6650 may merge search results obtained by the plurality of search methods 6620” ). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of processing multimodal plurality of queries into different search system and merging of search results of Kang into the feature of content retrieval by multimodal queries of Zhu, Popat and Ki to produce an expected result of extracting relevant contents. The modification would be obvious because one of ordinary skill in the art would be motivated to improve the user search query by clarifying the search context using multimodal input(Kang, para 0240). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Zhu, Xinliang et al(US Patent No. 12210516), hereafter referred as to “Zhu”, in view of Popat, Dalsukhbhai et al (PGPUB Document No. 20230342411), hereafter, referred to as “Popat”, in view of KI, Sehwan et al (PGPUB Document No. 20250157213), hereafter, referred to as “KI”, in further view of Wang, Zijia et al (PGPUB Document No. 20250335497), hereafter, referred to as “Wang”. Regarding Claim 6 (Original), Zhu, Popat and Ki teach all the limitations of 5 and Zhu further teaches wherein image search system is configured to: generate, by the image search system, an image query embedding based on an image content included in the multimodal input query(Zhu, element 506 (“Image Feature Vector”) of Fig. 5 discloses embeddings/vectorization of query image in response to multimodal query); But Zhu, Popat and Ki don’t explicitly teach access, by the image search system, a plurality of embedded images in an image database; generate, by the image search system, a similarity score for each embedded image in plurality of embedded images based on a calculated similarity to the image query embedding; and select, by the image search system, one or more search results to return based on the similarity scores for the plurality of embedded images. However, in the same field of endeavor of image query processing Wang teaches access, by the image search system, a plurality of embedded images in an image database; generate, by the image search system, a similarity score for each embedded image in plurality of embedded images based on a calculated similarity to the image query embedding(Wang, para 0023 discloses an embedded image databased “The method further includes encoding the representation as an image vector in a high-dimensional vector space and storing it into an image vector database. When retrieval is performed, the method may further include receiving a query that includes text information and that is for the image vector database, and determining, from the image vector database, an image associated with the text information”; para 0057 image vector similarities are measured by similarity matching scores ); and select, by the image search system, one or more search results to return based on the similarity scores for the plurality of embedded images(Wang, para 0057 discloses retrieving result by matching scores of vectors “the image vectors are sorted in an order of similarity scores of respective image vectors, and the image vector that best matches (the highest matching degree (or similarity)) the query vector Q.sub.v is retrieved. The image corresponding to the best matching image vector is used as a query result corresponding to the query Q”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the feature of retrieving images by its similarity matching scores of Wang into the feature of content retrieval by multimodal queries of Zhu, Popat and Ki to produce an expected result of extracting relevant contents. The modification would be obvious because one of ordinary skill in the art would be motivated to improve accuracy and efficiency of image retrieval in a search system(Wang, abstract). Response to Arguments I. 35 U.S.C §101 In light of applicant argument consideration, §101 abstract idea rejection to claim 1-20 has been withdrawn. II. 35 U.S.C §103 The crux of the applicant arguments regarding independent claim 1, 19 and 20 is “Applicant respectfully submits that Zhu and Popat, alone or in combination, do not describe, disclose, or suggest all the limitations of amended independent claim 1, even when taken in view of those of ordinary skill in the art. Specifically, Applicant respectfully submits that the cited art fails to disclose or suggest "generating, by the computing system, a model input comprising the multimodal input query and the selected subset of search results as context for responding to the multimodal input query; processing, by the computing system, the model input with a response generation model to generate a model output based on the model input, wherein the model output comprises a natural language response to the multimodal input query; and transmitting, by the computing system, the natural language response to the multimodal input query for display at a user computing device" as recited by amended independent claim 1.”. Applicant’s above mentioned arguments have been fully considered but the examiner respectfully disagrees for following reasons; Firstly, prior art Popat in paragraph 0060-0061 disclose inputting query and selected result set to a machine learning model for output generation. Secondly, Popat further in same paragraph 0060 discloses that inputted search result sub-set is being used as context to the query as following ”the inputs may include additional titles (e.g., corresponding to the additional passages). The additional passages are referred to as context passages”. For the argued last two limitations, new prior art KI in paragraph 58 discloses outputting result in natural language as following “the IQA apparatus may output the IQA result based on an actual natural language rather than the IQA score 135……” and further, fig. 10 and paragraph 0118 disclose transmitting natural language response to user device. Therefore, the examiner maintains the prior art rejection to claim 1-20. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDULLAH A DAUD whose telephone number is (469)295-9283. The examiner can normally be reached M~F: 9:30 am~6:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amy Ng can be reached at 571-270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ABDULLAH A DAUD/Examiner, Art Unit 2164 /AMY NG/Supervisory Patent Examiner, Art Unit 2164
Read full office action

Prosecution Timeline

Nov 25, 2024
Application Filed
Nov 19, 2025
Non-Final Rejection mailed — §103
Feb 19, 2026
Examiner Interview Summary
Feb 19, 2026
Response Filed
Feb 19, 2026
Applicant Interview (Telephonic)
Jun 10, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12602292
TENANT COPY USING INCREMENTAL DATABASE RECOVERY
2y 9m to grant Granted Apr 14, 2026
Patent 12566809
GRAPH LEARNING AND AUTOMATED BEHAVIOR COORDINATION PLATFORM
3y 11m to grant Granted Mar 03, 2026
Patent 12487887
FILESET PARTITIONING FOR DATA STORAGE AND MANAGEMENT
2y 10m to grant Granted Dec 02, 2025
Patent 12299037
GRAPH-BASED FEATURE ENGINEERING FOR MACHINE LEARNING MODELS
2y 9m to grant Granted May 13, 2025
Patent 12293262
ADAPTIVE MACHINE LEARNING TRAINING VIA IN-FLIGHT FEATURE MODIFICATION
5y 7m to grant Granted May 06, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
55%
Grant Probability
86%
With Interview (+31.0%)
3y 9m (~2y 0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 174 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month