Prosecution Insights
Last updated: October 02, 2026
Application No. 18/674,137

AESTHETIC IMAGE RETRIEVAL SYSTEM AND METHOD

Final Rejection §101§102§103
Filed
May 24, 2024
Examiner
SORRIN, AARON JOSEPH
Art Unit
2672
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
57 granted / 75 resolved
+14.0% vs TC avg
Strong +42% interview lift
Without
With
+42.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
32 currently pending
Career history
102
Total Applications
across all art units

Statute-Specific Performance

§101
20.0%
-20.0% vs TC avg
§103
37.1%
-2.9% vs TC avg
§102
14.1%
-25.9% vs TC avg
§112
28.0%
-12.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 75 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Claim objections and rejections under 35 USC 112(b) are withdrawn. Applicant's arguments regarding rejections under 35 USC 101 have been fully considered but they are not persuasive. On Pages 9-13, Applicant argues that the amended claim 1 provides an improvement in the technical field of visual content retrieval systems. The amended claim 1 does not disclose the alleged improvements described in Paragraph 19, which is cited by the Applicant as describing the improvement in the technical field. For example, claim 1 does not disclose the ANN search technique “to enable fast and efficient content selection”. Claim 1 further does not disclose the offline indexing system, the encoder, or the embeddings. While some of these features are described in dependent claims, they are highly generic in the technical field of visual content retrieval, failing to amount to an improvement. Applicant's arguments regarding rejections under 35 USC 102 and 35 USC 103 have been fully considered but they are not persuasive. On Pages 13-14 of Remarks, Applicant argues that Son does not each amended claim 1 for the following reasons: “Son teaches that the system is designed to help users express their search intent. (See, e.g., Abstract). Search intent is identified by generating visuals based on the user's search. The user selects visuals which reflect the user's intent. The suggested search queries can then be refined based on the selections. (See, e.g., Section 4.1.2). However, Son does not disclose or suggest a system that uses a model to analyze the text input by a user to determine search intent. Further, Son also does not disclose a system that automatically modifies a user query. As noted above, Son discloses a system that generates suggestions for a modified prompt which the user can then select as a prompt to use for additional searches.” First, Son does explicitly disclose ‘a system that uses a model to analyze the text input by a user to determine search intent’. As cited in the Non-Final Rejection, Section 4.2.1. recites, “GenQuery provides five concretized queries based on the text query entered by the user in Figure 3 (a), about one second later. To ensure a smooth user experience with GenQuery, it was critical to deliver suggestions at near-real time speed; hence, we used the gpt-3.5-turbo-16k-0613 LLM model. To deliver concretized queries as output, we applied zero-shot prompting with the [Current Search Query] entered by the user, following the prompting instruction in the Appendix A. For more accurate results, we prompted the LLM to explain the reasoning behind the concretized queries. Using these concretized queries, Gen Query performs image-based searches and presents the top eight results as shown in Figure 4-(b).” As shown in Appendix A prompts 1-2, the model is prompted to analyze the text and determine search intent: “prompt1: We would like to request you to ideate search queries to help designers explore and find useful reference images. The designer has now entered one text query into the image search system. However, there is currently an unspecified part of this query. If the designer looks for search results with this query, he/she can get too many different search results, so the designer wants to be recommended a more specific search query in the query they enter. These are described below. prompt2: Please suggest five search queries by following the steps. First, explain the non-specific parts of the current search query and how to specify them. Second, complete the current search query by adding more details to the end regarding color, shape, style, etc. Please add at least three words. Avoid changing the entire meaning of the query, but focus on specifying the unspecified parts in various aspects.”. Secondly, Son also explicitly discloses ‘a system that automatically modifies a user query’. In the excerpt above, the system adds details (at least three words). This is a clear modification of the user query. On Pages 14-15 of Remarks, Applicant argues that the additional references do not disclose the limitations of claim 1, and that the dependent claims are allowable in view of the independent claims being allowable. In view of the above arguments, Son fully discloses the limitations of claim 1, rendering the Applicant’s arguments here moot. Claim Objections Claims 1, 10, and 19 are objected to because of the following informalities: Claim 1 (and analogous claims 10 and 19) recite “automatically delivering the refined search query from [[to]] the visual content retrieval model”. Firstly, based on the claims filed 5/24/2024, the “from” should be underlined. Secondly, the revised limitation is unclear in view of the full Specification. For example, see Figure 2 which shows that the refined query is delivered to the visual content retrieval component. Appropriate correction is required. Claim Interpretation Regarding claims 1, 10, and 19, the limitation “automatically delivering the refined search query from [[to]] the visual content retrieval model” is being interpreted in accordance with the delivery workflow of Figure 2 (refined query is delivered to the visual content retrieval component), as described above under Claim Objections. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101. Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of query-based searching, without significantly more. The claim recites: “A data processing system comprising: a processor; and a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor alone or in combination with other processors, cause the data processing system to perform functions of: receiving user input defining an initial search query for a visual content retrieval system transmitted over a network from a client application, the initial search query describing at least one characteristic of visual content to be retrieved by the visual content retrieval system; automatically delivering the initial search query and a meta prompt to a refined query generating model as natural language inputs; automatically generating a refined search query using the refined query generating model, the refined query generating model being trained to analyze text of the initial search query to determine user intent based on terms or a combination of the terms found in the initial search query and to generate the refined search query with wording selected to cause a visual content retrieval model of the visual content retrieval system to retrieve aesthetic visual content based on the initial search query, the user intent, and the meta prompt; automatically delivering the refined search query from the visual content retrieval model using the refined query generating model; automatically retrieving the aesthetic visual content using the visual content retrieval model, the visual content retrieval model being trained to retrieve the aesthetic visual content with a reference to a visual content index; indexing, by the visual content index, retrievable visual content for the visual content retrieval system the visual content retrieval model being trained to compare the refined search query and the visual content index to identify the aesthetic visual content to retrieve; and automatically returning the aesthetic visual content to the client application via the network.” The limitations, as drafted, are processes that, under their broadest reasonable interpretation, cover performance of the limitation in the mind. A person can be given a search input describing a desired visual content, refine the search input based on a meta prompt (i.e. prompt: “add descriptive adjectives to the search input”), compare the refined search input to an image catalogue (visual content index), retrieve matched visual content from the catalogue and present it. This judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of a processor in communication with a memory, a network, models for query refinement and content retrieval, and a client application. The processor in communication with a memory and network amounts to generic computer equipment. The models are recited at a level of generality such that they amount to generic models. The client application is recited generically such that it amounts to a generic application. Accordingly, the additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are recited at a high-level of generality. It is therefore a judicial exception that is not integrated into a practical application, and does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This claim is not patent eligible. Claims 2-3 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of describing the content of the meta prompt, which can be done mentally, especially because the meta prompt is described in the context of natural language. The claims are not patent eligible. Claims 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of describing the refined query generating model as an LLM, which is recited generically such that it amounts to a generic LLM. The claim is not patent eligible. Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of creating embeddings for the retrievable visual content and the refined search query, then comparing the embeddings. These amount to mental processes. The claim is not patent eligible. Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of describing the visual content index as an ANN index, which is an additional element that is recited generically such that it does not meaningfully limit the abstract idea. The claim is not patent eligible. Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of describing a model as a VLM, which is an additional element that is recited generically such that it does not meaningfully limit the claim. The claim is not patent eligible. Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of introducing a meta prompt generating model, which is an additional element that is recited generically such that it does not meaningfully limit the claim. The claim is not patent eligible. Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of collecting data (human memory) and performing reinforcement training using the data, which amounts to a mental process of using previous data to contextualize a current query. In the context of the model, the reinforcement training is recited generically such that it fails to impose meaningful limits on practicing the abstract idea. The claim is not patent eligible. Claims 10-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of a method analogous to the abstract idea of claims 1-9. The claims are not patent eligible. Claims 19-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of a non-transitory computer readable medium analogous to the abstract idea of claims 1 and 9. The non-transitory computer readable medium is recited at a level of generality such that it amounts to generic data storage. The claims are not patent eligible. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-7, 10-16, and 19 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Son (GenQuery: Supporting Expressive Visual Search with Generative Models) Regarding claim 1, Son teaches “A data processing system comprising: a processor; and a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor alone or in combination with other processors,” (Son, Section 4.2.4., “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.”) “cause the data processing system to perform functions of: receiving user input defining an initial search query for a visual content retrieval system transmitted over a network from a client application, the initial search query describing at least one characteristic of visual content to be retrieved by the visual content retrieval system;” (Son, Figure 4 top left use query, ‘hiking poster design’. See Section 4.1.1., “Concretize the Vague Search Query. Firstly, Lily begins her search with the keyword “hiking poster design” to explore various hiking poster designs. After entering the prompt in the top text search bar, GenQuery provides Lily with five suggestions on how to concretize her prompt (Figure 4-a), along with expected search results for each suggestion (Figure 4-b) in Figure 3-a. Lily probes the concretized prompts and corresponding search results by pressing the up or down arrow keys (Figure 4-c) to decide what sort of text prompt to input.”). “automatically delivering the initial search query and a meta prompt to a refined query generating model as natural language inputs; automatically generating a refined search query using the refined query generating model, the refined query generating model being trained to analyze text of the initial search query to determine user intent based on terms or a combination of the terms found in the initial search query and to generate the refined search query with wording selected to cause a visual content retrieval model of the visual content retrieval system to retrieve aesthetic visual content based on the initial search query, the user intent, and the meta prompt;” (Son, Figure 4 and Section 4.2.1., “Query Concretization. GenQuery provides five concretized queries based on the text query entered by the user in Figure 3 (a), about one second later. To ensure a smooth user experience with GenQuery, it was critical to deliver suggestions at near-real time speed; hence, we used the gpt-3.5-turbo-16k-0613 LLM model. To deliver concretized queries as output, we applied zero-shot prompting with the [Current Search Query] entered by the user, following the prompting instruction in the Appendix A. For more accurate results, we prompted the LLM to explain the reasoning behind the concretized queries. Using these concretized queries, Gen Query performs image-based searches and presents the top eight results as shown in Figure 4-(b).” Note that the prompts 1 and 2 in Appendix A are mapped to the meta prompt. According to Appendix A, the prompt is provided along with the initial query as natural language inputs.) “automatically delivering the refined search query from the visual content retrieval model using the refined query generating model; automatically retrieving the aesthetic visual content using the visual content retrieval model, the visual content retrieval model being trained to retrieve the aesthetic visual content with a reference to a visual content index; indexing, by the visual content index, retrievable visual content for the visual content retrieval system the visual content retrieval model being trained to compare the refined search query and the visual content index to identify the aesthetic visual content to retrieve; and automatically returning the aesthetic visual content to the client application via the network.” (Son, Sections 4.2.1. and 4.2.4., and Figure 4, “Query Concretization. GenQuery provides five concretized queries based on the text query entered by the user in Figure 3 (a), about one second later. To ensure a smooth user experience with GenQuery, it was critical to deliver suggestions at near-real time speed; hence, we used the gpt-3.5-turbo-16k-0613 LLM model. To deliver concretized queries as output, we applied zero-shot prompting with the [Current Search Query] entered by the user, following the prompting instruction in the Appendix A. For more accurate results, we prompted the LLM to explain the reasoning behind the concretized queries. Using these concretized queries, GenQuery performs image-based searches and presents the top eight results as shown in Figure 4-(b).”; “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.” Figure 4 shows the returned aesthetic visual content to the client application.) Regarding claim 2, Son teaches “The data processing system of claim 1,” “wherein the meta prompt includes instructions for causing the refined query generating model to generate the refined search query in a manner that aligns with the user intent and that facilitates retrieval of the aesthetic visual content that is accurate.”(Son, Appendix A, prompts 1 and 2, “prompt1: We would like to request you to ideate search queries to help designers explore and find useful reference images. The designer has now entered one text query into the image search system. However, there is currently an unspecified part of this query. If the designer looks for search results with this query, he/she can get too many different search results, so the designer wants to be recommended a more specific search query in the query they enter. These are described below. prompt2: Please suggest five search queries by following the steps. First, explain the non-specific parts of the current search query and how to specify them. Second, complete the current search query by adding more details to the end regarding color, shape, style, etc. Please add at least three words. Avoid changing the entire meaning of the query, but focus on specifying the unspecified parts in various aspects.”) Regarding claim 3, Son teaches “The data processing system of claim 2,” “wherein the meta prompt also includes instructions for how to format the refined search query.” (Son, Appendix A, “Return your output as a valid JSON object of the following format: {"explanation": <explain how you generate the specified queries in the first and second steps>, "search_queries": [<list of five suggested queries that designer can use>]}”.) Regarding claim 4, Son teaches “The data processing system of claim 3,” “wherein the refined query generating model comprises a Large Language Model (LLM).” (Son, 4.2.1., “4.2.1 Query Concretization. GenQuery provides five concretized queries based on the text query entered by the user in Figure 3 (a), about one second later. To ensure a smooth user experience with GenQuery, it was critical to deliver suggestions at near-real time speed; hence, we used the gpt-3.5-turbo-16k-0613 LLM model. To deliver concretized queries as output, we applied zero-shot prompting with the [Current Search Query] entered by the user, following the prompting instruction in the Appendix A. For more accurate results, we prompted the LLM to explain the reasoning be hind the concretized queries. Using these concretized queries, Gen Query performs image-based searches and presents the top eight results as shown in Figure 4-(b).” Also, see Figure 4 caption, “Figure 4: Query concretization, which allows the user to concretize the user’s initial abstract search query through LLM zero shot prompting: (a) Suggested search query. User can swap their query by pressing the tab key; (b) Images searched by the suggested query are shown below in the search bar. Each image is clickable to process an image-based search; (c) GenQuery provides five suggestions at a time, and the user can explore other suggestions by pressing the up and down arrow keys.”) Regarding claim 5, Son teaches “The data processing system of claim 1,” “wherein: the visual content index is created by generating visual content embeddings for the retrievable visual content that maps the retrievable visual content to an embedding space, the visual content retrieval model includes an encoder for generating a query embedding that maps the refined search query to the embedding space, and the visual content retrieval model is trained to compare the query embedding to the visual content embeddings in the visual content index to identify a predetermined number of top visual content to retrieve in response to the refined search query.” (Son, Section 4.2.4., “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.” Note that the visual content index (LAION-5B dataset) contains computed embeddings for the visual content (i.e. maps the retrievable visual content to an embedding space), and the ‘clip-retrieval library’ maps text and images to a shared vector space for searching. Additionally, figure 4 shows the top eight search results (predetermined number of top visual content).) Regarding claim 6, Son teaches “The data processing system of claim 1,” “wherein: the visual content index is an Approximate k-Nearest Neighbors (ANN) index.” (Son, Section 4.2.4., “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.” Note that the visual content index (LAION-5B dataset) is an approximate k-nearest neighbors index.) Regarding claim 7, Son teaches “The data processing system of claim 1,” “wherein the visual content retrieval model comprises a vision language model.” (Son, Section 4.2.4., “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.” Note that the clip-retrieval uses CLIP, a visual language model).) Regarding claims 10-16, these claims recite a method with steps corresponding to the elements of the system recited in Claims 1-7. Therefore, the recited steps of these claims are mapped to the analogous elements in the corresponding system claims. Regarding claim 19, this claim recites a non-transitory computer readable medium storing instructions corresponding to the steps recited in Claim 10. Therefore, the recited programming instructions of this claim are mapped to the analogous steps in the corresponding method claim. Son teaches a non-transitory computer readable medium (Son, Section 4.2.4., “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.”) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 8 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Son in view of Lester (US20240378196A1). Regarding claim 8, Son teaches “The data processing system of claim 1,” While Son discloses the meta prompt, (Son, Appendix A, prompts 1 and 2), Son does not disclose a model for generating the meta prompt. Lester discloses a meta prompt generation model in the context of generating a search query (Lester, Paragraphs 79-80 and 140, “Alternatively and/or additionally, the meta-prompt can be tuned or refined by obtaining an aggregated example, in which the aggregated dataset can include a task description. The systems and methods can process the task description and the meta-prompt with a prompt tuning model to generate a task-specific prompt for the task description. The task-specific prompt, an example, and the task description can be processed with a prediction model to generate a prediction. The prediction can then be used in order to evaluate a loss function (e.g., the loss function may be evaluated by comparing the prediction and a respective label for the example.). One or more parameters of the meta-prompt can then be adjusted based on the loss function. Once the meta-prompt is generated, the meta-prompt can be stored on a server computing system to be utilized for prompt generation and refinement. The systems and methods for prompt generation can include receiving a prompt request from a user computing device and generating a requested prompt based on the prompt request and the meta-prompt. The requested prompt can then be sent back to the user computing device.”; “Alternatively and/or additionally, the prompt tuning model can be utilized to process a plurality of datasets and prompts to generate a meta-prompt. The meta-prompt can then be refined by processing aggregated examples and the meta-prompt with the prompt tuning model.” Note that the meta-prompt is described as used for query generation (see Paragraph 47). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to use the meta-prompt generation model of Lester to refine the initial meta-prompt of Son, thereby generating a refined meta-prompt for use in the searching of Son. The motivation for doing so would have been to generate an improved (refined) meta-prompt for more accurate searching. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Son with the above teaching of Lester to fully disclose, “wherein the meta prompt is generated using a meta prompt generating model.” Regarding claim 17, this claim recites a method with steps corresponding to the elements of the system recited in Claim 8. Therefore, the recited steps of this claim are mapped to the analogous elements in the corresponding system claim. Additionally, the rationale and motivation to combine the Son and Lester references apply here. Claim(s) 9, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Son in view of Lester further in view of Krishnamurthy (US20180218080A1). Regarding claim 9, Son in view of Lester teaches “The data processing system of claim 8,” “wherein the functions further comprise: collecting usage data pertaining to usage of the visual content retrieval system;” (Son, Figure 6 and Section 1 Paragraph 5, “Therefore, we propose GenQuery allows the users to concretize the abstract text query, express search intent through visuals, and diversify their search intent based on the search history. First, the user can concretize their vague search query through auto-complete suggestions with more specific search directions in text-based search (Figure 1 Query concretization). Second, the user can edit an image to generate an intent-aligned image as a search input through image-based image modification (Figure 1 Image-based modification). Third, the user can diversify their search intent through keyword-based image modification with the keywords suggested from the search history (e.g., saved visuals or inputted text queries). The modified image can also be used as a search input (Figure 1 Keyword-based modification).” Figure 6 also shows search terms from the search history.”) Son in view of Lester does not expressly disclose, “and performing reinforcement training of at least one of the meta prompt generating model and the visual content retrieval model using training data derived from the usage data.” Krishnamurthy discloses using reinforcement training of a visual content retrieval model using training data derived from usage data (Krishnamurthy, Paragraph 4, Embodiments of the present invention relate to, among other things, a conversational agent that can conduct conversations with users to assist with performing searches. In accordance with implementations of the present disclosure, the conversational agent is a reinforcement learning (RL) agent trained using a user model generated from existing session logs from a search engine. The user model is generated from the session logs by mapping entries from the session logs to user actions understandable by the RL agent and computing conditional probabilities of user actions occurring given previous user actions in the session logs. The RL agent is trained by conducting conversations with the user model in which the RL agent selects agent actions in response to user actions sampled using the conditional probabilities from the user model. The RL agent can subsequently be retrained by interacting with humans.” Note that the searching of Krishnamurthy is for visual content retrieval (image searching).) It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to incorporate the reinforcement training method of Krishnamurthy for training the visual content retrieval model of Son in view of Lester using, in part, the search history stored by Son in view of Lester. Note that one skilled in the art would be able to incorporate the saving of any additional usage data (i.e. session logs), other than the saved textual queries from Son in view of Lester, for the functionality of the reinforcement training of Krishnamurthy. The motivation for doing so would have been to improve search results by contextualizing a current search with respect to previous searches. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Son in view of Lester with the above teaching of Krishnamurthy to fully disclose, “and performing reinforcement training of at least one of the meta prompt generating model and the visual content retrieval model using training data derived from the usage data.” Regarding claim 18, this claim recites a method with steps corresponding to the elements of the system recited in Claim 9. Therefore, the recited steps of this claim are mapped to the analogous elements in the corresponding system claim. Additionally, the rationale and motivation to combine the Son, Lester, and Krishnamurthy references apply here. Regarding claim 20, this claim recites a non-transitory computer readable medium storing instructions corresponding to the steps recited in Claim 18. Therefore, the recited programming instructions of this claim are mapped to the analogous steps in the corresponding method claim. Additionally, the rationale and motivation to combine the Son, Lester, and Krishnamurthy references apply here. Son teaches a non-transitory computer readable medium (Son, Section 4.2.4., “Implementation Details. GenQuery is a web-based system, which is composed of a ReactJS front-end and a Python Flask server as the back-end. For the image dataset and basic search functions (e.g., text-based search and image-based search), we used the clip-retrieval library [1] on the hosted API provided by the library. Given an input query or image, the library processes the input into an embedding and then queries similar images from the LAION-5B dataset [47] through the API. We used a machine with an AMD Ryzen 9 5900X 12-Core Processor and NVIDIA GeForce RTX 3090 to implement GenQuery.”) Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AARON JOSEPH SORRIN whose telephone number is (703)756-1565. The examiner can normally be reached Monday - Friday 9am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AARON JOSEPH SORRIN/Examiner, Art Unit 2672 /SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672
Read full office action

Prosecution Timeline

May 24, 2024
Application Filed
Mar 24, 2026
Non-Final Rejection mailed — §101, §102, §103
Apr 22, 2026
Interview Requested
Apr 28, 2026
Examiner Interview Summary
Apr 28, 2026
Applicant Interview (Telephonic)
Jun 23, 2026
Response Filed
Sep 01, 2026
Final Rejection mailed — §101, §102, §103
Sep 30, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743886
COMPONENT MOUNTING SYSTEM, IMAGE PROCESSING DEVICE, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING SYSTEM
2y 10m to grant Granted Sep 22, 2026
Patent 12731269
ROTATION STATE ESTIMATION APPARATUS, METHOD THEREOF, AND PROGRAM
3y 1m to grant Granted Sep 08, 2026
Patent 12720096
IMAGE PROCESSING DEVICE, IMAGE DISPLAY SYSTEM, IMAGE PROCESSING METHOD, AND RECORDING MEDIUM
3y 0m to grant Granted Aug 25, 2026
Patent 12718945
A RADIOMIC-BASED MACHINE LEARNING ALGORITHM TO RELIABLY DIFFERENTIATE BENIGN RENAL MASSES FROM RENAL CELL CARCINOMA
2y 10m to grant Granted Aug 25, 2026
Patent 12705851
Method And System For Detecting, Quantifying, And Attributing Gas Emissions Of Industrial Assets
3y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+42.0%)
3y 0m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 75 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month