Prosecution Insights
Last updated: August 17, 2026
Application No. 18/627,340

CONTEXT-AWARE PROMPT MATCHING SYSTEM USING LARGE LANGUAGE MODELS

Non-Final OA §101§103
Filed
Apr 04, 2024
Examiner
CAMPOS, ALFREDO
Art Unit
Tech Center
Assignee
ORACLE INTERNATIONAL Corporation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
73%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
7 granted / 9 resolved
+17.8% vs TC avg
Minimal -5% lift
Without
With
+-5.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
19 currently pending
Career history
36
Total Applications
across all art units

Statute-Specific Performance

§101
35.2%
-4.8% vs TC avg
§103
43.0%
+3.0% vs TC avg
§102
3.6%
-36.4% vs TC avg
§112
18.2%
-21.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 9 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more. The claim(s) recite(s) significantly more. The subject matter eligibility test for products and process is describe below for claim 1 in view of dependent claims. Regarding claim 1: Step 1: Is the claim to a process machine manufacture or composition of matter? Yes – Claim 1 recites a method, which is a method that falls under the statutory categories. Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes – The claim recites the following: “based on the set of prompts and the prompt, identifying, [by the first LLM], a subset of the set of prompts;” - The limitation recites a mental process of identifying a set of prompts based on the input prompts and the prompt (see MPEP 2106.04(a)(2)III). “generating a particular embedding based on the prompt;”- The limitations recites a mental process of generating an embedding based on the prompt (see MPEP 2106.04(a)(2)III). “for each embedding in a set of embeddings, each of which corresponds to a different prompt in the subset of the set of prompts: generating a similarity score between said each embedding and the particular embedding;” - The limitations recites a mathematical process of generating a similarity score (see MPEP 2106.04(a)(2)I). “associating the similarity score with said each embedding;” - The limitations recites a mental process of associating the similarly score with said embedding (see MPEP 2106.04(a)(2)III). “adding the similarity score to a set of similarity scores;” - The limitations recites a mental process of adding the similarity score to a set of similarity scores (see MPEP 2106.04(a)(2)III). “ranking the set of embeddings based on the set of similarity scores;” - The limitations recites a mental process of ranking the set of embeddings based on the set of similarity scores (see MPEP 2106.04(a)(2)III). “identifying at least one highest ranked embedding, in the set of embeddings, that corresponds to a particular prompt in the subset;” - The limitations recites a mental process of identifying at least one highest ranked embedding that corresponds to a particular prompt in the subset (see MPEP 2106.04(a)(2)III). Step 2 Prong 2: Does the claim recite additional elements that integrate the judicial exception into a particular application? No – The claim includes the additional element(s): “A method comprising: receiving, by a first large language model (LLM), input that comprises a prompt for a second LLM;” The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by receiving class information created by an edge node. See MPEP 2106.5(g). “accessing, by the first LLM, a set of prompts;” The additional elements fall under “apply it” as using a generic computer to implement a first LLM to access a set of prompts. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)). “[based on the set of prompts and the prompt, identifying,] by the first LLM, [a subset of the set of prompts;]” The additional elements fall under “apply it” as using a generic computer to implement an LLM to identify a subset of the set of prompts (see MPEP 2106.05(f)). “wherein the method is performed by one or more computing devices.” The additional elements fall under “apply it” as using a generic computer to implement the method in one or more computing devices (see MPEP 2106.05(f)). Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No - The claim does not include additional elements that are sufficient to amount to a significantly more than the judicial exemption. As an order whole, the claim is directed to determining the best prompt based on a similarity score. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of receiving, accessing and using an LLM fall under using generic computer to apply an exemption and mere data gathering. The method does not improve on the function of a computer, transforms an article into another article, nor is it applied by a particular machine, making the claim not patent eligible. Regarding claim 2: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, wherein accessing, by the first LLM, a set of prompts comprises accessing a repository of stored prompts.” The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 3: Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes – The claim recites the following: “wherein identifying, [by the first LLM,] a subset of the set of prompts comprises performing keyword matching or performing semantic similarity analysis” - The limitation recites a mental process of identifying a set of prompts using keyword matching or performing sematic similarity analyst (see MPEP 2106.04(a)(2)III). Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 2, [wherein identifying,] by the first LLM, [a subset of the set of prompts comprises performing keyword matching or performing semantic similarity analysis.]” The additional elements fall under “apply it” as using a generic computer to implement a first LLM identify a subset of prompts using key word matching or sematic similarity analysis. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)). Regarding claim 4: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, wherein the first LLM has been pre-trained to perform contextual analysis.” The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 5: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, further comprising: causing the particular prompt to be presented on a screen of a computing device; receiving user selection of the particular prompt; and in response to receiving the user selection, inputting the particular prompt to the second LLM.” The additional elements fall under “apply it” as using a generic computer to present a screen of a computing device, receiving user selection, and input the particular prompt into the second LLM. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)). Regarding claim 6: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 5, further comprising: causing the second LLM to operate on the particular prompt to generate an output; and causing the output to be presented on the screen of the computing device.” The additional elements fall under “apply it” as using a generic computer to implement the second LLM to operate a particular prompt, generate out, and present the output on the screen of a computer (see MPEP 2106.05(f)). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 7: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, wherein the second LLM is different than the first LLM.” The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 8: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, wherein causing the particular prompt to be presented comprises causing multiple prompts to be presented on the screen of the computing device.” The additional elements fall under “apply it” as using a generic computer to preset multiple prompts of the computing device (see MPEP 2106.05(f)). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 9: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 8, wherein causing the multiple prompts to be presented comprises causing the multiple prompts to be presented based on their corresponding similarity scores.” The additional elements fall under “apply it” as using a generic computer to preset multiple prompts of the computing device based on their similarity scores (see MPEP 2106.05(f)). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 10: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, wherein identifying the subset comprises identifying a pre-determined number of prompts from the set of prompts.” The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Regarding claim 11: Step 2A Prong 2, Step 2B: The additional element(s): “The method of Claim 1, wherein the set of prompts comprises a plurality of categories of prompts, each category of the plurality of categories comprising multiple pre-defined prompts of a type belonging to said each category.” The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application. Claims 12-20 recite a computer readable medium product and are analogous to the method of claims 1, 3-9, and 11. Therefore, the rejections of claim 1, 3-9, and 11 above applies to claims 12-20. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 4, 7, 10, 12, 14, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Yuan et al. (US12675518B1) (“Yuan”) in view of Zhu, Hanlin, Banghua Zhu, and Jiantao Jiao. "Efficient prompt caching via embedding similarity." arXiv preprint arXiv:2402.01173 (2024). (“Zhu”). Regarding claim 1 and analogous claim 12, Yuan teaches A method comprising: receiving, by a first large language model (LLM), input that comprises a prompt for a second LLM; accessing, by the first LLM, a set of prompts (Yuan Col 8 line 38-42, It should also be noted that the size of the language model ( or LLM) included in the encoding generator subsystem 130 may be substantially less than the language model (or LLM) included in the summary generator subsystem 135. Col 10 line 17-25, At block 308, the user request and a first set of tokens are inputted into encoding generator subsystem 130. In some embodiments, the content summarization subsystem 120 may also cause a prompt (e.g., in natural language format), such as prompt 603 ("Identify the key topics in the transcription") to be inputted into the encoding generator subsystem 130 along with the tokens and/or the user request Such prompt 603 may, for example, include, be based on, or consist of the user request [receiving, by a first large language model (LLM),]. Col line 30-34, Alternatively, the prompt 603 may be one of a plurality of predetermined prompts stored in the content summarization subsystem 120, in which a predetermined prompt is selected based on the user request [accessing, by the first LLM, a set of prompts;]. As discussed previously, the encoding generator subsystem 130 may include a soft-prompt generator that is configured to generate an encoded output in the form of a soft prompt, which may provide refined instructions (based on the user request) to the summary generator subsystem 135 to generate the categorical description output and/or assists the summary generator subsystem 135 in classifying portions of the content 602 into the categorical description requested in the user instruction ( e.g., topics, issues, outcomes, or action items) [input that comprises a prompt for a second LLM;]); wherein the method is performed by one or more computing devices (Yuan Col 11 line 37-41, such as the computing device 700 shown in FIG. 7, and executed by one or more processors. In some embodiments, the routine 300, routine 400, and routine 500, or portions thereof may be implemented on multiple processors, serially or in parallel [wherein the method]. Col 14 line 58-67, FIG. 7 illustrates various components of an example computing device 700 configured to implement various functionality described herein. In some embodiments, the computing device 700 may be implemented using any of a variety of computing devices, such as server computing devices, desktop computing devices, personal computing devices, mobile computing devices, mainframe computing devices, midrange computing devices, host computing devices, or some combination thereof [is performed by one or more computing devices].). However Yuan does not explicitly teach based on the set of prompts and the prompt, identifying, by the first LLM, a subset of the set of prompts; generating a particular embedding based on the prompt; for each embedding in a set of embeddings, each of which corresponds to a different prompt in the subset of the set of prompts: generating a similarity score between said each embedding and the particular embedding; associating the similarity score with said each embedding; adding the similarity score to a set of similarity scores; ranking the set of embeddings based on the set of similarity scores; identifying at least one highest ranked embedding, in the set of embeddings, that corresponds to a particular prompt in the subset; However Zhu teaches based on the set of prompts and the prompt, identifying, by the first LLM, a subset of the set of prompts (Zhu page 3 Figure 2b, PNG media_image1.png 379 457 media_image1.png Greyscale page 3-4, 2 Preliminaries We introduce in this section some basic notations, definitions, and assumptions. Let Q denote the set of all possible prompts (queries) [based on the set of prompts]. For any prompt pair (q1, q2) ∈ Q × Q, we denote the ground-truth probability that (q1, q2) can be answered by the same response by P⋆(q1 = q2) [and the prompt]. Rigorously speaking, the ground-truth probability should be either 1 or 0, since for any two queries, we shall exactly know whether they can be answered by the same response. However, although the ground truth is always deterministic, we can still train a probabilistic predictor whose output represents the confidence of the prediction. On the other hand, even if a prompt pair can be answered by the same response in ground truth, the label (either by a human or a language model) might still be noisy at different levels. Therefore, we allow the ground-truth probability to be any real number in [0, 1]. Page 4 Definition 2 (Probability via embedding similarity). For any two prompts q1, q2, we denote the induced probability via embedding similarity that q1, q2 can be answered by the same response by (see Figure 2b) [identifying, by the first LLM, a subset of the set of prompts;]); generating a particular embedding based on the prompt (Zhu page 3 Figure 2a PNG media_image2.png 278 442 media_image2.png Greyscale page 4, Definition 1 (Embedding of prompts). For any prompt q, let vθ(q) ∈ Rd denote its vector embedding (see Figure 2a) where v can be viewed as the mapping of prompts to a specific layer of a language model, and θ ∈ Θ is the parameters of that model. [generating a particular embedding based on the prompt]); for each embedding in a set of embeddings, each of which corresponds to a different prompt in the subset of the set of prompts: generating a similarity score between said each embedding and the particular embedding; associating the similarity score with said each embedding; adding the similarity score to a set of similarity scores; ranking the set of embeddings based on the set of similarity scores; identifying at least one highest ranked embedding, in the set of embeddings, that corresponds to a particular prompt in the subset (Zhu page 3 Figure 2, PNG media_image3.png 481 1004 media_image3.png Greyscale Page 9 4.1 Construction of the dataset para 2, Now for each prompt, we search the three nearest neighbors using FAISS (Johnson et al., 2019), where the distance of a prompt pair is defined by the cosine similarity of the two corresponding embedding vectors as in Definition 2. We then sample 70k prompts uniformly at random, and for each prompt, we obtain three prompt pairs forming by the prompt itself and each of their three nearest neighbors [for each embedding in a set of embeddings, each of which corresponds to a different prompt in the subset of the set of prompts:]. Therefore, we get 210k prompt pairs in total, and we use GPT-4 (OpenAI, 2023) to label whether each prompt pair can be answered by the same response (0 or 1, where 1 means they can be answered by the same response). The exact prompt we use to label each pair is presented in Appendix B.1 [associating the similarity score with said each embedding;]. Page 10 4.1 Construction of the dataset para 4, In Algorithm 2, we first sort all data points by cosine similarity [ranking the set of embeddings based on the set of similarity scores;]. Then, we add data points in ascending order of similarity and ensure the added prompt pairs have alternate labels. By the construction procedure, the AUC of intfloat/e5-large-v2 on the constructed hard dataset D is approximately 0.5. In our experiments, the size of D is 37382, which is large enough for our training and implies that the existing embedding intfloat/e5-large-v2 might not be an ideal model for caching prediction. Finally, we show two examples of prompt pairs in our dataset, where the first one has label 0 but high cosine similarity according to intfloat/e5-large-v2, and the second one has label 1 but low cosine similarity. PNG media_image4.png 356 976 media_image4.png Greyscale 4.3 Simulation of the prompt streaming with caching para 2, The cache is initialized to be empty. We maintain two counters: nCorrectHit and nFalseHit, which represent the number of correct caching hits and the number of false caching hits, respectively. At each time, a prompt q from the streaming dataset arrives, and its embedding v(q) ∈ Rd (in our experiments, d = 1024) is calculated. We first find the nearest neighbor in the cache, which is a tuple (q0, v(q0), r0), where q0 is a prompt, r0 is the response, and v(q0) is the embedding. The distance is measured by cosine similarity. If sim(v(q1), v(q2)) > τ where τ ∈ [0, 1] is a threshold, we view it as a caching hit and use r0 as the response of q; otherwise, we view it as a caching miss, directly query the LLM to get the response r, and add the tuple (q, v(q), r) to the cache [identifying at least one highest ranked embedding, in the set of embeddings, that corresponds to a particular prompt in the subset;]); Yuan and Zhu are considered to be analogous to the claim invention because they are in the same field of large language models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Yuan to incorporate the teachings of Zhu to include identifying similar prompts. Doing so to improve inference effciency by prompt caching and improve the accuracy of the model (Zhu Abstract line 1-7, Large language models (LLMs) have achieved huge success in numerous natural language process (NLP) tasks. However, it faces the challenge of significant resource consumption during inference. In this paper, we aim to improve the inference efficiency of LLMs by prompt caching, i.e., if the current prompt can be answered by the same response of a previous prompt, one can directly utilize that previous response without calling the LLM. Specifically, we focus on the prediction accuracy of prompt caching for single-round question-answering tasks via embedding similarity. Abstract line 13-16, We then fine-tune the above embedding model, which significantly improves the AUC of caching prediction from 0.51 to 0.81. We also conduct simulations demonstrating that our trained models achieve better caching efficiency than the previous embedding model.). Regarding claim 4 and analogous claim 14, Yuan in view of Zhu teach the method of claim 1. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan teaches wherein the first LLM has been pre-trained to perform contextual analysis (Yuan Col 4 line 1-5, A Large Language Model ("LLM") is any type of language model that has been trained on a larger data set and has a larger number of training parameters compared to a regular language model. An LLM can understand more intricate patterns and generate text that is more coherent and contextually relevant due to its extensive training. Thus, an LLM may perform well on a wide range of topics and tasks. An LLM may comprise a NN trained using self-supervised learning. An LLM may be of any type, including a Question Answer ("QA") LLM that may be optimized for generating answers from a context, a multimodel LLM/model, and/or the like. An LLM (and/or other models of the present disclosure), may include, for example, attention-based and/ or transformer architecture or functionality [has been pre-trained to perform contextual analysis]. Col 6 line 37-39, The encoding generator subsystem 130 may include a least one of a language model and a large language model (LLM) [the first LLM]). Regarding claim 7 and analogous claim 17, Yuan in view of Zhu teach the method of claim 1. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan further teaches wherein the second LLM is different than the first LLM ((Yuan Col 4 line 18-28, While certain aspects and implementations are discussed herein with reference to use of a language model, LLM, and/or AI, those aspects and implementations may be performed by any other language model, LLM, AI model, generative AI model, generative model, ML model, NN, multimodel model, and/or other algorithmic processes. Similarly, while certain aspects and implementations are discussed herein with reference to use of a ML model, those aspects and implementations may be performed by any other AI model, generative AI model, generative model, NN, multimodel model, and/or other algorithmic processes. Col 4 line 50-60, Examples of models, language models, and/or LLMs that may be used in various implementations of the present disclosure include, for example, Bidirectional Encoder Representations from Transformers (BERT), LaMDA (Language Model for Dialogue Applications), PaLM (Pathways Language Model), PaLM 2 (Pathways Language Model 2), Generative Pre-trained Transformer 2 (GPT-2), Generative Pre-trained Transformer 3 (GPT-3), Generative Pre-trained Transformer 4 (GPT-4), LLAMA (Large Language Model Meta AI), and BigScience Large Open-science Open-access Multilingual Language Model (BLOOM) Col 6 line 37-41, The encoding generator subsystem 130 may include a least one of a language model and a large language model (LLM). In some embodiments, there may be two or more encoding generator subsystems 130 in the content summarization system 120. Col 8 line 50-57, In some embodiments, the content summarization system 120 comprises a large language model (LLM). The LLM may be stored in the language prograniming store 140 and requested by the at least one of the encoding generator subsystem 130 and the s=ary generator subsystem 135, when the LLM is needed. In some embodiments, the content summarization system 120 may request an LLM based on the type of content in the received request. Col 9 line 29-33, The content summarization module 215 may have access to one or more language models or LLMs that allow the content summarization module 215 to generate the content summary according to the particular categorical description requested by the user [second LLM is different than the first LLM]. (Examiner Note: The system selects an LLM to perform the encoding generator subsystems function and the content summarization module can be one or more LLMs models. The second LLM is different than the different LLM)). Regarding claim 10, Yuan in view of Zhu teach the method as recited in claim 1. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Zhu further teaches wherein identifying the subset comprises identifying a pre-determined number of prompts from the set of prompts (Zhu 4.1 Construction of the dataset para 2 line 1-5, Now for each prompt, we search the three nearest neighbors using FAISS (Johnson et al., 2019), where the distance of a prompt pair is defined by the cosine similarity of the two corresponding embedding vectors as in Definition 2. We then sample 70k prompts uniformly at random, and for each prompt, we obtain three prompt pairs forming by the prompt itself and each of their three nearest neighbors Now for each prompt, we search the three nearest neighbors using FAISS (Johnson et al., 2019), where the distance of a prompt pair is defined by the cosine similarity of the two corresponding embedding vectors as in Definition 2. We then sample 70k prompts uniformly at random, and for each prompt, we obtain three prompt pairs forming by the prompt itself and each of their three nearest neighbors [comprises identifying a pre-determined number of prompts from the set of prompts]). Claim(s) 2, 3, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of Zhu and further in view of Vu et al. (US20240020546A1) (“Vu”). Regarding claim 2, Yuan in view of Zhu teach the method of claim 1. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan does not explicitly teach wherein accessing, by the first LLM, a set of prompts comprises accessing a repository of stored prompts. However Vu teaches wherein accessing, by the first LLM, a set of prompts comprises accessing a repository of stored prompts (Vu para 107 line 1-10, In some implementations, the second training dataset 220 can be utilized to generate a second embedding 208 which can be utilized to query the prompt database 240 for a prompt associated with a similar embedding to the second embedding 208. For example, a second training example of the plurality of second training examples 222 and a set of parameters (e.g., the initial prompt) can be processed with a pre-trained machine-learned model 230 to generate a second output 226.). Yuan and Vu are considered to be analogous to the claim invention because they are in the same field of large language models prompt generation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Yuan to incorporate the teachings of Vu to use a database for accessing prompts. Doing so to query the prompt databased for a prompt associated with a similar embedding to the second embedding and use the output of the model to adjust one or more parameters of the initial prompt (Vu para 0107 line 9-17, The second output 226 and a respective second training label ( e.g., a second training label of the plurality of second training labels 224) associated with the second training example can be utilized to evaluate a loss function 250 to generate a prompt gradient. The prompt gradient can be backpropagated to adjust one or more parameters of the set of parameters ( e.g., the initial prompt). The training loop can be repeated for a portion of the second training dataset 220 to generate the second embedding 208.). Regarding claim 3 and analogous 13, Yuan in view of Zhu teach the method of claim 1. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan does not explicitly teach wherein identifying, by the first LLM, a subset of the set of prompts comprises performing keyword matching or performing semantic similarity analysis. Zhu teaches wherein identifying, by the first LLM, a subset of the set of prompts comprises performing keyword matching or performing semantic similarity analysis (Zhu page 2, 1 Introductions para 4 line 3 - 10 We propose a distillation-based method, which aims to learn whether a prompt pair can be answered by the same response via cosine similarity of the embeddings of the prompt pair, to fine-tune an existing semantic vector embedding from Wang et al. (2022). Theoretically, we provide finite sample guarantees for the learning error under mild assumptions using two different loss functions: binary cross entropy (BCE) and squared log difference (SLD), respectively (Section 3). Empirically, we carefully construct a hard dataset based on Kwiatkowski et al. (2019) (Section 4.1) and fine-tune the embedding from Wang et al. (2022) , which significantly improves the AUC of caching prediction from 0.51 to 0.81 (Section 4.2). 4.1 Construction of the dataset page 9 para 2 5-8, Therefore, we get 210k prompt pairs in total, and we use GPT-4 (OpenAI, 2023) to label whether each prompt pair can be answered by the same response (0 or 1, where 1 means they can be answered by the same response). The exact prompt we use to label each pair is presented in Appendix B.1 [performing semantic similarity analysis].). Claim(s) 5, 6, 8, 15, 16, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of Zhu and further in view of Luzhnica (US11516158B1) (“Luzhnica”). Regarding claim 5 and analogous 15, Yuan in view of Zhu teach the method of claim 1. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan teaches further comprising: causing the particular prompt to be presented on a screen of a computing device (Yuan Fig. 2 PNG media_image5.png 509 1060 media_image5.png Greyscale Col 9 line 15-29, The content summarization system 210 is a computing device that includes a content summarization module 215 which provides content summarization services. For example, the content summarization module 215 may be an application with which the user interacts with using the computing device 205. The content summarization module 215 may communicate with other components of the system 200 according to protocols defined by the network 240. The content summarization module 215 may be configured to receive a user request. For example, the user request may include a request to generate a content summary that is tailored to a particular categorical description of the content which includes (but is not limited to), topics, issues, outcomes, and action items ) Fig 6. 603 PNG media_image6.png 51 801 media_image6.png Greyscale Col 8 line 7-11, Alternatively, the prompt 603 may be one of a plurality of predetermined prompts stored in the content summarization subsystem 120, in which a predetermined prompt is selected based on the user request [causing the particular prompt to be presented on a screen of a computing device]); [and in response to receiving the user selection,] inputting the particular prompt to the second LLM (Col 8 line 25-32, The summary generator subsystem 135 may include a least one of a language model and a large language model (LLM). In some embodiments, there may be two or more summary generator subsystems 135 in the content summarization system 120. The summary generator subsystem 135 generates, based on the content provided by the user and the encoded outputs provided by the encoding generator sub system 130 [inputting the particular prompt to the second LLM]). Yuan does not explicitly teach receiving user selection of the particular prompt. Luzhnica teaches receiving user selection of the particular prompt (Luzhnica Col 120 line 19-24, FIG. 5 FIG. 5 illustrates the layout of an exemplary user interface, #500, for use in an exemplary system of the invention. Such a user interface can be presented on, e.g., a system interface such as a digital screen, for example a computer, laptop, pad, or phone screen. Col 125 line 47-58, FIG. 14A illustrates output provided by an exemplary system of the invention upon selectin of one or more specific setting(s), module(s), and/or input selection(s). For simplicity, the same type of interface #501 as shown in FIG. 13 is also presented here, along with a number of setting selections, inputs, etc., as described in FIG. 13. Three exemplary messages, #536, #538, and #540, are generated by the system and presented to the system user. Each of the three exemplary messages are presented with options, #1410 and #1420, for the user to use to indicate to the system their evaluation of the messages, transmit the messages, or otherwise. Col 126 line 11-17, FIG. 15 illustrates user evaluation/selection of system output provided in FIG. 14A (see the description of elements labeled in FIG. 14A but not of primary focus here within the discussion of, e.g., FIG. 13), but further illustrating how users can use the interface to make a section of messages to work with or rate, for application and for further training of the system. Two draft messages, #1520, are identified by a user as either disliked or not selected as indicated by shading (e.g., by using/selecting a rating indicator, #1420 for each message). One message, #1510 is selected for use (modification, transmission) or is rated as liked, e.g., using/selecting indicator/selector/control, #1410 [receiving user selection of the particular prompt;]); and in response to receiving the user selection, [inputting the particular prompt to the second LLM] (Luzhnica col 127, 63-67, User is then permitted, via the interface, to interact with automatically displayed message drafts, #1713, typically/ ideally resulting in identification of one or more positive interaction messages (PIMs) (via transmission, storage, editing, rating, or combination of any or all thereof). Col 128 line 7-13, Also, or alternatively, system-generated messages are subjected to analysis by further neural network( s) of a system. E.g., as shown in the Figure, messages 10 can be subject to a message-variation/message optimization neural network, #1715, which generates messages from system-generated messages, which can be used as, e.g., additional TSD [and in response to receiving the user selection]). Yuan and Luzhnica are considered to be analogous to the claim invention because they are in the same field of machine learning prompt generation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Yuan to incorporate the teachings of Luzhnica use user input. Doing so to allow the user to provide input to the system allowing control of the content generated messages to tailor messages as directed by the user to the system (Luzhnica col 122 line 12-19, Such user interfaces, providing a number, but a limited number, of structured input options and system/method settings, provide for an efficient way for users to control the content of system-generated messages and to efficiently provide instructional prompt and situational prompt direction to the system for tailoring/personalization messages to recipients, thereby increasing the chances of reading, impact, and desired response.) Regarding claim 6 and analogous 16, Yuan in view of Zhu and Luzhnica teach the method of claim 5 and analogous 15. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan and Luzhnica are combine in the same rational as set forth above with respect to claim 5 and analogous claim 15. Yuan further teaches wherein the instructions, when executed by the one or more computing devices, further cause: causing the second LLM to operate on the particular prompt to generate an output; and causing the output to be presented on the screen of the computing device (Yuan PNG media_image5.png 509 1060 media_image5.png Greyscale Col 9 line 15-29, The content summarization system 210 is a computing device that includes a content summarization module 215 which provides content summarization services. For example, the content summarization module 215 may be an application with which the user interacts with using the computing device 205. The content summarization module 215 may communicate with other components of the system 200 according to protocols defined by the network 240. The content summarization module 215 may be configured to receive a user request. For example, the user request may include a request to generate a content summary that is tailored to a particular categorical description of the content which includes (but is not limited to), topics, issues, outcomes, and action items [and causing the output to be presented on the screen of the computing device] (i.e. user 250 interacts with computing device 205). Col 12 line 60-65, As such, the transmission of the second encoded output causes the summary generator subsystem 135 to retrieve the content from, for example, memory or a database. At block 414, the summary generator subsystem 135 generates a content summary 606 based on the second encoded output and the content 602 [causing the second LLM to operate on the particular prompt to generate an output]). Regarding claim 8 and analogous 18, Yuan in view of Zhu and Luzhnica teach the method of claim 1 and analogous 2. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan and Luzhnica are combine in the same rational as set forth above with respect to claim 5 and analogous claim 15. Luzhnica teaches wherein causing the particular prompt to be presented comprises causing multiple prompts to be presented on the screen of the computing device (Luzhnica Col 120 line 19-24, FIG. 5 FIG. 5 illustrates the layout of an exemplary user interface, #500, for use in an exemplary system of the invention. Such a user interface can be presented on, e.g., a system interface such as a digital screen, for example a computer, laptop, pad, or phone screen. Col 125 line 47-58, FIG. 14A illustrates output provided by an exemplary system of the invention upon selectin of one or more specific setting(s), module(s), and/or input selection(s). For simplicity, the same type of interface #501 as shown in FIG. 13 is also presented here, along with a number of setting selections, inputs, etc., as described in FIG. 13. Three exemplary messages, #536, #538, and #540, are generated by the system and presented to the system user. Each of the three exemplary messages are presented with options, #1410 and #1420, for the user to use to indicate to the system their evaluation of the messages, transmit the messages, or otherwise [causing multiple prompts to be presented on the screen of the computing device]). Claim(s) 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of Zhu and Luzhnica and further in view of Gruenstein et al. (US20150279354A1) (“Gruenstein”). Regarding claim 9 and analogous 19, Yuan in view of Zhu and Luzhnica teach the method of claim 8 and analogous 18. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan and Luzhnica are combine in the same rational as set forth above with respect to claim 5 and analogous claim 15. Yuan does not teach wherein causing the multiple prompts to be presented comprises causing the multiple prompts to be presented based on their corresponding similarity scores. However Gruenstein teaches wherein causing the multiple prompts to be presented comprises causing the multiple prompts to be presented based on their corresponding similarity scores (Gruenstein para 0041, In FIG. 4A in an embodiment based on the example above, user interface 250 presents a list of generated speech recognition candidates 420 and prompts the user to choose one. In one embodiment, these choices are recognition results generated by embedded speech recognizer 210, while in another embodiment, these are stored speech tags that have been chosen based on their similarity to the audio stream [their corresponding similarity], and in an additional embodiment, these are speech recognition candidate results generated by a network-based speech recognizer [causing the multiple prompts to be presented]. When a user selects a result, the chosen result is then used to perform a query, and as described above, in an embodiment, a speech tag is generated and stored linking the chosen result to the audio stream.). Yuan and Gruenstein are considered to be analogous to the claim invention because they are in the same field of large language models prompt generation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Yuan to incorporate the teachings of Gruenstein present similar prompt candidates. Doing so to provide personalization to the user selection (Gruenstein para 0038, With generated speech tags stored on client device 110, in an embodiment, whenever a user performs a voice search, embedded speech recognizer 210 generates recognition candidates, as described above with the description of FIG. 2, and also, to provide personalization and resolve ambiguities in favor of past user selections, embedded speech recognizer 210 can use tag comparator 260 to compare the generated recognition candidates with speech tags stored in client database 240. In an embodiment, this comparison can influence the selection of a recognition candidate and thus have the benefits described above.). Claim(s) 11 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Yuan in view of Zhu and further in view of Pu et al. (US20210042662A1) (“Pu”). Regarding claim 11 and analogous 20, Yuan in view of Zhu teach the method of claim 1 and analogous 12. Yuan and Zhu are combine in the same rational as set forth above with respect to claim 1 and analogous claim 12. Yuan does not explicitly teach wherein the set of prompts comprises a plurality of categories of prompts, each category of the plurality of categories comprising multiple pre-defined prompts of a type belonging to said each category. However Pu teaches wherein the set of prompts comprises a plurality of categories of prompts, each category of the plurality of categories comprising multiple pre-defined prompts of a type belonging to said each category (Pu para 0014, One embodiment of the present invention discloses a method for interactive information capture via a set of predefined prompts grouped and configured for specific instances or types of information such as lecture material from a student's course, deals from nearby restaurants, or print ads from local. Para 0106, The disclosed method enables user-desired and tailored prompts and prompt-set templates to be created or modified by users and/or augmented by machine intelligence. The method offers a convenient way for individual users to personalize their information capture needs for any raw information types they see, hear, or may have access to, aided with the prompts collection (121), prompt-set templates collection (122) whenever and wherever available, as well as with heuristic and ML models (124) and data/ semantic/cognitive models with AI/ML/Data products and services (160) whenever and wherever applicable businesses [, each category of the plurality of categories comprising multiple pre-defined prompts of a type belonging to said each category]. Para 0130, Existing prompt-set template(s) on device or on server-side (112 or 122) may need to be modified to better suit the user's information capture needs for the applicable raw information contents and sources Para 0155, Related information, per the capture template set that the desired information is associated with and/or defined by the user's preference, can also be fetched and organized per the dimensions that the application may encompass. Examples include information related to a category or subject area, related to the class the desired information is associated with, within the capture template's structure as a sibling, parent, or child node information, related to other associated templates or prompts, or related to the location and timeline the information is captured, etc [wherein the set of prompts comprises a plurality of categories of prompts,]). Yuan and Pu are considered to be analogous to the claim invention because they are in the same field of machine learning and using prompts. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Yuan to incorporate the teachings of Pu to use set of prompt categories. Doing so to provide faster and more accurate information extraction via user instructed or directed prompts (Pu Abstract line 4-11, The disclosed system and methods effectively leverage human actions and cognitive processes with integration with market available AI/ML technologies and models for faster and more accurate information extraction via user instructed and/or directed prompts. It also leverages the unique information elements captured via the prompt based processing for enhanced information retrieval and rendering experiences.). Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure Zhou, Yongchao, et al. "Large language models are human-level prompt engineers." The eleventh international conference on learning representations. 2022 (2022) – teaches a method for receiving a prompt from a user with information that requires an answers and scoring the responses. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALFREDO CAMPOS whose telephone number is (571)272-4504. The examiner can normally be reached 7:00 - 4:00 pm M - F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALFREDO CAMPOS/Examiner, Art Unit 2129 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Apr 04, 2024
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682285
PROVIDING A SECURE AND COLLABORATIVE FEEDBACK MECHANISM FOR MACHINE LEARNING MODELS
3y 4m to grant Granted Jul 14, 2026
Patent 12651086
METHOD AND SERVER FOR DEFENDING SERVICE FROM PERSONAL PRIVACY INFERENCE ATTACK
3y 7m to grant Granted Jun 09, 2026
Patent 12561407
ONE-PASS APPROACH TO AUTOMATED TIMESERIES FORECASTING
4y 3m to grant Granted Feb 24, 2026
Patent 12561559
Neural Network Training Method and Apparatus, Electronic Device, Medium and Program Product
4y 2m to grant Granted Feb 24, 2026
Patent 12554973
HIERARCHICAL DATA LABELING FOR MACHINE LEARNING USING SEMI-SUPERVISED MULTI-LEVEL LABELING FRAMEWORK
3y 6m to grant Granted Feb 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
73%
With Interview (-5.0%)
3y 6m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 9 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month