Prosecution Insights
Last updated: August 17, 2026
Application No. 18/627,230

MAPPING DISPARATE LANGUAGE-BASED DATASETS USING A LANGUAGE MODEL

Non-Final OA §101§103
Filed
Apr 04, 2024
Examiner
LE, HUNG VAN
Art Unit
Tech Center
Assignee
Intuit Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
19 currently pending
Career history
3
Total Applications
across all art units

Statute-Specific Performance

§101
35.9%
-4.1% vs TC avg
§103
64.1%
+24.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1–20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (an abstract idea) without reciting significantly more. Regarding independent claims 1, 15, and 20 Step 1 — whether the claim falls within any statutory category. See MPEP § 2106.03. Independent claim 1 is drawn to a method claim. Independent claim 15 is drawn to a system claim comprising a processor and a data repository. Independent claim 20 is drawn to a method claim. Therefore, each of independent claims 1, 15, and 20 falls within at least one of the four categories of statutory subject matter: claim 1 as a process, claim 15 as a machine, and claim 20 as a process. Accordingly, independent claims 1, 15, and 20 satisfy Step 1 of the subject matter eligibility analysis. Step 2A Prong 1 — whether the claim recites a judicial exception. See MPEP § 2106.04, subsection II; MPEP § 2106.04(a)(2), subsection III. Regarding independent claim 1, the claim recites the limitations of: "generating an updated prompt by at least injecting a feature into the prompt"; "combining the set of term embedding vectors with the set of prompt embedding vectors to generate a set of combined vectors"; "applying the language model to the set of combined vectors to generate a mapping between the feature and a term in the plurality of terms"; and "presenting the mapping." These limitations recite an abstract idea. Determining a correspondence between a feature and a term — i.e., mapping a feature to a term in a set of terms — is a concept that can be practically performed in the human mind, including the observation, evaluation, and judgment of associating one item of information (a feature) with a related item of information (a term). A person could mentally, or with the assistance of pen and paper, review a feature and a list of terms and select the term that corresponds to the feature. These limitations therefore fall within the mental processes grouping of abstract ideas. See MPEP § 2106.04(a)(2), subsection III. To the extent the limitations of "applying a vector generation controller to the updated prompt to generate a set of prompt embedding vectors," "applying the vector generation controller to a language dataset to generate a set of term embedding vectors," and "combining the set of term embedding vectors with the set of prompt embedding vectors to generate a set of combined vectors" encompass, under their broadest reasonable interpretation, the generation and combination of numeric vector representations, these limitations additionally recite mathematical concepts, namely mathematical relationships and calculations performed on vectors. See MPEP § 2106.04(a)(2), subsection I. Where a claim recites multiple abstract ideas, they are considered together as a single abstract idea. See MPEP § 2106.04(a)(2). Step 2A Prong 2 — whether the claim recites additional elements that integrate the exception into a practical application. See MPEP §§ 2106.04(d), 2106.05(a)–(c), (e)–(h). Regarding independent claim 1, the claim recites the additional elements of: "a language model trained to process natural language data," "a vector generation controller," and the receiving and presenting of the prompt and mapping. These additional elements, considered individually and in combination, do not integrate the recited judicial exception into a practical application. The recitation of performing the mapping "using" a language model and a vector generation controller amounts to mere instructions to apply the abstract idea using generic computer components that act as a tool. The claim recites only the outcome — generating a mapping — without any detail of how the language model or vector generation controller is improved or functions in an unconventional way; the models are used to generally apply the abstract idea without placing any limits on how they operate. See MPEP §§ 2106.05(f). Consistent with the 2024 AI SME Update (Example 47), reciting the use of a trained model to achieve a result, without more, is mere instructions to apply the exception. Furthermore, confining the mapping to the environment of language models and embedding vectors merely links the use of the abstract idea to a particular technological environment or field of use. See MPEP § 2106.05(h). The steps of "receiving a prompt" and "presenting the mapping" constitute insignificant extra-solution activity — necessary data gathering (receiving an input) and data outputting (presenting a result) that are incidental to the abstract idea. See MPEP § 2106.05(g). The claim does not recite any improvement to the functioning of a computer or to another technology or technical field (MPEP § 2106.05(a)); the specification describes the language model (e.g., ChatGPT) and the vector generation controller (e.g., an ADA-002 embedding model) as pre-existing, generic models used as tools, and the claimed advance lies in the abstract mapping itself rather than in any improvement to the models or to computer functionality. Accordingly, the additional elements, individually and in combination, do not integrate the judicial exception into a practical application. The claim is therefore directed to the abstract idea (Step 2A: YES). Step 2B — whether the claim amounts to significantly more than the judicial exception. See MPEP § 2106.05. Regarding independent claim 1, the additional elements identified in Step 2A Prong Two — the language model trained to process natural language data, the vector generation controller, and the receiving and presenting steps — are re-evaluated, individually and in combination, and do not amount to significantly more than the judicial exception. The language model and vector generation controller are recited at a high level of generality and constitute generic computer components used as a tool to apply the abstract idea; this is well-understood, routine, and conventional activity. See MPEP §§ 2106.05(d), 2106.05(f). The receiving and presenting steps, previously identified as insignificant extra-solution activity, are re-evaluated at Step 2B and are found to be well-understood, routine, and conventional activities — receiving data as an input to a model and outputting/displaying a result are conventional computer functions long prevalent in the art. See MPEP § 2106.05(d), (g). Limiting the abstract idea to the technological environment of language models does not supply an inventive concept. See MPEP § 2106.05(h). Considered as an ordered combination, the additional elements add nothing that is not already present when the elements are considered individually; the combination merely applies the abstract mapping using generic models as tools. Accordingly, the claim does not include additional elements that, individually or in combination, amount to significantly more than the judicial exception (Step 2B: NO). Independent claim 1 is therefore rejected under 35 U.S.C. § 101. Regarding independent claim 15, the claim recites the same mapping limitations identified in claim 1 and therefore recites the same abstract idea (Step 2A Prong 1: YES) for the reasons set forth above. In addition to the abstract idea, claim 15 recites the additional elements of "a processor," "a data repository in communication with the processor," "a language model executable by the processor," "a vector generation controller executable by the processor," and "a mapping controller." These additional elements are generic computer components recited at a high level of generality that merely act as tools on which the abstract idea is performed, amounting to mere instructions to apply the exception and generally linking it to a technological environment (Step 2A Prong 2: NO; MPEP §§ 2106.05(f), (h)). At Step 2B, these generic components — a processor, a data repository/memory, and software controllers — are well-understood, routine, and conventional, and neither individually nor in combination amount to significantly more than the judicial exception (Step 2B: NO; MPEP § 2106.05(d)). Independent claim 15 is therefore rejected under 35 U.S.C. § 101. Regarding independent claim 20, the claim recites the limitations of "transforming the list of terms, the verbiage, and the metadata into a language dataset comprising a set of data structure files" and the underlying association of each term with its corresponding verbiage and metadata, which recite the abstract idea of organizing, associating, and evaluating information — a mental process that could be performed in the human mind or with pen and paper — and, to the extent the "applying a vector generation controller … to generate a plurality of term embedding vectors" limitation encompasses generating numeric vector representations, a mathematical concept (Step 2A Prong 1: YES; MPEP § 2106.04(a)(2), subsections I and III). The additional elements — "a vector generation controller," "a non-transitory computer readable storage medium," and the "receiving," "storing," steps — are generic computer components and conventional data gathering/storage operations that apply the abstract idea, generally link it to a technological environment, and constitute insignificant extra-solution activity, and thus do not integrate the exception into a practical application (Step 2A Prong 2: NO; MPEP §§ 2106.05(f), (g), (h)). At Step 2B, storing data on a computer-readable medium and generating embeddings using a generic controller are well-understood, routine, and conventional, and do not, individually or in combination, amount to significantly more (Step 2B: NO; MPEP § 2106.05(d)). Independent claim 20 is therefore rejected under 35 U.S.C. § 101. Regarding dependent claims 2–14 and 16–19 Regarding dependent claims 2–14 and 16–19, these claims merely narrow the previously identified abstract idea limitations recited in independent claims 1, 15, and 20. Step 1 — whether the claim falls within any statutory category. See MPEP § 2106.03. Dependent claims 2–14 depend from independent claim 1, which is drawn to a method, and therefore each falls within the process category. Dependent claims 16–19 depend from independent claim 15, which is drawn to a system, and therefore each falls within the machine category. Step 2A Prong 1, Step 2A Prong 2, and Step 2B. For the reasons described above with respect to the independent claims, the judicial exceptions recited in these dependent claims are not meaningfully integrated into a practical application, nor do they amount to significantly more than the abstract idea. The additional limitations introduced in the dependent claims further define the abstract idea and constitute mental processes and mathematical concepts — including generating the language dataset by adding terms and metadata (claims 2, 16); applying a prediction model and determining the feature that contributed most to a prediction (claims 3, 4, 17, 18); retrieving a prompt template (claim 5); testing a mapping against a list of allowed mappings and adjusting and re-generating the mapping (claims 6, 7); expressing the mapping as an object notation data structure (claim 8); generating and transmitting messages conveying the mapping (claims 9, 10, 19); converting between data-structure formats (claim 11); determining a semantic distance and generating and substituting a new term (claims 12, 13, 14) — which are practically capable of being performed in the human mind or with the assistance of pen and paper, or recite mathematical relationships and calculations. The additional elements recited in these dependent claims (a prediction model, a language model, a computing device, a storage medium) are generic computer components that apply the abstract idea, generally link it to a technological environment, or add insignificant extra-solution activity, and are well-understood, routine, and conventional. See MPEP §§ 2106.05(d), (f), (g), (h). Accordingly, dependent claims 2–14 and 16–19 also recite abstract ideas that do not integrate into a practical application and do not amount to significantly more than the judicial exception. Therefore, claims 2–14 and 16–19 are rejected under 35 U.S.C. § 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 2, 5, 9-11, 15, 16, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (Lewis), Non-Patent Literature, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," NeurIPS, published December 2020, in view of Saxena (Saxena), U.S. Patent Application Publication No. US 2024/0330579 A1, and further in view of Hernandez et al. (Hernandez), U.S. Patent No. US 11,922,495 B1. Regarding Claim 1, (Lewis) teaches a method comprising: "receiving a prompt" "generating an updated prompt" "applying a vector generation controller to the updated prompt to generate a set of prompt embedding vectors;" "applying the vector generation controller to a language dataset to generate a set of term embedding vectors," "combining the set of term embedding vectors with the set of prompt embedding vectors to generate a set of combined vectors;" "applying the language model to the set of combined vectors to generate a mapping between the feature and a term in the plurality of terms; and" "presenting the mapping." As to "receiving a prompt," (Lewis) discloses receiving an input query x that is provided to the model: (Lewis) discloses models that "use the input sequence x to retrieve text documents z and use them as additional context when generating the target sequence y" (Lewis, p. 2, § 2, Methods; p. 2, Fig. 1, "Query"). The received input query x reads on the recited prompt. Thus (Lewis) teaches receiving a prompt. As to "generating an updated prompt," (Lewis) discloses updating the received input through fine-tuning, wherein the retriever and generator components are fine-tuned jointly on the input-output pairs such that the input provided to the model is updated during end-to-end fine-tuning (Lewis, p. 2, § 2, Methods; p. 4, § 2.4, Training). The input as updated through fine-tuning reads on the recited updated prompt. Thus (Lewis) teaches generating an updated prompt. As to "applying a vector generation controller to the updated prompt to generate a set of prompt embedding vectors," (Lewis) discloses a query encoder that transforms the input into a dense vector representation: "q(x) = BERTq(x)" (Lewis, p. 3, § 2.2). This query encoder is the recited vector generation controller, and the dense query representation q(x) that it produces from the input is the recited set of prompt embedding vectors. Thus (Lewis) teaches applying a vector generation controller to the updated prompt to generate a set of prompt embedding vectors. As to "applying the vector generation controller to a language dataset to generate a set of term embedding vectors," (Lewis) discloses that the same encoder architecture is applied to the corpus, wherein a document encoder transforms each document of the language dataset into a dense vector representation: "d(z) = BERTd(z)" (Lewis, p. 3, § 2.2), and further discloses that "we use the document encoder to compute an embedding for each document, and build a single MIPS index" (Lewis, p. 4, § 3). The dense document representations d(z) so produced are the recited set of term embedding vectors. Thus (Lewis) teaches applying the vector generation controller to a language dataset to generate a set of term embedding vectors. As to "combining the set of term embedding vectors with the set of prompt embedding vectors to generate a set of combined vectors," (Lewis) discloses combining the very same two sets of vectors generated in the two immediately preceding steps. Specifically, having generated the prompt embedding vectors q(x) by applying the query encoder to the input (Lewis, p. 3, § 2.2) and the term embedding vectors d(z) by applying the document encoder to the language dataset (Lewis, p. 3, § 2.2; p. 4, § 3), (Lewis) combines those two sets of vectors by computing the inner product of the term embedding vectors with the prompt embedding vectors: "pη(z|x) ∝ exp(d(z)ᵀ q(x))," where "d(z) = BERTd(z)" and "q(x) = BERTq(x)" (Lewis, p. 3, § 2.2). The expression d(z)ᵀ q(x) is a direct combination of the set of term embedding vectors d(z) with the set of prompt embedding vectors q(x), and the resulting combination is used to select, by Maximum Inner Product Search (MIPS), the documents whose term embedding vectors are combined with the prompt embedding vectors (Lewis, p. 3, § 2.2). (Lewis) further discloses that the input and the retrieved content so selected are then combined and supplied together as the input to the generator, wherein the input x and the retrieved content z are concatenated for input to the BART generator (Lewis, p. 4, § 2.3). The result of combining the term embedding vectors with the prompt embedding vectors, which is thereafter supplied to the language model, reads on the recited set of combined vectors. This construction is consistent with the specification, which states that combining the vectors includes "providing each of the set of term embedding vectors and the set of prompt embedding vectors directly to the language model" (Specification, ¶ [0081]). Thus (Lewis) teaches combining the set of term embedding vectors with the set of prompt embedding vectors to generate a set of combined vectors. As to "applying the language model to the set of combined vectors to generate a mapping between the feature and a term in the plurality of terms," (Lewis) discloses applying a pre-trained seq2seq language model (BART) to the combined vectors resulting from the preceding combining step, wherein the generator produces its output conditioned jointly on the input and the retrieved document, "pθ(yi | x, z, y1:i−1)" (Lewis, p. 3, § 2.1; p. 2, Fig. 1), and further discloses that the model may generate a mapping of an input to a target term because "RAG can be used for sequence classification tasks by considering the target class as a target sequence of length one" (Lewis, p. 3, § 2.1). Thus (Lewis) teaches applying the language model to the set of combined vectors to generate a mapping between the input and a term. As to "presenting the mapping," (Lewis) discloses producing and outputting the final prediction y (Lewis, p. 2, Fig. 1, "Output"; p. 3, § 2.1). Thus (Lewis) teaches presenting the mapping. (Lewis) teaches the language model receiving and being directed by the input, namely that the input query x is supplied to the seq2seq generator, which generates the output conditioned on that input (Lewis, p. 2, § 2, Methods; p. 3, § 2.1). However, (Lewis) does not teach: receiving a prompt for a language model trained to process natural language data, wherein the prompt commands the language model; generating an updated prompt by at least injecting a feature into the prompt; In the same field of endeavor, (Saxena) teaches "for a language model trained to process natural language data, wherein the prompt commands the language model." Specifically, (Saxena) discloses a large language model trained to process natural language data, to which a prompt is provided as a command directing the model to generate text: "The prompt templates store 215 is a repository of prompt templates that can be used by other components … to generate prompts to the LLM 130" (Saxena, p. 14, [0028]), and the system, "[u]sing the prompt template, the first value, and the second value, … generates a prompt to a large language model" that causes the model to generate text (Saxena, Abstract; p. 13, [0018]). Thus (Saxena) teaches that the prompt is for a language model trained to process natural language data and that the prompt commands the language model. (Saxena) further teaches "by at least injecting a feature into the prompt." Specifically, (Saxena) discloses injecting a determined value into a prompt template to form the prompt: "Each prompt template includes at least one parameter, a value for which is selected to generate a unique prompt to the LLM 130" (Saxena, p. 14, [0028]), and a first value for a first parameter is determined and injected, and a second value is requested and injected, into the template to generate the prompt (Saxena, Abstract; p. 13, [0018]). The injection of the determined parameter value into the prompt reads on injecting a feature into the prompt. Thus (Saxena) teaches generating the updated prompt by at least injecting a feature into the prompt. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the retrieval-augmented generation pipeline of (Lewis) with the prompt-command and feature-injection of (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) recognizes that "it can be difficult to generate prompts that will cause the text produced by a large language model to be appropriate for the context … especially difficult to generate these prompts programmatically" (Saxena, p. 13, [0017]), such that one would provide the input of (Lewis) as a prompt that commands the language model and would inject the feature into that prompt as taught by (Saxena), in order to programmatically generate a prompt that directs and controls the mapping produced by the language model. The combination of (Lewis) and (Saxena), however, does not teach: "wherein the language dataset comprises a plurality of terms disparate from the feature." In the same field of endeavor, (Hernandez) teaches "wherein the language dataset comprises a plurality of terms disparate from the feature." Specifically, (Hernandez) discloses an intelligent lending platform in which a machine-learned model analyzes a query and, when the query is denied, outputs the reasons for the denial to the user (Hernandez, Abstract), and further discloses that the model generates output signals wherein a denied query is accompanied by "one or more respective reasons why the new lending query is denied," which are sent "for presentation on the user device" (Hernandez, col. 63–64, claim). The model's contributing signals (the feature) and the corresponding human-readable reasons (the terms of the language dataset) are distinct, disparate representations lacking a fixed one-to-one correspondence — the model feature is an opaque signal, whereas the term is a human-readable reason. Thus (Hernandez) teaches that the language dataset comprises a plurality of terms disparate from the feature. (Lewis), (Saxena), and (Hernandez) are analogous to the claimed invention as all are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the retrieval-augmented generation pipeline of (Lewis) and the prompt-command and feature-injection of (Saxena) with the disparate feature-to-reason mapping of (Hernandez). The motivation to combine (Lewis), (Saxena), and (Hernandez) is that (Hernandez) recognizes the need to explain an opaque model's decision by mapping the model's contributing feature to a human-understandable reason for presentation to the user (Hernandez, Abstract; col. 63–64), such that one would apply the pipeline of (Lewis) and (Saxena) to map a feature to a term in a disparate language dataset, thereby producing a human-understandable mapping where no fixed correspondence between the feature and the term otherwise exists. Regarding Claim 2, (Lewis) teaches the method of claim 1, further comprising: "generating the language dataset by: adding the plurality of terms to the language dataset, and" As to "generating the language dataset by: adding the plurality of terms to the language dataset," (Lewis) discloses generating the language dataset by adding a plurality of terms (documents) to it: "Each Wikipedia article is split into disjoint 100-word chunks, to make a total of 21M documents. We use the document encoder to compute an embedding for each document, and build a single MIPS index" (Lewis, p. 4, § 3). Thus (Lewis) teaches generating the language dataset by adding the plurality of terms to the language dataset. The combination of (Lewis) and (Saxena), however, does not teach: "adding metadata related to the plurality of terms to the language dataset." In the same field of endeavor, (Hernandez) teaches "adding metadata related to the plurality of terms to the language dataset." Specifically, (Hernandez) discloses that each term (signal) of the dataset is added together with related metadata, namely a risk-metric value and a direction indication, wherein each signal is "associated with a value of a risk metric and an indication of whether the signal increases or decreases a corresponding risk metric" (Hernandez, col. 63–64, claim). This risk-metric value and direction indication constitute metadata related to the plurality of terms that is added to the dataset. Thus (Hernandez) teaches adding metadata related to the plurality of terms to the language dataset. (Lewis) and (Hernandez) are analogous to the claimed invention as both are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the generation of the language dataset by adding the plurality of terms, as taught by (Lewis), with the addition of metadata related to those terms, as taught by (Hernandez). The motivation to combine (Lewis) and (Hernandez) is that (Hernandez) associates each term with related metadata such as a risk-metric value and direction indication (Hernandez, col. 63–64), such that one would add such metadata to the language dataset of (Lewis) in order to provide the language model with additional context corresponding to each term, thereby improving the accuracy of the mapping between a feature and a term in the disparate language dataset. Regarding Claim 5, (Lewis) teaches obtaining the prompt, namely receiving an input query x that is provided to the seq2seq generator (Lewis, p. 2, Fig. 1; p. 3, § 2.1). However, (Lewis) does not teach: "generating, prior to receiving, the prompt by retrieving a prompt template stored as natural language text," "wherein the prompt comprises the prompt template." In the same field of endeavor, (Saxena) teaches "generating, prior to receiving, the prompt by retrieving a prompt template stored as natural language text." Specifically, (Saxena) discloses a repository of prompt templates from which a prompt template is retrieved and used to generate the prompt: "The prompt templates store 215 is a repository of prompt templates that can be used by other components … to generate prompts to the LLM 130" (Saxena, p. 14, [0028]), and the system "obtains a prompt template associated with a section of the webpage" and, using the prompt template, "generates a prompt to a large language model" (Saxena, Abstract; see also p. 13, [0018]). Each prompt template is stored as natural-language text that includes at least one parameter (Saxena, p. 14, [0028]). Because the prompt is generated from the retrieved prompt template before it is provided to the language model, (Saxena) teaches generating, prior to receiving, the prompt by retrieving a prompt template stored as natural language text. (Saxena) further teaches "wherein the prompt comprises the prompt template." Specifically, (Saxena) discloses that the generated prompt is formed from the prompt template together with the selected parameter values: "Using the prompt template, the first value, and the second value, the system generates a prompt to a large language model" (Saxena, Abstract). Because the generated prompt is formed from and incorporates the prompt template, the prompt comprises the prompt template. Thus (Saxena) teaches wherein the prompt comprises the prompt template. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the generation of the prompt by retrieving a stored prompt template as taught by (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) retrieves a stored, reusable prompt template to programmatically generate a context-appropriate prompt for the language model (Saxena, p. 14, [0028]; Abstract), such that one would generate the prompt of (Lewis) by retrieving a stored prompt template in order to standardize and efficiently produce prompts that reliably command the language model to generate the mapping. Regarding Claim 9, (Lewis) teaches presenting the mapping, namely generating and outputting the final prediction y (Lewis, p. 2, Fig. 1, "Output"; p. 3, § 2.1). However, (Lewis) does not teach: "generating a response text based on the term," "generating an electronic message using the response text, and" "transmitting the electronic message to a user." In the same field of endeavor, (Hernandez) teaches "generating a response text based on the term." Specifically, (Hernandez) discloses that, when a lending query is denied, the model provides "one or more respective reasons why the new lending query is denied" (Hernandez, col. 63–64, claim), and outputs the reason(s) for the denial to the user (Hernandez, Abstract). The reason text produced from the denial reason (the term) reads on generating a response text based on the term. Thus (Hernandez) teaches generating a response text based on the term. (Hernandez) further teaches "generating an electronic message using the response text, and transmitting the electronic message to a user." Specifically, (Hernandez) discloses generating an indication conveying the reason and "sending, from the payment service and for presentation on the user device, an indication" of the query result to the user (Hernandez, col. 63–64, claim; see also Abstract). The indication conveyed for presentation on the user device is an electronic message generated using the response text and transmitted to a user. Thus (Hernandez) teaches generating an electronic message using the response text and transmitting the electronic message to a user. (Lewis) and (Hernandez) are analogous to the claimed invention as both are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the generation and transmission of a message to a user as taught by (Hernandez). The motivation to combine (Lewis) and (Hernandez) is that (Hernandez) generates and transmits to the user an electronic message conveying the reason for a model's decision (Hernandez, Abstract; col. 63–64), such that one would present the mapping generated by (Lewis) by generating a response text based on the term, generating an electronic message using the response text, and transmitting the electronic message to the user, thereby conveying to the user, in a human-understandable form, the term to which the feature was mapped. Regarding Claim 10, (Lewis) teaches presenting the mapping, namely generating and outputting the final prediction y (Lewis, p. 2, Fig. 1, "Output"; p. 3, § 2.1). However, (Lewis) does not teach: "generating an electronic message comprising the mapping;" "transmitting the mapping to a computing device;" "receiving, from the computing device, a presentation electronic message; and" "transmitting the presentation electronic message to a user." In the same field of endeavor, (Hernandez) teaches "generating an electronic message comprising the mapping" and "transmitting the mapping to a computing device." Specifically, (Hernandez) discloses generating an indication that conveys the model's result, including the reason(s) mapped to a denied query, and "sending, from the payment service and for presentation on the user device, an indication" of that result (Hernandez, col. 63–64, claim; see also Abstract). Generating the indication conveying the mapped reason reads on generating an electronic message comprising the mapping, and sending that indication to the user device reads on transmitting the mapping to a computing device. Thus (Hernandez) teaches these limitations. (Hernandez) further teaches "receiving, from the computing device, a presentation electronic message; and transmitting the presentation electronic message to a user." Specifically, (Hernandez) discloses a payment service that communicates with user devices and other systems over one or more networks and returns communications for presentation to buyers and users (Hernandez, col. 59–60; Abstract). In such an arrangement, a presentation message determined at the computing device is received back at the payment service and transmitted for presentation to the user, which reads on receiving, from the computing device, a presentation electronic message and transmitting the presentation electronic message to a user. Thus (Hernandez) teaches these limitations. (Lewis) and (Hernandez) are analogous to the claimed invention as both are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the generation, transmission, and return of messages between the service and a computing device for presentation to a user as taught by (Hernandez). The motivation to combine (Lewis) and (Hernandez) is that (Hernandez) exchanges messages with a computing device over a network to present a model's result to a user (Hernandez, col. 59–60; Abstract), such that one would present the mapping generated by (Lewis) by generating an electronic message comprising the mapping, transmitting the mapping to a computing device, receiving a presentation electronic message from that device, and transmitting the presentation electronic message to the user, thereby enabling the mapping to be formatted and delivered to the user for presentation. Regarding Claim 11, (Lewis) teaches: "wherein applying the vector generation controller to the language dataset comprises applying the vector generation controller to the plurality of second object notation data structures." As to "wherein applying the vector generation controller to the language dataset comprises applying the vector generation controller to the plurality of second object notation data structures," (Lewis) discloses applying the document encoder "d(z) = BERTd(z)" to each data structure (document) of the language dataset to generate the term embedding vectors, and further discloses that "we use the document encoder to compute an embedding for each document, and build a single MIPS index" (Lewis, p. 3, § 2.2; p. 4, § 3). Applying the document encoder to each of the plurality of data structures of the language dataset reads on applying the vector generation controller to the plurality of second object notation data structures. Thus (Lewis) teaches this limitation. (Lewis) teaches forming the updated prompt, namely providing an input to the seq2seq generator (Lewis, p. 2, Fig. 1). However, (Lewis) does not teach: "wherein the feature comprises a first tab separated value data structure that further includes a variable prompt," "wherein generating the updated prompt comprises injecting the first object notation data structure into the prompt." In the same field of endeavor, (Saxena) teaches "wherein the feature comprises a first tab separated value data structure that further includes a variable prompt" and "wherein generating the updated prompt comprises injecting the first object notation data structure into the prompt." Specifically, (Saxena) discloses that a prompt template includes at least one parameter whose value is injected into the template to form the prompt, and that a second, variable parameter value is added to the prompt: "Each prompt template includes at least one parameter, a value for which is selected to generate a unique prompt to the LLM 130" (Saxena, p. 14, [0028]), and "A request to provide input for a second value of a second parameter is sent for display to a user. Using the prompt template, the first value, and the second value, the system generates a prompt to a large language model" (Saxena, Abstract; p. 13, [0018]). Injecting the feature's data structure and the second, variable parameter value into the prompt reads on the feature including a variable prompt and on generating the updated prompt by injecting the first data structure into the prompt. Thus (Saxena) teaches these limitations. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the injection of a data structure and a variable prompt into the prompt as taught by (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) injects parameter values, including a variable value, into a prompt template to programmatically form a context-appropriate prompt (Saxena, p. 14, [0028]; Abstract), such that one would inject the feature's data structure and a variable prompt into the prompt of (Lewis) in order to controllably form the updated prompt that commands the language model. The combination of (Lewis) and (Saxena), however, does not teach: "wherein the language dataset comprises a plurality of second tab separated value data structures," "converting, prior to applying the vector generation controller to the feature, the first tab separated value data structure into a first object notation data structure," "converting, prior to applying the vector generation controller to the language dataset, the plurality of second tab separated value data structures into a plurality of second object notation data structures." A person of ordinary skill in the art would readily infer representing the feature and the language dataset as tab separated value data structures and converting those tab separated value data structures into object notation data structures prior to applying the vector generation controller. Representing data as a tab separated value data structure, and converting a tab separated value data structure into an object notation data structure (e.g., converting a TSV structure into a JSON structure), are notoriously well-known and conventional in the art of data processing; tab separated value and object notation formats are standardized formats for storing tabular and structured data, respectively, and converting data from one such format to the other prior to further processing is a routine and well-understood operation. A person of ordinary skill in the art would readily infer performing such a conversion prior to vectorizing the feature and the language dataset of the combination of (Lewis) and (Saxena), as this amounts to no more than the application of a known, standardized data-format conversion to known data to yield predictable results. It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to represent the feature and the language dataset of the combination of (Lewis) and (Saxena) as tab separated value data structures, and to convert those tab separated value data structures into object notation data structures prior to applying the vector generation controller. The motivation to do so is that converting tabular data into a standardized object notation format yields a structured, machine-readable representation that is reliably parsed and processed by downstream components, which is a recognized benefit of such formats and a predictable use of a known technique to yield predictable results. Regarding Claim 15, (Lewis) teaches a system comprising: "a processor;" "a data repository in communication with the processor, the data repository storing:" "a prompt and an updated prompt," "a language dataset comprising a plurality of terms, which includes a term," "a set of prompt embedding vectors, a set of term embedding vectors, and a set of combined vectors, and" "a feature and a mapping between the feature and the term," "a language model executable by the processor" "a vector generation controller executable by the processor; and" "a mapping controller programmed, when executed by the processor, to perform a computer-implemented method comprising:" "receiving the prompt," "generating the updated prompt" "applying the vector generation controller to the updated prompt to generate a set of prompt embedding vectors," "applying the vector generation controller to the language dataset to generate the set of term embedding vectors," "combining the set of term embedding vectors with the set of prompt embedding vectors to generate the set of combined vectors," "applying the language model to the set of combined vectors to generate the mapping, and" "returning the mapping." As to "a processor" and "a data repository in communication with the processor," (Lewis) discloses a computer-implemented system in which the models are executed and in which the computed document representations are stored as a searchable index: "we use the document encoder to compute an embedding for each document, and build a single MIPS index using FAISS" (Lewis, p. 4, § 3), and the document encoder "(and index)" is kept fixed during operation (Lewis, p. 4, § 2.4). The computing system that executes the encoders and the generator is the recited processor, and the stored index in communication with that system is the recited data repository. Thus (Lewis) teaches a processor and a data repository in communication with the processor. As to the data repository storing "a prompt and an updated prompt," (Lewis) discloses receiving and holding an input query x that is provided to the model, wherein the models "use the input sequence x to retrieve text documents z and use them as additional context when generating the target sequence y" (Lewis, p. 2, § 2, Methods; p. 2, Fig. 1, "Query"), and further discloses updating that input through fine-tuning, wherein the retriever and generator are fine-tuned jointly on the input-output pairs (Lewis, p. 2, § 2, Methods; p. 4, § 2.4, Training). The received input query x is the recited prompt, and the input as updated through fine-tuning is the recited updated prompt. Thus (Lewis) teaches the data repository storing a prompt and an updated prompt. As to the data repository storing "a language dataset comprising a plurality of terms, which includes a term," (Lewis) discloses a stored corpus of documents constituting the language dataset: "Each Wikipedia article is split into disjoint 100-word chunks, to make a total of 21M documents" (Lewis, p. 4, § 3). The stored documents are the recited plurality of terms, which includes a term. Thus (Lewis) teaches this limitation. As to the data repository storing "a set of prompt embedding vectors, a set of term embedding vectors, and a set of combined vectors," and "a feature and a mapping between the feature and the term," (Lewis) discloses generating and storing the document representations as the index (Lewis, p. 3, § 2.2; p. 4, § 3), and generating the query representation, the combination of the two representations, and the resulting output during operation of the system (Lewis, p. 3, §§ 2.1–2.2; p. 2, Fig. 1). Thus (Lewis) teaches the data repository storing these items. As to "a language model executable by the processor," (Lewis) discloses a pre-trained seq2seq generator (BART) that is executed by the computing system to generate the output (Lewis, p. 2, Fig. 1; p. 3, § 2.1). Thus (Lewis) teaches a language model executable by the processor. As to "a vector generation controller executable by the processor," (Lewis) discloses encoders executed by the computing system to transform text into dense vector representations, namely "q(x) = BERTq(x)" and "d(z) = BERTd(z)" (Lewis, p. 3, § 2.2). Thus (Lewis) teaches a vector generation controller executable by the processor. As to "a mapping controller programmed, when executed by the processor, to perform a computer-implemented method comprising: receiving the prompt," (Lewis) discloses receiving the input query x that is provided to the model (Lewis, p. 2, § 2, Methods; p. 2, Fig. 1, "Query"). Thus (Lewis) teaches receiving the prompt. As to "generating the updated prompt," (Lewis) discloses updating the received input through fine-tuning, wherein the retriever and generator components are fine-tuned jointly such that the input provided to the model is updated during end-to-end fine-tuning (Lewis, p. 2, § 2, Methods; p. 4, § 2.4, Training). Thus (Lewis) teaches generating the updated prompt. As to "applying the vector generation controller to the updated prompt to generate a set of prompt embedding vectors," (Lewis) discloses applying the query encoder to the input to produce a dense query representation: "q(x) = BERTq(x)" (Lewis, p. 3, § 2.2). The query encoder is the recited vector generation controller, and the representation q(x) is the recited set of prompt embedding vectors. Thus (Lewis) teaches this limitation. As to "applying the vector generation controller to the language dataset to generate the set of term embedding vectors," (Lewis) discloses applying the same encoder architecture to the corpus, wherein the document encoder produces a dense representation of each document, "d(z) = BERTd(z)" (Lewis, p. 3, § 2.2), and "we use the document encoder to compute an embedding for each document, and build a single MIPS index" (Lewis, p. 4, § 3). The representations d(z) are the recited set of term embedding vectors. Thus (Lewis) teaches this limitation. As to "combining the set of term embedding vectors with the set of prompt embedding vectors to generate the set of combined vectors," (Lewis) discloses combining the very same two sets of vectors generated in the two immediately preceding steps. Specifically, having generated the prompt embedding vectors q(x) by applying the query encoder to the input (Lewis, p. 3, § 2.2) and the term embedding vectors d(z) by applying the document encoder to the language dataset (Lewis, p. 3, § 2.2; p. 4, § 3), (Lewis) combines those two sets of vectors by computing the inner product of the term embedding vectors with the prompt embedding vectors: "pη(z|x) ∝ exp(d(z)ᵀ q(x))," where "d(z) = BERTd(z)" and "q(x) = BERTq(x)" (Lewis, p. 3, § 2.2). The expression d(z)ᵀ q(x) is a direct combination of the set of term embedding vectors d(z) with the set of prompt embedding vectors q(x), and the resulting combination is used to select, by Maximum Inner Product Search (MIPS), the documents whose term embedding vectors are combined with the prompt embedding vectors (Lewis, p. 3, § 2.2). (Lewis) further discloses that the input and the retrieved content so selected are then combined and supplied together as the input to the generator, wherein the input x and the retrieved content z are concatenated for input to the BART generator (Lewis, p. 4, § 2.3). The result of combining the term embedding vectors with the prompt embedding vectors, which is thereafter supplied to the language model, reads on the recited set of combined vectors. This construction is consistent with the specification, which states that combining the vectors includes "providing each of the set of term embedding vectors and the set of prompt embedding vectors directly to the language model" (Specification, ¶ [0081]). Thus (Lewis) teaches combining the set of term embedding vectors with the set of prompt embedding vectors to generate the set of combined vectors. As to "applying the language model to the set of combined vectors to generate the mapping," (Lewis) discloses applying the seq2seq language model to the combined vectors resulting from the preceding combining step, wherein the generator produces its output conditioned jointly on the input and the retrieved document, "pθ(yi | x, z, y1:i−1)" (Lewis, p. 3, § 2.1; p. 2, Fig. 1), and further discloses that the model may generate a mapping of an input to a target term because "RAG can be used for sequence classification tasks by considering the target class as a target sequence of length one" (Lewis, p. 3, § 2.1). Thus (Lewis) teaches this limitation. As to "returning the mapping," (Lewis) discloses producing and outputting the final prediction y (Lewis, p. 2, Fig. 1, "Output"; p. 3, § 2.1). Thus (Lewis) teaches returning the mapping. (Lewis) teaches the language model receiving and being directed by the input, namely that the input query x is supplied to the seq2seq generator, which generates the output conditioned on that input (Lewis, p. 2, § 2, Methods; p. 3, § 2.1). However, (Lewis) does not teach: "and trained to process natural language data, wherein the prompt commands the language model;" "by injecting the feature into the prompt," In the same field of endeavor, (Saxena) teaches "and trained to process natural language data, wherein the prompt commands the language model." Specifically, (Saxena) discloses a large language model trained to process natural language data, to which a prompt is provided as a command directing the model to generate text: "The prompt templates store 215 is a repository of prompt templates that can be used by other components … to generate prompts to the LLM 130" (Saxena, p. 14, [0028]), and the system, "[u]sing the prompt template, the first value, and the second value, … generates a prompt to a large language model" that causes the model to generate text (Saxena, Abstract; p. 13, [0018]). Thus (Saxena) teaches a language model trained to process natural language data, wherein the prompt commands the language model. (Saxena) further teaches "by injecting the feature into the prompt." Specifically, (Saxena) discloses injecting a determined value into a prompt template to form the prompt: "Each prompt template includes at least one parameter, a value for which is selected to generate a unique prompt to the LLM 130" (Saxena, p. 14, [0028]), and a first value for a first parameter is determined and injected, and a second value is requested and injected, into the template to generate the prompt (Saxena, Abstract; p. 13, [0018]). The injection of the determined parameter value into the prompt reads on injecting the feature into the prompt. Thus (Saxena) teaches generating the updated prompt by injecting the feature into the prompt. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the retrieval-augmented generation system of (Lewis) with the prompt-command and feature-injection of (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) recognizes that "it can be difficult to generate prompts that will cause the text produced by a large language model to be appropriate for the context … especially difficult to generate these prompts programmatically" (Saxena, p. 13, [0017]), such that one would provide the input of (Lewis) as a prompt that commands the language model and would inject the feature into that prompt as taught by (Saxena), in order to programmatically generate a prompt that directs and controls the mapping produced by the language model. The combination of (Lewis) and (Saxena), however, does not teach: "wherein the feature is disparate from the plurality of terms;" In the same field of endeavor, (Hernandez) teaches "wherein the feature is disparate from the plurality of terms." Specifically, (Hernandez) discloses an intelligent lending platform in which a machine-learned model analyzes a query and, when the query is denied, outputs the reasons for the denial to the user (Hernandez, Abstract), and further discloses that the model generates output signals wherein a denied query is accompanied by "one or more respective reasons why the new lending query is denied," which are sent "for presentation on the user device" (Hernandez, col. 63–64, claim). The model's contributing signals (the feature) and the corresponding human-readable reasons (the terms of the language dataset) are distinct, disparate representations lacking a fixed one-to-one correspondence — the model feature is an opaque signal, whereas the term is a human-readable reason. Thus (Hernandez) teaches that the feature is disparate from the plurality of terms. (Lewis), (Saxena), and (Hernandez) are analogous to the claimed invention as all are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the retrieval-augmented generation system of (Lewis) and the prompt-command and feature-injection of (Saxena) with the disparate feature-to-reason mapping of (Hernandez). The motivation to combine (Lewis), (Saxena), and (Hernandez) is that (Hernandez) recognizes the need to explain an opaque model's decision by mapping the model's contributing feature to a human-understandable reason for presentation to the user (Hernandez, Abstract; col. 63–64), such that one would configure the system of (Lewis) and (Saxena) to map a feature that is disparate from the plurality of terms, thereby producing a human-understandable mapping where no fixed correspondence between the feature and the terms otherwise exists. Regarding Claim 16, the limitations of claim 16 are commensurate in scope with the limitations of claim 2. Accordingly, claim 16 is rejected under 35 U.S.C. 103 for the same reasons and under the same rationale set forth in the rejection of claim 2 above. Regarding Claim 19, the limitations of claim 19 are commensurate in scope with the limitations of claim 10. Accordingly, claim 19 is rejected under 35 U.S.C. 103 for the same reasons and under the same rationale set forth in the rejection of claim 10 above. Regarding Claim 20, (Lewis) teaches a method comprising: "receiving a list of terms;" "transforming the list of terms … into a language dataset comprising a set of data structure files," "applying a vector generation controller to the language dataset to generate a plurality of term embedding vectors, wherein each of the plurality of term embedding vectors comprises one or more embedding vectors corresponding to a file in the set of data structure files;" "storing the plurality of term embedding vectors in a non-transitory computer readable storage medium;" "storing … in the non-transitory computer readable storage medium," As to "receiving a list of terms," (Lewis) discloses receiving a corpus of text that constitutes the list of terms used by the model, wherein the non-parametric memory is "a dense vector index of Wikipedia" (Lewis, p. 1, Abstract) and "we use the December 2018 Wikipedia dump" as the document corpus (Lewis, p. 4, § 3). The received corpus of text entries is the recited list of terms. Thus (Lewis) teaches receiving a list of terms. As to "transforming the list of terms … into a language dataset comprising a set of data structure files," (Lewis) discloses transforming the received corpus into a set of discrete files constituting the language dataset: "Each Wikipedia article is split into disjoint 100-word chunks, to make a total of 21M documents" (Lewis, p. 4, § 3). The received corpus is thereby transformed into a set of 21M discrete documents, each of which is a data structure file of the language dataset. Thus (Lewis) teaches transforming the list of terms into a language dataset comprising a set of data structure files. As to "applying a vector generation controller to the language dataset to generate a plurality of term embedding vectors, wherein each of the plurality of term embedding vectors comprises one or more embedding vectors corresponding to a file in the set of data structure files," (Lewis) discloses applying the document encoder to the very language dataset produced by the immediately preceding transforming step, and generating one embedding per file thereof. Specifically, having split the corpus into the set of documents (files) (Lewis, p. 4, § 3), (Lewis) applies a document encoder that produces a dense vector representation of each such document, "d(z) = BERTd(z)" (Lewis, p. 3, § 2.2), and expressly states that "we use the document encoder to compute an embedding for each document" (Lewis, p. 4, § 3). The document encoder is the recited vector generation controller, and the per-document embeddings d(z) so generated are the recited plurality of term embedding vectors, each of which comprises one or more embedding vectors corresponding to a file in the set of data structure files because each embedding is computed for, and corresponds to, one document of the set. Thus (Lewis) teaches this limitation. As to "storing the plurality of term embedding vectors in a non-transitory computer readable storage medium," (Lewis) discloses storing the very term embedding vectors generated in the immediately preceding step as a persistent, searchable index: having computed an embedding for each document, (Lewis) "build[s] a single MIPS index using FAISS" from those embeddings (Lewis, p. 4, § 3), and the document encoder "(and index)" is thereafter kept fixed and accessed during operation (Lewis, p. 4, § 2.4). The stored index of term embedding vectors resides in a non-transitory computer readable storage medium of the computing system. Thus (Lewis) teaches storing the plurality of term embedding vectors in a non-transitory computer readable storage medium. As to "storing … in the non-transitory computer readable storage medium," as recited in the final two limitations, (Lewis) discloses storing the data used by the model in the same non-transitory computer readable storage medium in which the term embedding vectors are stored, namely the stored index that is retained and accessed during operation (Lewis, p. 4, § 3; p. 4, § 2.4). Thus (Lewis) teaches storing data in the non-transitory computer readable storage medium. (Lewis) teaches directing the language model, namely that the input query x is supplied to the seq2seq generator, which generates the output conditioned on that input (Lewis, p. 2, § 2, Methods; p. 3, § 2.1). However, (Lewis) does not teach: "a plurality of prompt templates … wherein each of the plurality of prompt templates includes a command to map a feature to one or more terms in the list of terms; and" "a plurality of variable prompts …, the plurality of variable prompts configured to modify the plurality of prompt templates." In the same field of endeavor, (Saxena) teaches "a plurality of prompt templates … wherein each of the plurality of prompt templates includes a command to map a feature to one or more terms in the list of terms." Specifically, (Saxena) discloses a stored repository containing a plurality of prompt templates, each of which includes at least one parameter that is used to command the language model by way of the generated prompt: "The prompt templates store 215 is a repository of prompt templates that can be used by other components … to generate prompts to the LLM 130," and "Each prompt template includes at least one parameter, a value for which is selected to generate a unique prompt to the LLM 130" (Saxena, p. 14, [0028]). Each such stored prompt template, which commands the language model to produce an output associating an injected value with the terms of the dataset, reads on a prompt template that includes a command to map a feature to one or more terms in the list of terms. Thus (Saxena) teaches this limitation. (Saxena) further teaches "a plurality of variable prompts …, the plurality of variable prompts configured to modify the plurality of prompt templates." Specifically, (Saxena) discloses that a second, variable parameter value is obtained and applied to the prompt template to form the prompt: "A request to provide input for a second value of a second parameter is sent for display to a user. Using the prompt template, the first value, and the second value, the system generates a prompt to a large language model" (Saxena, Abstract; p. 13, [0018]). These second, variable parameter values, which are applied to and thereby alter the stored prompt templates to produce a unique prompt, read on a plurality of variable prompts configured to modify the plurality of prompt templates. Thus (Saxena) teaches this limitation. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the term-embedding generation and storage of (Lewis) with the prompt templates and variable prompts of (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) recognizes that "it can be difficult to generate prompts that will cause the text produced by a large language model to be appropriate for the context … especially difficult to generate these prompts programmatically" (Saxena, p. 13, [0017]), such that one would store, together with the term embedding vectors of (Lewis), a plurality of prompt templates each including a mapping command and a plurality of variable prompts that modify those templates, in order to programmatically generate the prompts that direct the language model to map a feature to the stored terms. The combination of (Lewis) and (Saxena), however, does not teach: "receiving verbiage corresponding to the list of terms, wherein each term in the list of terms corresponds to a set of verbiage in the verbiage;" "receiving metadata, wherein each term in the list of terms corresponds to a set of metadata in the metadata;" "wherein each file in the set of data structure files comprises a triplet of a term, the set of verbiage for the term, and the set of metadata for the term;" In the same field of endeavor, (Hernandez) teaches "receiving verbiage corresponding to the list of terms, wherein each term in the list of terms corresponds to a set of verbiage in the verbiage." Specifically, (Hernandez) discloses that each term of the list of terms, in the form of a signal used by the lending model, has a corresponding human-readable reason, wherein the model provides "one or more respective reasons why the new lending query is denied" (Hernandez, col. 63–64, claim; Abstract). The human-readable reason corresponding to each signal is the recited set of verbiage corresponding to each term. Thus (Hernandez) teaches this limitation. (Hernandez) further teaches "receiving metadata, wherein each term in the list of terms corresponds to a set of metadata in the metadata." Specifically, (Hernandez) discloses that each term (signal) is received together with corresponding metadata, namely a risk-metric value and a direction indication, wherein each signal is "associated with a value of a risk metric and an indication of whether the signal increases or decreases a corresponding risk metric" (Hernandez, col. 63–64, claim). The risk-metric value and direction indication corresponding to each signal are the recited set of metadata corresponding to each term. Thus (Hernandez) teaches this limitation. (Hernandez) further teaches "wherein each file in the set of data structure files comprises a triplet of a term, the set of verbiage for the term, and the set of metadata for the term." Specifically, (Hernandez) discloses maintaining, for each term (signal), the association of that signal with both its corresponding human-readable reason and its corresponding risk-metric value and direction indication (Hernandez, col. 63–64, claim; Abstract). Each such association of a signal, its reason, and its risk-metric data constitutes a triplet of a term, the set of verbiage for the term, and the set of metadata for the term. Thus (Hernandez) teaches this limitation. (Lewis), (Saxena), and (Hernandez) are analogous to the claimed invention as all are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the term-embedding generation and storage of (Lewis) and the prompt templates and variable prompts of (Saxena) with the receiving of verbiage and metadata corresponding to each term and their association into triplets as taught by (Hernandez). The motivation to combine (Lewis), (Saxena), and (Hernandez) is that (Hernandez) associates each term with its corresponding human-readable reason and its corresponding risk-metric metadata for use by the model (Hernandez, Abstract; col. 63–64), such that one would receive the verbiage and metadata corresponding to each term and transform them, together with the list of terms, into the language dataset of triplet files that is vectorized and stored by (Lewis) and used with the prompt templates of (Saxena), thereby preparing a language dataset that provides the language model with each term together with its associated descriptive text and context for accurate mapping. Claims 3, 4, 17, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (Lewis), Non-Patent Literature, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," NeurIPS, published December 2020, in view of Saxena (Saxena), U.S. Patent Application Publication No. US 2024/0330579 A1, in view of Hernandez et al. (Hernandez), U.S. Patent No. US 11,922,495 B1, and further in view of Nourian et al. (Nourian), U.S. Patent No. US 11,645,581 B2. Regarding Claim 3, (Lewis) teaches using a feature as an input to the language model, namely an input query x that is provided to the seq2seq generator to produce an output (Lewis, p. 2, Fig. 1; p. 3, § 2.1). However, (Lewis) does not teach: "generating the feature by: applying a prediction model to user data to generate a prediction regarding a user, wherein the user data comprises a plurality of features," In the same field of endeavor, (Hernandez) teaches "generating the feature by: applying a prediction model to user data to generate a prediction regarding a user, wherein the user data comprises a plurality of features." Specifically, (Hernandez) discloses applying a machine-learned model to a user's lending query to obtain a prediction that the query is approved or denied for the user: "applying, at the payment service, the model to the … lending query to obtain an indication that the … lending query is approved" or denied (Hernandez, col. 63–64, claim; see also Abstract), wherein the query comprises a plurality of signals, each "associated with a value of a risk metric" (Hernandez, col. 63–64, claim). Thus (Hernandez) teaches generating the feature by applying a prediction model to user data, comprising a plurality of features, to generate a prediction regarding a user. (Lewis) and (Hernandez) are analogous to the claimed invention as both are from the same field of endeavor of applying machine-learned and language models to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the prediction model of (Hernandez). The motivation to combine (Lewis) and (Hernandez) is that (Hernandez) applies a prediction model to a user's data to generate a prediction whose reason must be provided to the user (Hernandez, Abstract; col. 63–64), such that one would generate the feature using the prediction model of (Hernandez) to supply the feature that is mapped to a human-understandable term by the pipeline of (Lewis), thereby explaining the prediction generated by the model. The combination of (Lewis) and (Hernandez), however, does not teach: "determining a first feature in the plurality of features that contributed, during application of the prediction model to the user data, to the prediction more than at least one other feature in the plurality of features, and" "specifying the first feature as the feature." In the same field of endeavor, (Nourian) teaches "determining a first feature in the plurality of features that contributed, during application of the prediction model to the user data, to the prediction more than at least one other feature in the plurality of features." Specifically, (Nourian) discloses that the model analyzes features associated with an applicant's profile applying for credit, based upon which the applicant is approved or denied (Nourian, col. 3–4), and that the contribution of each feature to the model's output is determined and ranked by importance: "average importance rankings for different features in the target model may be listed across one or more importance measures" (Nourian, col. 7–8, Fig. 7), and "feature contributions … show the average relevance of each feature with respect to the target output" (Nourian, col. 7–8). By ranking the features according to their contribution to the prediction, (Nourian) determines a first feature that contributed to the prediction more than at least one other feature. Thus (Nourian) teaches determining a first feature in the plurality of features that contributed to the prediction more than at least one other feature in the plurality of features. (Nourian) further teaches "specifying the first feature as the feature." Specifically, (Nourian) discloses that the features deemed most important, i.e., the highest-ranked feature according to the importance rankings and feature contributions, are selected and analyzed (Nourian, col. 7–8, Figs. 7–8). Specifying the highest-ranked, most-contributing feature reads on specifying the first feature as the feature. Thus (Nourian) teaches specifying the first feature as the feature. (Lewis), (Hernandez), and (Nourian) are analogous to the claimed invention as all are from the same field of endeavor of applying machine-learned and language models to user data and explaining the resulting predictions. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) and the prediction model of (Hernandez) with the feature-contribution determination of (Nourian). The motivation to combine (Lewis), (Hernandez), and (Nourian) is that (Nourian) determines and ranks the contribution of each feature to a credit model's output to meaningfully explain the model's decision (Nourian, col. 7–8), such that one would apply (Nourian)'s feature-contribution determination to identify the specific feature that most contributed to the prediction generated by (Hernandez)'s model, and provide that feature to the mapping pipeline of (Lewis), thereby mapping the most-contributing feature to a human-understandable term and improving the explainability of the model's decision. Regarding Claim 4, the limitations: "generating the feature by: applying a prediction model to user data to generate a prediction regarding a user, wherein the user data comprises a plurality of features," "determining a first feature in the plurality of features that contributed, during application of the prediction model to the user data, to the prediction more than at least one other feature in the plurality of features, and" "specifying the first feature as the feature," are taught by the combination of (Lewis), (Hernandez), and (Nourian) for the same reasons set forth in the rejection of claim 3 above — namely, (Lewis) teaches using a feature as an input to the language model but does not teach generating the feature by applying a prediction model to user data; (Hernandez) teaches applying a prediction model to a user's data comprising a plurality of features to generate a prediction (Hernandez, col. 63–64, claim; Abstract); and (Nourian) teaches determining and ranking the feature that contributed most to the prediction and specifying that feature (Nourian, col. 3–4; col. 7–8, Figs. 7–8). The combination of (Lewis), (Hernandez), and (Nourian), however, does not teach: "wherein generating the updated prompt further comprises adding a variable prompt to the prompt." In the same field of endeavor, (Saxena) teaches "wherein generating the updated prompt further comprises adding a variable prompt to the prompt." Specifically, (Saxena) discloses that, in addition to a first value determined for a first parameter, a second value for a second parameter is requested from a user and added to the prompt template to generate the prompt: "A request to provide input for a second value of a second parameter is sent for display to a user. Using the prompt template, the first value, and the second value, the system generates a prompt to a large language model" (Saxena, Abstract; see also p. 13, [0018]). This second, user-provided parameter value that is added to the prompt reads on a variable prompt added to the prompt. Thus (Saxena) teaches that generating the updated prompt further comprises adding a variable prompt to the prompt. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the addition of a variable prompt taught by (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) adds a second, variable parameter value to the prompt template to generate a unique prompt appropriate for the context (Saxena, Abstract; p. 13, [0018]), such that one would add a variable prompt to the prompt of (Lewis) in order to customize the prompt at runtime and thereby control and adapt the mapping generated by the language model. Regarding Claim 17, the limitations of claim 17 are commensurate in scope with the limitations of claim 3. Accordingly, claim 17 is rejected under 35 U.S.C. 103 for the same reasons and under the same rationale set forth in the rejection of claim 3 above. Regarding Claim 18, the limitations of claim 18 are commensurate in scope with the limitations of claim 4. Accordingly, claim 18 is rejected under 35 U.S.C. 103 for the same reasons and under the same rationale set forth in the rejection of claim 4 above. Claims 6 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (Lewis), Non-Patent Literature, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," NeurIPS, published December 2020, in view of Saxena (Saxena), U.S. Patent Application Publication No. US 2024/0330579 A1, in view of Hernandez et al. (Hernandez), U.S. Patent No. US 11,922,495 B1, in view of Kuchnik et al. (Kuchnik), Non-Patent Literature, "Validating Large Language Models with ReLM," MLSys 2023, and further in view of Madaan et al. (Madaan), Non-Patent Literature, "Self-Refine: Iterative Refinement with Self-Feedback," NeurIPS, published December 2023. Regarding Claim 6, (Lewis) teaches the method of claim 1, further comprising: "reapplying the vector generation controller to the adjusted prompt to generate an adjusted set of prompt embedding vectors;" "combining the set of term embedding vectors with the adjusted set of prompt embedding vectors to generate an adjusted set of combined vectors; and" "reapplying the language model to the adjusted set of combined vectors to generate a new mapping." As to "reapplying the vector generation controller to the adjusted prompt to generate an adjusted set of prompt embedding vectors," (Lewis) discloses applying the query encoder "q(x) = BERTq(x)" to an input to generate prompt embedding vectors (Lewis, p. 3, § 2.2); applying that same query encoder again to an adjusted input to generate an adjusted set of prompt embedding vectors reads on the recited reapplying. As to "combining the set of term embedding vectors with the adjusted set of prompt embedding vectors to generate an adjusted set of combined vectors," (Lewis) discloses combining the input representation with the retrieved-document representation for the generator (Lewis, p. 3, § 2.1; p. 4, § 2.3); performing that same combination with the adjusted prompt embedding vectors reads on generating an adjusted set of combined vectors. As to "reapplying the language model to the adjusted set of combined vectors to generate a new mapping," (Lewis) discloses applying the seq2seq language model to the combined vectors to generate a mapping (Lewis, p. 3, § 2.1; p. 2, Fig. 1); applying that same language model again to the adjusted set of combined vectors reads on generating a new mapping. Thus (Lewis) teaches the reapplying, combining, and reapplying limitations. (Lewis) teaches returning the mapping, namely outputting the prediction y (Lewis, p. 2, Fig. 1, "Output"). However, (Lewis) does not teach: "testing the mapping against a list of allowed mappings, wherein returning the mapping is performed responsive to a positive result of the testing;" In the same field of endeavor, (Kuchnik) teaches "testing the mapping against a list of allowed mappings, wherein returning the mapping is performed responsive to a positive result of the testing." Specifically, (Kuchnik) discloses ReLM, "a system for validating and querying LLMs using standard regular expressions," which reduces evaluation rules to a defined set of allowed strings and tests the language model's output against that set (Kuchnik, p. 1, Abstract). (Kuchnik) further discloses testing an output against a structured query defining all allowed results — for example, "A structured query over all dates of the form <Month> <Day>, <Year>" — such that the output is evaluated against, and accepted when it conforms to, the allowed set (Kuchnik, p. 2, Fig. 1(c)). Testing the model's output against the defined set of allowed results, and accepting (returning) the output upon a conforming (positive) result, reads on testing the mapping against a list of allowed mappings and returning the mapping responsive to a positive result of the testing. Thus (Kuchnik) teaches this limitation. (Lewis) and (Kuchnik) are analogous to the claimed invention as both are from the same field of endeavor of generating and evaluating outputs of language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the validation of the model's output against a list of allowed results as taught by (Kuchnik). The motivation to combine (Lewis) and (Kuchnik) is that (Kuchnik) validates a language model's output against a defined set of allowed results to ensure the output conforms to permissible values (Kuchnik, p. 1, Abstract; p. 2, Fig. 1), such that one would test the mapping generated by (Lewis) against a list of allowed mappings and return the mapping only upon a positive result, thereby ensuring that the returned mapping is a permissible one. The combination of (Lewis) and (Kuchnik), however, does not teach: "adjusting, responsive to a negative result of the testing, the updated prompt to generate an adjusted prompt;" In the same field of endeavor, (Madaan) teaches "adjusting, responsive to a negative result of the testing, the updated prompt to generate an adjusted prompt." Specifically, (Madaan) discloses Self-Refine, in which an initial output is generated, feedback is obtained on that output, and the feedback is used to refine and regenerate the output iteratively: the model is used "to get feedback" on its output, "the feedback is passed back to M, which refines the previously generated output," and "Steps … iterate until a stopping condition is met" (Madaan, p. 2, Fig. 1 and § 1). Responsive to feedback indicating the output is not satisfactory (a negative result), (Madaan) adjusts the input provided to the model to generate a refined output, which reads on adjusting the updated prompt to generate an adjusted prompt responsive to a negative result. Thus (Madaan) teaches this limitation. (Lewis), (Kuchnik), and (Madaan) are analogous to the claimed invention as all are from the same field of endeavor of generating and evaluating outputs of language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) and the output validation of (Kuchnik) with the iterative feedback-and-refinement of (Madaan). The motivation to combine (Lewis), (Kuchnik), and (Madaan) is that (Madaan) refines a language model's input responsive to feedback in order to improve the output (Madaan, p. 2, § 1), such that one would, responsive to a negative result of testing the mapping of (Lewis) against the allowed mappings of (Kuchnik), adjust the updated prompt to generate an adjusted prompt and re-run the pipeline, thereby producing a new mapping that conforms to the list of allowed mappings. Regarding Claim 7, (Lewis) teaches the method of claim 1, further comprising: "reapplying the vector generation controller to the adjusted prompt to generate an adjusted set of prompt embedding vectors;" "combining the set of term embedding vectors with the adjusted set of prompt embedding vectors to generate an adjusted set of combined vectors;" "reapplying the language model to the adjusted set of combined vectors to generate a new mapping;" As to "reapplying the vector generation controller to the adjusted prompt to generate an adjusted set of prompt embedding vectors," (Lewis) discloses applying the query encoder "q(x) = BERTq(x)" to an input to generate prompt embedding vectors (Lewis, p. 3, § 2.2); applying that same query encoder again to an adjusted input to generate an adjusted set of prompt embedding vectors reads on the recited reapplying. As to "combining the set of term embedding vectors with the adjusted set of prompt embedding vectors to generate an adjusted set of combined vectors," (Lewis) discloses combining the input representation with the retrieved-document representation for the generator (Lewis, p. 3, § 2.1; p. 4, § 2.3); performing that same combination with the adjusted prompt embedding vectors reads on generating an adjusted set of combined vectors. As to "reapplying the language model to the adjusted set of combined vectors to generate a new mapping," (Lewis) discloses applying the seq2seq language model to the combined vectors to generate a mapping (Lewis, p. 3, § 2.1; p. 2, Fig. 1); applying that same language model again to the adjusted set of combined vectors reads on generating a new mapping. Thus (Lewis) teaches the reapplying, combining, and reapplying limitations. (Lewis) teaches returning the mapping, namely outputting the prediction y (Lewis, p. 2, Fig. 1, "Output"). However, (Lewis) does not teach: "testing the mapping against a list of allowed mappings, wherein returning the mapping is performed responsive to a positive result of the testing;" "retesting the new mapping against the list of allowed mappings, wherein returning the mapping is performed using the new mapping responsive to a second positive result of retesting; and" In the same field of endeavor, (Kuchnik) teaches "testing the mapping against a list of allowed mappings, wherein returning the mapping is performed responsive to a positive result of the testing" and "retesting the new mapping against the list of allowed mappings, wherein returning the mapping is performed using the new mapping responsive to a second positive result of retesting." Specifically, (Kuchnik) discloses ReLM, "a system for validating and querying LLMs using standard regular expressions," which reduces evaluation rules to a defined set of allowed strings and tests the language model's output against that set (Kuchnik, p. 1, Abstract), and discloses testing an output against a structured query defining all allowed results — for example, "A structured query over all dates of the form <Month> <Day>, <Year>" — such that the output is evaluated against, and accepted when it conforms to, the allowed set (Kuchnik, p. 2, Fig. 1(c)). Testing the model's output against the defined set of allowed results and accepting (returning) the output upon a conforming (positive) result reads on testing the mapping against a list of allowed mappings and returning the mapping responsive to a positive result; performing that same validation again upon a subsequently generated output reads on retesting the new mapping against the list of allowed mappings and returning the mapping using the new mapping responsive to a second positive result of retesting. Thus (Kuchnik) teaches these limitations. (Lewis) and (Kuchnik) are analogous to the claimed invention as both are from the same field of endeavor of generating and evaluating outputs of language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the validation and re-validation of the model's output against a list of allowed results as taught by (Kuchnik). The motivation to combine (Lewis) and (Kuchnik) is that (Kuchnik) validates a language model's output against a defined set of allowed results to ensure the output conforms to permissible values (Kuchnik, p. 1, Abstract; p. 2, Fig. 1), such that one would test, and retest, the mapping generated by (Lewis) against a list of allowed mappings and return the mapping only upon a positive result, thereby ensuring that the returned mapping is a permissible one. The combination of (Lewis) and (Kuchnik), however, does not teach: "adjusting, responsive to a negative result of the testing, the updated prompt to generate an adjusted prompt;" "returning, responsive to the retesting exceeding a threshold number of negative results, the mapping as a failure result." In the same field of endeavor, (Madaan) teaches "adjusting, responsive to a negative result of the testing, the updated prompt to generate an adjusted prompt" and "returning, responsive to the retesting exceeding a threshold number of negative results, the mapping as a failure result." Specifically, (Madaan) discloses Self-Refine, in which an initial output is generated, feedback is obtained on that output, and the feedback is used to refine and regenerate the output iteratively: the model is used "to get feedback" on its output, "the feedback is passed back to M, which refines the previously generated output," and the process is "repeated either for a specified number of iterations or until M determines that no further refinement is necessary" (Madaan, p. 2, Fig. 1 and § 1). Responsive to feedback indicating the output is not satisfactory (a negative result), (Madaan) adjusts the input provided to the model to generate a refined output, which reads on adjusting the updated prompt to generate an adjusted prompt responsive to a negative result; and terminating the iterative process once the specified number of iterations is reached without a satisfactory result reads on returning, responsive to the retesting exceeding a threshold number of negative results, the mapping as a failure result. Thus (Madaan) teaches these limitations. (Lewis), (Kuchnik), and (Madaan) are analogous to the claimed invention as all are from the same field of endeavor of generating and evaluating outputs of language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) and the output validation of (Kuchnik) with the iterative feedback-and-refinement of (Madaan). The motivation to combine (Lewis), (Kuchnik), and (Madaan) is that (Madaan) refines a language model's input responsive to feedback and terminates after a specified number of iterations (Madaan, p. 2, § 1), such that one would, responsive to a negative result of testing the mapping of (Lewis) against the allowed mappings of (Kuchnik), adjust the updated prompt to generate an adjusted prompt and re-run the pipeline, and, upon the retesting exceeding a threshold number of negative results, return the mapping as a failure result, thereby ensuring that the process either produces a mapping conforming to the list of allowed mappings or terminates with a defined failure result rather than iterating indefinitely. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over (Lewis), Non-Patent Literature, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," NeurIPS, published December 2020, in view of (Saxena), U.S. Patent Application Publication No. US 2024/0330579 A1, in view of (Hernandez), U.S. Patent No. US 11,922,495 B1, and further in view of (Willard), Non-Patent Literature, "Efficient Guided Generation for Large Language Models," arXiv:2307.09702v4, published August 19, 2023. Regarding Claim 8, (Lewis) teaches the form of the mapping, namely that the language model generates the mapping as the output of the generator and that the output is produced and returned (Lewis, p. 2, Fig. 1, "Output"; p. 3, § 2.1). However, (Lewis) does not teach: "the mapping comprises an object notation language data structure specifying the feature and the term." In the same field of endeavor, (Willard) teaches "the mapping comprises an object notation language data structure specifying the feature and the term." Specifically, (Willard) discloses guiding the generation of a large language model so that the model's generated output is constrained to conform to a specified structure, wherein the approach "allows one to enforce domain-specific knowledge and constraints, and enables the construction of reliable interfaces by guaranteeing the structure of the generated text" (Willard, p. 1, Abstract). (Willard) further discloses that such guided generation is directed to "generating sequences of tokens from a large language model (LLM) … that conform to regular expressions or context-free grammars (CFGs)," and that "[t]his kind of guided LLM generation is used to make LLM model output usable under rigid formatting requirements" (Willard, p. 1, § 1, Introduction). (Willard) expressly discloses that the generated output is constrained to conform to an object notation language data structure, stating that "our indexing approach can also be extended to CFGs and LALR(1) parsers to allow for efficient guided generation according to popular data formats and programming languages (e.g. JSON, Python, SQL, etc.)" (Willard, p. 2, § 1, Introduction). JSON, i.e., JavaScript Object Notation, is an object notation language data structure. (Willard) further discloses the mechanism by which the generated output is so constrained, wherein a boolean mask restricts the support of the next-token distribution such that the generated sequences represent "strings that parse according to a specified grammar" (Willard, p. 4, § 2.2, Guiding generation), and provides an implementation in which a language model is guided to produce output conforming to the specified structure (Willard, p. 8, § 3.1, Examples). Accordingly, when the language model of the combination generates the mapping between the feature and the term, guiding that generation according to the JSON data format as taught by (Willard) causes the generated mapping to be emitted as an object notation language data structure, and, because that generated mapping is the mapping between the feature and the term, the resulting object notation language data structure specifies the feature and the term. Thus (Willard) teaches that the mapping comprises an object notation language data structure specifying the feature and the term. (Lewis) and (Willard) are analogous to the claimed invention as both are from the same field of endeavor of generating and processing the outputs of language models applied to natural-language data. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping generated by (Lewis) with the guided generation of an object notation language data structure as taught by (Willard). The motivation to combine (Lewis) and (Willard) is as recited by (Willard), which teaches that guided generation "enables the construction of reliable interfaces by guaranteeing the structure of the generated text" (Willard, p. 1, Abstract) and is "used to make LLM model output usable under rigid formatting requirements that are either hard or costly to capture through fine-tuning alone" (Willard, p. 1, § 1, Introduction), such that one would guide the language model of (Lewis) to emit the generated mapping as a JSON object notation language data structure, thereby guaranteeing that the mapping is produced in a reliably parsable, machine-readable form suitable for consumption by downstream processes. Claims 12, 13, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Lewis et al. (Lewis), Non-Patent Literature, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," NeurIPS, published December 2020, in view of Saxena (Saxena), U.S. Patent Application Publication No. US 2024/0330579 A1, in view of Hernandez et al. (Hernandez), U.S. Patent No. US 11,922,495 B1, and further in view of Luitel et al. (Luitel), Non-Patent Literature, "Improving Requirements Completeness: Automated Assistance through Large Language Models.", published February 14, 2024. Regarding Claim 12, (Lewis) teaches: "determining the semantic distance between the feature and the term;" As to "determining the semantic distance between the feature and the term," (Lewis) discloses computing a measure of similarity between the query representation and each document representation using the inner product of their embeddings, "pn(z|x) ∝ exp(d(z)ᵀq(x))," and selecting documents by "Maximum Inner Product Search (MIPS)" (Lewis, p. 3, § 2.2). This inner-product similarity between the feature (query) representation and the term (document) representation is a determination of the semantic distance between the feature and the term. Thus (Lewis) teaches determining the semantic distance between the feature and the term. (Lewis) teaches providing a prompt to and generating an output from the language model (Lewis, p. 2, Fig. 1; p. 3, § 2.1). However, (Lewis) does not teach: "wherein the prompt includes a second command to suggest a new term if the term is unsuitable," "generating, by the language model, the new term; and" "replacing, by the language model, the term with the new term such that the mapping that is generated is from the feature to the new term." In the same field of endeavor, (Saxena) teaches "wherein the prompt includes a second command to suggest a new term if the term is unsuitable," "generating, by the language model, the new term," and "replacing, by the language model, the term with the new term such that the mapping that is generated is from the feature to the new term." Specifically, (Saxena) discloses that a prompt is constructed from a template that includes one or more parameters commanding the language model, and that, using the prompt, the language model generates text: "Each prompt template includes at least one parameter, a value for which is selected to generate a unique prompt to the LLM 130" (Saxena, p. 14, [0028]), and "Using the prompt template, the first value, and the second value, the system generates a prompt to a large language model" to generate text (Saxena, Abstract; p. 13, [0018]). Including a further command in the prompt directing the language model to generate a new term, having the language model generate that new term, and having the language model substitute the new term such that the mapping is from the feature to the new term, read on these limitations. Thus (Saxena) teaches wherein the prompt includes a second command to suggest a new term if the term is unsuitable, generating by the language model the new term, and replacing by the language model the term with the new term such that the mapping that is generated is from the feature to the new term. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the prompt-command-directed text generation of (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) commands the language model, via the prompt, to generate text appropriate to the context (Saxena, p. 14, [0028]; Abstract), such that one would include in the prompt of (Lewis) a command directing the language model to generate and substitute a new term, in order to produce a suitable mapping when an existing term is inadequate. The combination of (Lewis) and (Saxena), however, does not teach: "wherein unsuitable is defined as the term having a semantic distance to the feature that fails to satisfy a predetermined threshold, and" "determining that the new term is unsuitable because the semantic distance fails to satisfy the predetermined threshold." In the same field of endeavor, (Luitel) teaches "wherein unsuitable is defined as the term having a semantic distance to the feature that fails to satisfy a predetermined threshold" and "determining that the new term is unsuitable because the semantic distance fails to satisfy the predetermined threshold." Specifically, (Luitel) discloses employing a predetermined cosine-similarity threshold on word embeddings to determine whether a predicted term is a good match for a term, wherein a prediction whose semantic similarity fails to meet the predetermined threshold is not treated as a valid match (Luitel, § 6, Discussion — an 85% cosine-similarity threshold on word embeddings used to assess whether predictions are good matches for the novel terms). Defining a term as unsuitable when its embedding-based semantic distance to the feature fails to satisfy the predetermined threshold, and determining unsuitability on that basis, read on these limitations. Thus (Luitel) teaches wherein unsuitable is defined as the term having a semantic distance to the feature that fails to satisfy a predetermined threshold, and determining that the term is unsuitable because the semantic distance fails to satisfy the predetermined threshold. (Lewis), (Saxena), and (Luitel) are analogous to the claimed invention as all are from the same field of endeavor of processing natural-language data with language and embedding models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the semantic-distance determination of (Lewis) and the prompt-command-directed generation of (Saxena) with the predetermined-threshold suitability determination of (Luitel). The motivation to combine (Lewis), (Saxena), and (Luitel) is that (Luitel) applies a predetermined semantic-similarity threshold to decide whether a predicted term is an acceptable match (Luitel, § 6), such that one would determine, using the semantic distance of (Lewis), that a term is unsuitable when it fails the predetermined threshold, and command the language model of (Saxena) to generate and substitute a new term, thereby ensuring that the resulting mapping from the feature to the term satisfies a required degree of semantic correspondence. Regarding Claim 13, (Lewis) teaches the method of claim 12, further comprising: "updating, after applying the language model to the set of combined vectors, the language dataset to include the new term." As to "updating, after applying the language model to the set of combined vectors, the language dataset to include the new term," (Lewis) discloses that the language dataset — the non-parametric memory — can be revised and expanded to update the knowledge available to the model: (Lewis) identifies "updating their world knowledge" as an objective (Lewis, p. 1, Abstract) and discloses that, by combining the model with a non-parametric memory, "knowledge can be directly revised and expanded, and accessed knowledge can be inspected and interpreted" (Lewis, p. 1, § 1). Revising and expanding the language dataset to add a term reads on updating the language dataset to include the new term, and doing so following the model's operation reads on performing the update after applying the language model to the set of combined vectors. Thus (Lewis) teaches updating, after applying the language model to the set of combined vectors, the language dataset to include the new term. Regarding Claim 14, (Lewis) teaches generating an output from the language model, namely generating the prediction y from the seq2seq generator (Lewis, p. 2, Fig. 1; p. 3, § 2.1). However, (Lewis) does not teach: "adding a third command to the prompt, wherein the third prompt commands the language model to generate verbiage for the new term." In the same field of endeavor, (Saxena) teaches "adding a third command to the prompt, wherein the third prompt commands the language model to generate verbiage for the new term." Specifically, (Saxena) discloses that a prompt is constructed from a template that includes one or more parameters, each of which commands the language model, and that the language model generates text in response to the prompt: "Each prompt template includes at least one parameter, a value for which is selected to generate a unique prompt to the LLM 130" (Saxena, p. 14, [0028]), and "Using the prompt template, the first value, and the second value, the system generates a prompt to a large language model" to generate text (Saxena, Abstract; p. 13, [0018]). Adding a further command to the prompt that directs the language model to generate text (verbiage) for the new term reads on adding a third command to the prompt that commands the language model to generate verbiage for the new term. Thus (Saxena) teaches adding a third command to the prompt, wherein the third prompt commands the language model to generate verbiage for the new term. (Lewis) and (Saxena) are analogous to the claimed invention as both are from the same field of endeavor of processing natural-language data with language models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the mapping pipeline of (Lewis) with the prompt-command-directed text generation of (Saxena). The motivation to combine (Lewis) and (Saxena) is that (Saxena) commands the language model, via commands added to the prompt, to generate text appropriate to the context (Saxena, p. 14, [0028]; Abstract), such that one would add a third command to the prompt of (Lewis) directing the language model to generate verbiage for the new term, in order to produce descriptive text corresponding to the newly generated term for use in subsequent presentation and processing. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG VAN LE whose telephone number is (571)270-0164. The examiner can normally be reached 8 a.m. - 5 p.m.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HUNG VAN LE/Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Apr 04, 2024
Application Filed
Aug 05, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month