Prosecution Insights
Last updated: August 16, 2026
Application No. 18/454,262

PROMPT-BASED SEQUENTIAL LEARNING

Non-Final OA §101§103§112
Filed
Aug 23, 2023
Priority
Aug 25, 2022 — provisional 63/400,767
Examiner
ZENG, WENWEI
Art Unit
2146
Tech Center
2100 — Computer Architecture & Software
Assignee
NEC Laboratories America Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
21 currently pending
Career history
18
Total Applications
across all art units

Statute-Specific Performance

§101
44.1%
+4.1% vs TC avg
§103
49.2%
+9.2% vs TC avg
§102
3.4%
-36.6% vs TC avg
§112
3.4%
-36.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103 §112
CTNF 18/454,262 CTNF 101543 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 8 and 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claims 8 and 18, the limitation “ ⋅ : ⋅ denotes a concatenation operation”, where the symbol ‘ ⋅ : ⋅ ’ is considered indefinite since this symbol is not found in the equation or elaborated in the specification other than a brief mention in [0026]. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mathematical concept or mental process) without significantly more. Claim 1: Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “ A computer-implemented method for training a language model, comprising: retrieving a knowledge sentence, related to an input sentence, from a knowledge base; encoding the input sentence, the knowledge sentence, and a prompt into an intermediate representation; decoding the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt; and fine-tuning a language model based on the named entity ,” and a method is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components: encoding the input sentence, the knowledge sentence, and a prompt into an intermediate representation; ( This is considered a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in paragraphs [0024] state “the knowledge vector may be aggregated using the equation K=Aggregate(HI,Hk), where Hk is the knowledge vector and HI is the intermediate representation. The prompt vector is used to calculate HI, as above. The prompt may be generated manually, and the prompt sentence is encoded to initialize prompt parameters ϕ.” , and [0026] state “The attention function of the encoder, with the prompt, may be written as: PNG media_image1.png 9 120 media_image1.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), decoding the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt; ( This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in [0029] state “The decoding uses the output of the encoder and previous decoder output tokens to decode a next token,” where the decoding process involves a calculation on the output of the encoder, see MPEP 2106.04(a)(2), subsection I), and fine-tuning a language model based on the named entity ( This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in [0029] state “The index y i ~ is the position of each token in the input sentence, and is used to select tokens from the input sentence. Using the ground truth value pi, a categorical cross-entropy loss may be determined between pi and y i and may be used to fine-tune the parameters of the model in 212,” PNG media_image2.png 87 129 media_image2.png Greyscale Here, a loss function is used to fine- tune the model, and is considered a math calculation, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A computer-implemented method for training a language model, comprising: retrieving a knowledge sentence, related to an input sentence, from a knowledge base ; ( In step 2A, prong 2, this recites mere data receiving or gathering, and is considered insignificant extra-solution activity – see MPEP 2106.05(g)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element iv recites mere data gathering, and is considered insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is a well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea . Therefore, the claim is not patent eligible. Claim 2: Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 2 recites the following additional element: The method of claim 1, wherein retrieving the knowledge sentence includes searching a knowledge base for entities in the input sentence , ( In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g),). In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) – see MPEP 2106.05(d) (II)(i). Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 3: Regarding claim 3, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis applied to claim 2. Further, claim 3 recites the following additional element: The method of claim 2, wherein retrieving the knowledge sentence further includes retrieving relations from the knowledge base , ( In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g),). In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) – see MPEP 2106.05(d) (II)(i). Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 4: Regarding claim 4, it is dependent upon claim 3, and thereby incorporates the limitations of, and corresponding analysis applied to claim 3. Further, claim 4 recites the following abstract idea: The method of claim 3, wherein retrieving the knowledge sentence further includes generating a set of knowledge sentences from the entities and the relations and selecting a percentage with a highest relevance to the input sentence , ( This recites a mental process, since a person can mentally generate sentences, and then select a percentage that is highly relevant to an input sentence, see MPEP 2106.04(a)(2)(III)), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 5: Regarding claim 5, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Further, claim 5 recites the following additional element: The method of claim 4, wherein the set of knowledges sentences include sentences of the form <entity> is a <type> , (In step 2A, prong 2, this recites an indication to a field of use or technological environment – see MPEP 2106.05(h)), (In step 2B, this also recites a field of use or technological environment – see MPEP 2106.05(h)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 6: Regarding claim 6, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis applied to claim 2. Further, claim 6 recites the following additional element: The method of claim 2, wherein the knowledge base is a multilingual knowledge graph , (In step 2A, prong 2, this recites an indication to a field of use or technological environment – see MPEP 2106.05(h)), (In step 2B, this also recites a field of use or technological environment – see MPEP 2106.05(h)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 7: Regarding claim 7, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 7 recites the following additional element: The method of claim 1, further comprising pre-training the model with labeled training data from a source domain, wherein the input sentence includes labeled entities from a second domain , (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 8: Regarding claim 8, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 8 recites the following abstract idea: The method of claim 1, wherein encoding uses an attention function : PNG media_image1.png 9 120 media_image1.png Greyscale where l designates an attention layer, Q, K, and V are query, key, and value parameters of the attention layer, respectively, ϕk and ϕv are prompt parameters corresponding to K and V, and [ ⋅ : ⋅] denotes a concatenation operation, and d is a dimension size , ( This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in [0026] state “two trainable embedding matrices may be defined as trainable prompt parameters. The attention function of the encoder, with the prompt PNG media_image1.png 9 120 media_image1.png Greyscale The input sequence representation may be projected to the Q , K and V values. For each ϕv l and ϕk l , they may be initialized with the word embedding of a manually defined prompt. The hidden representation HI of the input sentence may be determined using the encoder,” where the attention function takes in prompt parameters and project input sequence to Q ,K and V , see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 9: Regarding claim 9, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 9 recites the following additional element: The method of claim 1, wherein the prompt specifies a type of entity to be identified , (In step 2A, prong 2, this recites an indication to a field of use or technological environment – see MPEP 2106.05(h)), (In step 2B, this also recites a field of use or technological environment – see MPEP 2106.05(h)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 10: Regarding claim 10, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 10 recites the following abstract idea: The method of claim 1, wherein encoding includes aggregating a representation of the input sentence with a representation of the knowledge sentence, based on the prompt , ( This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in [0027] state “a scaled dot-product attention may be used to aggregate the knowledge sentences and input sentence,” where a scaled dot-product attention is used to aggregate the sentences, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 11: Regarding claim 11, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “ A system for training a language model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to: retrieve a knowledge sentence, related to an input sentence, from a knowledge base; encode the input sentence, the knowledge sentence, and a prompt into an intermediate representation; decode the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt; and fine-tune a language model based on the named entity ,” and a system is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math process but for recitation of generic computer components: encode the input sentence, the knowledge sentence, and a prompt into an intermediate representation; ( This is considered a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in paragraphs [0024] state “the knowledge vector may be aggregated using the equation K=Aggregate(HI,Hk), where Hk is the knowledge vector and HI is the intermediate representation. The prompt vector is used to calculate HI, as above. The prompt may be generated manually, and the prompt sentence is encoded to initialize prompt parameters ϕ.” , and [0026] state “The attention function of the encoder, with the prompt, may be written as: PNG media_image1.png 9 120 media_image1.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), decode the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ; ( This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in [0029] state “The decoding uses the output of the encoder and previous decoder output tokens to decode a next token,” where the decoding process involves a calculation on the output of the encoder, see MPEP 2106.04(a)(2), subsection I), and fine-tune a language model based on the named entity ( This recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see specification in [0029] state “The index y i ~ is the position of each token in the input sentence, and is used to select tokens from the input sentence. Using the ground truth value pi, a categorical cross-entropy loss may be determined between pi and y i and may be used to fine-tune the parameters of the model in 212,” PNG media_image2.png 87 129 media_image2.png Greyscale Here, a loss function is used to fine-tune the model, and is considered a math calculation, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mathematical concept but for the recitation of generic computer components, then it falls within the mathematical concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A system for training a language model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor , ( This recites mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), to: retrieve a knowledge sentence, related to an input sentence, from a knowledge base; ( In step 2A, prong 2, this recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element iv recites mere instructions to apply the judicial exception using generic computer components, which is not indicative of significantly more. The additional element v recites mere data gathering, and is considered insignificant extra-solution activity. In step 2B, this insignificant extra-solution activity is a well understood routine and conventional activity, which includes receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea . Therefore, the claim is not patent eligible. Claims 12 - 20: Regarding claim 11, all of claim 11’s dependent claims follow the deficiencies of their parent claim. Since claims 12-20 recite similar limitations as corresponding claims 2-10 listed above, and are rejected for similar reasons under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 2, 3, 9, 11, 12, 13, and 19 are rejected under 35 U.S.C. 103 over Gong, B. et al., (Pub. No. CN 113868392A), published on December 31, 2021, (hereafter, GONG), in view of Overell, S. et al., (US PG Pub. No. US20110307435A), published on December 15, 2011, (hereafter, OVERELL), further in view of Jiang, W. et al., (Pub. No. CN113641830A), published on November 12, 2021, (hereafter, JIANG). Claim 1: Regarding claim 1, GONG teaches “ A computer-implemented method for training a language model, comprising: retrieving a knowledge sentence, related to an input sentence, from a knowledge base; ” See GONG in [0003] describe “The general information retrieval-based question-answering system works in two steps: question resolution and answer retrieval. Problem analysis is a natural language understanding task , and the main work of the problem analysis is to extract useful information from a question set by a user so as to guide subsequent retrieval . The answer retrieval is to find the answer from the constructed domain knowledge base.” Here, GONG explains retrieving an answer which is in a form of a natural language. Further, see GONG in [0004] specify “The problem analysis method adopts two sequence tagging technologies of named entity identification and part of speech tagging, which are also called slot filling. The sequence tagging is to regard an input sentence as an input sequence , and tag each word in the input sequence, so as to tag important elements in the sentence.” Here, GONG describes the input sequence is similar to a question, and using the answer retrieval step involves finding an answer also in a form of a sentence from a knowledge base. This relates to retrieving a knowledge sentence, where the knowledge sentence is construed to mean the sentence, where after retrieving the relevant information from a knowledge base, generates an answer to a user prompt to a language model. Further, GONG teaches “ encoding the input sentence, the knowledge sentence, and a prompt into an intermediate representation; ” See GONG in [0039] describe “the method for realizing the question-answering system is used in a specific field, a field knowledge base is designed by considering the characteristics of knowledge in the field, the structure of the knowledge base is considered , two sequence tagging tasks of named body recognition and part of speech tagging are completed based on a bidirectional transducer encoder representation technology (BERT), sentence information is extracted in a targeted manner ,” Here, GONG describes BERT, a model that is specifically designed to encode two sequence tasks: 1. the extracted sentence information (i.e. input sentence) as well as the 2. sequence task extracted from a knowledge base (i.e. knowledge sentence), into contextualized numeric representations called embeddings or label text sequence with tags in this case. GONG mentions BERT model has an encoder and transducer, and performs the process of encoding. Further, see GONG in paragraphs [0103-0106], describe "The specific training method comprises the following steps: (1) and performing participlization on the corpus data training set labeled above, namely dividing the question text into separate participles (tokens)... (2) And then converted to word embedding . Processing NLP task by means of neural network model, it is often necessary to map words into a vector in a high-dimensional dense space , and express semantic relation between corresponding words by cosine distance between each vector, which is word embedding. " Here, GONG describes labeling the question text into word embeddings, which shows encoding the prompt into an intermediate representation. Note the examiner construes intermediate representation to mean any vector representation in space, which includes embeddings, from the specification [0023] stating “the encoder encodes the inputs together and generates an intermediate representation, for example as a vector in a latent space.” Further, see GONG in paragraphs [0110-0111, 0114-0116] describe “ (1) the method includes inputting question texts, simply screening the text lengths, ... (2) And performing participlization and converting into word embedding. …(1) Extracting relation elements contained in the question from the labeled result; (2) converting relational elements into database query statements (3) And giving corresponding prompts for the conditions of inquiring or not inquiring answers.” Here, GONG describes using the process of viewing the question text, converting this into word embeddings (i.e. intermediate representation), extracting relation elements from the question (this step implies using information from a knowledge base), then giving prompts that correspond with the content of the question. The term corresponding prompts means prompts that relate to the original text content, which was converted into word embeddings. Since the prompt is originally part of the text, which is converted to word embeddings, this means the prompt also was converted into intermediate representation. Further, see GONG in [0020 - 0023] describe “step 2-2: obtaining a question sentence of a user , screening according to the length of the text, and then preprocessing the screened text; step 2-3: inputting the preprocessed data into a sequence labeling model to obtain a label sequence; step 2-4: extracting relationship elements from the tag sequence; step 2-5: converting the relationship elements into database query statements ;” Here, GONG describes the process after getting a question sentence from a user, the method helps preprocess data into sequence labeling model, which refers to the BERT model, and labels the sequence (which includes input sentence). See GONG in [0025-0030] for details. Further, GONG teaches “decoding the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ; ” See GONG in [0090 - 0091] describe "The encoder representation model (BERT) of a bidirectional Transformer is composed of a number of bidirectional Transformer modules, each of which comprises a concatenation of a number of encoders and decoders. A single decoder or encoder has one attribute layer, and inputs one fully connected layer after residual chaining and normalization," Further, see GONG in [0105-0106] describe "(2) And then converted to word embedding. Processing NLP task by means of neural network model, it is often necessary to map words into a vector in a high-dimensional dense space, and express semantic relation between corresponding words by cosine distance between each vector, which is word embedding." See [0020-0023] where GONG describes “step 2-3: inputting the preprocessed data into a sequence labeling model to obtain a label sequence; step 2-4: extracting relationship elements from the tag sequence; step 2-5: converting the relationship elements into database query statements .” Here, GONG teaches the BERT model contains decoders in [0090-0091], and later GONG describes converting the text sentence from word embedding (i.e. intermediate representation) into expressing semantic relation between corresponding words, and later labeling those relationship elements by tags which relates to named entity. From the specification note in [0018] describe “ In the specific example of FIG. 1 , the input text 100 includes certain named entities, illustrated with bold text, including a personal name, “Ms. Green,” an organization's name, “LIRR,” and a location name, “Manhattan.” Each of these named entities has an associated type: person, organization, or location. A given corpus of training data may be labeled, such that each named entity may be associated with one or more such types . ” Note the examiner construes ‘named’ to mean any label or category of a text or word (viewed as ‘entity’). Further, see GONG in [0112] describe “(3) And calculating in the network, performing softmax transformation …, taking the obtained maximum probability, and taking the corresponding sequence tag as the tag corresponding to the token. And obtaining a label sequence corresponding to the sentence .” Using named entity means to label an item or text with a name by a tag or other labels. Here, GONG has named or labeled the data containing words into a labeled sequence of the words from [0105-0106], and later obtain a label sequence that corresponds to a sentence in [0112], which relates to named entity from the input sentence. Further, GONG teaches “ and fine-tuning a language model …” See GONG in [0028] describe " the training method comprises the following steps: performing word segmentation processing on data in the training data set, converting the data into word embedding, and inputting the word embedding into a model for training; [0029] finally, fine tuning is carried out on the encoder representation model of the bidirectional Transformer ." Here, GONG shows details in fine-tuning a model of the bidirectional transformer model for word segmentation in language processing. Further details, see GONG in [0018-0019]. However, GONG did not teach “decoding the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ;” “and fine-tuning a language model based on the named entity,” In an analogous field, OVERELL teaches “decoding the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ;” See OVERELL in [0169, 0171] describe "One method for finding the principal class of an object is first to identify the classes of which the object is a member, i.e. a query is done looking for objects to which the entity has the relation [is an instance of]. The resulting class objects are then ordered using the [is a subclass of] relation and the most specific class labelled as a principal class is then considered the PC for the object … A similar check is done while adding a new object when prompting the user entity for a class of which the object is a member. After prompting the user entity for a class, both this class and the classes to which this class is on the right in the relation [is a subclass of] are retrieved from the knowledge base and again they are ordered. The most specific class labelled as a PC is taken as the object's PC. If one is not found using this method the user entity is prompted for a more specific class , e.g. the string “policeman” will find [human being] as the principal class (the class of policemen is a subclass of the PC [human being]) but “living thing” will result in the user being prompted to be more specific." Here, OVERELL describes class as a type of the entity, where class is a certain category of the text the user has input into their query (i.e. input sentence). OVERELL shows that the prompt asks the user to label a text by a class (i.e. type). For example, if the user enters policeman, the prompt will specify the class (i.e. type), which is a human being. When OVERELL mentions “ The most specific class labelled as a PC is taken as the object's PC. If one is not found using this method the user entity is prompted for a more specific class ,” this means that if a prompt does not recognize the term entered, then the prompt will ask to specify what is the type of item or entity. If the user enters a living thing, then the prompt asks the user to specify the term. For example, living thing is a broad term, whereas policeman is a type of human being. In this case, OVERELL shows the prompt ask or clarify the type of entity the user wants to know about. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of GONG and incorporate into the teachings of OVERELL because both references teach retrieving knowledge sentences based on a knowledge base, and encoding and decoding steps for a language model. One of ordinary skill in the art would be motivated to do so because by integrating OVERELL’s framework into the methods of GONG, one with ordinary skill in the art would achieve a goal of “many advantages in terms of efficient processing of the query,” (OVERELL, [0239]), and has “instructions thus providing the user with a mechanism to correct and improve the problem for all users,” (OVERELL, [0948]). However, GONG in view of OVERELL did not teach “and fine-tuning a language model based on the named entity,” In an analogous field, JIANG teaches “ and fine-tuning a language model based on the named entity ,” See JIANG in [n0117] describe “ The labeled data required for encoder-decoder learning can be large-scale collected text data and related subgraphs retrieved from the knowledge graph based on it. Text data can be sentences, phrases , or passages, or combinations of these linguistic units, to be compatible with different higher-level artificial intelligence tasks. The process of retrieving knowledge subgraphs from a knowledge graph based on text data can employ simple character matching or utilize mature basic tools such as entity recognition and entity chaining.” Here, JIANG mentions that “ The labeled data required for encoder-decoder learning can be large-scale collected text data”, and text data includes sentences or phrases. See JIANG in [n0034] describe “ Knowledge graphs are also widely used to improve higher-level artificial intelligence tasks or higher-level natural language processing tasks (referred to as higher-level tasks). These higher-level tasks typically use a pre-trained language model as a basis and fine-tune the parameters on their own labeled data .” Here, JIANG describes fine tuning parameters of a language model with knowledge graphs using labeled data in [n0034], where labeled data refers to labeled text data in the form of labeled sentences mentioned in [n0117], and relates to named entity . Note the examiner construes the term ‘named entity’ to mean labeled item, such as classifying text into a category, such as identifying Manhattan as a location. Note the examiner also construes ‘named’ to mean any label or category of a text or word (viewed as ‘entity’). From the specification note in [0018] describe “ In the specific example of FIG. 1 , the input text 100 includes certain named entities, illustrated with bold text, including a personal name, “Ms. Green,” an organization's name, “LIRR,” and a location name, “Manhattan.” Each of these named entities has an associated type: person, organization, or location. A given corpus of training data may be labeled, such that each named entity may be associated with one or more such types . ” Note the examiner construes ‘named’ to mean any label or category of a text or word (viewed as ‘entity’). JIANG shows using labeled data in form of labeled sentences or labeled text (i.e. named entity) to fine tune a model and its parameters based on this labeled text data. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references GONG and OVERELL, and incorporate with the teachings of JIANG by using the teachings of GONG and OVERELL of retrieving knowledge sentences based on a knowledge base, and encoding and decoding steps for a language model, with JIANG’s teaching of fine-tuning a language model based on the named entity. One of ordinary skill in the art would be motivated to do so because by integrating JIANG’s framework into the methods of GONG and OVERELL, one with ordinary skill in the art would achieve a method that “provides a pre trained model for knowledge enhancement that achieves effective knowledge utilization while maintaining ease of use,” (JIANG, [n0120]). Claim 2: Regarding claim 2, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 1. Further, GONG teaches “ The method of claim 1, wherein retrieving the knowledge sentence includes searching a knowledge base for entities in the input sentence ,” See GONG describe in [0013 - 0015] in "step 1-2: describing a relationship between two entities using the relationship based on a knowledge graph ; step 1-3: and establishing a relational database. More preferably, the relationship between the two entities in step 1-2 is described by two elements, namely, the occurrence time of the relationship and the type of the relationship." Here, GONG shows using a knowledge graph to describe a relationship between two entities. See GONG in [0039] for details. Further, see GONG in [0003] describe “The general information retrieval-based question-answering system works in two steps: question resolution and answer retrieval. Problem analysis is a natural language understanding task, and the main work of the problem analysis is to extract useful information from a question set by a user so as to guide subsequent retrieval. The answer retrieval is to find the answer from the constructed domain knowledge base.” Here, GONG describes the method in which the question-answering system works. First, the system obtains a question (i.e. input sentence) from a user, then uses a knowledge base to find the response to the user’s question, which relates to searching a knowledge base for entities in input sentence. See GONG in [0109-0112] for more information. Claim 3: Regarding claim 3, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 2. Further, JIANG teaches “ The method of claim 2, wherein retrieving the knowledge sentence further includes retrieving relations from the knowledge base. ” See JIANG in [n0117-n0118] describe "The process of retrieving knowledge subgraphs from a knowledge graph based on text data can employ simple character matching or utilize mature basic tools such as entity recognition and entity chaining . Furthermore, the retrieved knowledge subgraph can be expanded into a larger knowledge subgraph along the edges connecting the edge nodes. This extension enables the learned knowledge to enhance the pre-trained model, allowing it to not only characterize relevant knowledge for the input information, but also related knowledge . Whether to perform subgraph expansion and what kind of subgraph expansion to perform depends on the specific upper-level artificial intelligence task being addressed. " JIANG refers to retrieving a knowledge sentence when describing the “ characterizing relevant knowledge for input information ”. Further, JIANG describes in [n0058] “For example, the input information can be sentences, phrases, or passages, or combinations of these language units, to be compatible with different higher-level artificial intelligence tasks.” Here, JIANG describes the process of retrieving information based on text data or input information (which includes sentences). When JIANG describes “ related knowledge”, JIANG refers to retrieving relations from a knowledge base to help classify input sentences. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references GONG and OVERELL, and incorporate with the teachings of JIANG by using the teachings of GONG and OVERELL of retrieving knowledge sentences based on a knowledge base, encoding and decoding steps to generate a named entity, to later fine-tune a language model, with JIANG’s teaching of retrieving the knowledge sentence further includes retrieving relations from the knowledge base. One of ordinary skill in the art would be motivated to do so because by integrating JIANG’s framework into the methods of GONG and OVERELL, one with ordinary skill in the art would achieve a method that “provides a pre trained model for knowledge enhancement that achieves effective knowledge utilization while maintaining ease of use,” (JIANG, [n0120]). Claim 9: Regarding claim 9, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 1. Further, OVERELL teaches “ The method of claim 1, wherein the prompt specifies a type of entity to be identified ,” See OVERELL in [0169, 0171] describe "One method for finding the principal class of an object is first to identify the classes of which the object is a member, i.e. a query is done looking for objects to which the entity has the relation [is an instance of]. The resulting class objects are then ordered using the [is a subclass of] relation and the most specific class labelled as a principal class is then considered the PC for the object … A similar check is done while adding a new object when prompting the user entity for a class of which the object is a member. After prompting the user entity for a class, both this class and the classes to which this class is on the right in the relation [is a subclass of] are retrieved from the knowledge base and again they are ordered. The most specific class labelled as a PC is taken as the object's PC. If one is not found using this method the user entity is prompted for a more specific class , e.g. the string “policeman” will find [human being] as the principal class (the class of policemen is a subclass of the PC [human being]) but “living thing” will result in the user being prompted to be more specific. " OVERELL describes class as a type of the entity. Here, OVERELL describes that the system prompts a user to be more specific when entering a specific class or type for an object (which is viewed as an entity). For example, OVERELL classifies the policeman entity to be a human being type, however, the system will prompt the user to be more specific in the input when the input is “living thing”. OVERELL shows that the prompt asks the user to label a text by a class (i.e. type). For example, if the user enters policeman, the prompt will specify the class (i.e. type), which is a human being. When OVERELL mentions “ The most specific class labelled as a PC is taken as the object's PC. If one is not found using this method the user entity is prompted for a more specific class ,” this means that if a prompt does not recognize the item entered, then the prompt will ask for more details on what is the type of item. If the user enters a living thing, then the prompt asks the user to specify the item. The prompt specifies the class (i.e. type) of an item to be identified illustrated by identifying a policeman as a type of human being. The prompt, similar to the query, identifies ‘ the entity has the relation [is an instance of] ’, where [is an instance of] also let the prompt specify an instance (i.e. type) of entity to be identified later, and follows a similar format as the above policeman example. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of GONG, and incorporate with the reference of OVERELL, by using the teachings of GONG of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with OVERELL’s teaching of having a prompt specifying a type of entity. One of ordinary skill in the art would be motivated to do so because by integrating OVERELL’s framework into the methods of GONG, one with ordinary skill in the art would achieve the goal of “many advantages in terms of efficient processing of the query,” (OVERELL, [0239]), and has “instructions thus providing the user with a mechanism to correct and improve the problem for all users,” (OVERELL, [0948]). Claim 11: Regarding claim 11, GONG teaches “ A system for training a language model,… ” See GONG in [0003] describe "This presents two tasks to the question and-answer system: natural language understanding and the construction of knowledge bases. “ Here, GONG describes a system for understanding natural language. See GONG in [0106] describe "Processing NLP task by means of neural network model, 768-dimensional word embedding codes for each token are generated by means of a BERT pre-training model. " Here, GONG describes pre-training a BERT model for NLP task. Further, GONG teaches “ retrieve a knowledge sentence, related to an input sentence, from a knowledge base ;” See GONG in [0003] describe “The general information retrieval-based question-answering system works in two steps: question resolution and answer retrieval. Problem analysis is a natural language understanding task , and the main work of the problem analysis is to extract useful information from a question set by a user so as to guide subsequent retrieval . The answer retrieval is to find the answer from the constructed domain knowledge base.” Here, GONG explains retrieving an answer which is in a form of a natural language. Further, see GONG in [0004] specify “The problem analysis method adopts two sequence tagging technologies of named entity identification and part of speech tagging, which are also called slot filling. The sequence tagging is to regard an input sentence as an input sequence , and tag each word in the input sequence, so as to tag important elements in the sentence.” Here, GONG describes the input sequence is similar to a question, and using the answer retrieval step involves finding an answer also in a form of a sentence from a knowledge base. This relates to retrieving a knowledge sentence, where the knowledge sentence is construed to mean the sentence, where after retrieving the relevant information from a knowledge base, generates an answer to a user prompt to a language model. Further, GONG teaches “ encode the input sentence, the knowledge sentence, and a prompt into an intermediate representation; ” See GONG in [0039] describe “the method for realizing the question-answering system is used in a specific field, a field knowledge base is designed by considering the characteristics of knowledge in the field, the structure of the knowledge base is considered , two sequence tagging tasks of named body recognition and part of speech tagging are completed based on a bidirectional transducer encoder representation technology (BERT), sentence information is extracted in a targeted manner ,” Here, GONG describes BERT, a model that is specifically designed to encode two sequence tasks: 1. the extracted sentence information (i.e. input sentence) as well as the 2. sequence task extracted from a knowledge base (i.e. knowledge sentence), into contextualized numeric representations called embeddings or label text sequence with tags in this case. GONG mentions BERT model has an encoder and transducer, and performs the process of encoding. Further, see GONG in paragraphs [0103-0106], describe "The specific training method comprises the following steps: (1) and performing participlization on the corpus data training set labeled above, namely dividing the question text into separate participles (tokens)... (2) And then converted to word embedding . Processing NLP task by means of neural network model, it is often necessary to map words into a vector in a high-dimensional dense space , and express semantic relation between corresponding words by cosine distance between each vector, which is word embedding. " Here, GONG describes labeling the question text or prompt into word embeddings, which shows encoding the prompt into an intermediate representation. Note the examiner construes intermediate representation to mean any vector representation in space, which includes embeddings, from the specification [0023] stating “the encoder encodes the inputs together and generates an intermediate representation, for example as a vector in a latent space.” Further, see GONG in paragraphs [0110-0111, 0114-0116] describe “ (1) the method includes inputting question texts , simply screening the text lengths, ... (2) And performing participlization and converting into word embedding . …(1) Extracting relation elements contained in the question from the labeled result; (2) converting relational elements into database query statements (3) And giving corresponding prompts for the conditions of inquiring or not inquiring answers.” Here, GONG describes using the process of viewing the question text, converting this into word embeddings (i.e. intermediate representation), extracting relation elements from the question (this step implies using information from a knowledge base), then giving prompts that correspond with the content of the question. The term corresponding prompts means prompts that relate to the original text content, which was converted into word embeddings. Since the prompt is originally part of the text, which is converted to word embeddings, this means the prompt also was converted into intermediate representation. Further, see GONG in [0020 - 0023] describe “step 2-2: obtaining a question sentence of a user , screening according to the length of the text, and then preprocessing the screened text; step 2-3: inputting the preprocessed data into a sequence labeling model to obtain a label sequence; step 2-4: extracting relationship elements from the tag sequence; step 2-5: converting the relationship elements into database query statements ;” Here, GONG describes the process after getting a question sentence from a user, the method helps preprocess data into sequence labeling model, which refers to the BERT model, and labels the sequence (which includes input sentence). See GONG in [0025-0030] for details. Further, GONG teaches “decode the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ; ” See GONG in [0090 - 0091] describe "The encoder representation model (BERT) of a bidirectional Transformer is composed of a number of bidirectional Transformer modules, each of which comprises a concatenation of a number of encoders and decoders. A single decoder or encoder has one attribute layer, and inputs one fully connected layer after residual chaining and normalization," Further, see GONG in [0105-0106] describe "(2) And then converted to word embedding. Processing NLP task by means of neural network model, it is often necessary to map words into a vector in a high-dimensional dense space, and express semantic relation between corresponding words by cosine distance between each vector, which is word embedding." See [0020-0023] where GONG describes “step 2-3: inputting the preprocessed data into a sequence labeling model to obtain a label sequence; step 2-4: extracting relationship elements from the tag sequence; step 2-5: converting the relationship elements into database query statements .” Here, GONG teaches the BERT model contains decoders in [0090-0091], and later GONG describes converting the text sentence from word embedding (i.e. intermediate representation) into expressing semantic relation between corresponding words, and later labeling those relationship elements by tags which relates to named entity. From the specification note in [0018] describe “ In the specific example of FIG. 1 , the input text 100 includes certain named entities, illustrated with bold text, including a personal name, “Ms. Green,” an organization's name, “LIRR,” and a location name, “Manhattan.” Each of these named entities has an associated type: person, organization, or location. A given corpus of training data may be labeled, such that each named entity may be associated with one or more such types . ” Note the examiner construes ‘named’ to mean any label or category of a text or word (viewed as ‘entity’). Further, see GONG in [0112] describe “(3) And calculating in the network, performing softmax transformation …, taking the obtained maximum probability, and taking the corresponding sequence tag as the tag corresponding to the token. And obtaining a label sequence corresponding to the sentence .” Using named entity means to label an item or text with a name by a tag or other labels. Here, GONG has named or labeled the data containing words into a labeled sequence of the words from [0105-0106], and later obtain a label sequence that corresponds to a sentence in [0112], which relates to named entity from the input sentence. Further, GONG teaches “ and fine-tune a language model …” See GONG in [0028] describe " the training method comprises the following steps: performing word segmentation processing on data in the training data set, converting the data into word embedding, and inputting the word embedding into a model for training; [0029] finally, fine tuning is carried out on the encoder representation model of the bidirectional Transformer;" Here, GONG shows details in fine-tuning a model of the bidirectional transformer model for word segmentation in language processing. Further details, see GONG in [0018-0019]. However, GONG did not teach “decode the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ;” “and fine-tuning a language model based on the named entity,” “ a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to…” In an analogous field, OVERELL teaches “decode the intermediate representation to generate a named entity from the input sentence that is of a type specified by the prompt ;” See OVERELL in [0169, 0171] describe "One method for finding the principal class of an object is first to identify the classes of which the object is a member, i.e. a query is done looking for objects to which the entity has the relation [is an instance of]. The resulting class objects are then ordered using the [is a subclass of] relation and the most specific class labelled as a principal class is then considered the PC for the object … A similar check is done while adding a new object when prompting the user entity for a class of which the object is a member. After prompting the user entity for a class, both this class and the classes to which this class is on the right in the relation [is a subclass of] are retrieved from the knowledge base and again they are ordered. The most specific class labelled as a PC is taken as the object's PC. If one is not found using this method the user entity is prompted for a more specific class, e.g. the string “policeman” will find [human being] as the principal class (the class of policemen is a subclass of the PC [human being]) but “living thing” will result in the user being prompted to be more specific ." Here, OVERELL describes class as a type of the entity, where class is a certain category of the text the user has input into their query (i.e. input sentence). OVERELL shows that the prompt asks the user to label a text by a class (i.e. type). For example, if the user enters policeman, the prompt will specify the class (i.e. type), which is a human being. When OVERELL mentions “ The most specific class labelled as a PC is taken as the object's PC. If one is not found using this method the user entity is prompted for a more specific class ,” this means that if a prompt does not recognize the term entered, then the prompt will ask to specify what is the type of item or entity. If the user enters a living thing, then the prompt asks the user to specify the term. For example, living thing is a broad term, whereas policeman is a type of human being. In this case, OVERELL shows the prompt ask or clarify the type of entity the user wants to know about. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of GONG and incorporate into the teachings of OVERELL because both references teach retrieving knowledge sentences based on a knowledge base, and encoding and decoding steps for a language model. One of ordinary skill in the art would be motivated to do so because by integrating OVERELL’s framework into the methods of GONG, one with ordinary skill in the art would achieve a goal of “many advantages in terms of efficient processing of the query,” (OVERELL, [0239]), and has “instructions thus providing the user with a mechanism to correct and improve the problem for all users,” (OVERELL, [0948]). However, GONG in view of OVERELL did not teach “and fine-tuning a language model based on the named entity,” In an analogous field, JIANG teaches “ and fine-tune a language model based on the named entity ,” See JIANG in [n0117] describe “ The labeled data required for encoder-decoder learning can be large-scale collected text data and related subgraphs retrieved from the knowledge graph based on it. Text data can be sentences, phrases , or passages, or combinations of these linguistic units, to be compatible with different higher-level artificial intelligence tasks. The process of retrieving knowledge subgraphs from a knowledge graph based on text data can employ simple character matching or utilize mature basic tools such as entity recognition and entity chaining.” Here, JIANG mentions that “ The labeled data required for encoder-decoder learning can be large-scale collected text data”, and text data includes sentences or phrases. See JIANG in [n0034] describe “ Knowledge graphs are also widely used to improve higher-level artificial intelligence tasks or higher-level natural language processing tasks (referred to as higher-level tasks). These higher-level tasks typically use a pre-trained language model as a basis and fine-tune the parameters on their own labeled data .” Here, JIANG describes fine tuning parameters of a language model with knowledge graphs using labeled data in [n0034], where labeled data refers to labeled text data in the form of labeled sentences mentioned in [n0117], and relates to named entity . Note the examiner construes the term ‘named entity’ to mean labeled item, such as classifying text into a category, such as identifying Manhattan as a location. Note the examiner also construes ‘named’ to mean any label or category of a text or word (viewed as ‘entity’). From the specification note in [0018] describe “ In the specific example of FIG. 1 , the input text 100 includes certain named entities, illustrated with bold text, including a personal name, “Ms. Green,” an organization's name, “LIRR,” and a location name, “Manhattan.” Each of these named entities has an associated type: person, organization, or location. A given corpus of training data may be labeled, such that each named entity may be associated with one or more such types . ” Note the examiner construes ‘named’ to mean any label or category of a text or word (viewed as ‘entity’). JIANG shows using labeled data in form of labeled sentences or labeled text (i.e. named entity) to fine tune a model and its parameters based on this labeled text data. Further, JIANG teaches “comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to…” See JIANG in [n0054] describe "electronic devices can be any device with computing capabilities, such as personal computers (PCs), mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and/or displays, such as mobile phones, tablets." Here, JIANG mentions a hardware device. Further, see JIANG in [n0141] describe "this disclosure also provides an electronic device , which may include the broadcaster client or server in the above embodiments. The electronic device may include at least one processor and a memory communicatively connected to the at least one processor . The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the model pre training method..." Here, JIANG shows a device (that includes hardware) and a memory that enables at least one processor to run a pre-training method on a model . See [n0019] in JIANG for details. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references GONG and OVERELL, and incorporate with the teachings of JIANG by using the teachings of GONG and OVERELL of retrieving knowledge sentences based on a knowledge base, and encoding and decoding steps for a language model, with JIANG’s teaching of fine-tuning a language model based on the named entity. One of ordinary skill in the art would be motivated to do so because by integrating JIANG’s framework into the methods of GONG and OVERELL, one with ordinary skill in the art would achieve a method that “provides a pre trained model for knowledge enhancement that achieves effective knowledge utilization while maintaining ease of use,” (JIANG, [n0120]). Claim 12: Regarding claim 12, the claim recites similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Claim 13: Regarding claim 13, the claim recites similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Claim 19: Regarding claim 19, the claim recites similar limitations as corresponding claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale. Claims 4, 10, 14, and 20 are rejected under 35 U.S.C. 103 over GONG, in view of OVERELL, further in view of JIANG, and further in view of Ya, J. et al., in “A compare-aggregate model with external knowledge for query-focused summarization”, published on October 21, 2020, available at https://link.springer.com/chapter/10.1007/978-3-030-62008-0_5 , (hereafter, YA). Claim 4: Regarding claim 4, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 3. Further, OVERELL teaches “ The method of claim 3, wherein retrieving the knowledge sentence further includes generating a set of knowledge sentences from the entities and the relations and selecting a percentage with a highest relevance to the input sentence,” See OVERELL in [0088] describe “Knowledge generation is implemented in the preferred embodiment via a collection of “generators” which comprise a pattern of the facts which they can generate in combination with one or more mechanisms to generate facts which match this pattern. Some generators achieve this by providing a query linked to the pattern which if answered provides values for unknowns in the pattern thus enabling the generation of the facts.” Here, OVERELL shows that a collection of generators shows a set of items are created from a knowledge base that has entities and relations. Further, see OVERELL in [1192] describe "The Sentence Splitter (4804) takes unstructured text as input and generates a list of sentences." Later, see OVERELL in [1193] describe "The Sentence Markup may … mark up entities of the following type: dates, currencies, quantities, named entities. Named Entities may not only be identified and/or linked to an external Knowledge Base but can also be classified and/or linked within the document being processed either within sentences or across multiple sentences." Here, OVERELL explains that the set contain sentences that has labeled entities and relations, classified from a knowledge base. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of GONG, and incorporate with the reference of OVERELL, by using the teachings of GONG of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with OVERELL’s teaching of generating a set of knowledge sentences from the entities and the relations. One of ordinary skill in the art would be motivated to do so because by integrating OVERELL’s framework into the methods of GONG, one with ordinary skill in the art would achieve the goal of “many advantages in terms of efficient processing of the query,” (OVERELL, [0239]), and has “instructions thus providing the user with a mechanism to correct and improve the problem for all users,” (OVERELL, [0948]). However, GONG in view of OVERELL, further in view of JIANG, did not teach “The method of claim 3, wherein retrieving the knowledge sentence further includes generating a set of knowledge sentences from the entities and the relations and selecting a percentage with a highest relevance to the input sentence.” In an analogous field, YA teaches “The method of claim 3, wherein retrieving the knowledge sentence further includes generating a set of knowledge sentences from the entities and the relations and selecting a percentage with a highest relevance to the input sentence ,” See YA in page 72, section 3.1 Problem Formulation, describe “Sentence Tokenizer decomposes the input document into sentences… Sentence Relevance Scorer quantifies the relevance matching between query and each sentence with a score Score(Q, S i ) . Sentence Selector selects sentences with high relevance scores and little redundancy to output a query-focused summary that meets the requirement.” Also, see YA in page 76, section 3.3 Sentence selection, describe “A summary is expected to provide both relevant and non-redeemable information. Our proposed approach focuses on sentence scoring . We employ a simple greedy algorithm used in previous work [4] to select sentences which can be chosen as summary. We sort the sentences in descending order according to the derived relevance matching scores . And, we iteratively dequeue the top-ranked sentence, and append it to the current summary if it is non-redundant. A sentence is considered non-redundant if it contains significantly new bi-grams compared with the current summary content. We set the threshold of the new bi-gram ratio to 0.5.” Here, YA describes calculating relevance scores between a query and a sentence from an input document (i.e. input sentence), and select sentences that have high relevance scores, and provides an example of 0.5 as a relevance score threshold, showing a percentage relevance. See page 79, section 5.1 Unsupervised Methods where YA describes “Graph-based models played a leading role in the extractive summarization area, due to its ability to reflect various sentence relationships. For example, Wan et al. [21] adopted manifold ranking to make use of the within-document sentence relationships, the cross-document sentence relationships and the sentence-to-query relationships. In a graph, nodes are sentences and the edge scores reflect the similarity between sentences, each node is given a relevance weight based on its relevance to the query.” Here, YA shows that scores and relevance weights are calculated for sentences to determine how a particular sentence is relevant to the query. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of GONG, OVERELL, and JIANG, and incorporate with the teachings of YA by using the teachings of GONG, OVERELL, and JIANG, of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with YA’s teaching of selecting a percentage with a highest relevance to the input sentence. One of ordinary skill in the art would be motivated to do so because by integrating YA’s framework into the methods of GONG, OVERELL, and JIANG, one with ordinary skill in the art would achieve the goal of “finding out such words and enhancing the weights of them can help to make a more effective summary related to the query,” (see YA in page 69, paragraph 5, Introduction section). Claim 10: Regarding claim 10, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 1. Further, OVERELL teaches “ wherein encoding includes aggregating a representation of the input sentence with a representation of the knowledge sentence, based on the prompt ,” See OVERELL in [1135] describe “The preferred embodiment is operable to receive the request data via an HTTP (or HTTPS) request where the request data is encoded using HTTP request variables.” Under broadest reasonable interpretation, e ncoding is construed to be inputting information from one format into another format, which OVERELL shows by encoding the information. Further, see OVERELL in [0747] mention “For example, when adding a new object to the knowledge base , the user may be prompted for the name of a class to which this object belongs. If the user tries to specify a class which does not yet appear in the system, they may choose to add the class, opening the “add class” process as a sub-process.” Here, OVERELL shows this method relates to using a knowledge base to apply information for a prompt. Further, OVERELL shows in [1195] “The POS Tagger (4808), takes sentences (marked up or otherwise) as input, and generates lists of Tokens.” Here, OVERELL shows a method that takes sentences as input. From [0747], OVERELL shows taking input sentence and extracting information from a knowledge base, and adding information that was not part of the knowledge base. Later, see OVERELL in [0094] for more details. Later, OVERELL shows in [0687] "In an example authentication interaction with the system, the system first prompts the user to say who he/she is (step 1902). The user responds by entering his/her name (e.g. “Michael Smith”). The system then looks up this natural language string in the knowledge base [“michael smith”] to see which entities it could denote . If it only denotes one entity, the system moves immediately on to prompting for a password". Here, OVERELL illustrates the process of first using a prompt to ask information from a user (that has been encoded in http request format from [1135]) that contains input sentence (from the example is entering the user’s name) with a knowledge sentence (from the example of looking up in natural language string in the knowledge base), then based on the prompt, the system asks the user to confirm credentials with entering a password. However, OVERELL did not teach “ encoding includes aggregating a representation of the input sentence with a representation of the knowledge sentence, based on the prompt” In an analogous art, YA teaches “The method of claim 1, wherein encoding includes aggregating a representation of the input sentence with a representation of the knowledge sentence, based on the prompt,” See YA in page 75, section 3.2 Sentence relevance scorer, External Conceptual Knowledge describe “construct a subgraph which is derived from ConceptNet based on the concepts mentioned in the input query and document by mapping words to concepts in the external knowledge graph . This subgraph includes connections between the concepts mentioned in the query and each sentence in document, via concepts available in ConceptNet. Such subgraph enhances the concepts mentioned in the query and document with additional information and can improve the performance of relevance matching in our task.” Further, see YA in page 74, section 3.2, subsection Comparison , describe “ . The goal of the comparison layer is to match each s - j (the j -th word in sentence S and its context in document) with h j s ( a weighted version of query Q that best matches s - j ). A comparison function is used to match each word in the query and sentence to a corresponding attention-applied vector representation : PNG media_image3.png 132 1152 media_image3.png Greyscale where denotes element-wise multiplication.” Here, YA shows that the s - j in sentence S shows a representation of the input sentence , and h j s shows a representation of the knowledge sentence. Later, YA describes in page 75, section 3.2, sub section Aggregation that “ In the aggregation layer, we apply a comparison function to each pair of s - j and h j s to obtain a series of vectors. Finally, we aggregate these vectors using a one-layer CNN [18] with n-types of filters and calculate the relevance score for each sentence S .” Later, YA shows in section 3.2 that the s - j and h j s , are both aggregated to obtain a series of vectors. These vectors are combining both the representations from input sentence and knowledge sentence. Further, see YA in page 69, Introduction section, paragraphs 6-7, describe “To address the above problem caused by sentence-level attention, we introduce a fine-grained and interactive word-by-word attention to the query-focused extractive summarization system. In that way, we can capture the real intent of query. Specifically, we utilize a Compare-Aggregate framework to implement the idea. Furthermore, to fill the expression gap between query and document, we leverage conceptual knowledge to enrich the model.” See figure 2 in YA for details. PNG media_image4.png 346 1417 media_image4.png Greyscale Here, YA shows using a query-focused (i.e. based on prompt) feature and incorporate into an aggregation method called compare-aggregate framework. See YA in page 76, section 4.1 Experimental Setup, paragraph 2, for more details. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine reference GONG, OVERELL, and JIANG, and incorporate with the teachings of YA by using the teachings of GONG, OVERELL, and JIANG, of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with YA’s teaching of the method of aggregating a representation of the input sentence with a representation of the knowledge sentence. One of ordinary skill in the art would be motivated to do so because by integrating YA’s framework into the methods of GONG, OVERELL, and JIANG, one with ordinary skill in the art would achieve the goal of “finding out such words and enhancing the weights of them can help to make a more effective summary related to the query,” (see YA in page 69, paragraph 5, Introduction section). Claim 14: Regarding claim 14, the claim recites similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Claim 20: Regarding claim 20, the claim recites similar limitations as corresponding claim 10 and is rejected for similar reasons as claim 10 using similar teachings and rationale. Claims 5 and 15 are rejected under 35 U.S.C. 103 over GONG, in view of OVERELL, further in view of JIANG, and further in view of YA, and further in view of Li, J. et al., in “A Survey on Deep Learning for Named Entity Recognition,” available at https://ieeexplore.ieee.org/abstract/document/9039685 , (hereafter, LI). Claim 5: Regarding claim 5, GONG in view of OVERELL, further in view of JIANG, and further in view of YA, teach the limitations in claim 4. However, GONG in view of OVERELL, further in view of JIANG, and further in view of YA, did not teach “ In an analogous field, LI teaches “ The method of claim 4, wherein the set of knowledges sentences include sentences of the form <entity> is a <type> . ” See LI in page 51, section 2 Background, subsection 2.1 What is NER? describe "given a sequence of tokens s= ⟨ w1,w2,…,wN ⟩ , NER is to output a list of tuples ⟨ Is,Ie,t ⟩ , each of which is a named entity mentioned in s. Here, Is ∈ [1,N] and Ie ∈ [1,N] are the start and the end indexes of a named entity mention; t is the entity type from a predefined category set . Fig. 1 shows an example where a NER system recognizes three named entities from the given sentence. When NER was first defined in MUC-6 [10], the task is to recognize names of people, organizations, locations, and time, currency, percentage expressions in text. Note that the task focuses on a small set of coarse entity types and one type per named entity . We call this kind of NER tasks as coarse-grained NER [10], [11]." Here, LI mentions a model that recognizes entities by their type or if that entity belongs to a predefined category. See figure 1 in LI for details. PNG media_image5.png 600 1440 media_image5.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of GONG, OVERELL, JIANG, and YA and incorporate with the teachings of LI by using the teachings of GONG, OVERELL, JIANG, and YA, of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with LI’s teaching of the knowledge sentences being in the form of <entity> is a <type>. One of ordinary skill in the art would be motivated to do so because by integrating LI’s framework into the methods of GONG, OVERELL, JIANG, and YA, one with ordinary skill in the art would achieve the goal of “integrating or fine-tuning pre-trained language model embeddings is becoming a new paradigm for neural NER. When leveraging these language model embeddings, there are significant performance improvements,” (see LI in section 3.5 Summary of DL-based NER, page 61 ). Claim 15: Regarding claim 15, the claim recites similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale. Claims 6 and 16 are rejected under 35 U.S.C. 103 over GONG, in view of OVERELL, further in view of JIANG, and further in view of Zhu, C. et al, (US PG Pub. No. US20220230625A1), published on July 21, 2022, (hereafter, ZHU). Claim 6: Regarding claim 6, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 2. However, GONG in view of OVERELL, further in view of JIANG, did not teach “ The method of claim 2, wherein the knowledge base is a multilingual knowledge graph .” In an analogous art, ZHU teaches “ The method of claim 2, wherein the knowledge base is a multilingual knowledge graph ,” See ZHU in [0181-0182] describe "Additionally, … the knowledge module 1312 is also configured to generate knowledge-based entity representations (e.g., knowledge information 1318) in the first language based on the second knowledge graph. In such embodiments, the multi-lingual integrated knowledge-language module 1330 is trained to perform semantic analysis in the first language based on knowledge learned from entities and entity relations in the second language. The language module 1314 is also configured to provide context information 1316 to the knowledge module 1312. The first language knowledge graph 1320 and second language knowledge graph 1322 are optionally language or textual-based knowledge graphs corresponding to the same or different languages . Accordingly, methods and systems are provided for obtaining electronic content comprising a first set of speech transcriptions in the first language and applying the electronic content as input to the knowledge-language module, wherein the first set of speech transcriptions is translated into a second set of speech transcriptions in the second language using the integrated knowledge-language module.” Here, ZHU teaches using knowledge graphs that correspond to either the same or different languages, and they are translated from one language to another language. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of GONG, OVERELL, and JIANG, and incorporate with the teachings of ZHU by using the teachings of GONG, OVERELL, and JIANG, of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with ZHU’s teaching of a knowledge base is a multilingual knowledge graph. One of ordinary skill in the art would be motivated to do so because by integrating ZHU’s framework into the methods of GONG, OVERELL, and JIANG, one with ordinary skill in the art would achieve a method to “train a model (e.g., for a specific query or specific task) to increase accuracy, efficiency, and efficacy of that model in the desired natural language understanding application,” (ZHU, [0049]). Claim 16: Regarding claim 16, the claim recites similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale. Claims 7 and 17 are rejected under 35 U.S.C. 103 over GONG, in view of OVERELL, further in view of JIANG, and further in view of Meftah, S. et al., in “Neural Transfer Learning for Domain Adaptation in Natural Language Processing,” available at https://theses.hal.science/tel-03206378/, (hereafter, MEFTAH), and further in view of Sedinkina, M. in “Domain adaptation in Natural Language Processing,”, published on 2021, available at https://edoc.ub.uni-muenchen.de/27831/7/Sedinkina_Marina.pdf, (hereafter, SEDINKINA). Claim 7: Regarding claim 7, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 1. However, GONG in view of OVERELL, further in view of JIANG, did not teach “The method of claim 1, further comprising pre-training the model with labeled training data from a source domain, wherein the input sentence includes labeled entities from a second domain .” In an analogous art, MEFTAH teaches “The method of claim 1, further comprising pre-training the model with labeled training data from a source domain , wherein the input sentence includes labeled entities from a second domain.” See MEFTAH in page iii, abstract describe the study “propose two methods to transfer the knowledge encoded in the neural representations of a source model– pretrained on large labelled datasets from the source domain – to the target model.” Here, MEFTAH shows that the pre-training process in the source domain involves using labelled datasets. Further, see MEFTAH in page 12, section 2.2 Formalisation, describe “The aim behind using transfer learning is to improve the learning of the predictive target function f T by leveraging the knowledge gained from D S and T S . Generally, in a transfer learning scheme, labelled training examples from the source domain D S = {(x i S ,y i S ) ∈ X S ×Y S : i ∈ (1,...,n S )} are abundant.” Here, MEFTAH teaches that training examples from a source domain are labelled. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references GONG, OVERELL, and JIANG, and incorporate with the teachings of MEFTAH by using the teachings of GONG, OVERELL, and JIANG, of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with MEFTAH’s teaching of pre-training the model with labeled training data from a source domain. One of ordinary skill in the art would be motivated to do so because by integrating MEFTAH’s framework into the methods of GONG, OVERELL, and JIANG, one with ordinary skill in the art would achieve a method to “detecting proper nouns in the POS tagging task would hopefully help to better identify named entities in Named Entity Recognition (NER) task... when tasks are related, the efficiency of MTL is significant for many evidences. First, MTL allows augmenting training data, implicitly, which begets a better regularisation and thus avoids over-fitting. Second, MTL allows an “eavesdropping” process, which means that when two tasks ‘A’ and ‘B’ are jointly trained, in some cases, a set of features that are important for task ‘A’ can be easier to learn by task ‘B’, ” (See MEFTAH in page 22, section 2.5.2. Multi-task learning). However, GONG in view of OVERELL, further in view of JIANG, further in view of MEFTAH, did not teach “The method of claim 1, further comprising pre-training the model with labeled training data from a source domain, wherein the input sentence includes labeled entities from a second domain .” In an analogous art, SEDINKINA teaches “The method of claim 1, further comprising pre-training the model with labeled training data from a source domain, wherein the input sentence includes labeled entities from a second domain .” See SEDINKINA in page 63, section 5.2, Overview of methods, describe “During pre-training, the model learns a general-purpose representation of inputs, and during fine-tuning (adaptation), the representation is transferred to a new task. Diverse context-sensitive models have been successfully applied on various downstream NLP tasks, ranging from question-answering to sentiment analysis... That is why this model is beneficial for sequence labeling problems such as NER and POS tagging ... Due to this architecture, it achieves state-of-the-art performance on a large suite of sentence-level and token-level tasks , such as question-answering, paraphrasing, POS tagging and sentiment analysis.” Here, SEDINKINA mentions the model learns raw data such as input from text at a sentence-level (i.e. input sentence) that captures essential features usable across many different, often unseen, downstream tasks. Further, SEDINKINA in page 59, section 5.2, describes “domain adaptation (DA) techniques have been proposed to facilitate the transfer of knowledge from one domain (source) to a specialized domain (target) . They have been successfully applied for different NLP tasks such as sentiment analysis (Glorot, Bordes, and Bengio, 2011; Pan et al., 2010b; Sarma, Liang, and Sethares, 2019), POS tagging (Astudillo et al., 2015; Schnabel and Schütze, 2014) or Named Entity Recognition (NER) (Newman-Griffis and Zirikly, 2018).” Here, SEDINKINA mentions that using domain adaptation methods apply to named entity recognition, which classifies unstructured text into pre-defined categories (i.e. labeled entities) like person names, organizations, locations, and dates. This classification, as SEDINKINA describes, is a transfer of knowledge from one source domain to a specialized target domain (i.e. second domain) . PNG media_image6.png 392 914 media_image6.png Greyscale From table 13, under the supervised domain adaptation setting, the method SEDINKINA discusses use labeled data for the target domain (i.e. second domain). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of GONG, OVERELL, JIANG, and MEFTAH and incorporate with the teachings of SEDINKINA by using the teachings of GONG, OVERELL, JIANG, and MEFTAH of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model, with SEDINKINA’s teaching of labeled entities from a second domain. One of ordinary skill in the art would be motivated to do so because by integrating SEDINKINA’s framework into the methods of GONG, OVERELL, JIANG, and MEFTAH, one with ordinary skill in the art would achieve a method that “improve the performance especially for tasks involving a longer text sequence. Hence, the authors demonstrate consistent improvement over BERT on a wide spectrum of problems including language understanding, reading comprehension text classification, and document ranking tasks,” (see SEDINKINA in page 16, third full paragraph, from section 2.2.2 Contextualized Embedding-based Methods, Transformer-based). Claim 17: Regarding claim 17, the claim recites similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Claims 8 and 18 are rejected under 35 U.S.C. 103 over GONG, in view of OVERELL, further in view of JIANG, and further in view of Liu, X. et al., (Pub. No. CN114860915A), published on August 5, 2022, (hereafter, LIU). Claim 8: Regarding claim 8, GONG in view of OVERELL, further in view of JIANG, teach the limitations in claim 1. Further, GONG teaches “The method of claim 1, wherein encoding uses an attention function: PNG media_image1.png 9 120 media_image1.png Greyscale where l designates an attention layer, Q, K, and V are query, key, and value parameters of the attention layer, respectively … and d is a dimension size ” See GONG in paragraphs [0092-0095, 0097] describe “For a single attribute, represented by self-attribute, the input sequence (represented by matrix X) is cross-multiplied with the weight matrix to obtain three matrices of Q (query), Key (Key) and Value, and finally, for each row of the V matrix, softmax weighted average of inner product of Q and K is taken as output Z of the attribute layer. [0093] Q=X x WQ [0094] K=X x WK [0095] V=X x WV … The Attention calculation process largely uses matrix operation, and can maximally utilize the computer to optimize the matrix operation.” See GONG in [0096] show equation PNG media_image7.png 110 316 media_image7.png Greyscale Here, GONG shows the softmax attention function where output Z includes the query Q, key K, and value V parameters of the function, and includes dimension size d k . However, GONG in view of OVERELL, further in view of JIANG did not teach “ ϕk and ϕv are prompt parameters corresponding to K and V, and ⋅ : ⋅ denotes a concatenation operation ,” In an analogous field, LIU teaches “ ϕk and ϕv are prompt parameters corresponding to K and V, and ⋅ : ⋅ denotes a concatenation operation ...” See LIU in [n0083-n0086] describe " the pre-trained language model is subjected to cue learning using the parameter update amount in this round, including: Step S131: Modify the attention matrix in the Transformer class model according to the parameter update amount in this round to obtain the modified attention matrix… implementation of step S131 above is as follows: The basic calculation operation process of the Transformer class model has been described above. Since the size of each head of each layer of K<sup>(h)</sup>(x<sub>i</sub>) and V<sup>(h)</sup>(x<sub>i</sub>) is exactly the same, for ease of understanding, the following explanation uses one head of one layer as an example. Assume that the prompt matrix is represented as P<sub>k</sub>, P<sub>q</sub> and/or P<sub>v</sub> (the prompt matrix can be randomly initialized at the beginning, and then the prompt matrix is determined by the parameter update amount in the current round ). Here, the prompt matrix P<sub>k</sub> is used to modify the key matrix (K) in the attention matrix, the prompt matrix P<sub>q</sub> is used to modify the query matrix (Q) in the attention matrix, and the prompt matrix P<sub>v</sub> is used to modify the numerical matrix (V) in the attention matrix. In practice, the prompt matrix can be used to modify any one matrix (Q, K, or V), any two matrices (Q and K, or K and V, or Q and V), or all three matrices (Q, K, and V)." Here, LIU shows the parameters are updated within the attention function of the model, where " prompt matrix is determined by the parameter update amount " shows updating prompt parameters in the function. Note the prompt parameters are construed by examiner to include any parameters that modifies the attention matrix for a language model. Further, see LIU in [n0057] describe "then, the first attention matrix and the second attention matrix are concatenated along the h dimension to obtain the concatenated matrix. Finally, the concatenated matrix is linearly projected again to obtain the output data of multi-head attention. " Here, LIU mentions a concatenation operation on the attention function. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references GONG, OVERELL, and JIANG, and incorporate with the teachings of LIU by using the teachings of GONG, OVERELL, and JIANG, of retrieving knowledge sentences based on a knowledge base to later fine-tune a language model using an attention function, with LIU’s teaching of prompt parameters and using a concatenation operation on the attention function. One of ordinary skill in the art would be motivated to do so because by integrating LIU’s framework into the methods of GONG, OVERELL, and JIANG, one with ordinary skill in the art would achieve a method to “optimize the prompt learning process of the pre-trained language model, and effectively accelerating the prompt learning efficiency of the pre-trained language model,” ([n0005], LIU). Claim 18: Regarding claim 18, the claim recites similar limitations as corresponding claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WenWei Zeng/Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146 Application/Control Number: 18/454,262 Page 2 Art Unit: 2146 Application/Control Number: 18/454,262 Page 3 Art Unit: 2146 Application/Control Number: 18/454,262 Page 4 Art Unit: 2146 Application/Control Number: 18/454,262 Page 5 Art Unit: 2146 Application/Control Number: 18/454,262 Page 6 Art Unit: 2146 Application/Control Number: 18/454,262 Page 7 Art Unit: 2146 Application/Control Number: 18/454,262 Page 8 Art Unit: 2146 Application/Control Number: 18/454,262 Page 9 Art Unit: 2146 Application/Control Number: 18/454,262 Page 10 Art Unit: 2146 Application/Control Number: 18/454,262 Page 11 Art Unit: 2146 Application/Control Number: 18/454,262 Page 12 Art Unit: 2146 Application/Control Number: 18/454,262 Page 13 Art Unit: 2146 Application/Control Number: 18/454,262 Page 14 Art Unit: 2146 Application/Control Number: 18/454,262 Page 15 Art Unit: 2146 Application/Control Number: 18/454,262 Page 16 Art Unit: 2146 Application/Control Number: 18/454,262 Page 17 Art Unit: 2146 Application/Control Number: 18/454,262 Page 18 Art Unit: 2146 Application/Control Number: 18/454,262 Page 19 Art Unit: 2146 Application/Control Number: 18/454,262 Page 20 Art Unit: 2146 Application/Control Number: 18/454,262 Page 21 Art Unit: 2146 Application/Control Number: 18/454,262 Page 22 Art Unit: 2146 Application/Control Number: 18/454,262 Page 23 Art Unit: 2146 Application/Control Number: 18/454,262 Page 24 Art Unit: 2146 Application/Control Number: 18/454,262 Page 25 Art Unit: 2146 Application/Control Number: 18/454,262 Page 26 Art Unit: 2146 Application/Control Number: 18/454,262 Page 27 Art Unit: 2146 Application/Control Number: 18/454,262 Page 28 Art Unit: 2146 Application/Control Number: 18/454,262 Page 29 Art Unit: 2146 Application/Control Number: 18/454,262 Page 30 Art Unit: 2146 Application/Control Number: 18/454,262 Page 31 Art Unit: 2146 Application/Control Number: 18/454,262 Page 32 Art Unit: 2146 Application/Control Number: 18/454,262 Page 33 Art Unit: 2146 Application/Control Number: 18/454,262 Page 34 Art Unit: 2146 Application/Control Number: 18/454,262 Page 35 Art Unit: 2146 Application/Control Number: 18/454,262 Page 36 Art Unit: 2146 Application/Control Number: 18/454,262 Page 37 Art Unit: 2146 Application/Control Number: 18/454,262 Page 38 Art Unit: 2146 Application/Control Number: 18/454,262 Page 39 Art Unit: 2146 Application/Control Number: 18/454,262 Page 40 Art Unit: 2146 Application/Control Number: 18/454,262 Page 41 Art Unit: 2146 Application/Control Number: 18/454,262 Page 42 Art Unit: 2146 Application/Control Number: 18/454,262 Page 43 Art Unit: 2146 Application/Control Number: 18/454,262 Page 44 Art Unit: 2146 Application/Control Number: 18/454,262 Page 45 Art Unit: 2146 Application/Control Number: 18/454,262 Page 46 Art Unit: 2146 Application/Control Number: 18/454,262 Page 47 Art Unit: 2146 Application/Control Number: 18/454,262 Page 48 Art Unit: 2146 Application/Control Number: 18/454,262 Page 49 Art Unit: 2146 Application/Control Number: 18/454,262 Page 50 Art Unit: 2146 Application/Control Number: 18/454,262 Page 51 Art Unit: 2146 Application/Control Number: 18/454,262 Page 52 Art Unit: 2146 Application/Control Number: 18/454,262 Page 53 Art Unit: 2146 Application/Control Number: 18/454,262 Page 54 Art Unit: 2146 Application/Control Number: 18/454,262 Page 55 Art Unit: 2146 Application/Control Number: 18/454,262 Page 56 Art Unit: 2146 Application/Control Number: 18/454,262 Page 57 Art Unit: 2146 Application/Control Number: 18/454,262 Page 58 Art Unit: 2146 Application/Control Number: 18/454,262 Page 59 Art Unit: 2146 Application/Control Number: 18/454,262 Page 60 Art Unit: 2146 Application/Control Number: 18/454,262 Page 61 Art Unit: 2146
Read full office action

Prosecution Timeline

Aug 23, 2023
Application Filed
May 14, 2026
Non-Final Rejection mailed — §101, §103, §112
Jul 31, 2026
Interview Requested
Aug 06, 2026
Examiner Interview Summary

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month