DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on 12/30/2025.
Claims 1-20 are pending and have been examined.
All previous objections / rejections not mentioned in this Office Action have been withdrawn by the examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments / Amendments
Regarding the Applicant’s arguments for the rejections under 35 U.S.C. § 101, applicant asserts that independent claim limitations cannot be performed in the mind and the amended claims integrate the alleged abstract idea into a practical application by implementing a cross-lingual evaluation metric to evaluate the language capability of a large language model without human annotations. Examiner respectfully disagrees. During patent examination, pending claims must be “given their broadest reasonable interpretation consistent with the specification.” MPEP 2111. Also, claims should not be interpreted by reading limitations of the specification into the claim, to narrow the scope of the claim, by implicitly adding disclosed limitations that have no express basis in the claim language. In re Prater, 415 F.2d 1393. Here, the steps in the claim language are broad and examiner interprets the claim broadly.
First, the steps recited in the claim limitation can be performed in the mind. Specifically, the human mind can think of a question in two languages, thinking of an answer to the question in both languages in the mind, thinking of a similarity in the answers, and adjusting rules in the thought process in answering the question in one language. The large language model is not specifically defined in the specification and is interpreted as a model that has a natural language input and a natural language output. The human mind can provide answers to questions. The large language model is interpreted as an additional element that does not integrate the judicial exception into a practical application. A use of a model that has a natural language input and natural language output does not improve technology to an extent to cover a particular solution to a problem. The claim encompasses mental observations or evaluations that can be practically performed in the human mind.
Second, amended claims does not integrate the alleged abstract idea into a practical application by implementing a cross-lingual evaluation metric to evaluate the language capability of a large language model without human annotations. MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. Here, the claim recites “evaluating a cross-lingual similarity between the target output in the target language and the reference output in the reference language to obtain an evaluation score for the target output”. The evaluation step, which provides the solution to the problem of evaluation of a low-resource LLM “more accurately in a non-biased way” (Spec. P0017), does not describe any specific technological improvement or provide a solution to the problem. Specifically, claim recites only the idea of a solution or outcome and fails to recite details of how a solution to a problem is accomplished. Therefore, the claims as currently recited does not overcome the 35 U.S.C. § 101 abstract idea rejection.
Regarding the Applicant’s arguments for the rejections under 35 U.S.C. § 103, applicant asserts that prior art reference does not teach “reference large language model” and “target large language model” by asserting that the models are different models and prior art reference utilizes the same model. Examiner respectfully disagrees. Prior art reference Ranaldi utilizes different models for different languages. Table 1 of Ranaldi illustrates that the prior art reference utilizes language specific versions of Alpaca. The models are distinct models with distinct model names as illustrated. The use of distinct models reads on the claim limitation of reference large language model and and target large language model.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 2, 5-11, 13-18, and 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1, 10, and 17 the limitations of “accessing a pair of parallel inputs comprising a reference input in a reference language and a target input in a target language, wherein the reference input in the reference language corresponds to the target input in the target language, the target language different from the reference language”, “executing a reference large language model for a generative task to obtain a reference output in the reference language based on the reference input”, “executing a target large language model for the generative task to obtain a target output in the target language based on the target input”, “evaluating a cross-lingual similarity between the target output in the target language and the reference output in the reference language to obtain an evaluation score for the target output”, and “fine-tuning the target large language model based on the evaluation score for the target output using a reinforcement learning algorithm”, as drafted, are processes that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. More specifically, the mental process of a human thinking of a question in two languages, thinking of an answer to the question in both languages in the mind, thinking of a similarity in the answers, and adjusting rules in the thought process in answering the question in one language. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the --Mental Processes-- grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
This judicial exception is not integrated into a practical application because the recitation of a system in claim 10 and a non-transitory computer readable storage medium in claims 17, reads to generalized computer components, based upon the claim interpretation wherein the structure is interpreted using P0089-P0092 in the specification. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using generalized computer components to think of a question in two languages, think of an answer to the question in both languages in the mind, think of a similarity in the answers, and adjust rules in the thought process in answering the question in one language amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible.
With respect to claim 2, 11, and 18, the claim recites “wherein the target input comprises communication data in the target language on a communication platform, comprising meeting transcripts, chat messages, audio messages, and emails”, which reads on a human reading chat messages in the mind. No additional limitations are present.
With respect to claim 5 and 13, the claim recites “further comprising translating the target input in the target language to obtain the reference input in the reference language using a machine translation model”, which reads on a human translating a question from one language to another language in the mind. No additional limitations are present.
With respect to claim 6 and 14, the claim recites “wherein the generative task comprises summarization, paraphrasing, or question-answer generation”, which reads on a human answering a question in the mind. No additional limitations are present.
With respect to claim 7, the claim recites “wherein the evaluation score comprises a reference- less machine translation metric, a similarity metric, or a predicted estimate for human judgment”, which reads on a human determining a similarity in responses to questions in the mind. No additional limitations are present.
With respect to claim 8, the claim recites “deploying multiple large language models for the generative task to generate multiple outputs based on the target input”, “evaluating a similarity between each output of the multiple outputs and the reference output to generate multiple evaluation scores”, “ranking the multiple large language models based on the multiple evaluation scores for the multiple large language models”, and “selecting a large language model corresponding to a highest evaluation score for the generative task in the target language”, which reads on a human utilizing multiple rules or instructions to generate responses to questions, determine similarity score between the responses in the mind, ranking the rules based on the scores, and selecting the rules with the highest score in the mind. No additional limitations are present.
With respect to claim 9 and 20, the claim recites “obtaining a first output generated from a target input by using the target large language model for the generative task at a first time”, “evaluating a first similarity between the first output and the reference output to generate a first evaluation score”, “obtaining a second output generated from the target input by using the target large language model for the generative task at a second time later than the first time”, “evaluating a second similarity between the second output and the reference output to generate a second evaluation score”, and “detecting a model drift of the target large language model based on a different between the first evaluation score and the second evaluation score”, which reads on a human determining a response to a question, determining a response to a question at a later time, detecting a similarity in the response, and detecting a change in the response in the mind. No additional limitations are present.
With respect to claim 15, the claim recites “wherein the evaluation score comprises a reference- less machine translation metric, a vector-based similarity metric, or a semantic similarity metric”, which reads on a human determining a similarity in responses to questions in the mind. No additional limitations are present.
With respect to claim 16, the claim recites “deploy multiple large language models for the generative task to generate multiple outputs based on the target input”, “evaluate a similarity between each output of the multiple outputs and the reference output to generate multiple evaluation scores for the multiple large language models”, “rank the multiple large language models based on the multiple evaluation scores”, and “select a large language model corresponding to a highest evaluation score for the generative task in the target language”, which reads on a human utilizing multiple rules or instructions to generate responses to questions, determine similarity score between the responses in the mind, ranking the rules based on the scores, and selecting the rules with the highest score in the mind. No additional limitations are present.
These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-7, 10, 12-15, 17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over "Empowering Cross-lingual Abilities of Instruction-tuned Large Language Models by Translation-following demonstrations" by Ranaldi et al., hereinafter Ranaldi, in view of Foley et al. (U.S. PG Pub No. 20250028992), hereinafter Foley.
Regarding claim 1, 10, and 17 Ranaldi teaches:
(Claim 1) A method comprising: (2.2 Instruction-tuning Paradigm, Instruction-tuning method.)
(Claim 10) A system comprising: a communications interface; a non-transitory computer-readable medium; and one or more processors communicatively coupled to the communications interface and the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non- transitory computer-readable medium to: (4.3.2 Models Setup, Running our experiments on a workstation equipped with two Nvidia RTX A6000 with 48 GB of VRAM.)
(Claim 17) A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to: (4.3.2 Models Setup, Running our experiments on a workstation equipped with two Nvidia RTX A6000 with 48 GB of VRAM.)
accessing a pair of parallel inputs comprising a reference input in a reference language and a target input in a target language, wherein the reference input in the reference language corresponds to the target input in the target language, the target language different from the reference language; (Figure 1, Parallel inputs are fed into Stanford Alpaca LLM and x-CrossAlpaca LLM. The instruction and input are in English for Stanford Alpaca LLM and foreign language for x-CrossAlpaca LLM.)
executing a reference large language model for a generative task to obtain a reference output in the reference language based on the reference input; (Figure 1, Original Stanford Alpaca LLM outputs text.)
executing a target large language model for the generative task to obtain a target output in the target language based on the target input; (Figure 1, x-CrossAlpaca LLM outputs text.)
evaluating a cross-lingual similarity between the target output in the target language and the reference output in the reference language to obtain an evaluation score for the target output; and (4 Experiments, Our approach is based on instruction-tuning on language-specific data augmented with a crosslingual semantic alignment. Hence, we set several baseline models (Section 4.1), which we augmented with our CrossAlpaca approach (Section 4.2). Finally, we performed a series of systematic evaluations (Section 4.3.1) to observe the impact of the proposed intervention.; 4.3.3 Evaluation, We estimate accuracy by measuring exact match values in the zero-shot setting.)
fine-tuning the target large language model based on the evaluation score for the target output using a reinforcement learning algorithm. (4.3.2 Models Setup, In order to align the results with the state-of-the-art models, we used the alpaca_LoRA (Hu et al., 2021b) code, adopting the same hyperparameters. We performed the fine-tuning with a single epoch and a batch-size of 128 examples. [Alpaca LoRAa code uses context target pairs in its training data to adjust model parameters.]; 2.3 Instruction-tuning is at hand, Parameter-Efficient Tuning (PEFT) is an efficient technique to adjust a small part of the model parameters and freeze the others. The main goal is to significantly reduce computational and storage costs while maintaining the performance of the original models. The standard techniques for PEFT are LoRA (Hu et al., 2021a), Prefix Tuning (Li and Liang, 2021), P-Tuning (Liu et al., 2022).)
Ranaldi teaches high level concepts of instruction tuning on language specific data and does not specifically teach:
fine-tuning the target large language model based on the evaluation score for the target output using a reinforcement learning algorithm.
Foley, however, teaches:
fine-tuning the target large language model based on the evaluation score for the target output using a reinforcement learning algorithm. (P0020, An embodiment uses one or more training prompt responses, or combinations of training prompts and their corresponding responses, to train an attribution model to attribute a fine-tuned model to a source foundation model. In some embodiments, the attribution model is a classifier model that outputs a classification of a fine-tuned model to a source foundation model, and a confidence score in the classification.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize an evaluation score between responses of large language models. It would have been obvious to combine the references because a fine-tuned model may generate false information and generating an evaluation score enables the fine-tuned model to link or attribute to the foundation model. (Foley, P0015)
Regarding claim 3 Ranaldi in view of Foley teaches claim 1.
Ranaldi further teaches:
wherein the reference large language model is a Generative Pre-trained Transformer 4 (GPT-4), and wherein the reference language is English. (Figure 1, Original LLM language in English.)
Ranaldi does not specifically teach:
wherein the reference large language model is a Generative Pre-trained Transformer 4 (GPT-4), and wherein the reference language is English.
Foley, however, teaches:
wherein the reference large language model is a Generative Pre-trained Transformer 4 (GPT-4), and wherein the reference language is English. (P0002, A foundation model, or base model, is a machine learning model that is trained, generally using self-supervision or semi-supervised learning. … Some examples of presently available LLMs are the Generative Pre-trained Transformer (GPT) family of models, such as GPT-3 and GPT-4.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have reference large language model be GPT-4. It would have been obvious to combine the references because utilizing GPT-4 as a reference LLM is a known technique to yield a predictable result of providing a generative output to an input.
Regarding claim 4 Ranaldi in view of Foley teaches claim 1.
Ranaldi further teaches:
wherein the target language is non-English, wherein the target large language model comprises a GPT model or a non-GPT model. (Figure 1, x-CrossAlpaca LLM using foreign language.)
Regarding claim 5 and 13 Ranaldi in view of Foley teaches claim 1 and 10.
Ranaldi further teaches:
further comprising translating the target input in the target language to obtain the reference input in the reference language using a machine translation model. (3.2 Cross-Lingual Instruction-tuning, Instruction-following demonstrations. The original version of Alpaca is in English. However, as described in Section 3.1, different open-source translations are produced with a translation engine. In our experiments, we propose the Instruction tuning phase with the original English version (enAlpaca) and the translated language-specific versions.)
Regarding claim 6 and 14 Ranaldi in view of Foley teaches claim 1 and 10.
Ranaldi further teaches:
wherein the generative task comprises summarization, paraphrasing, or question-answer generation. (4.3 Experimental Setup, Cross-lingual Question Answering Dataset (XQUAD) (Artetxeet al., 2019) consists of a subset of 240 paragraphs and 1190 question-answer pairs from the development set of SQuAD v1.1 (Rajpurkar et al., 2016) with their manual translations into several languages. Consequently, the dataset is entirely parallel across 11 languages.)
Regarding claim 7 Ranaldi in view of Foley teaches claim 1.
Ranaldi does not specifically teach:
wherein the evaluation score comprises a reference- less machine translation metric, a similarity metric, or a predicted estimate for human judgment.
Foley, however, teaches:
wherein the evaluation score comprises a reference- less machine translation metric, a similarity metric, or a predicted estimate for human judgment. (P0021, Each binary prediction is a classification of a training input as either attributed to a particular foundation model or not attributed to that particular foundation model, with a confidence score in the attribution.; P0022, The attribution model is implemented using TripletNet-based classifiers that use a margin-based loss function using the separate embeddings of the base and fine-tuned model responses. TripletNet is a presently available technique that is able to make predictions by taking in a single sentence or other portion of text, computing an output embedding (a numerical representation of a prompt response), and finding the closest embedding from the prompt responses in the training set and using the label of the training sentence as a prediction. The cosine distance between the anchor input (a reference input used in the loss function), positive example (a prompt pair or a prompt-response pair with correct attribution), and negative example (a prompt pair or a prompt-response pair with incorrect attribution is computed as the loss function (the value that the machine learning training procedure tries to minimize—e.g., cos (anchor, negative) and -cos (anchor, positive) for TripleNet).)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize an evaluation score of a similarity metric between responses of large language models. It would have been obvious to combine the references because a fine-tuned model may generate false information and generating an evaluation score enables the fine-tuned model to link or attribute to the foundation model. (Foley, P0015)
Regarding claim 15 Ranaldi in view of Foley teaches claim 10.
Ranaldi does not specifically teach:
wherein the evaluation score comprises a reference- less machine translation metric, a vector-based similarity metric, or a semantic similarity metric.
Foley, however, teaches:
wherein the evaluation score comprises a reference- less machine translation metric, a vector-based similarity metric, or a semantic similarity metric. (P0021, Each binary prediction is a classification of a training input as either attributed to a particular foundation model or not attributed to that particular foundation model, with a confidence score in the attribution.; P0022, The attribution model is implemented using TripletNet-based classifiers that use a margin-based loss function using the separate embeddings of the base and fine-tuned model responses. TripletNet is a presently available technique that is able to make predictions by taking in a single sentence or other portion of text, computing an output embedding (a numerical representation of a prompt response), and finding the closest embedding from the prompt responses in the training set and using the label of the training sentence as a prediction. The cosine distance between the anchor input (a reference input used in the loss function), positive example (a prompt pair or a prompt-response pair with correct attribution), and negative example (a prompt pair or a prompt-response pair with incorrect attribution is computed as the loss function (the value that the machine learning training procedure tries to minimize—e.g., cos (anchor, negative) and -cos (anchor, positive) for TripleNet).)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize an evaluation score of a similarity metric between responses of large language models. It would have been obvious to combine the references because a fine-tuned model may generate false information and generating an evaluation score enables the fine-tuned model to link or attribute to the foundation model. (Foley, P0015)
Regarding claim 12 and 19 Ranaldi in view of Foley teaches claim 10 and 17.
Ranaldi further teaches:
wherein the reference large language model is a Generative Pre-trained Transformer 4 (GPT-4), wherein the target large language model comprises a GPT model or a non-GPT model, wherein the reference language is English, and wherein the target language is non- English. (Figure 1, Original LLM language in English.; Figure 1, x-CrossAlpaca LLM using foreign language.)
Ranaldi does not specifically teach:
wherein the reference large language model is a Generative Pre-trained Transformer 4 (GPT-4), wherein the target large language model comprises a GPT model or a non-GPT model, wherein the reference language is English, and wherein the target language is non- English.
Foley, however, teaches:
wherein the reference large language model is a Generative Pre-trained Transformer 4 (GPT-4), wherein the target large language model comprises a GPT model or a non-GPT model, wherein the reference language is English, and wherein the target language is non- English. (P0002, A foundation model, or base model, is a machine learning model that is trained, generally using self-supervision or semi-supervised learning. … Some examples of presently available LLMs are the Generative Pre-trained Transformer (GPT) family of models, such as GPT-3 and GPT-4.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have reference large language model be GPT-4. It would have been obvious to combine the references because utilizing GPT-4 as a reference LLM is a known technique to yield a predictable result of providing a generative output to an input.
Claims 2, 11, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ranaldi in view of Foley and further view of Ben-David et al. (U.S. PG Pub No. 20240346238), hereinafter Ben-David.
Regarding claim 2, 11, and 18 Ranaldi in view of Foley teach claim 1, 10, and 17.
Ranaldi in view of Foley does not specifically teach:
wherein the target input comprises communication data in the target language on a communication platform, comprising meeting transcripts, chat messages, audio messages, and emails.
Ben-David, however, teaches:
wherein the target input comprises communication data in the target language on a communication platform, comprising meeting transcripts, chat messages, audio messages, and emails. (P0024, Generating sales call summaries and/or deal summaries using a specific-trained language model.; P0034, The summary generator is configured to process the data in the database.; P0032, A database stores message data that may be related to conversations. Such message data may include but is not limited to, email messages, chat logs, instant messages, or other text, or data otherwise, including contents of communications or content related to communications between and among individuals.; P0033, A database also includes transcript data related to audio/video conversations.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have target input as communication data. It would have been obvious to combine the references because the input of communication data to a an LLM is a known technique to yield a predictable result of providing a generative output to an input.
Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ranaldi in view of Foley and further view of Srinivasan et al. (U.S. PG Pub No. 20250111169), hereinafter Srinivasan.
Regarding claim 8 and 16 Ranaldi in view of Foley teach claim 1, 10.
Ranaldi further teaches:
deploying multiple large language models for the generative task to generate multiple outputs based on the target input; (Figure 1, Utilizing multiple LLMs (Stanford Alpaca, x-CrossAlpaca) to answer question and providing a response based on the input.)
Ranaldi does not specifically teach:
evaluating a similarity between each output of the multiple outputs and the reference output to generate multiple evaluation scores;
Foley, however, teaches:
evaluating a similarity between each output of the multiple outputs and the reference output to generate multiple evaluation scores; (P0020, An embodiment uses one or more training prompt responses, or combinations of training prompts and their corresponding responses, to train an attribution model to attribute a fine-tuned model to a source foundation model. In some embodiments, the attribution model is a classifier model that outputs a classification of a fine-tuned model to a source foundation model, and a confidence score in the classification.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize an evaluation score between responses of large language models. It would have been obvious to combine the references because a fine-tuned model may generate false information and generating an evaluation score enables the fine-tuned model to link or attribute to the foundation model. (Foley, P0015)
Ranaldi in view of Foley does not specifically teach:
ranking the multiple large language models based on the multiple evaluation scores for the multiple large language models; and
selecting a large language model corresponding to a highest evaluation score for the generative task in the target language.
Srinivasan, however, teaches:
ranking the multiple large language models based on the multiple evaluation scores for the multiple large language models; and (P0014, Providing the prompt to multiple large language models (LLMs), the multiple LLMs being trained on different datasets and having different knowledge and capabilities.; P0015, Receiving multiple responses from the multiple LLMs.; P0016, Determining a rank for each of the multiple responses, the rank indicating a level of confidence.; P0059, The ranking engine can also compare and assign a level of pair-wise similarity or dissimilarity between selected pairs of responses from the LLM array. This can be done by comparing the embedding vectors generated from the selected pair of responses to determine a level of agreement or disagreement between the responses and source LLMs. This can be done for all possible pairings of LLM responses. As will be appreciated, cosine similarity is a mathematical metric that measures how similar two vectors are in a multi-dimensional space.)
selecting a large language model corresponding to a highest evaluation score for the generative task in the target language. (P0018, Selecting the response having the best rank.; P0026, Various outputs of the LLMs are analyzed and ranked by an AI-based ranking and reasoning engine, with the highest ranked output indicating the most probable best fit or optimal solution for the prompt.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to rank large language models based on evaluation scores and select the highest evaluation score. It would have been obvious to combine the references because by efficiently ranking and/or comparing the similarity of the output to yield an output having the highest likelihood of being the ground truth, LLM hallucination incidents can be reduced. (Srinivasan P0029)
Claims 9 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ranaldi in view of Foley and further view of Deepak et al. (U.S. PG Pub No. 20250259011), hereinafter Deepak.
Regarding claim 9 and 20 Ranaldi in view of Foley teach claim 1, 17.
Ranaldi further teaches:
obtaining a first output generated from a target input by using the target large language model for the generative task at a first time; (Figure 1, Utilizing multiple LLMs (Stanford Alpaca, x-CrossAlpaca) to answer question and providing a response based on the input.)
Ranaldi does not specifically teach:
evaluating a first similarity between the first output and the reference output to generate a first evaluation score;
Foley, however, teaches:
evaluating a first similarity between the first output and the reference output to generate a first evaluation score; (P0020, An embodiment uses one or more training prompt responses, or combinations of training prompts and their corresponding responses, to train an attribution model to attribute a fine-tuned model to a source foundation model. In some embodiments, the attribution model is a classifier model that outputs a classification of a fine-tuned model to a source foundation model, and a confidence score in the classification.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize an evaluation score between responses of large language models. It would have been obvious to combine the references because a fine-tuned model may generate false information and generating an evaluation score enables the fine-tuned model to link or attribute to the foundation model. (Foley, P0015)
Ranaldi in view of Foley does not specifically teach:
obtaining a second output generated from the target input by using the target large language model for the generative task at a second time later than the first time; evaluating a second similarity between the second output and the reference output to generate a second evaluation score; and detecting a model drift of the target large language model based on a different between the first evaluation score and the second evaluation score.
Deepak, however, teaches:
obtaining a second output generated from the target input by using the target large language model for the generative task at a second time later than the first time; evaluating a second similarity between the second output and the reference output to generate a second evaluation score; and detecting a model drift of the target large language model based on a different between the first evaluation score and the second evaluation score. (P0043, In certain embodiments, a drift monitoring module may analyze LLM outputs over time, such as by using statistical methods and machine learning algorithms to identify patterns of uncertainty or errors in the model's responses (e.g., assessing the confidence level of the LLM in its responses, comparing LLM outputs with trusted external data sources for factual accuracy, or preforming longitudinal analysis to track the consistency of LLM responses over time in various domains, among others). Drift detection may then focus on identifying changes in the LLM's performance over time.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to obtain a second output at a later time to determine model drift of the target large language model. It would have been obvious to combine the references because LLM model can produce uncertainty or incorrectness over time with loss of knowledge and detecting the drift can allow for LLM updates to counteract the drift. (Srinivasan P0043, P0049)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Taori et al. (“Alpaca: A Strong, Replicable Instruction-Following Model”): Supervised finetuning of an LLM model utilizing the output of a second LLM model.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL WONSUK CHUNG whose telephone number is (571)272-1345. The examiner can normally be reached Monday - Friday (7am-4pm)[PT].
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571)272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL W CHUNG/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659