DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 09/02/2026 is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-14 and 17-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Independent claims 1, 18, and 19 recite a method, a computer-readable medium (CRM), and a system, respectively. These claims therefore invoke a statutory category (machine and process) in Step 1 of the Subject Matter Eligibility Test.
Step 2A, Prong One: Independent claims 1, 18, and 19, under their broadest reasonable interpretation, recite creating a prompt for input to a large language model (LLM), input the prompt to an LLM to generate an out, and evaluating that output by comparing it to an expected output. This is an abstract idea in the form of certain methods of organizing human activity (i.e. mental processes such as observation, evaluation, judgement, and opinion). The steps of accessing data (i.e. an LLM, a context, a question, an answer), processing the data, inputting the data into a system, and analyzing the outputted data could be performed by a human using pen and paper or by purely mental reasoning, save for the recitation of generic computer components
Step 2A, Prong Two: The claims do not integrate the judicial exception into a practical application. The recitation of “a large language model (LLM)” are generic instructions to perform the abstract idea on/using a computer and do not impose a meaningful limit on the judicial exception. The LLM is recited at such a high-level of generality and is merely used as a tool to perform the abstract idea faster and more efficiently. The data receiving, a pre-solution activity, and outputting, a post-solution activity, steps required to perform the method do not add a meaningful limitation. Mere data gathering, analysis, and output do not provide an inventive concept. There is no improvement to the functioning of the model itself, evaluation of model outputs, the functioning of computers, or to any other technology or technical field.
Step 2B: The claims do not include any additional elements that amount to significantly more than the judicial exception. The only additional elements beyond the abstract idea is the LLM, which performs generic computational functions such as receiving, analyzing, and outputting data. Such elements are well-understood, routing, and conventional within the field.
Accordingly, claims 1, 18, and 19 are directed to an abstract idea and do not include significantly more than the abstract idea itself.
With respect to claim 2, the claim relates to including instructions in the prompt to respond to the prompt with an indication of if the context answers the question. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 3, the claim relates to the context comprising a title and a text in a second language, and the question being related to a topic associated with the title. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 4, the claim relates to the question being in a second language, the context being in a specified language, and the prompt including an instruction to respond in a third language. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 5, the claim relates to the prompt including instructions to indicate if the context answers the question based on a comparison of the context and a ground truth. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 6, the claim relates to the indication of whether a language of the output matches a ground truth language. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 7, the claim relates to the prompt including instructions to respond in a specified language, determining a language of the output, and labelling the output as inaccurate. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 8, the claim relates to the prompt including instructions to respond in the same language, determining a language of the output, and labelling the output as inaccurate. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 9, the claim relates to evaluating the output using recall-oriented understudy for gisting evaluation metrics. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 10, the claim relates to the question being a first question of a plurality in a second language and the context being a first context of a plurality in the second language. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 11, the claim relates to the question being a first question of a plurality in different languages and the context being a first context of a plurality in different languages. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 12, the claim relates to the plurality of questions comprising a same question translated, and the plurality of contexts comprising a same context translated. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 13, the claim relates to the prompt including a multilingual consistency task comprising answering deterministic questions. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 14, the claim relates to generating a score for the ability of the LLM to respond in the specified language, a score for the ability of the LLM to reason correctly, and a score for the ability of the LLM to response consistently. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. The only additional element is “the LLM”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present.
With respect to claim 17, the claim relates to calculating a token score for the LLM based on vocabulary size and language distribution. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 19, the claim relates to various limitations similar to those of dependent claims 2-8, and thus is similarly rejected under the same rationale.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, 9-16, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kahana et al. ("EVALUATION METHODOLOGY FOR LARGE LANGUAGE MODELS FOR MULTILINGUAL DOCUMENT QUESTION AND ANSWER", 02/01/2024), hereinafter referred to as Kahana, in view of Lai et al. ("LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback", 08/16/2024), hereinafter referred to as Lai.
Regarding claim 1, Kahana discloses a method for evaluating one or more large language models for multilingual performance ("Evaluating multilingual model performance is also an area of active research since most of the popular model performance benchmarks are also predominately for the English language," Kahana Section 1 pg. 1), comprising: accessing a large language model (LLM), a context, a question, and an answer ("The flow of the tests starts with randomly selecting a subset from these datasets, where each sample has a context, a question, and one or more answers," Kahana Section 2.1 pg. 2);
generating a prompt, the prompt comprising a first instruction to generate a specified response in a specified language based on the context and the question ("The second part of the flow involves querying the LLM to answer the question. We use simple prompts like ‘Here is context for the question: ’ and ‘Please, given the context, answer the following question:’, followed by the context and question respectively. We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2);
generating an output by providing the prompt as input into the LLM, wherein at least one of the question, first instruction, context, and answer are in a second language different from the specified language ("The second part of the flow involves querying the LLM to answer the question. We use simple prompts like ‘Here is context for the question: ’ and ‘Please, given the context, answer the following question:’, followed by the context and question respectively. We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2 and "Lastly, we ask the LLM to verify if the generated answer is correct," Kahana Section 2.1 pg. 2);
evaluating, based on a comparison of a contents of the output to the answer ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2), at least one of: (a) an ability of the LLM to respond in the specified language for a multilingual task ("The first experiment involves the XQuAD dataset, where the context-question-answers samples are originally in English and about the North American culture, industry, etc. The translated versions include translated contexts, as well as the questions and true answers, which means we do not need to translate anything in the process. This sets the benchmark for the performance. In addition, we run the English tests 10 times and compute the mean accuracy and the standard deviation across all 10 tests, to ensure that the proposed testing framework is consistent," Kahana Section 2.3 pg. 3 and Kahana Table 1 pg. 4 shows results);
or (c) an ability of the LLM to respond consistently for a plurality of languages ("The first experiment involves the XQuAD dataset, where the context-question-answers samples are originally in English and about the North American culture, industry, etc. The translated versions include translated contexts, as well as the questions and true answers, which means we do not need to translate anything in the process. This sets the benchmark for the performance. In addition, we run the English tests 10 times and compute the mean accuracy and the standard deviation across all 10 tests, to ensure that the proposed testing framework is consistent," Kahana Section 2.3 pg. 3 and Kahana Table 1 pg. 4 shows results);
wherein the method is performed by at least one device including a hardware processor ("In this work, we present results when using GPT-4-32K and GPT-3.5-Turbo models," Kahana Section 2.1 pg. 2, large language models inherently require hardware processors).
However, Kahana fails to disclose (b) an ability of the LLM to reason correctly for a multilingual task. Lai teaches a scalable multilingual capability for large language models (LLMs) with cross-lingual feedback.
Lai teaches (b) an ability of the LLM to reason correctly for a multilingual task (Lai Section 4.1 pg. 6 shows the reasoning task).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Lai’s teaching of evaluating a model on its reasoning abilities. This would ensure that the model being evaluated does not suffer from any performance gaps or deficits. It would also ensure that these deficits do not propagate to the other languages that the model possesses. Evaluating a model on tasks, such as reasoning, is a very common practice within the art of language models.
Regarding claim 2, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the prompt includes an instruction following task comprising: instructions, in a first language, to respond to the prompt with an indication of whether the context answers the question ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2).
Regarding claim 3, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the context comprises a title in the second language and a text in the second language (Kahana Section 2.2 pg. 2 shows datasets XQuAD and SQuAD, which include titles, paragraphs (context), questions, and answers);
and the question is in the second language and is related to a topic associated with the title (Kahana Section 2.2 pg. 2 shows datasets XQuAD and SQuAD, which include titles, paragraphs (context), questions, and answers).
Regarding claim 4, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the question is in the second language, the context is in the specified language, and the prompt includes a second instruction to respond in a third language ("We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2 and "In the evaluation flow, we add a step to translate the context, the question and the true answers, but using the system prompt ‘Please translate the following questions to LANGUAGE:’, as an example for translation of the questions to language LANGUAGE," Kahana Section 2.2 pg. 3).
Regarding claim 5, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the prompt includes a second instruction to indicate whether the context answers the question based on correlation between the context and a ground truth ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2).
Regarding claim 6, Kahana, in view of Lai, discloses all of the limitations of claim 5. Kahana further discloses wherein: the indication comprises an indication whether a language of the output matches a language of a ground truth ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2).
Regarding claim 9, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses comprising: evaluating the output by comparing the output to a ground-truth using recall-oriented understudy for gisting evaluation (ROUGE) metrics ("For the XL Sum and Self-Instruct* benchmark, we report the multilingual ROUGE-1 score," Lai Section 4.5 pg. 6).
Regarding claim 10, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the question is a first question of a plurality of questions in the second language and the context is a first context of a plurality of contexts in the second language ("This dataset includes 12 languages and 1,190 context-question-answers samples. The topics are generic and vary between many fields of interest. An important aspect of this dataset is that each question has been translated from English to all these languages (Arabic, German, Greek, Spanish, Hindi, Russian, Thai, Turkish, Vietnamese, Chinese, Romanian) with a high level of confidence in the translation, as it is translated by human translators," Kahana Section 2.2 pg. 2).
Regarding claim 11, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the question is a first question of a plurality of questions in different languages and the context is a first context of a plurality of contexts in different languages ("This dataset includes 12 languages and 1,190 context-question-answers samples. The topics are generic and vary between many fields of interest. An important aspect of this dataset is that each question has been translated from English to all these languages (Arabic, German, Greek, Spanish, Hindi, Russian, Thai, Turkish, Vietnamese, Chinese, Romanian) with a high level of confidence in the translation, as it is translated by human translators," Kahana Section 2.2 pg. 2).
Regarding claim 12, Kahana, in view of Lai, discloses all of the limitations of claim 11. Kahana further discloses wherein: the plurality of questions in different languages comprises a same question translated into different languages ("This dataset includes 12 languages and 1,190 context-question-answers samples. The topics are generic and vary between many fields of interest. An important aspect of this dataset is that each question has been translated from English to all these languages (Arabic, German, Greek, Spanish, Hindi, Russian, Thai, Turkish, Vietnamese, Chinese, Romanian) with a high level of confidence in the translation, as it is translated by human translators," Kahana Section 2.2 pg. 2);
and the plurality of contexts in different languages comprises one or more same contexts translated into different languages ("This dataset includes 12 languages and 1,190 context-question-answers samples. The topics are generic and vary between many fields of interest. An important aspect of this dataset is that each question has been translated from English to all these languages (Arabic, German, Greek, Spanish, Hindi, Russian, Thai, Turkish, Vietnamese, Chinese, Romanian) with a high level of confidence in the translation, as it is translated by human translators," Kahana Section 2.2 pg. 2).
Regarding claim 13, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the prompt includes a multilingual consistency task prompt comprising: answering a plurality of deterministic questions in a plurality of languages ("The second part of the flow involves querying the LLM to answer the question. We use simple prompts like ‘Here is context for the question: ’ and ‘Please, given the context, answer the following question:’, followed by the context and question respectively. We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2) and an instruction to indicate, in the output, a plurality of yes or no answers for the plurality of deterministic questions based solely on the context ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2);
and comparing the plurality of yes or no answers to a ground truth to determine an accuracy of the output ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2).
Regarding claim 14, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses comprising: generating a total score as weighted sum of a score for the ability of the LLM to respond in the specified language for the multilingual task (Kahana Table 1 pg. 4 shows accuracy scores), and a score for the ability of the LLM to respond consistently for the plurality of languages (Kahana Table 1 pg. 4 shows accuracy scores).
However, Kahana fails to disclose a score for the ability of the LLM to reason correctly for the multilingual task.
Lai teaches a score for the ability of the LLM to reason correctly for the multilingual task ("For the XCOPA and PAWS-X benchmarks, we utilize the accuracy score for evaluation," Lai Section 4.5 pg. 6).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Lai’s teaching of utilizing a metric score for how well an LLM is able to reason. Using scores to define metrics is a well-known technique within the art of large language model training, optimization, and evaluation, and it is an obvious progression here.
Regarding claim 15, Kahana, in view of Lai, discloses all of the limitations of claim 14. Kahana further discloses comprising: the standard data set comprising one or more pairs of question, context, answer triplets, the one or more pairs comprising a first triplet in the specified language paired with a second triplet in the second language (Kahana Section 2.2 pg. 2 XQuAD dataset).
However, Kahana fails to disclose training the LLM using a standard data set to generate a trained LLM; generating a baseline score for the LLM; computing a feedback score based on comparing a second score for the trained LLM to the baseline score for the LLM; and computing a final score based on the feedback score and the total score.
Lai discloses training the LLM using a standard data set to generate a trained LLM (Lai Section 3.4 pg. 5 shows the two-step training process: supervised fine-tuning and aligning LLMs with human feedback);
generating a baseline score for the LLM (Lai Section 4.2 pg. 6 details the baselines for the comparison);
computing a feedback score based on comparing a second score for the trained LLM to the baseline score for the LLM (Lai Table 3 shows the scores for the models);
and computing a final score based on the feedback score and the total score (Lai Table 3 shows the scores for the models).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Lai’s teaching of training a multilingual LLM and optimizing it. Training a model with a dataset, using baseline scoring to represent an initial point, and comparing further scores to the baseline to see how well the model has improved are all well-known techniques within the art of model training.
Regarding claim 16, Kahana, in view of Lai, discloses all of the limitations of claim 15. However, Kahana fails to disclose comprising: training the LLM using the standard data set to generate a trained LLM using at least one training method selected from: parameter efficient fine-tuning, vocabulary extension tuning, and instructional tuning.
Lai teaches comprising: training the LLM using the standard data set to generate a trained LLM using at least one training method selected from: parameter efficient fine-tuning, vocabulary extension tuning, and instructional tuning (Lai Section 3.4 pg. 5 shows the two-step training process: supervised fine-tuning and aligning LLMs with human feedback).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Lai’s teaching of using fine-tuning to train an LLM. Parameter fine-tuning is a well-known technique within the art of large language models and their respective training and would have been an obvious inclusion.
As to claim 18, computer-readable medium (CRM) claim 18 and method claim 1 are related as method and CRM of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 18 is similarly rejected under the same rationale as applied above with respect to the method claim.
As to claim 20, system claim 20 and method claim 1 are related as method and system of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 20 is similarly rejected under the same rationale as applied above with respect to the method claim.
Claim(s) 7-8 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kahana, in view of Lai, and further in view of Marchisio et al. ("Understanding and Mitigating Language Confusion in LLMs", 11/16/2024), hereinafter referred to as Marchisio.
Regarding claim 7, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the prompt includes a second instruction to respond to the prompt in the specified language ("In the evaluation flow, we add a step to translate the context, the question and the true answers, but using the system prompt ‘Please translate the following questions to LANGUAGE:’, as an example for translation of the questions to language LANGUAGE," Kahana Section 2.2 pg. 3).
However, Kahana fails to disclose the method comprising: determining a language of the output; and labelling the output as inaccurate responsive to determining that the output is not in the specified language. Marchisio teaches a method for understanding and mitigating language confusion in LLMs.
Marchisio teaches the method comprising: determining a language of the output (Marchisio Section 2.2 pg. 2 checks the response against the desired language both by line-level and word-level);
and labelling the output as inaccurate responsive to determining that the output is not in the specified language (Marchisio Section 2.2 pg. 3 shows the line-level pass rate and the word-level pass rate).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Marchisio’s teaching of determining the output language and labelling the output as incorrect. This is a well-known technique within the art of multilingual large language models, and it’s an obvious progression to check and conclude that the model is responding in the correct and desired language.
Regarding claim 8, Kahana, in view of Lai, discloses all of the limitations of claim 1. Kahana further discloses wherein: the prompt includes a second instruction to respond to the prompt in a same language as the question ("The second part of the flow involves querying the LLM to answer the question. We use simple prompts like ‘Here is context for the question: ’ and ‘Please, given the context, answer the following question:’, followed by the context and question respectively. We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2).
However, Kahana fails to disclose the method comprising: determining a language of the output; and labelling the output as inaccurate responsive to determining that the output is not in the same language as the question.
Marchisio teaches the method comprising: determining a language of the output (Marchisio Section 2.2 pg. 2 checks the response against the desired language both by line-level and word-level);
and labelling the output as inaccurate responsive to determining that the output is not in the same language as the question (Marchisio Section 2.2 pg. 3 shows the line-level pass rate and the word-level pass rate).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Marchisio’s teaching of determining the output language and labelling the output as incorrect. This is a well-known technique within the art of multilingual large language models, and it’s an obvious progression to check and conclude that the model is responding in the correct and desired language.
Regarding claim 19, Kahana, in view of Lai, discloses all of the limitations of claim 18. Kahana further discloses wherein: the prompt includes an instruction following task comprising instructions, in a first language, to respond to the prompt with an indication of whether the context answers the question ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2);
the context comprises a title in the second language and a text in the second language (Kahana Section 2.2 pg. 2 shows datasets XQuAD and SQuAD, which include titles, paragraphs (context), questions, and answers);
the question is in the second language and is related to a topic associated with the title (Kahana Section 2.2 pg. 2 shows datasets XQuAD and SQuAD, which include titles, paragraphs (context), questions, and answers);
the question is in the second language, the context is in the specified language, and the prompt includes a second instruction to respond in a third language ("We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2 and "In the evaluation flow, we add a step to translate the context, the question and the true answers, but using the system prompt ‘Please translate the following questions to LANGUAGE:’, as an example for translation of the questions to language LANGUAGE," Kahana Section 2.2 pg. 3);
the prompt includes a second instruction to indicate whether the context answers the question based on correlation between the context and a ground truth ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2);
the indication comprises an indication whether a language of the output matches a language of a ground truth ("Lastly, we ask the LLM to verify if the generated answer is correct. To do this, we use a system prompt like ‘Is your answer correct? The ‘true’ and correct answer is: X. Your answer is: Y. Reply Only ‘TRUE’ IF ‘YES’ OR ‘FALSE’ IF ‘NO”. We concatenate the ‘true’ answer from the dataset in the place of X and the inferred answer by the LLM in the place of Y. Then we look for the word ‘True’ (checking for upper/lower case as well), as well as ‘Yes’ which interestingly sometimes gets returned instead of ‘True’ (and ‘No’ instead of ‘False’). In the case of multiple answers, we iterate over all the answers and if a correct answer has been supplied by the LLM, we advance the counter and continue to the next question," Kahana Section 2.1 pg. 2);
the prompt includes a second instruction to respond to the prompt in the specified language ("In the evaluation flow, we add a step to translate the context, the question and the true answers, but using the system prompt ‘Please translate the following questions to LANGUAGE:’, as an example for translation of the questions to language LANGUAGE," Kahana Section 2.2 pg. 3);
and the prompt includes a second instruction to respond to the prompt in a same language as the question ("The second part of the flow involves querying the LLM to answer the question. We use simple prompts like ‘Here is context for the question: ’ and ‘Please, given the context, answer the following question:’, followed by the context and question respectively. We emphasize that even though the system prompts are in English, the contexts and questions are injected in any foreign language," Kahana Section 2.1 pg. 2).
However Kahana fails to disclose the operations further comprising: determining a language of the output; labelling the output as inaccurate responsive to determining that the output is not in the specified language; determining a language of the output; and labelling the output as inaccurate responsive to determining that the output is not in the same language as the question.
Marchisio teaches the operations further comprising: determining a language of the output (Marchisio Section 2.2 pg. 2 checks the response against the desired langauge both by line-level and word-level);
labelling the output as inaccurate responsive to determining that the output is not in the specified language (Marchisio Section 2.2 pg. 3 shows the line-level pass rate and the word-level pass rate);
determining a language of the output (Marchisio Section 2.2 pg. 2 checks the response against the desired langauge both by line-level and word-level);
and labelling the output as inaccurate responsive to determining that the output is not in the same language as the question (Marchisio Section 2.2 pg. 3 shows the line-level pass rate and the word-level pass rate).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Marchisio’s teaching of determining the output language and labelling the output as incorrect. This is a well-known technique within the art of multilingual large language models, and it’s an obvious progression to check and conclude that the model is responding in the correct and desired language.
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kahana, in view of Lai, and further in view of Chelombitko et al. ("QTOK: A COMPREHENSIVE FRAMEWORK FOR EVALUATING MULTILINGUAL TOKENIZER QUALITY IN LARGE LANGUAGE MODELS", 10/16/2024), hereinafter referred to as Chelombitko.
Regarding claim 17, Kahana, in view of Lai, discloses all of the limitations of claim 14. However, Kahana fails to disclose comprising: calculating a token score for the LLM based on a size of a vocabulary of the LLM and on a language distribution of tokens in the vocabulary, wherein the total score is based at least in part on a weighting of the token score.
Chelombitko teaches a method for evaluating multilingual tokenizer quality in LLMs.
Chelombitko teaches calculating a token score for the LLM (Chelombitko Sections 2.8 and 2.9 pg. 16-17 and Table 6) based on a size of a vocabulary (Chelombitko Section 2.7.1 pg. 15-16) of the LLM and on a language distribution of tokens in the vocabulary (Chelombitko Section 2.6 pg. 12-14), wherein the total score is based at least in part on a weighting of the token score (Chelombitko Sections 2.8 and 2.9 pg. 16-17 and Table 6).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kahana’s disclosure of evaluating a large language model for multilingual performance by including Chelombitko’s teaching of utilizing token scores to ensure that a model’s vocabulary is sufficient. Ensuring that the vocabulary known to a model is sufficient by calculating the vocabulary size and language distribution ensures that the model knows enough to answer most potential queries. This would ensure that the model is well-trained and equipped. This is an obvious progression.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Huang et al., “BenchMAX:AComprehensiveMultilingual Evaluation Suite for Large Language Models”, 02/11/2025
Liu et al., “Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models”, 06/20/2024
Singh et al., “INDICGENBENCH: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages”, 08/16/2024
US Patent Application Publication No. 2026/0064994
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM MICHAEL WEAVER whose telephone number is (571)272-7062. The examiner can normally be reached Monday-Friday, 8AM-5PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ADAM MICHAEL WEAVER/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658