Prosecution Insights
Last updated: October 02, 2026
Application No. 18/627,828

PERFORMANCE EVALUATION OF GENERATIVE QUESTION-ANSWERING SYSTEMS

Non-Final OA §102§103
Filed
Apr 05, 2024
Examiner
PULLIAS, JESSE SCOTT
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
3 (Non-Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
885 granted / 1072 resolved
+20.6% vs TC avg
Moderate +12% lift
Without
With
+12.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
31 currently pending
Career history
1110
Total Applications
across all art units

Statute-Specific Performance

§101
15.5%
-24.5% vs TC avg
§103
53.1%
+13.1% vs TC avg
§102
19.8%
-20.2% vs TC avg
§112
4.6%
-35.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1072 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/16/26 has been entered. This office action is in response to correspondence 06/16/26 regarding application 18/627,828, in which claims 1-4, 6-9, 12-15, and 17-21 were amended, claim 5 was cancelled, and new claim 22 was added. Claims 1-4, 6-10, and 12-22 are pending in the application and have been considered. Response to Arguments The examiner agrees with Applicant on page 8 that the amendments to claims 1-4, 6-9, 12-15, and 17-21 and addition of new claim 22 do not introduce new matter. The examiner has reconsidered the previous finding that Zhang does not disclose a regression model configured to output an evaluation for an input question-answer pair. Zhang clearly discloses that classification model may be a regression model that generates a score ([0072]). The new 35 U.S.C. 102(a)(2) rejections based on Zhang are made based in part on this finding. The examiner has reviewed Zhang in depth to prepare this office action, and believes it may be helpful to discuss the three main types of machine learning models in Zhang, what role(s) they serve, and how they correspond to the machine learning models claimed according to the examiner’s analysis of the reference. “machine-learned model” of model serving system – this ML model receives a query from the user and generates an answer set of output tokens, [0032], based on knowledge sources, [0039]. This corresponds to the claimed “first artificial intelligence (AI) model”. This can be, but is not necessarily a language model or LLM. “classification model” – this is a machine learning model that receives a pair of question answer pair and provides a score indicating how well the output answers the input. [0065], [0039], [0072], [0143]. The classification model may be a second LLM, but for the purposes of mapping Zhang to claim 1, is best thought of as a regression model that provides a score [0072]. This corresponds to the claimed “evaluation model”. “machine-learned language model” – this is an LLM that receives a question/answer pair along with the provided score from classification model and provides an evaluation label based on how well the score scored the answer to the question. See [0070-0072], [0143], and [0034]. This corresponds to the claimed “first large language model (LLM) different from the first AI model”. Because the claim language does not require a “numerical” score from the LLM, and Applicant’s specification describes “evaluation score” as containing “a value indicative of the quality of the answer to the question” (see Applicant’s Abstract), the “evaluation label” provided by the LLM of Zhang corresponds to the claimed “first evaluation score”. Zhang contains numerous embodiments and overlapping terminology that can mean different things. For purposes of this office action, the examiner is considering the embodiments of Zhang that operate as follows, based on the passages cited above: a. user provides a question or query b. “machine-learned model” of model serving system provides an answer to user’s question or query c. “classification model” generates a score for the Q/A pair d. “machine-learned language model” is prompted with the scored Q/A pair and provides an evaluation label e. evaluation label is used to retrain “classification model” f. user provides another question or query g. “machine-learned model” of model serving system provides another answer to user’s question or query h. “classification model” generates another score for the Q/A pair Applicant’s arguments on pages 9-11 regarding the 35 U.S.C. 103 rejections of claims 1, 3-9, 12, 15, and 17 are essentially that the claims as amended distinguish from Zhang because the claim language requires using a second AI model different from the first AI model to provide the evaluation score. However, as seen above, Zhang uses a second, different model, namely “machine-learned language model” to provide the evaluation score, different from the “machine learned model” used by model serving system 150 to answer the user’s queries, see [0031]. “[0031] The model serving system 150 receives requests from the online concierge system 140 to perform tasks using machine-learned models. The tasks include, but are not limited to, natural language processing (NLP) tasks, audio processing tasks, image processing tasks, video processing tasks, and the like. In one or more embodiments, the machine-learned models deployed by the model serving system 150 are models configured to perform one or more NLP tasks. The NLP tasks include, but are not limited to, text generation, query processing, machine translation, chatbots, and the like.” Applicant’s arguments on page 11-13 are similar to those addressed above, and are not persuasive for similar reasons. After a search, the examiner agrees with Applicant that new claim 22 contains allowable subject matter, for the reasons below. Claim Objections In claim 17, line 6, should “receive” be “receiving”? Claim Interpretation Claims 17-21 are directed to a ”computer-readable storage medium”. The specification explicitly states that this term does not encompass communication media, propagating signals, and signals per se (paragraph [0146], page 45). Claims 17-20 are interpreted as only encompassing eligible storage medium types. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 3, 4, 6, 7, 9, 12, and 17 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Zhang et al. (US 20250200356). Consider claim 1, Zhang discloses a system for evaluating the performance of a question-answering model (machine learning model that evaluates results of a classification model by generating a score evaluating an input-output pair of a question-answering task, [0070-0072], [0039]), the system comprising: a processor (computer processor, [0171]); and a memory device that stores program code structured to cause the processor to (a non-transitory, tangible, computer-readable medium stores instructions executed by a processor, [0171]-[0172]; “memory” is inherent in this computer architecture): obtain a prior question-answer pair comprising a question and an associated answer associated with a first artificial intelligence (AI) model (online concierge system 140 performs question-answering using “machine-learned model” of model serving system, i.e. “a first artificial intelligence model”, based on knowledge sources, [0039]); receive a first evaluation score for the prior question-answer pair from a first large language model (LLM) different from the first AI model (online system provides a batch of evaluation request prompts to machine-learned “language” model, which is different from “machine-learned model” of model service system, with each input-output pair, i.e. question-answer pair, [0143], which is an LLM, [0034], and generates an evaluation label, [0143], which may be a generated score, [0072]); train an evaluation model based on features that comprise information from the prior question-answer pair and labels based on the first evaluation score (online system generates training samples of labeled historical input-output, i.e. scored question-answer pairs for the classification model and trains classification model using a supervised training technique, [0143], classification model trained by comparing the classification score to the LLM label score with a loss function, [0067], [0072]), the evaluation model comprising a regression model configured to output an evaluation for an input question-answer pair (each machine learning model may be a linear regression model trained to use the set of parameters to transform inputs into outputs, [0065]; in the context of QA, this would transform a QA pair into a scored QA pair, [0039], [0072]); obtain a current question-answer pair (online concierge system receives an input query provided by a user, [0145], and provides an answer in the form of images, descriptions, as an input-output pair, [0146]; although this context refers to the input-output pair as a query-item pair in this context, for a question-answering task, it would be a question-answer pair, [0039]); and generate a current evaluation score for the current question-answer pair by applying the current question-answer pair to the evaluation model (the next query-item pair, which in the context of question-answering is a question-answer pair, see above, is supplied to classification model for generating an evaluation score, [0071-0072]). Consider claim 12, Zhang discloses a method for evaluating the performance of a question-answering model (evaluating results of a classification model by generating a score evaluating an input-output pair of a question-answering task, [0070-0072], [0039]), comprising: obtaining a prior question-answer pair comprising a question and an associated answer generated by a first artificial intelligence (AI) model (online concierge system 140 performs question-answering using “machine-learned model” of model serving system, i.e. “a first artificial intelligence model”, based on knowledge sources, [0039]); receiving a first evaluation score for the prior question-answer pair from a first large language model (LLM) different from the first AI model (online system provides a batch of evaluation request prompts to machine-learned “language” model, which is different from “machine-learned model” of model service system, with each input-output pair, i.e. question-answer pair, [0143], which is an LLM, [0034], and generates an evaluation label, [0143], which may be a generated score, [0072]); training an evaluation model based on features that comprise information from the prior question-answer pair and labels based on the first evaluation score (online system generates training samples of labeled historical input-output, i.e. scored question-answer pairs for the classification model and trains classification model using a supervised training technique, [0143], classification model trained by comparing the classification score to the LLM label score with a loss function, [0067], [0072]), the evaluation model comprising a regression model configured to output an evaluation for an input question-answer pair (each machine learning model may be a linear regression model trained to use the set of parameters to transform inputs into outputs, [0065]; in the context of QA, this would transform a QA pair into a scored QA pair, [0039], [0072]); obtaining a current question-answer pair (online concierge system receives an input query provided by a user, [0145], and provides an answer in the form of images, descriptions, as an input-output pair, [0146]; although this context refers to the input-output pair as a query-item pair in this context, for a question-answering task, it would be a question-answer pair, [0039]); and applying the evaluation model to the current question-answer pair, resulting in a second evaluation score for the current question-answer pair (the next query-item pair, which in the context of question-answering is a question-answer pair, see above, is supplied to classification model for generating an evaluation score, [0071-0072]). Consider claim 17, Zhang discloses a computer-readable storage medium having computer program code recorded thereon that when executed by at least one processor causes the at least one processor to perform a method (non-transitory computer readable medium storing code executed by a processor, [0171], [0172]) comprising: obtaining a prior question-answer pair comprising a question and an associated answer generated by a first artificial intelligence (AI) model (online concierge system 140 performs question-answering using “machine-learned model” of model serving system, i.e. “a first artificial intelligence model”, based on knowledge sources, [0039]); receiving a first evaluation score for the prior question-answer pair from a first large language model (LLM) different from the first AI model (online system provides a batch of evaluation request prompts to machine-learned “language” model, which is different from “machine-learned model” of model service system, with each input-output pair, i.e. question-answer pair, [0143], which is an LLM, [0034], and generates an evaluation label, [0143], which may be a generated score, [0072]); training an evaluation model based on features that comprise information from the prior question-answer pair and labels based on the first evaluation score (online system generates training samples of labeled historical input-output, i.e. scored question-answer pairs for the classification model and trains classification model using a supervised training technique, [0143], classification model trained by comparing the classification score to the LLM label score with a loss function, [0067], [0072]), the evaluation model comprising a regression model configured to output an evaluation for an input question-answer pair (each machine learning model may be a linear regression model trained to use the set of parameters to transform inputs into outputs, [0065]; in the context of QA, this would transform a QA pair into a scored QA pair, [0039], [0072]); obtaining a current question-answer pair (online concierge system receives an input query provided by a user, [0145], and provides an answer in the form of images, descriptions, as an input-output pair, [0146]; although this context refers to the input-output pair as a query-item pair in this context, for a question-answering task, it would be a question-answer pair, [0039]); and applying the evaluation model to the current question-answer pair, resulting in a second evaluation score for the current question-answer pair (the next query-item pair, which in the context of question-answering is a question-answer pair, see above, is supplied to classification model for generating an evaluation score, [0071-0072]). Consider claim 3, Zhang discloses the current question-answer pair comprises a current question and a current answer, the current question provided to the first AI model and the current answer returned by the first AI model (online concierge system 140 performs question-answering using “machine-learned model” of model serving system, i.e. “a first artificial intelligence model”, based on knowledge sources, [0039]). Consider claim 4, Zhang discloses the second evaluation score is indicative of a quality of the current answer to the current question (classification model generates an evaluation score for an input/output pair, [0071-0072]; for a question-answering task, this would be indicative of the quality of the answer to the question in the question-answer pair, [0039]). Consider claim 6, Zhang discloses the first AI model is a different type of model than a type of the first LLM (online concierge system 140 performs question-answering using “machine-learned model” of model serving system, i.e. “a first artificial intelligence model”, based on knowledge sources, [0039], which is different from machine-learned “language” model, an LLM, [0034]). Consider claim 7, Zhang discloses the program code is structured to cause the processor to receive the evaluation score for the prior question-answer pair by: generating a prompt that includes the prior question and prior answer of the prior question-answer pair to the first LLM (instruction prompt instructs machine learned language model to evaluate the quality of the result relative to the query, [0075]; for a question-answering task, this would be quality of the answer to the question in the question-answer pair, [0039]). Consider claim 9, Zhang discloses the program code is further structured to cause the processor to: provide, to a user interface, a rating based on the second evaluation score and the current answer (the evaluation label and reasoning are output, [0096], in this instance, a rating of exact match, [0102], displayed to the user, [0075]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 13, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20250200356) in view of Jiang et al. (“LLM-BLENDER: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion” arXiv:2306.02561v3 [cs.CL] 30 Jun 2023). Consider claim 2, Zhang discloses the program code is further structured to cause the processor to: provide each prior question-answer pair to the first LLM and a second LLM causing the first LLM to return a first evaluation score and the second LLM to return a third evaluation score for the prior question-answer pair (“the language models are LLMs”, [0034], [0064]; receiving scores from both machine-learned language model in response to instruction prompt, [0071], and when classification model is a second machine-learned language model different from the first machine-learned language model, a score from a second different LLM, [0072]); and train the evaluation model based on the scores (online system generates training samples of labeled historical input-output, i.e. scored question-answer pairs for the classification model and trains classification model using a supervised training technique, [0143], classification model trained by comparing the classification score to the label score with a loss function, [0067], [0072]). Zhang does not specifically mention a combination of the evaluation scores. Jiang discloses a combination of the evaluation scores (aggregation of scores, Section 3.3, pages 5-6). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by utilizing a combination of the first and third evaluation scores in order to address the known strengths and weaknesses of different LLMs, as suggested by Jiang (Section 1, page 1). Doing so would have led to predictable results of alleviating biases, errors, and uncertainties in individual LLMs, as suggested by Jiang (Section 1, page 2). The references cited are analogous art in the same field of natural language processing. Consider claim 13, Zhang discloses providing each prior question-answer pair to the first LLM and a second LLM causing the first LLM to return a first evaluation score and the second LLM to return a third evaluation score for the prior question-answer pair (“the language models are LLMs”, [0034], [0064]; receiving scores from both machine-learned language model in response to instruction prompt, [0071], and when classification model is a second machine-learned language model different from the first machine-learned language model, a score from a second different LLM, [0072]); and training the evaluation model based on the scores (online system generates training samples of labeled historical input-output, i.e. scored question-answer pairs for the classification model and trains classification model using a supervised training technique, [0143], classification model trained by comparing the classification score to the label score with a loss function, [0067], [0072]). Zhang does not specifically mention a combination of the evaluation scores. Jiang discloses a combination of the evaluation scores (aggregation of scores, Section 3.3, pages 5-6). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by utilizing a combination of the first and third evaluation scores for reasons similar to those for claim 2. Consider claim 18, Zhang discloses providing each prior question-answer pair to the first LLM and a second LLM causing the first LLM to return a first evaluation score and the second LLM to return a third evaluation score for the prior question-answer pair (“the language models are LLMs”, [0034], [0064]; receiving scores from both machine-learned language model in response to instruction prompt, [0071], and when classification model is a second machine-learned language model different from the first machine-learned language model, a score from a second different LLM, [0072]); and training the evaluation model based on the scores (online system generates training samples of labeled historical input-output, i.e. scored question-answer pairs for the classification model and trains classification model using a supervised training technique, [0143], classification model trained by comparing the classification score to the label score with a loss function, [0067], [0072]). Zhang does not specifically mention a combination of the evaluation scores. Jiang discloses a combination of the evaluation scores (aggregation of scores, Section 3.3, pages 5-6). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by utilizing a combination of the first and third evaluation scores for reasons similar to those for claim 2. Claims 10, 16, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20250200356) in view of Chopra et al. (US 10445745). Consider claim 10, Zhang does not, but Chopra discloses the program code is further structured to cause the processor (program code, Col 15 lines 54-67) to: obtain a chat history that identifies a conversation between a user and a question-answering system (questions from chat a historical questions asked on the platform, Col 3 lines 3-9); and select the prior question-answer pair from the conversation based on a filtering criteria (highest scoring question-answer pairs, Col 13-14 lines 51-5). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by obtaining a chat history that identifies a conversation between a user and a question-answering system and selecting the prior question-answer pair from the conversations based on a filtering criteria in order to make question answering more precise and less cumbersome to use, as suggested by Chopra, (Col 1 lines 39-40), predictably addressing customer queries more efficiently, as suggested by Chopra, (Col 1 lines 11-12). The references cited are analogous art in the same field of natural language processing. Consider claim 16, Zhang does not, but Chopra discloses obtaining a chat history that identifies a conversation between a user and a question-answering system (questions from chat a historical questions asked on the platform, Col 3 lines 3-9); and selecting the prior question-answer pair from the conversations based on a filtering criteria (highest scoring question-answer pairs, Col 13-14 lines 51-5). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by obtaining a chat history that identifies a conversation between a user and a question-answering system and selecting the prior question-answer pairs from the conversations based on a filtering criteria for reasons similar to those for claim 10. Consider claim 21, Zhang does not, but Chopra discloses obtaining a chat history that identifies a conversation between a user and a question-answering system (questions from chat a historical questions asked on the platform, Col 3 lines 3-9); and selecting the prior question-answer pair from the conversations based on a filtering criteria (highest scoring question-answer pairs, Col 13-14 lines 51-5). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by obtaining a chat history that identifies a conversation between a user and a question-answering system and selecting the prior question-answer pairs from the conversations based on a filtering criteria for reasons similar to those for claim 10. Claims 14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20250200356) in view of Barbetta et al. (US 20150161106). Consider claim 14, Zhang does not, but Barbetta discloses the second evaluation score is indicative of a quality of a current question of the current question-answer pair to a current answer of the current question-answer pair (coercion analyzer generates a score that reflects ease of coercion of the answer to the type of question that the question seeks, [0049], indicative of question quality in the question-answer pair, [0035]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang such that the second evaluation score is indicative of a quality of a current question of the current question-answer pair to a current answer of the current question-answer pair in order to reduce errors in the question-answer sets, as suggested by Barbetta ([0003]), predictably improving accuracy, as suggested by Barbetta ([0001]). The references cited are analogous art in the same field of natural language processing. Consider claim 19, Zhang does not, but Barbetta discloses the second evaluation score is indicative of a quality of a current question of the current question-answer pair to a current answer of the current question-answer pair (coercion analyzer generates a score that reflects ease of coercion of the answer to the type of question that the question seeks, [0049], indicative of question quality in the question-answer pair, [0035]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang such that the second evaluation score is indicative of a quality of a current question of the current question-answer pair to a current answer of the current question-answer pair for reasons similar to those for claim 14. Claims 8, 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20250200356) in view of Shen et al. (“Joint Generator-Ranker Learning for Natural Language Generation”. Findings of the Association for Computational Linguistics: ACL 2023, pages 7681–7699 July 9-14, 2023). Consider claim 8, Zhang discloses the program code is further structured to cause the processor to perform an action in response to applying the evaluation model to the current question answer pair, the action comprising at least one of: providing an indication relating to a quality of the current answer (the evaluation label is stored with historical input-output pair and provided to classification model for supervised training, [0143]); or providing an indication relating to the quality of the current answer to a planner of a question-answering system that selected the trained model to generate the current answer (noting the claim language “at least one of” only requires this limitation in the alternative). Zhang does not specifically mention providing an indication related to a quality of the current answer to a second AI model that generated the current answer. Shen discloses providing an indication related to a quality of the current answer to a second AI model that generated the current answer (ranker gives a reward to each candidate, which is used to train the generator, Section 3, page 7683). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by providing an indication related to a quality of the current answer to a second AI model that generated the current answer by using evaluation score from classifier as a reward signal similarly to Shen in order to improve quality of output text, predictably improving natural language task performance, as suggested by Shen (Section 1, page 7681). The references cited are analogous art in the same field of natural language processing. Consider claim 15, Zhang discloses performing an action in response to applying the current question answer pair to the evaluation model, the action comprising at least one of: providing an indication relating to a quality of the current answer (the evaluation label is stored with historical input-output pair and provided to classification model for supervised training, [0143]); or providing an indication relating to the quality of the current answer to a planner of a question-answering system that selected the trained model to generate the current answer (noting the claim language “at least one of” only requires this limitation in the alternative). Zhang does not specifically mention providing an indication related to a quality of the current answer to a second AI model that generated the current answer. Shen discloses providing an indication related to a quality of the current answer to a second AI model that generated the current answer (ranker gives a reward to each candidate, which is used to train the generator, Section 3, page 7683). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by providing an indication related to a quality of the current answer to a second AI model that generated the current answer by using evaluation score from classifier as a reward signal similarly to Shen for reasons similar to those for claim 8. Consider claim 20, Zhang discloses performing an action in response to applying the current question answer pair to the evaluation model, the action comprising at least one of: providing an indication relating to a quality of the current answer (the evaluation label is stored with historical input-output pair and provided to classification model for supervised training, [0143]); or providing an indication relating to the quality of the current answer to a planner of a question-answering system that selected the trained model to generate the current answer (noting the claim language “at least one of” only requires this limitation in the alternative). Zhang does not specifically mention providing an indication related to a quality of the current answer to a second AI model that generated the current answer. Shen discloses providing an indication related to a quality of the current answer to a second AI model that generated the current answer (ranker gives a reward to each candidate, which is used to train the generator, Section 3, page 7683). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Zhang by providing an indication related to a quality of the current answer to a second AI model that generated the current answer by using evaluation score from classifier as a reward signal similarly to Shen for reasons similar to those for claim 8. Allowable Subject Matter Claim 22 is objected to as being dependent on a rejected base claim, but would be allowable if rewritten in independent form including all limitations of its base claim 1 and intervening claim 2. Claim 22 recites: “apply a first weight to the first evaluation score based on the first LLM being the first type of LLM; and apply a second weight to the second evaluation score based on the second LLM being the second type of LLM.” Techniques for weighted ensembles of LLMs were known in the art, see Yang et al. (“One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering”. medRxiv preprint doi: https://doi.org/10.1101/2023.12.21.23300380 ; December 24, 2023). Yang discloses answering medical questions with a weighted majority vote of multiple LLMs. Yang learns the weights based on how the LLMs perform on certain question types. This is helpful in answering questions correctly, but not necessarily applicable to the claimed training the evaluation model based on a combination of LLM generated evaluation scores weighted based on LLM type. Coste et al. (“Reward Model Ensembles Help Mitigate Overoptimization”. arXiv:2310.02743v2 [cs.LG] 10 Mar 2024) is more relevant in spirit to the claimed training the evaluation model based on a combination of LLM generated evaluation scores weighted based on LLM type since a reward corresponds more closely to the claimed evaluation score, but applying the reward ensemble of Coste to Zhang still would not have resulted in the claimed training the evaluation model based on a combination of LLM generated evaluation scores weighted based on LLM type. The subject matter of claim 22 is therefore considered novel and non-obvious over the prior art of record. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Jesse S Pullias/ Primary Examiner, Art Unit 2655 09/03/26
Read full office action

Prosecution Timeline

Show 1 earlier event
Nov 18, 2025
Non-Final Rejection mailed — §102, §103
Feb 18, 2026
Response Filed
Mar 16, 2026
Final Rejection mailed — §102, §103
May 04, 2026
Applicant Interview (Telephonic)
May 05, 2026
Examiner Interview Summary
Jun 16, 2026
Request for Continued Examination
Jun 18, 2026
Response after Non-Final Action
Sep 08, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749483
SYSTEM, APPARATUS, AND METHOD FOR PROCESSING NATURAL LANGUAGE, AND NON-TRANSITORY COMPUTER READABLE RECORDING MEDIUM
3y 2m to grant Granted Sep 29, 2026
Patent 12738286
POST-PROCESSOR, AUDIO DECODER AND RELATED METHODS FOR ENHANCING TRANSIENT PROCESSING
6y 3m to grant Granted Sep 15, 2026
Patent 12730980
REAL-TIME ADAPTATION OF MACHINE LEARNING MODELS USING LARGE LANGUAGE MODELS
2y 9m to grant Granted Sep 08, 2026
Patent 12694234
IMAGE-BASED TEXT TRANSLATION AND PRESENTATION
3y 10m to grant Granted Jul 28, 2026
Patent 12694224
Detecting Random and/or Algorithmically-Generated Character Sequences in Domain Names
2y 2m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
95%
With Interview (+12.5%)
2y 7m (~1m remaining)
Median Time to Grant
High
PTA Risk
Based on 1072 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month