Prosecution Insights
Last updated: October 02, 2026
Application No. 18/664,125

Measuring The Efficacy Of Large Language Models On Classification Tasks

Non-Final OA §101§103
Filed
May 14, 2024
Examiner
MARU, MATIYAS T
Art Unit
Tech Center
Assignee
ORACLE INTERNATIONAL Corporation
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
1y 9m
Est. Remaining
67%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
36 granted / 59 resolved
+1.0% vs TC avg
Moderate +6% lift
Without
With
+5.9%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
22 currently pending
Career history
83
Total Applications
across all art units

Statute-Specific Performance

§101
33.9%
-6.1% vs TC avg
§103
53.7%
+13.7% vs TC avg
§102
2.2%
-37.8% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 59 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1 – 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e. an abstract idea) without significantly more. In step 1, of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, falls within one or more statutory categories (processes). In step 2A prong 1, of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following limitations recite a process that, under broadest reasonable interpretation, recites abstract idea but for the recitation of generic computer components: Regarding claim(s) 1 and analogous claim 13: comparing each particular label of the first plurality of labels to an expected label for the first prompt (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves reviewing generated labels and comparing each label to an expected label. See (MPEP 2106.04)). to compute a distance value for each particular label from the expected label to generate a first plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to quantify the difference between generated labels and expected labels. See (MPEP 2106.04)). generating a first evaluation based at least in part on the first plurality of distance values. (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves analyzing calculated distance to generate a first evaluation. See (MPEP 2106.04)). If the claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process, but for the recitation of generic computer components, then it falls within the mental process. Accordingly, the claim recites an abstract idea. Step 2A Prong 2 of the 101-analysis, set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: As evaluated below: One or more non-transitory computer readable media comprising instructions that, when executed by one or more hardware processors, cause performance of operations comprising: (i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation which does not amount to more than a recitation of the words "apply it" (or an equivalent), such as mere instructions to implement an abstract idea on a computer. See MPEP 2106.05(f)). inputting a first plurality of submissions of a same first prompt to a first generative model to generate a corresponding first plurality of labels by the first generative model; (i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the claim recites limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity, See MPEP (2106.05(g))). wherein the first prompt comprises: a) a first instruction, and b) a first content item; (i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h)). wherein each particular label of the first plurality of labels is one of a set of two or more candidate labels; (i.e.: deemed insufficient to transform the judicial exception to a patentable invention because the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h)). In Step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception: Regarding limitation (I), recite mere application of the abstract idea or mere instructions to implement an abstract idea on a computer are deemed insufficient to transform the judicial exception to a patentable invention because the limitations generally apply the use of a generic computer and/or process with the judicial exception, see MPEP 2106.05(f). Regarding limitation (II), additional elements considered extra/post solution activity, as analyzed above, are activity that are well-understood routine and conventional, specifically: the courts have recognized the computer functions as well‐understood, routine, and conventional functions. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). Regarding limitation (III and IV), additional elements are deemed insufficient to transform the judicial exception to a patentable invention to a patentable invention because they generally link the judicial exception to the technology environment, see MPEP 2106.05(h). As analyzed above, the additional elements, analyzed above, do not integrate the noted judicial exception into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Therefore, the claim is directed to an abstract idea. Regarding claim 2, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: inputting a second plurality of submissions of a same second prompt to the first generative model to generate a corresponding second plurality of labels by the first generative model; The recitation in the additional limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. wherein the second prompt comprises: a) a second instruction, and b) the first content item; The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels; The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. comparing each particular label of the second plurality of labels to an expected label for the second prompt (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves reviewing generated labels and comparing each label to an expected label. See (MPEP 2106.04)). to compute a distance value for each particular label from the expected label to generate a second plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to quantify the difference between generated labels and expected labels. See (MPEP 2106.04)). generating a second evaluation of the second instruction based on the second plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves analyzing calculated distance to generate a second evaluation. See (MPEP 2106.04)). selecting one of the first instruction or the second instruction based at least in part on respective first and second evaluations. (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves comparing evaluation results and choosing one of multiple alternatives based on the results. See (MPEP 2106.04)). Claim(s) 14 and 18, recite similar subject matter as claim 2, so are rejected under the same rationale. Regarding claim 3, dependent upon claim 2, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: inputting a third plurality of submissions to the first generative model, wherein each of the third plurality of submissions includes: a) the selected instruction, and b) a target content item of a plurality of content items; receiving a label for each submission of the third plurality of submissions. The recitation in the additional limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. Regarding claim 4, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: inputting a second plurality of submissions of the first prompt to a second generative model to generate a corresponding second plurality of labels by the second generative model; The recitation in the additional limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels; The recitation in the additional limitation simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. comparing each particular label of the second plurality of labels to an expected label for the second prompt (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves reviewing generated labels and comparing each label to an expected label. See (MPEP 2106.04)). to compute a distance value for each particular label from the expected label to generate a second plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to quantify the difference between generated labels and expected labels. See (MPEP 2106.04)). generating a second evaluation of the second generative model based on the second plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves analyzing calculated distance to generate a second evaluation. See (MPEP 2106.04)). selecting one of the first instruction or the second instruction based at least in part on respective first and second evaluations. (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves comparing evaluation results and choosing one of multiple alternatives based on those results. See (MPEP 2106.04)). Claim(s) 15 and 19, recite similar subject matter as claim 4, so are rejected under the same rationale. Regarding claim 5, dependent upon claim 4, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: inputting a third plurality of submissions to the selected generative model, wherein each of the third plurality of submissions includes: a) the first instruction, and b) a target content item of a plurality of content items; receiving a label for each submission of the third plurality of submissions. The recitation in the additional limitation directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. Regarding claim 6, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: inputting a second plurality of submissions of a same second prompt to the first generative model to generate a corresponding second plurality of labels by the first generative model; inputting a third plurality of submissions of the first prompt to a second generative model to generate a corresponding third plurality of labels by the second generative model; inputting a fourth plurality of submissions of the second prompt to the second generative model to generate a corresponding fourth plurality of labels by the second generative model; The recitation in the additional limitations directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. wherein the second prompt comprises: a) a second instruction, and b) the first content item; wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels; wherein the second evaluation corresponds to the combination of the first generative model and the second instruction; … wherein the third evaluation corresponds to the combination of the second generative model and the first instruction; … wherein the fourth evaluation corresponds to the combination of the second generative model and the second instruction; based at least in part on respective first, second, third, and fourth evaluations, selecting one of: a) the combination of the first generative model and the first instruction; b) the combination of the first generative model and the second instruction; c) the combination of the second generative model and the first instruction; or d) the combination of the second generative model and the second instruction. The recitation in the additional limitations simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. comparing each particular label of the second plurality of labels to an expected label for the second prompt (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves reviewing generated labels and comparing each label to an expected label. See (MPEP 2106.04)). to compute a distance value for each particular label from the expected label to generate a second plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to quantify the difference between generated labels and expected labels. See (MPEP 2106.04)). generating a second evaluation based on the second plurality of distance values, (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves analyzing calculated distance values to evaluate the performance of the second generative model. See (MPEP 2106.04)). comparing each particular label of the third plurality of labels to the expected label for the first prompt (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves reviewing generated labels and comparing each label to an expected label. See (MPEP 2106.04)). to compute a distance value for each particular label from the expected label to generate a third plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to quantify the difference between generated labels and expected labels. See (MPEP 2106.04)). generating a third evaluation based on the third plurality of distance values, (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves analyzing calculated distance values to evaluate the performance of the third generative model. See (MPEP 2106.04)). comparing each particular label of the fourth plurality of labels to the expected label for the second prompt (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves reviewing generated labels and comparing each label to an expected label. See (MPEP 2106.04)). to compute a distance value for each particular label from the expected label to generate a fourth plurality of distance values; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to quantify the difference between generated labels and expected labels. See (MPEP 2106.04)). generating a fourth evaluation based on the fourth plurality of distance values, (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves analyzing calculated distance values to evaluate the performance of the fourth generative model. See (MPEP 2106.04)). Claim(s) 16 and 20, recite similar subject matter as claim 6, so are rejected under the same rationale. Regarding claim 7, dependent upon claim 6, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: inputting a fifth plurality of submissions to the selected generative model, wherein each of the fifth plurality of submissions includes: a) the selected instruction, and b) a target content item of a plurality of content items; receiving a label for each submission of the fifth plurality of submissions. The recitation in the additional limitations directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above. Regarding claim 8, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: wherein the operations further comprise: instructing the first generative model to classify content items using the set of three or more candidate labels by at least one of: a) submitting a content item; b) submitting a content item in conjunction with explicit commands to classify the content item; or c) submitting a content item in conjunction with one or more characters that imply the content item is to be classified by the first generative model. The recitation in the additional limitations simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Regarding claim 9, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: wherein the first prompt further comprises commands that indicate a requested output format. The recitation in the additional limitations simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Regarding claim 10, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: wherein a first distance value of the first plurality of distance values is a non-binary distance value. The recitation in the additional limitations simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Regarding claim 11, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: wherein the first evaluation comprises one or more numeric metrics. The recitation in the additional limitations simply links the judicial exception to a field of use and/or technology environment, see MPEP 2106.05(h). Limitations directed to field of use cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Regarding claim 12, dependent upon claim 1, and fail to resolve the deficiencies identified above by integrating the judicial exception into a practical application, or introducing significantly more than the judicial exception. The claim recites: selecting a number of submissions in the plurality of submissions (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves choosing a quantity of submissions from a larger set based on specified criteria. See (MPEP 2106.04)). performing a set of two or more runs of each test prompt of a plurality of test prompts, Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. wherein a run comprises: inputting a test prompt to the first generative model to generate a corresponding test label; The recitation in the additional limitations directed to mere data gathering as deemed insufficient to transform the judicial exception because claimed elements are considered insignificant extra-solution activity and well-understood routine and conventional (2106.05(d)). Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TL| Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network). See MPEP 2106.05(d)(II). The additional limitations as analyze failed to integrate a judicial exception into a practical application at Step 2A and provide an inventive concept in Step 2B, per the analysis above based on the set of two or more runs for each test prompt: computing a mean correctness metric for each test prompt; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to compute an average correctness metric from multiple runs. See (MPEP 2106.04)). computing a variance metric for each test prompt; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves performing mathematical calculation to compute the variance of information associated with each test prompt. See (MPEP 2106.04)). based at least in part on the mean correctness metrics and the variance metrics for each test prompt, computing a confidence value; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves analyzing statistical metrics and performing mathematical calculation to compute a confidence value. See (MPEP 2106.04)). performing additional runs of each test prompt Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. re-computing the confidence value until the confidence value meets a predefined threshold; (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mathematical concept: It involves repeatedly performing mathematical calculations until the result meets a predetermined threshold. See (MPEP 2106.04)). in response to determining that the confidence value meets the predefined threshold, identifying the number of runs as the number of submissions to be used when inputting the first plurality of submissions. (i.e.: the broadest reasonable interpretation, the claim recites abstract idea: mental process: It involves evaluating whether a condition is satisfied and selecting a corresponding number of runs based on the evaluation. See (MPEP 2106.04)). Regarding claim 17, The rest of the limitations are analogous to claim 1, so are rejected under similar rationale. A system, comprising: at least one device including a hardware processor; the system being configured to perform operations comprising: Deemed insufficient to transform the judicial exception to a patentable invention because the limitation is directed to mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and are considered to adding the words “apply it” (or an equivalent) with the judicial exception, See MPEP 2106.05(f). Limitations directed to using the computer as a tool for implementing an abstract idea cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1 – 5, 8, 11, 13 – 15 and 17 – 19 are rejected under 35 U.S.C. 103 as being unpatentable over Hearty et al., Pub. No.: US20240330655A1, in view of Atlan et al., Pub. No.: US20240185137A1. Regarding claim 1, Hearty teaches: One or more non-transitory computer readable media comprising instructions that, when executed by one or more hardware processors, cause performance of operations comprising: (Hearty, “[0082] Example 6. A system comprising: memory hardware storing instructions and one or more electronic processors configured to execute the instructions [One or more non-transitory computer readable media comprising instructions that, when executed by one or more hardware processors, cause performance of operations], wherein the instructions include: providing a plurality of training inputs to a first artificial intelligence model to generate a plurality of training outputs,...”) comparing each particular label of the first plurality of labels to an expected label for the first prompt to compute a distance value for each particular label from the expected label to generate a first plurality of distance values; (Hearty, “[0070] … In the example process 1000, the generative artificial intelligence test module 176 computes a first metric corresponding to a count of the selected labeled features (at block 1036). In the example process 1000, the generative artificial intelligence test module 176 computes a second metric corresponding to distances between the selected labeled features and the test feature (at block 1040). In some embodiments, the second metric may be computed as a Manhattan distance (or the distance measure representing the sum of distances between the test feature and each selected label feature) [comparing each particular label of the first plurality of labels to an expected label for the first prompt to compute a distance value for each particular label from the expected label to generate a first plurality of distance values].”) generating a first evaluation based at least in part on the first plurality of distance values. (Hearty, “[0082] … computing a first metric corresponding to a count of the selected labeled features, computing a second metric corresponding to distances between the selected labeled features and the test feature [generating a first evaluation based at least in part on the first plurality of distance values], computing a risk score based on the first metric and the second metric,...”) Hearty does not teach: inputting a first plurality of submissions of a same first prompt to a first generative model to generate a corresponding first plurality of labels by the first generative model; wherein the first prompt comprises: a) a first instruction, and b) a first content item; wherein each particular label of the first plurality of labels is one of a set of two or more candidate labels; Atlan teaches: inputting a first plurality of submissions of a same first prompt to a first generative model to generate a corresponding first plurality of labels by the first generative model; (Atlan, “[0008] In some embodiments, the graph augmentation system receives the name of an application. The graph augmentation system generates a prompt for an LLM based on the name of the application. The prompt includes a request for one or more tags associated with the application. The one or more tags describe user interests associated with the application. The graph augmentation system provides the prompt to the LLM for execution [inputting a first plurality of submissions of a same first prompt to a first generative model] and receives, as output from the LLM, a plurality of candidate tags [to generate a corresponding first plurality of labels by the first generative model]. The graph augmentation system inputs the plurality of candidate tags into a classifier. The classifier is trained to classify candidate tags into known tags. Known tags are tags that already exist in a graph.”) wherein the first prompt comprises: a) a first instruction, and b) a first content item; wherein each particular label of the first plurality of labels is one of a set of two or more candidate labels; (Atlan, “[0008] In some embodiments, the graph augmentation system receives the name of an application. The graph augmentation system generates a prompt for an LLM based on the name of the application. The prompt includes a request for one or more tags associated with the application. The one or more tags describe user interests associated with the application [wherein the first prompt comprises: a) a first instruction, and b) a first content item]. The graph augmentation system provides the prompt to the LLM for execution and receives, as output from the LLM, a plurality of candidate tags [wherein each particular label of the first plurality of labels is one of a set of two or more candidate labels]. The graph augmentation system inputs the plurality of candidate tags into a classifier. The classifier is trained to classify candidate tags into known tags. Known tags are tags that already exist in a graph.”) Atlan and Hearty are related to the same field of endeavor (i.e.: knowledge-based neural networks). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Atlan with teachings of Hearty to add a graph based mechanism for organizing machine learning outputs with confidence weighted relationships to improve the organization and retrieval of labeled features (Atlan, Abstract). Claim 13, recites limitations analogous to claim 1, so is rejected under the same rationale. Regarding claim 2, Hearty in view of Atlan teach the method of claim 1. Hearty further teaches: wherein the first evaluation corresponds to the first instruction and wherein the operations further comprise: inputting a second plurality of submissions of a same second prompt to the first generative model to generate a corresponding second plurality of labels by the first generative model; (Hearty, “[0005] In some aspects, the techniques described herein relate to a non-transitory computer-readable storage medium including executable instructions, wherein the executable instructions cause an electronic processor to: provide a plurality of training inputs to a first artificial intelligence model to generate a plurality of training outputs [inputting a second plurality of submissions of a same second prompt to the first generative model to generate a corresponding second plurality of labels by the first generative model]; preprocess the plurality of training inputs and/or training outputs; label one or more of the plurality of preprocessed training inputs and/or training outputs;”) comparing each particular label of the second plurality of labels to an expected label for the second prompt to compute a distance value for each particular label from the expected label to generate a second plurality of distance values; (Hearty, “ [0070] …In the example process 1000, the generative artificial intelligence test module 176 computes a first metric corresponding to a count of the selected labeled features (at block 1036). In the example process 1000, the generative artificial intelligence test module 176 computes a second metric corresponding to distances between the selected labeled features and the test feature (at block 1040). In some embodiments, the second metric may be computed as a Manhattan distance (or the distance measure representing the sum of distances between the test feature and each selected label feature) [comparing each particular label of the second plurality of labels to an expected label for the second prompt to compute a distance value for each particular label from the expected label to generate a second plurality of distance values]…”) generating a second evaluation of the second instruction based on the second plurality of distance values; (Hearty, “[0005] … compute a first metric corresponding to a count of the selected labeled features; compute a second metric corresponding to distances between the selected labeled features [generating a second evaluation of the second instruction based on the second plurality of distance values] and the test feature; compute a risk score based on the first metric and the second metric; in response to the risk score being above a threshold, assign a first label to the test output, wherein the first label is indicative of an erroneous output from the first artificial intelligence model;…”) selecting one of the first instruction or the second instruction based at least in part on respective first and second evaluations. (Hearty, “[0071] In the example process 1000, the generative artificial intelligence test module 176 computes a risk score based on the first metric and the second metric (at block 1044). [selecting one of the first instruction or the second instruction based at least in part on respective first and second evaluations] In the example process 1000, the generative artificial intelligence test module 176 determines whether the risk score is above a threshold (at decision block 1048).”) Atlan further teaches: wherein the second prompt comprises: a) a second instruction, and b) the first content item; wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels; (Atlan, “[0008] In some embodiments, the graph augmentation system receives the name of an application. The graph augmentation system generates a prompt for an LLM based on the name of the application. The prompt includes a request for one or more tags associated with the application. The one or more tags describe user interests associated with the application [wherein the first prompt comprises: a) a first instruction, and b) a first content item]. The graph augmentation system provides the prompt to the LLM for execution and receives, as output from the LLM, a plurality of candidate tags [wherein each particular label of the first plurality of labels is one of a set of two or more candidate labels]. The graph augmentation system inputs the plurality of candidate tags into a classifier. The classifier is trained to classify candidate tags into known tags. Known tags are tags that already exist in a graph.”) It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Atlan with teachings of Hearty for the same reasons disclosed for claim 1. Claim(s) 14 and 18, recite limitations analogous to claim 2, so are rejected under the same rationale. Regarding claim 3, Hearty in view of Atlan teach the method of claim 2. Hearty further teaches: wherein the operations further comprise: inputting a third plurality of submissions to the first generative model, wherein each of the third plurality of submissions includes: a) the selected instruction, and b) a target content item of a plurality of content items; (Hearty, “[0065] … In some examples, the generative artificial intelligence test module 176 applies a second label to each preprocessed training input and/or preprocessed training output that has been marked as problematic by users. In various implementations, the generative artificial intelligence test module 176 applies a third label to each preprocessed training input and/or preprocessed training output [inputting a third plurality of submissions to the first generative model] that are designed to conform to a one or more criteria (for example, to be offensive and/or insulting). [wherein each of the third plurality of submissions includes: a) the selected instruction, and b) a target content item of a plurality of content items]”) receiving a label for each submission of the third plurality of submissions. (Hearty, “[0085] Example 9. The system of example 8, wherein labeling one or more of the plurality of preprocessed training inputs and/or training outputs includes at least one of: assigning a third label to each preprocessed training input and/or training output that contains a term present in a list [receiving a label for each submission of the third plurality of submissions]; assigning a fourth label to each preprocessed training input and/or training output marked by a user; and assigning a fifth label to each preprocessed training input and/or training output conforming to a criteria.”) Regarding claim 4, Hearty in view of Atlan teach the method of claim 1. Hearty further teaches: wherein the first evaluation corresponds to the first generative model and wherein the operations further comprise: inputting a second plurality of submissions of the first prompt to a second generative model to generate a corresponding second plurality of labels by the second generative model; (Hearty, “[0005] …provide a test input to the first artificial intelligence model to generate a test output; preprocess the test output; add the preprocessed test output to the feature space as a test feature using the second artificial intelligence model [inputting a second plurality of submissions of the first prompt to a second generative model to generate a corresponding second plurality of labels by the second generative model]; select labeled features in the feature space within a radius of the test feature; compute a first metric corresponding to a count of the selected labeled features;”) wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels; (Hearty, “[0026] … A training protocol for generating such a corpus or corpora in a time-efficient way includes: (1) generating labels [wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels] in a single domain by retrieving statements on one topic, (2) generating labels in an additional domain by retrieving statements on another topic, (3) training a machine learning model that evaluates statements across multiple domains, (4) continuing to add additional domains while holding out a multi-domain dataset to test generalization abilities of the machine learning model, and (5) continuing to train the machine learning model until generalization ability reaches a threshold.”) comparing each particular label of the second plurality of labels to an expected label for the second prompt to compute a distance value for each particular label from the expected label to generate a second plurality of distance values; (Hearty, “[0005] … using a second artificial intelligence model; provide a test input to the first artificial intelligence model to generate a test output; preprocess the test output; add the preprocessed test output to the feature space as a test feature using the second artificial intelligence model; select labeled features in the feature space within a radius of the test feature; compute a first metric corresponding to a count of the selected labeled features; compute a second metric corresponding to distances between the selected labeled features and the test feature; [comparing each particular label of the second plurality of labels to an expected label for the second prompt to compute a distance value for each particular label from the expected label to generate a second plurality of distance values]”) generating a second evaluation of the second generative model based on the second plurality of distance values; (Hearty, “[0044] …In various implementations, the secondary processing [generating a second evaluation of the second generative model] may produce a labeled set with labels indicating whether responses or portions of responses from the first artificial intelligence model are valid or invalid. In various implementations, subnetworks of the embedded data may be evaluated [based on the second plurality of distance values]. For example, subnetworks may be evaluated based on their proximity and associations with other subnetworks.”) selecting one of the first generative model or the second generative model based at least in part on respective first and second evaluations. (Hearty, “[0045] In the example process 500, the generative artificial intelligence test module 176 may provide the labeled data to a second artificial intelligence model to generate a classification (at block 520). In various implementations, the second artificial intelligence model may be a binary classification model that receives the labeled data and generates a classification.”) Claim(s) 15 and 19, recite limitations analogous to claim 4, so are rejected under the same rationale. Regarding claim 5, Hearty in view of Atlan teach the method of claim 4. Hearty further teaches: wherein the operations further comprise: inputting a third plurality of submissions to the selected generative model, (Hearty, “[0085] Example 9. The system of example 8, wherein labeling one or more of the plurality of preprocessed training inputs and/or training outputs includes at least one of: assigning a third label to each preprocessed training input and/or training output that contains a term present in a list [inputting a third plurality of submissions to the selected generative model]; assigning a fourth label to each preprocessed training input and/or training output marked by a user; and assigning a fifth label to each preprocessed training input and/or training output conforming to a criteria.”) wherein each of the third plurality of submissions includes: a) the first instruction, and b) a target content item of a plurality of content items; (Hearty, “[0065] … In some examples, the generative artificial intelligence test module 176 applies a second label to each preprocessed training input and/or preprocessed training output that has been marked as problematic by users. In various implementations, the generative artificial intelligence test module 176 applies a third label to each preprocessed training input and/or preprocessed training output [wherein each of the third plurality of submissions] that are designed to conform to a one or more criteria (for example, to be offensive and/or insulting) [includes: a) the first instruction, and b) a target content item of a plurality of content items].”) receiving a label for each submission of the third plurality of submissions. (Hearty, “[0085] Example 9. The system of example 8, wherein labeling one or more of the plurality of preprocessed training inputs and/or training outputs includes at least one of: assigning a third label to each preprocessed training input and/or training output [receiving a label for each submission of the third plurality of submissions] that contains a term present in a list; assigning a fourth label to each preprocessed training input and/or training output marked by a user; and assigning a fifth label to each preprocessed training input and/or training output conforming to a criteria.”) Regarding claim 8, Hearty in view of Atlan teach the method of claim 1. Atlan further teaches: wherein the operations further comprise: instructing the first generative model to classify content items using the set of three or more candidate labels by at least one of: (Atlan, “[0007] The graph augmentation system may generate tags for applications using a large language model (LLM). The graph augmentation system may prompt the LLM with a name of an application and a request to generate tags (e.g., descriptions, labels, interests) for the application. The LLM may output candidate tags describing the application, and the graph augmentation system may use a classifier to classify the candidate tags into tags in the graph [instructing the first generative model to classify content items using the set of three or more candidate labels by at least one of:].”) a) submitting a content item; b) submitting a content item in conjunction with explicit commands to classify the content item; or c) submitting a content item in conjunction with one or more characters that imply the content item is to be classified by the first generative model. (Atlan, “[0049] … The tagging module 215 inputs each of the candidate tags into the tag classifier [a) submitting a content item] and receives, as output from the tag classifier, a known tag that corresponds to each of the candidate categories. In some embodiments, the tagging module 215 may determine that a candidate tag has no matching known tag.”) It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Atlan with teachings of Hearty for the same reasons disclosed for claim 1. Regarding claim 11, Hearty in view of Atlan teach the method of claim 1. Hearty further teaches: wherein the first evaluation comprises one or more numeric metrics. (Hearty, “[0082] … adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model, selecting labeled features in the feature space within a radius of the test feature, computing a first metric corresponding to a count of the selected labeled features [wherein the first evaluation comprises one or more numeric metrics], computing a second metric corresponding to distances between the selected labeled features and the test feature, computing a risk score based on the first metric and the second metric, in response to the risk score being above a threshold, assigning a first label to the test output,…”) Regarding claim 17, Hearty teaches: A system, comprising: at least one device including a hardware processor; the system being configured to perform operations comprising: (Hearty, “[0082] Example 6. A system comprising: memory hardware storing instructions and one or more electronic processors configured to execute the instructions [A system, comprising: at least one device including a hardware processor; the system being configured to perform operations comprising], wherein the instructions include: providing a plurality of training inputs to a first artificial intelligence model to generate a plurality of training outputs,...”) The rest of the limitations are analogous to claim 1, so are rejected under similar rationale. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Hearty in view of Atlan and further view of MATAMOROS et al., Pub. No.: US20250209309A1. Regarding claim 9, Hearty in view of Atlan teach the method of claim 1. Hearty in view of Atlan do not teach: wherein the first prompt further comprises commands that indicate a requested output format. MATAMOROS teaches: wherein the first prompt further comprises commands that indicate a requested output format. (MATAMOROS, “[0002] A large language model (LLM) is a type of machine learning (ML) model that may generate text output, including natural language text output. A LLM may receive a natural language input, for example, a LLM may be provided with a prompt, which may be a natural language instruction that instructs the LLM to generate a desired output, including natural language text or other generative output in various desired formats [wherein the first prompt further comprises commands that indicate a requested output format].”) MATAMOROS, Hearty and Atlan are related to the same field of endeavor (i.e.: knowledge-based neural networks). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of MATAMOROS with teachings of Hearty and Atlan to automatically generating labeled training data using similarity based candidate retrieval and LLM annotation to improve the efficiency of producing the labeled training inputs. (MATAMOROS, Abstract). Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Hearty in view of Atlan and further view of Ardhanari et al., Pub. No.: US20210248268A1. Regarding claim 10, Hearty in view of Atlan teach the method of claim 1. Hearty in view of Atlan do not teach: wherein a first distance value of the first plurality of distance values is a non-binary distance value Ardhanari teaches: wherein a first distance value of the first plurality of distance values is a non-binary distance value. (Ardhanari, “[0253] … In a non-binary or “smooth” context, terms in the corpus may be weighted (e.g., assigned a value between 0 and 1) based on factors such as the distance from the query term. For example, the weight assigned to a term in a non-binary context may attenuate exponentially based on distance of the term from the query term [wherein a first distance value of the first plurality of distance values is a non-binary distance value].”) Ardhanari, Hearty and Atlan are related to the same field of endeavor (i.e.: knowledge-based neural networks). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Ardhanari with teachings of Hearty and Atlan to add preprocessing training data by identifying sensitive entities to improve the privacy and security of the training inputs. (Ardhanari, Abstract). Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Hearty in view of Atlan and further view of Yoshikawa et al., Pub. No.: US20220391596A1 and MUTHU et al., Pub. No.: US20250209372A1. Regarding claim 12, Hearty in view of Atlan teach the method of claim 1. Hearty further teaches: wherein the operations further comprise: selecting a number of submissions in the plurality of submissions at least by: performing a set of two or more runs of each test prompt of a plurality of test prompts, wherein a run comprises: inputting a test prompt to the first generative model to generate a corresponding test label; based on the set of two or more runs for each test prompt: (Hearty, “[0004] … training outputs, organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model, providing a test input to the first artificial intelligence model [selecting a number of submissions in the plurality of submissions at least by: performing a set of two or more runs of each test prompt of a plurality of test prompts, wherein a run comprises: inputting a test prompt to the first generative model to generate a corresponding test label] to generate a test output, preprocessing the test output [based on the set of two or more runs for each test prompt], adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model…”) performing additional runs of each test prompt and re-computing the confidence value until the confidence value meets a predefined threshold; (Hearty, “[0026] … A training protocol for generating such a corpus or corpora in a time-efficient way includes: (1) generating labels in a single domain by retrieving statements on one topic, (2) generating labels in an additional domain by retrieving statements on another topic, (3) training a machine learning model that evaluates statements across multiple domains, (4) continuing to add additional domains while holding out a multi-domain dataset to test generalization abilities of the machine learning model, and (5) continuing to train the machine learning model until generalization ability reaches a threshold [performing additional runs of each test prompt and re-computing the confidence value until the confidence value meets a predefined threshold].”) Hearty in view of Atlan do not teach: computing a mean correctness metric for each test prompt; computing a variance metric for each test prompt; based at least in part on the mean correctness metrics and the variance metrics for each test prompt, computing a confidence value; in response to determining that the confidence value meets the predefined threshold, identifying the number of runs as the number of submissions to be used when inputting the first plurality of submissions. Yoshikawa teaches: computing a mean correctness metric for each test prompt; (Yoshikawa, “[0059] Furthermore, in the case R3 where the difference in the probability distribution is small and the context-dependency of the dummy contexts (c.sub.1, c.sub.2, c.sub.3) for the output result when the input sentence x is input into the language model M1 is low, the value of the confidence C for the correct [computing a mean correctness metric for each test prompt] response becomes high (0.9 in the illustrated example). Thus, in the case R3, the response by the language model M1 is output, assuming that it is likely to be the correct response. In this way, the information processing apparatus 1 according to the embodiment can assist optimization of the output of the language model M1.”) computing a variance metric for each test prompt; (Yoshikawa, “[0062] In addition, the information processing apparatus 1 calculates the variance based on each distribution of the output [computing a variance metric for each test prompt] result when each of the combined sentences is input into the language model M1, and assumes the calculated variance to be the index value of the confidence C. This allows the information processing apparatus 1 to assume the variance based on each distribution of the output result in the combined sentences to be the index value of the confidence C and to obtain the confidence C in consideration of the context-dependency.”) based at least in part on the mean correctness metrics and the variance metrics for each test prompt, computing a confidence value; (Yoshikawa, “[0059] Furthermore, in the case R3 where the difference in the probability distribution is small and the context-dependency of the dummy contexts (c.sub.1, c.sub.2, c.sub.3) for the output result when the input sentence x is input into the language model M1 is low, the value of the confidence C for the correct response becomes high (0.9 in the illustrated example) [based at least in part on the mean correctness metrics and the variance metrics for each test prompt, computing a confidence value]. Thus, in the case R3, the response by the language model M1 is output, assuming that it is likely to be the correct response. In this way, the information processing apparatus 1 according to the embodiment can assist optimization of the output of the language model M1.”) Yoshikawa, Hearty and Atlan are related to the same field of endeavor (i.e.: knowledge-based neural networks). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of Yoshikawa with teachings of Hearty and Atlan to evaluate confidence of outputs based on distributions to improve the reliability of the risk score. (Yoshikawa, Abstract). Hearty in view of Atlan and Yoshikawa do not teach: in response to determining that the confidence value meets the predefined threshold, identifying the number of runs as the number of submissions to be used when inputting the first plurality of submissions. MUTHU teaches: in response to determining that the confidence value meets the predefined threshold, identifying the number of runs as the number of submissions to be used when inputting the first plurality of submissions. (MUTHU, “[0029] Some embodiments provide for repeatedly optimizing an input prompt for a number of iterations. For example, an optimization configuration (e.g., specified by a user) may specify that an input prompt should be optimized for a given number of iterations or until some other condition is met (e.g., a score that meets a threshold condition, a number of successive iterations without improvement, and/or the like) [in response to determining that the confidence value meets the predefined threshold]. Accordingly, the optimization process may be repeated for a plurality of iterations (e.g., up to the given number of iterations or until some other condition is met) [identifying the number of runs as the number of submissions to be used when inputting the first plurality of submissions] using the optimized prompt from each completed iteration as the input prompt for the subsequent iteration. The optimization may stop before completing the one or more specified conditions if all scoring criteria are satisfied.”) MUTHU, Hearty, Atlan and Yoshikawa are related to the same field of endeavor (i.e.: knowledge-based neural networks). It would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the teaching of MUTHU with teachings of Hearty, Atlan and Yoshikawa to optimize prompts based on evaluation scores to improve the quality of the training and test outputs. (MUTHU, Abstract). Allowable subject matter Claim(s) 6 – 7, 16 and 20 would be allowable if rewritten or amended to overcome the rejection 35 U.S.C. 101 set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. Claim 6 recites: The non-transitory media of Claim 1, wherein the first evaluation corresponds to the combination of the first generative model and the first instruction, and wherein the operations further comprise: inputting a second plurality of submissions of a same second prompt to the first generative model to generate a corresponding second plurality of labels by the first generative model; wherein the second prompt comprises: a) a second instruction, and b) the first content item; wherein each particular label of the second plurality of labels is one of the set of three or more candidate labels; comparing each particular label of the second plurality of labels to an expected label for the second prompt to compute a distance value for each particular label from the expected label to generate a second plurality of distance values; generating a second evaluation based on the second plurality of distance values, wherein the second evaluation corresponds to the combination of the first generative model and the second instruction; inputting a third plurality of submissions of the first prompt to a second generative model to generate a corresponding third plurality of labels by the second generative model; comparing each particular label of the third plurality of labels to the expected label for the first prompt to compute a distance value for each particular label from the expected label to generate a third plurality of distance values; generating a third evaluation based on the third plurality of distance values, wherein the third evaluation corresponds to the combination of the second generative model and the first instruction; inputting a fourth plurality of submissions of the second prompt to the second generative model to generate a corresponding fourth plurality of labels by the second generative model; comparing each particular label of the fourth plurality of labels to the expected label for the second prompt to compute a distance value for each particular label from the expected label to generate a fourth plurality of distance values; generating a fourth evaluation based on the fourth plurality of distance values, wherein the fourth evaluation corresponds to the combination of the second generative model and the second instruction; based at least in part on respective first, second, third, and fourth evaluations, selecting one of: a) the combination of the first generative model and the first instruction; b) the combination of the first generative model and the second instruction; c) the combination of the second generative model and the first instruction; or d) the combination of the second generative model and the second instruction. Closest prior arts: Hearty et al., Pub. No.: US20240330655A1. Hearty teaches providing a plurality of training inputs to a first artificial intelligence model to generate a plurality of training outputs, organizing the preprocessed training inputs and/or training outputs in a feature space based on proximity using a second artificial intelligence model, providing a test input to the first artificial intelligence model to generate a test output, adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model, computing a first metric corresponding to a count of selected labeled features in the feature space, computing a second metric corresponding to distances between the selected labeled features and the test feature, computing a risk score based on the first metric and the second metric. However, Hearty does not teach inputting multiple submissions of a second prompt, including a second instruction and the same first content item, to a first generative model to generate corresponding labels selected from a set of three or more candidate labels. The generated labels are compared with an expected label to compute distance values, which are used to generate a second evaluation for the first generative model with the second instruction. The method further inputs multiple submissions of the first prompt to a second generative model, compares the resulting labels with the expected label for the first prompt to compute a third plurality of distance values and generate a third evaluation for the second generative model with the first instruction. Atlan et al., Pub. No.: US20240185137A1. Atlan teaches inputting signals into a machine learning model, and receives, as output from the model, tags that correspond to the new application and levels of confidence for each tag. The system updates the graph to include one or more nodes corresponding to the new application, with the tags linked to the one or more nodes with an edge that has a weight corresponding to the level of confidence. However, Atlan does not teach inputting multiple submissions of a second prompt, including a second instruction and the same first content item, to a first generative model to generate corresponding labels selected from a set of three or more candidate labels. The generated labels are compared with an expected label to compute distance values, which are used to generate a second evaluation for the first generative model with the second instruction. The method further inputs multiple submissions of the first prompt to a second generative model, compares the resulting labels with the expected label for the first prompt to compute a third plurality of distance values and generate a third evaluation for the second generative model with the first instruction. Claim(s) 16 and 20 recites limitations analogous to claim 6, so would be allowable for same rationale. Dependent claim 7, would be allowable because of their dependency to claim 6. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kumar et al., Pub. No.: US20250321848A1. Kumar teaches obtaining multiple assumptions and determines which of the multiple assumptions are satisfied by the first and the second history of the metric to obtain multiple satisfied assumptions. The system obtains multiple tests associated with the multiple satisfied assumptions. Lu et al., Pub. No.: US20240273286A1 Lu teaches a generative language model, a first version of a first document, generating a second version of the first document by dividing the first version of the first document into a plurality of segments, where a first segment of the plurality of segments includes a subset of the digital content generated by the generative language model; Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATIYAS T MARU whose telephone number is (571)270-0902 or via email: matiyas.maru@uspto.gov. The examiner can normally be reached Monday 8:00am - Friday 4:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached on (571)431-0762. The fax phone number for the organization were this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.M./ Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

May 14, 2024
Application Filed
Aug 19, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743611
HYBRID NEURAL NETWORK SYSTEM WITH MULTI-THREADED INPUTS
3y 11m to grant Granted Sep 22, 2026
Patent 12731068
SELECTING REPRESENTATIVE FEATURES FOR MACHINE LEARNING MODELS
5y 5m to grant Granted Sep 08, 2026
Patent 12725031
SYSTEM AND METHOD FOR IMPLEMENTING FEDERATED LEARNING ENGINE FOR INTEGRATION OF VERTICAL AND HORIZONTAL AI
5y 0m to grant Granted Sep 01, 2026
Patent 12725011
BAYESIAN NEURAL NETWORKS FOR RANSOMWARE INCIDENT DETECTION
4y 4m to grant Granted Sep 01, 2026
Patent 12705482
POINT PROCESS LEARNING METHOD, POINT PROCESS LEARNING APPARATUS AND PROGRAM
3y 3m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
67%
With Interview (+5.9%)
4y 1m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 59 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month