DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. Using the subject matter eligibility test from page 74621 of the Federal Register Notice titled “2014 Interim Guidance on Patent Subject Matter Eligibility,” a two-step process is performed. Under step 1, the claims are analyzed to determine if the claim is directed to a process, machine, article of manufacture, or composition of matter. In this case, claims 1-11 are directed to a method, which is a process; claims 12-16 are directed to a system, which is a machine or an article of manufacture; and claims 17-20 are directed to a computer program product, which is a machine or an article of manufacture. Step 2A (part 1 of the Mayo test), using the guidance from pages 50-57 of the Federal Register Vol. 84 No. 4 from Monday, January 7, 2019, requires applying a two-prong inquiry. In Prong One, examiners evaluate whether the claim recites a judicial exception, determining if the claim is directed to a law of nature, a natural phenomenon, or an abstract idea. In this case, claim 1 recites receiving a prompt, standardizing and adjusting the prompt by pre-processing, tokenizing, and cleaning, determining evaluation criteria, distributing the prompt to LLMs, generating and evaluating responses to the prompt, generating rankings, generating a top ranked response, aggregating rankings, and identifying a winning LLM, which are mental processes. In Prong Two, examiners evaluate whether the judicial exception is integrated into a practical application that imposes a meaningful limit on the judicial exception. In this case, sending data to a user is mere extrasolution activity, while structural elements of processor, memory, storage medium, and LLMs are generic computing components, and do not integrate the abstract ideas into a practical application.
Step 2B (part 2 of the Mayo test) requires analyzing the claims to determine if they recite additional elements that amount to significantly more than the judicial exception. In this case, the claims do not include additional elements that are sufficient to amount to significantly more than the abstract idea itself.
Regarding claims 1, 12, and 17, receiving a prompt, standardizing and adjusting the prompt by pre-processing, tokenizing, and cleaning, determining evaluation criteria, distributing the prompt to LLMs, generating and evaluating responses to the prompt, generating rankings, generating a top ranked response, aggregating rankings, and identifying a winning LLM are mental processes, which is an abstract idea. For example, a human could receive a prompt and process it for LLM compatibility, could decide upon criteria for evaluation, and could generate different responses to the prompt, and could rank the responses to find a winner. Additional limitations of sending data to a user is mere extrasolution activity, while structural elements of processor, memory, storage medium, and LLMs are generic computing components, and do not integrate the abstract ideas into a practical application or constitute significantly more.
Regarding claims 2, 13, and 18, determining a style, determining information is not in a database, and updating the database are mental processes, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claims 3, 14, and 19, determining a style, determining information is in a database, and selecting an LLM are mental processes, which is an abstract idea. Further limitations of receiving and sending data are mere extrasolution activity, and do not integrate the abstract idea into a practical application or constitute significantly more.
Regarding claims 4, 15, and 20, providing benchmarking, evaluation, and ranking are mental processes, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claims 5 and 16, requesting and generating criteria are mental processes, which is an abstract idea. Use of an LLM is simply having a computer perform the abstract idea, and does not integrate the abstract idea into a practical application or constitute significantly more.
Regarding claim 6, determining a context is a mental process, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claim 7, generating scoring criteria is a mental process, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claim 8, generating confidence levels and explanations for rankings is a mental process, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claim 9, determining a tie and the further iterations is a mental process, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claim 10, verifying compatibility is a mental process, which is an abstract idea without integration into a practical application and without significantly more.
Regarding claim 11, determining a prompt type and selecting an LLM are mental processes, which is an abstract idea. Maintaining LLMs is interpreted as storage, which is mere extrasolution activity, and does not integrate the abstract idea into a practical application or constitute significantly more.
The limitations of the claims, taken alone, do not amount to significantly more than the above-identified judicial exception (the abstract idea). Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements individually. Applicable case law cited in the Federal Register includes, but is not limited to: Alice Corp., 134 S. Ct. at 2355-56, Digitech Image Tech., LLC v. Electronics for Imaging, Inc., 758 F.3d 1344 (Fed. Cir. 2014), Benson, 409 U.S. at 63.
See "Preliminary Examination Instructions in view of the Supreme Court Decision in Alice Corporation Pty. Ltd. v. CLS Bank International, et al.," dated June 25, 2014, and the Federal Register notice titled "2014 Interim Guidance on Patent Subject Matter Eligibility" (79 FR 74618).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4-12, 15-17, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Srinivasan et al. (US 2025/0111169 A1), hereinafter referred to as Srinivasan, in view of Lucas (US 2025/0190801 A1), and further in view of Li et al. (Li, R., Patel, T., & Du, X. (2023). Prd: Peer rank and discussion improve large language model based evaluations. arXiv preprint arXiv:2307.02762.), hereinafter referred to as Li.
Regarding claim 1, Srinivasan teaches:
A computer-implemented method comprising:
determining one or more evaluation criteria for evaluating responses to the prompt (para [0055], where criteria for ranking are predetermined);
distributing, in parallel and simultaneously, the prompt and the one or more evaluation criteria to the contestant LLMs (Fig. 1 element 104, 108, para [0026], where an array of LLMs execute a prompt in parallel);
generating and evaluating the responses to the prompt and based on the one or more evaluation criteria, generating rankings of the responses (para [0054-55], where the interim outputs are provided and ranked based on the criteria),
determining a top-ranked response included in the responses by aggregating the rankings, and identifying a winning LLM among the contestant LLMs based on the winning LLM having generated the top-ranked response (para [0055], where a highest aggregate response score is determined, and para [0067], where feedback is provided to each LLM indicating why they were chosen or not as the most probably ground truth response); and
sending the top-ranked response and an identification of the winning LLM to the user (para [0028] where LLMs respond to the user, and [0073-77], where the output includes the response and choice of winning LLM).
Srinivasan does not teach:
receiving a prompt from a user and standardizing and adjusting the prompt by pre- processing, tokenizing, and cleaning the prompt, so that the standardized and adjusted prompt is compatible with contestant large language models (LLMs);
wherein the responses and the rankings are generated by the contestant LLMs, respectively;
Lucas teaches:
receiving a prompt from a user and standardizing and adjusting the prompt by pre- processing, tokenizing, and cleaning the prompt, so that the standardized and adjusted prompt is compatible with contestant large language models (LLMs) (para [0022], where a prompt from a user is received, and Fig. 3, para [0041-43], where a prompt is pre-processed including filtering and denoising, and tokenized for input into language models);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Srinivasan by using the preprocessing of Lucas (Lucas para [0041-43]) on the prompts of Srinivasan (Srinivasan Fig. 1 element 120), in order to improve quality and security of the outputs (Lucas para [0039]).
Li teaches:
wherein the responses and the rankings are generated by the contestant LLMs, respectively (Page 3 Figure 1, where each LLM model acts both as reviewers and contestants);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Srinivasan in view of Lucas by using the LLMs of Srinivasan in view of Lucas (Srinivasan para [0026]) as both contestants and reviewers as taught by Li (Li page 3 Figure 1), in order to mitigate biases in automated evaluations while still benefiting from strong capability in reading and writing reviews (Li page 2 first paragraph).
Regarding claim 4, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
providing ongoing benchmarking of the contestant LLMs by repeatedly using the determining the one or more evaluation criteria, the distributing the prompt and the one or more evaluation criteria, the generating and the evaluating the responses, the generating the rankings of the responses (Srinivasan Fig. 5, para [0091], where another prompt is received after processing the first prompt); and
based on the ongoing benchmarking, providing an evaluation and a ranking of subsequent responses while ensuring an accuracy of the evaluation and the ranking of the subsequent responses, even though one or more contestant LLMs have improved after an evaluation and a ranking of previous responses (Srinivasan Fig. 6, para [0100-101], where feedback and self learning is performed for each LLM, followed by subsequent processing and ranking of responses).
Regarding claim 5, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
requesting a contestant LLM included in the contestant LLMs to generate the one or more evaluation criteria (Lucas para [0017], where the evaluation metric is calculated by the LLM); and
generating the one or more evaluation criteria by the contestant LLM (Lucas para [0017], where the evaluation metric is calculated by the LLM),
wherein the responses and the rankings being generated by the contestant LLMs, and the one or more evaluation criteria being generated by the contestant LLM provides a benchmarking process for LLMs that eliminates human bias and enhances accuracy and fairness (Srinivasan para [0054-55], where the interim outputs are provided and ranked based on predetermined criteria, and Li page 3 Figure 1, where using an LLM as a reviewer eliminates human bias and enhances accuracy and fairness).
Regarding claim 6, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
determining a context of the prompt, wherein the generating and the evaluating the responses and the determining the top-ranked response are based on the context of the prompt, which ensures that the winning LLM is contextually relevant to the prompt (Lucas para [0040], where contextual information related to the prompt is used to augment the prompt).
Regarding claim 7, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
generating scoring criteria tailored to specifics of the prompt, wherein the evaluating the responses and the generating the rankings of the responses includes using the scoring criteria, which provides an accuracy in a comprehensive evaluation of performances of the contestant LLMs (Srinivasan para [0048-49], [0054-55], where the interim outputs are provided and ranked based on predetermined criteria relevant to the application).
Regarding claim 8, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
generating respective confidence levels and respective explanations for the rankings, wherein a confidence level included in the confidence levels and an explanation included in the explanations are associated with a ranking of the top-ranked response, and wherein the sending the top-ranked response and the identification of the winning LLM to the user includes sending to the user the confidence level and the explanation (Srinivasan para [0016], where the rank indicates a level of confidence of the response representing a ground truth, and para [0074-77], where the reasoning engine explains the results in the output).
Regarding claim 9, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
determining that an initial performance of the generating and the evaluating the responses to the prompt, the generating the rankings of the responses, and determining the top-ranked response results in a tie in top rankings of a first response and a second response included in the responses (Srinivasan para [0079-80], where the LLMs share the same ranking); and
performing one or more subsequent iterations of the generating and the evaluating the responses, the generating the rankings of the responses, and the determining the top-ranked response until the top-ranked response is determined without a ranking of another response being tied with a ranking of the top-ranked response (Srinivasan para [0073-77], where the GPT is ranked higher than BARD in a different iteration).
Regarding claim 10, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
verifying that the standardized and adjusted prompt is compatible with the contestant LLMs, wherein the determining the one or more evaluation criteria is performed in response to the verifying (Lucas para [0041-42], where the tokenizing and pre-processing are suitable for input to the LM, and Srinivasan para [0032], where routing a prompt to an appropriate LLM is interpreted as the verification, followed by the criteria determination in para [0055]).
Regarding claim 11, Srinivasan in view of Lucas and Li teaches:
The method of claim 1, further comprising:
maintaining an active register of LLMs which are available to be the contestant LLMs in an evaluation of responses to prompts, wherein the register includes model annotations that are continuously updated with classification information that specifies types of prompts associated with the LLMs in the active register (Srinivasan para [0032], [0065], where an array of LLMs is used and where routing of prompts is performed based on analysis of the prompt and which LLMs are most likely or capable of providing a most probably ground truth response, and para [0101], where the learning input is fed back to the prompt router);
determining a type of the prompt (Srinivasan para [0065], where the prompt is analyzed and classified); and
selecting the contestant LLMs from the active register of LLMs based on the active register specifying an association between the type of the prompt and each of the contestant LLMs in the active register (Srinivasan para [0065], where a subset of LLMs is selected based on the category of the prompt).
Regarding claim 12, Srinivasan teaches:
A computer system comprising:
a processor set (Fig. 2 element 212, para [0036], where a processor is used);
one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media (para [0005], [0103], where computer readable storage media is used to store instructions) to cause the processor set to perform operations comprising:
determining one or more evaluation criteria for evaluating responses to the prompt (para [0055], where criteria for ranking are predetermined);
distributing, in parallel and simultaneously, the prompt and the one or more evaluation criteria to the contestant LLMs (Fig. 1 element 104, 108, para [0026], where an array of LLMs execute a prompt in parallel);
generating and evaluating the responses to the prompt and based on the one or more evaluation criteria, generating rankings of the responses (para [0054-55], where the interim outputs are provided and ranked based on the criteria),
determining a top-ranked response included in the responses by aggregating the rankings, and identifying a winning LLM among the contestant LLMs based on the winning LLM having generated the top-ranked response (para [0055], where a highest aggregate response score is determined, and para [0067], where feedback is provided to each LLM indicating why they were chosen or not as the most probably ground truth response); and
sending the top-ranked response and an identification of the winning LLM to the user (para [0028] where LLMs respond to the user, and [0073-77], where the output includes the response and choice of winning LLM).
Srinivasan does not teach:
receiving a prompt from a user and standardizing and adjusting the prompt by pre-processing, tokenizing, and cleaning the prompt, so that the standardized and adjusted prompt is compatible with contestant large language models (LLMs);
wherein the responses and the rankings are generated by the contestant LLMs, respectively;
Lucas teaches:
receiving a prompt from a user and standardizing and adjusting the prompt by pre-processing, tokenizing, and cleaning the prompt, so that the standardized and adjusted prompt is compatible with contestant large language models (LLMs) (para [0022], where a prompt from a user is received, and Fig. 3, para [0041-43], where a prompt is pre-processed including filtering and denoising, and tokenized for input into language models);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Srinivasan by using the preprocessing of Lucas (Lucas para [0041-43]) on the prompts of Srinivasan (Srinivasan Fig. 1 element 120), in order to improve quality and security of the outputs (Lucas para [0039]).
Li teaches:
wherein the responses and the rankings are generated by the contestant LLMs, respectively (Page 3 Figure 1, where each LLM model acts both as reviewers and contestants);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Srinivasan in view of Lucas by using the LLMs of Srinivasan in view of Lucas (Srinivasan para [0026]) as both contestants and reviewers as taught by Li (Li page 3 Figure 1), in order to mitigate biases in automated evaluations while still benefiting from strong capability in reading and writing reviews (Li page 2 first paragraph).
Regarding claim 15, Srinivasan in view of Lucas and Li teaches:
The computer system of claim 12, wherein the operations further comprise:
providing ongoing benchmarking of the contestant LLMs by repeatedly using the determining the one or more evaluation criteria, the distributing the prompt and the one or more evaluation criteria, the generating and the evaluating the responses, the generating the rankings of the responses (Srinivasan Fig. 5, para [0091], where another prompt is received after processing the first prompt); and
based on the ongoing benchmarking, providing an evaluation and a ranking of subsequent responses while ensuring an accuracy of the evaluation and the ranking of the subsequent responses, even though one or more contestant LLMs have improved after an evaluation and a ranking of previous responses (Srinivasan Fig. 6, para [0100-101], where feedback and self learning is performed for each LLM, followed by subsequent processing and ranking of responses).
Regarding claim 16, Srinivasan in view of Lucas and Li teaches:
The computer system of claim 12, wherein the operations further comprise:
requesting a contestant LLM included in the contestant LLMs to generate the one or more evaluation criteria (Lucas para [0017], where the evaluation metric is calculated by the LLM); and
generating the one or more evaluation criteria by the contestant LLM (Lucas para [0017], where the evaluation metric is calculated by the LLM),
wherein the responses and the rankings being generated by the contestant LLMs, and the one or more evaluation criteria being generated by the contestant LLM provides a benchmarking process for LLMs that eliminates human bias and enhances accuracy and fairness (Srinivasan para [0054-55], where the interim outputs are provided and ranked based on predetermined criteria, and Li page 3 Figure 1, where using an LLM as a reviewer eliminates human bias and enhances accuracy and fairness).
Regarding claim 17, Srinivasan teaches:
A computer program product comprising:
one or more computer-readable storage media (para [0005], [0103], where computer readable storage media is used to store instructions); and
program instructions stored on the one or more computer-readable storage media (para [0005], [0103], where computer readable storage media is used to store instructions) to perform operations comprising:
determining one or more evaluation criteria for evaluating responses to the prompt (para [0055], where criteria for ranking are predetermined);
distributing, in parallel and simultaneously, the prompt and the one or more evaluation criteria to the contestant LLMs (Fig. 1 element 104, 108, para [0026], where an array of LLMs execute a prompt in parallel);
generating and evaluating the responses to the prompt and based on the one or more evaluation criteria, generating rankings of the responses (para [0054-55], where the interim outputs are provided and ranked based on the criteria),
determining a top-ranked response included in the responses by aggregating the rankings, and identifying a winning LLM among the contestant LLMs based on the winning LLM having generated the top-ranked response (para [0055], where a highest aggregate response score is determined, and para [0067], where feedback is provided to each LLM indicating why they were chosen or not as the most probably ground truth response); and
sending the top-ranked response and an identification of the winning LLM to the user (para [0028] where LLMs respond to the user, and [0073-77], where the output includes the response and choice of winning LLM).
Srinivasan does not teach:
receiving a prompt from a user and standardizing and adjusting the prompt by pre-processing, tokenizing, and cleaning the prompt, so that the standardized and adjusted prompt is compatible with contestant large language models (LLMs);
wherein the responses and the rankings are generated by the contestant LLMs, respectively;
Lucas teaches:
receiving a prompt from a user and standardizing and adjusting the prompt by pre-processing, tokenizing, and cleaning the prompt, so that the standardized and adjusted prompt is compatible with contestant large language models (LLMs) (para [0022], where a prompt from a user is received, and Fig. 3, para [0041-43], where a prompt is pre-processed including filtering and denoising, and tokenized for input into language models);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Srinivasan by using the preprocessing of Lucas (Lucas para [0041-43]) on the prompts of Srinivasan (Srinivasan Fig. 1 element 120), in order to improve quality and security of the outputs (Lucas para [0039]).
Li teaches:
wherein the responses and the rankings are generated by the contestant LLMs, respectively (Page 3 Figure 1, where each LLM model acts both as reviewers and contestants);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Srinivasan in view of Lucas by using the LLMs of Srinivasan in view of Lucas (Srinivasan para [0026]) as both contestants and reviewers as taught by Li (Li page 3 Figure 1), in order to mitigate biases in automated evaluations while still benefiting from strong capability in reading and writing reviews (Li page 2 first paragraph).
Regarding claim 20, Srinivasan in view of Lucas and Li teaches:
providing ongoing benchmarking of the contestant LLMs by repeatedly using the determining the one or more evaluation criteria, the distributing the prompt and the one or more evaluation criteria, the generating and the evaluating the responses, the generating the rankings of the responses (Srinivasan Fig. 5, para [0091], where another prompt is received after processing the first prompt); and
based on the ongoing benchmarking, providing an evaluation and a ranking of subsequent responses while ensuring an accuracy of the evaluation and the ranking of the subsequent responses, even though one or more contestant LLMs have improved after an evaluation and a ranking of previous responses (Srinivasan Fig. 6, para [0100-101], where feedback and self learning is performed for each LLM, followed by subsequent processing and ranking of responses).
Allowable Subject Matter
Claims 2-3, 13-14, and 18-19 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 101, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: the closest prior art of Srinivasan, Lucas, and Li do not teach the limitations of the claims. Specifically, none of the cited prior art teaches determining the style is not in a database by querying the database for the style, and updating the database with each of the cited pieces of information, and updating further over time, in combination with the other limitations. Hence, none of the cited prior art, either alone or in combination thereof, teaches the combination of limitations found in the claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 2025/0328786 para [0023], [0034] teaches determination of metrics, para [0044] teaches outputting the winning LLM, and para [0084] teaches ranking the models and displaying the best.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRYAN S BLANKENAGEL whose telephone number is (571)270-0685. The examiner can normally be reached 8:00am-5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRYAN S BLANKENAGEL/Primary Examiner, Art Unit 2658