Prosecution Insights
Last updated: August 18, 2026
Application No. 18/650,651

BENCHMARK CREATOR - AN ARTIFICIAL INTELLIGENCE-BASED APPROACH TO EVALUATING THE KNOWLEDGE OF A LANGUAGE MODEL FOR A DATASET

Non-Final OA §101§103§112
Filed
Apr 30, 2024
Examiner
ANDERSON, SCOTT C
Art Unit
Tech Center
Assignee
Intuit Inc.
OA Round
1 (Non-Final)
58%
Grant Probability
Moderate
1-2
OA Rounds
5m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 58% of resolved cases
58%
Career Allowance Rate
611 granted / 1044 resolved
-1.5% vs TC avg
Strong +31% interview lift
Without
With
+31.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
41 currently pending
Career history
1083
Total Applications
across all art units

Statute-Specific Performance

§101
36.8%
-3.2% vs TC avg
§103
28.8%
-11.2% vs TC avg
§102
14.1%
-25.9% vs TC avg
§112
18.6%
-21.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1044 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This Office action is in reply to application no. 18/650,651, filed 30 April 2024. Claims 1-20 are pending and are considered below. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being incomplete for omitting essential steps, such omission amounting to a gap between the steps. See MPEP § 2172.01. In each independent claim, benchmark questions are removed “in order to generate decontaminated benchmark data”, but it is not positively claimed that any such generating takes place. The following step uses “the decontaminated benchmark data”, but there isn’t any. To overcome this rejection the Examiner suggests bifurcating the “removing” step into two steps: one of removing and one of generating decontaminated benchmark data. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims lie within statutory categories of invention, as each is directed to a method (process) or system (machine). The claim(s) recite(s) creating a set of questions in no particular manner, comparing the questions to other questions, removing some questions in no particular manner but simply for a particular purpose, confirming that the remaining questions meet certain criterial, and measuring a level of performance in no particular manner. This recites human mental work, requiring nothing more than a pen and paper. A person can create a set of questions by writing them down, can compare them to other questions mentally, can remove some of the questions by striking through them, can ask another person the questions to see if the answers are correct, and can determine a score mentally by any means at all. None of this presents any practical difficulty and none requires any technology beyond pen and paper. This judicial exception is not integrated into a practical application because aside from the bare inclusion of a generic computer and nondescript use of AI, discussed below, nothing is done beyond what was set forth above, which does not go beyond generally linking the abstract idea to the technological environment of AI-enabled computers. See MPEP § 2106.05(h). As the claims only manipulate data pertaining to sets of questions, scores and the like, they do not improve the “functioning of a computer” or of “any other technology or technical field”. See MPEP § 2106.05(a). They do not apply the abstract idea “with, or by use of a particular machine”, MPEP § 2106.05(b), as the below-cited Guidance is clear that a generic computer is not the particular machine envisioned. They do not effect a “transformation or reduction of a particular article to a different state or thing”, MPEP § 2106.05(c). First, such data, being intangible, are not a particular article at all. Second, the claimed manipulation is neither transformative nor reductive; as the courts have pointed out, in the end, data are still data. They do not apply the abstract idea “in some other meaningful way beyond generally linking [it] to a particular technological environment”, MPEP § 2106.05(e), as the lack of technical and algorithmic detail in the claims is so as not to go beyond such a general linkage. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional claim limitations, considered individually and in ordered combination, are insufficient to elevate an otherwise-ineligible claim. Claim 11, which has the most, includes a processor and memory comprising instructions for execution by the processor. These elements are recited at a high degree of generality and the specification is clear, ¶ 62, that nothing more than a “general purpose processor” is required, which encompasses a generic computer. It only performs generic computer functions of nondescriptly manipulating information and sharing information with persons and/or other devices. Generic computers performing generic computer functions, without an inventive concept, do not amount to significantly more than the abstract idea. The type of information being manipulated does not impose meaningful limitations or render the idea less abstract. In light of Recentive1, simply making use of known machine learning techniques in a different data environment is insufficient in and of itself. The claim elements when considered in ordered combination – a generic computer performing a chronological sequence of abstract steps while making nondescript use of AI – do nothing more than when they are analyzed individually. The other independent claim is simply a different embodiment but is likewise directed to a generic computer performing, essentially, the same process. The dependent claims further do not amount to significantly more than the abstract idea: claims 2 and 12 simply recite iteration; claims 3, 13, 10 and 20 consist entirely of a mere duplication of parts. Claims 4-6, 8, 9, 14-16, 18 and 19 are simply further descriptive of the type of information being manipulated, and claims 7 and 17 simply recite additional, abstract manipulation of data. The claims are not patent eligible. For further guidance please see MPEP § 2106.03 – 2106.07(c) (formerly referred to as the “2019 Revised Patent Subject Matter Eligibility Guidance”, 84 Fed. Reg. 50, 55 (7 January 2019)). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 5, 6, 9-13, 15, 16, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. (U.S. Publication No. 2024/0160634, filed 8 November 2022) in view of Agrawal et al. (U.S. Publication No. 2021/0365643). In-line citations are to Kumar. Claims are examined as best understood. With regard to Claim 1: Kumar teaches: A method of automated evaluation of a language processing machine learning model, [0018; natural language processing is performed; 0035; machine learning is used] comprising: creating, using a validated language processing machine learning model, benchmark data comprising benchmark questions based on a dataset; [0054; models and algorithms may be created; 0023; they may be trained within a “question answering system” to provide “accurate determinations of answers to questions” from “heterogeneous data sources”; 0052; proposed answers are compared to known answers to questions, which reads on such known answers to questions being benchmark answers and questions] comparing the benchmark questions to training questions in a training data set used to train a target language processing machine learning model; [0052; it provides scores based on noise found based on the comparisons] removing one or more of the benchmark questions from the benchmark data based on the comparing in order to generate decontaminated benchmark data… [0033; the data may be “denoised”, which reads on this] testing the decontaminated benchmark data to determine whether a question-testing machine learning model can provide correct answers to input benchmark questions from the decontaminated benchmark data without being provided with the dataset as an input; [0057; testing is performed on “testing data that is near, but not identical to training data; 0023; it is “able to reliably and accurately determine a correct answer to a question over hybrid contexts and/or from heterogeneous data sources”] and measuring a level of performance of the target language processing machine learning model using the benchmark data. [0030; relevance scores are computed] Kumar does not explicitly teach confirming that the decontaminated benchmark data corresponds to a threshold proportion of information within the dataset, but it is known in the art. Agrawal teaches a natural language processing system [title] that uses “machine learning techniques”. [0049] A “low-pass filter” is provided to data “to remove as much noise as possible”. [0077] It may remove any data “that do not contribute more than a threshold minimum proportion to variation” in order to “reduce simulation time by orders of magnitude”. [0148] A measure of textual similarity such as a “moving average” may be employed. [0077] Agrawal and Kumar are analogous art as each is directed to electronic means for performing natural language processing using machine learning techniques. It would have been obvious to one of ordinary skill in the art just prior to the filing of the claimed invention to combine the teaching of Agrawal with that of Kumar in order to reduce processing time, as taught by Agrawal; further, it is simply a substitution of one known part for another with predictable results, simply characterizing data in the manner of Agrawal rather than, or in addition to, that of Kumar; the substitution produces no new and unexpected result. With regard to Claim 2: The method of Claim 1, further comprising retraining the target language processing machine learning model based on the measured level of performance of the target language processing machine learning model failing to meet a performance threshold. [Agrawal, as cited above in regard to claim 1; Kumar; 0049; retraining is performed; this is simply a matter of replacing the data of Kumar with that of Agrawal, with no new and unexpected result inherent or disclosed] With regard to Claim 3: The method of Claim 1, wherein multiple language processing machine learning models, including the target language processing machine learning model, are trained using the training data set, and wherein the target language processing machine learning model is selected from the multiple language processing machine learning models for use based on the measured level of performance of the target language processing machine learning model meeting a performance threshold. This claim is not patentably distinct from claim 1 as it consists entirely of a mere duplication of parts, simply performing steps on two sets of data rather than one, with no new and unexpected result inherent or disclosed. See MPEP § 2144.04(VI)(B). With regard to Claim 5: The method of Claim 1, wherein comparing the benchmark questions to training questions in a training data set comprises determining a level of textual similarity between a training question of the training questions and a benchmark question of the benchmark questions. [Agrawal, 0077 as cited above in regard to claim 1] With regard to Claim 6: The method of Claim 1, wherein comparing the benchmark questions to training questions in a training data set comprises determining a level of semantic similarity between a training question of the training questions and a benchmark question of the benchmark questions. [id.] With regard to Claim 9: The method of Claim 1, wherein measuring the level of performance of the target language processing machine learning model using the decontaminated benchmark data is based on using the target language processing machine learning model to generate answers to questions in the decontaminated benchmark data and scoring the generated answers. [0035; weighted scores are associated with candidate answer items which have been generated] With regard to Claim 10: The method of Claim 9, wherein measuring the level of performance of the target language processing machine learning model comprises generating multiple sets of answers to the questions in the decontaminated benchmark data, wherein the generated multiple sets of answers are used to determine a level of consistency of the target language processing machine learning model. This claim is not patentably distinct from claim 9. Generating multiple sets of answers, as opposed to generating answers (as in the earlier claim) is at most a mere duplication of parts which is considered but given no patentable weight. The purpose of the answers (what they are used for) consists entirely of manner-of-use language which is considered but given no patentable weight; there is no positively claimed step of determining a level of consistency. With regard to Claim 11: Kumar teaches: A system for automated evaluation of a language processing machine learning model, comprising: one or more processors; and a memory comprising instructions that, when executed by the one or more processors, cause the system to: create, using a validated language processing machine learning model, benchmark data comprising benchmark questions based on a dataset; [0054; models and algorithms may be created; 0023; they may be trained within a “question answering system” to provide “accurate determinations of answers to questions” from “heterogeneous data sources”; 0052; proposed answers are compared to known answers to questions, which reads on such known answers to questions being benchmark answers and questions] compare the benchmark questions to training questions in a training data set used to train a target language processing machine learning model; [0052; it provides scores based on noise found based on the comparisons] remove one or more of the benchmark questions from the benchmark data based on the comparing in order to generate decontaminated benchmark data… [0033; the data may be “denoised”, which reads on this] test the decontaminated benchmark data to determine whether a question-testing machine learning model can provide correct answers to input benchmark questions from the decontaminated benchmark data without being provided with the dataset as an input; [0057; testing is performed on “testing data that is near, but not identical to training data; 0023; it is “able to reliably and accurately determine a correct answer to a question over hybrid contexts and/or from heterogeneous data sources”] and measure a level of performance of the target language processing machine learning model using the benchmark data. [0030; relevance scores are computed] Kumar does not explicitly teach confirm that the decontaminated benchmark data corresponds to a threshold proportion of information within the dataset, but it is known in the art. Agrawal teaches a natural language processing system [title] that uses “machine learning techniques”. [0049] A “low-pass filter” is provided to data “to remove as much noise as possible”. [0077] It may remove any data “that do not contribute more than a threshold minimum proportion to variation” in order to “reduce simulation time by orders of magnitude”. [0148] A measure of textual similarity such as a “moving average” may be employed. [0077] Agrawal and Kumar are analogous art as each is directed to electronic means for performing natural language processing using machine learning techniques. It would have been obvious to one of ordinary skill in the art just prior to the filing of the claimed invention to combine the teaching of Agrawal with that of Kumar in order to reduce processing time, as taught by Agrawal; further, it is simply a substitution of one known part for another with predictable results, simply characterizing data in the manner of Agrawal rather than, or in addition to, that of Kumar; the substitution produces no new and unexpected result. With regard to Claim 12: The system of Claim 11, further comprising retraining the target language processing machine learning model based on the measured level of performance of the target language processing machine learning model failing to meet a performance threshold. [Agrawal, as cited above in regard to claim 1; Kumar; 0049; retraining is performed; this is simply a matter of replacing the data of Kumar with that of Agrawal, with no new and unexpected result inherent or disclosed] With regard to Claim 13: The system of Claim 11, wherein multiple language processing machine learning models, including the target language processing machine learning model, are trained using the training data set, and wherein the target language processing machine learning model is selected from the multiple language processing machine learning models for use based on the measured level of performance of the target language processing machine learning model meeting a performance threshold. This claim is not patentably distinct from claim 11 as it consists entirely of a mere duplication of parts, simply performing steps on two sets of data rather than one, with no new and unexpected result inherent or disclosed. See MPEP § 2144.04(VI)(B). With regard to Claim 15: The system of Claim 11, wherein comparing the benchmark questions to training questions in a training data set comprises determining a level of textual similarity between a training question of the training questions and a benchmark question of the benchmark questions. [Agrawal, 0077 as cited above in regard to claim 1] With regard to Claim 16: The system of Claim 11, wherein comparing the benchmark questions to training questions in a training data set comprises determining a level of semantic similarity between a training question of the training questions and a benchmark question of the benchmark questions. [id.] With regard to Claim 19: The system of Claim 11, wherein measuring the level of performance of the target language processing machine learning model using the decontaminated benchmark data is based on using the target language processing machine learning model to generate answers to questions in the decontaminated benchmark data and scoring the generated answers. [0035; weighted scores are associated with candidate answer items which have been generated] With regard to Claim 20: The system of Claim 19, wherein measuring the level of performance of the target language processing machine learning model comprises generating multiple sets of answers to the questions in the decontaminated benchmark data, wherein the generated multiple sets of answers are used to determine a level of consistency of the target language processing machine learning model. This claim is not patentably distinct from claim 19. Generating multiple sets of answers, as opposed to generating answers (as in the earlier claim) is at most a mere duplication of parts which is considered but given no patentable weight. The purpose of the answers (what they are used for) consists entirely of intended-use language which is considered but given no patentable weight; there is no positively claimed step of determining a level of consistency. Claim(s) 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. in view of Agrawal et al. further in view of Brisebois et al. (U.S. Patent No. 9,563,782). These claims are similar so are analyzed together. With regard to Claim 4: The method of Claim 1, wherein confirming that the decontaminated benchmark data corresponds to the threshold proportion of information within the dataset comprises using a question-evaluating machine learning model to determine that a threshold number of topics within the dataset are represented by the decontaminated benchmark data. With regard to Claim 14: The system of Claim 11, wherein confirming that the decontaminated benchmark data corresponds to the threshold proportion of information within the dataset comprises using a question-evaluating machine learning model to determine that a threshold number of topics within the dataset are represented by the decontaminated benchmark data. Kumar and Agrawal teach the method of claim 1 and system of claim 11, including the claimed determining of correspondence but do not explicitly teach determination of a number of topics, but it is known in the art. Brisebois teaches secure self-service access to content [title] that uses “machine learning techniques”. [Col. 11, line 63] Criteria for a knowledge manager may include a “threshold number” of “topic-relevant conversations” [Col. 56, lines 24-26] This may include determining a proportion of conversations relating to a specific number of topics, e.g. two or more. [Col. 15, lines 34-36] Brisebois and Kumar are analogous art as each is directed to electronic means for using machine learning to process textual communications. It would have been obvious to one of ordinary skill in the art just prior to the filing of the claimed invention to combine the teaching of Brisebois with that of Kumar and Agrawal in order to improve trust, as taught by Brisebois; [Col. 58, lines 54-55] further, it is simply a substitution of one known part for another with predictable results, simply using Brisebois’ basis for a determination rather than, or in addition to, those of Kumar; the substitution produces no new and unexpected result. Claim(s) 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kumar et al. in view of Agrawal et al. further in view of Toyama (U.S. Publication No. 2022/0309939). These claims are similar so are analyzed together. With regard to Claim 8: The method of Claim 1, wherein the training data set further comprises multiple choice answers comprising one correct answer and at least one incorrect answer to each of the training questions, wherein the benchmark data further comprises multiple choice answers comprising one correct answer and at least one incorrect answer to each of the benchmark questions. With regard to Claim 18: The system of Claim 11, wherein the training data set further comprises multiple choice answers comprising one correct answer and at least one incorrect answer to each of the training questions, wherein the benchmark data further comprises multiple choice answers comprising one correct answer and at least one incorrect answer to each of the benchmark questions. Kumar and Agrawal teach the method of claim 1 and system of claim 11, but do not explicitly teach using multiple choice answers, and though it is of no patentable significance as explained below, it is known in the art. Toyama teaches a text processing system [abstract] in which choices of answers are displayed “in a multiple choice question” [0026] in which the answers “include a correct word and similar words similar to the correct word”. [0041] Toyama and Kumar are analogous art as each is directed to electronic means for managing questions and answers. It would have been obvious to one of ordinary skill in the art just prior to the filing of the claimed invention to combine the teaching of Toyama with that of Kumar and Agrawal in order to present more relevant information, as taught by Toyama; [0006] further, it is simply a substitution of one known part for another with predictable results, simply using Toyama’s answer format in place of that of Kumar; the substitution produces no new and unexpected result. These claims are not patentably distinct from their respective parent claims as they consist entirely of nonfunctional, descriptive language, disclosing at most human interpretation of data, but which imparts neither structure nor functionality to any claimed embodiment. The reference is provided for the purpose of compact prosecution. Conclusion No “art rejection” is made herein under 35 U.S.C. § 103 of claims 7 or 17. The prior art cited does not teach the limitations of these claims, and no prior art has been found that teaches the claimed limitations and would reasonably serve to make an obvious combination with the art already cited. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCOTT C ANDERSON whose telephone number is (571)270-7442. The examiner can normally be reached M-F 9:00 to 5:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bennett Sigmond can be reached at (303) 297-4411. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SCOTT C ANDERSON/Primary Examiner, Art Unit 3694 1 Recentive Analytics, Inc. v. Fox Corp. et al., 134 F.4th 1205, 1216 (Fed. Cir. 2025)
Read full office action

Prosecution Timeline

Apr 30, 2024
Application Filed
Jul 16, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705597
POST-PURCHASE CREDIT OFFER AND TENDER SWITCH
1y 7m to grant Granted Aug 11, 2026
Patent 12700020
SYSTEM AND METHOD FOR FUNDING A VIRTUAL LOCATION
4y 1m to grant Granted Aug 04, 2026
Patent 12694456
AUTOMATIC GENERATION OF OPTIMIZED AGGREGATION METRICS FOR USAGE BASED INSURANCE
2y 5m to grant Granted Jul 28, 2026
Patent 12687931
Enhanced Systems and Methods for Multi-Platform Advertising Using Holographic Displays, Biometric Integration, Quantum Technologies, and Device Synchronization
1y 4m to grant Granted Jul 21, 2026
Patent 12682364
DETECTING UNAUTHORIZED ONLINE APPLICATIONS USING MACHINE LEARNING
1y 12m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
58%
Grant Probability
90%
With Interview (+31.4%)
2y 9m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1044 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month