Prosecution Insights
Last updated: August 18, 2026
Application No. 18/662,533

METHOD AND APPARATUS FOR SELF-CONSISTENCY BOOSTS CALIBRATION FOR MATH REASONING

Non-Final OA §101§103§112
Filed
May 13, 2024
Examiner
PATEL, SHREYANS A
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
3 (Non-Final)
89%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
364 granted / 411 resolved
+26.6% vs TC avg
Moderate +8% lift
Without
With
+8.5%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
34 currently pending
Career history
457
Total Applications
across all art units

Statute-Specific Performance

§101
26.4%
-13.6% vs TC avg
§103
41.1%
+1.1% vs TC avg
§102
21.4%
-18.6% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 411 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments with respect to 35 U.S.C. 101 Abstract Idea in regards to claims 1-2, 4-9, 11-12 and 14-20 have been considered, however are not found to be persuasive due to the following reasons. Examiner respectfully disagrees with Applicant’s arguments because Under Alice/Mayo Step 2A, the claims are directed to an abstract idea. The claims recites mental processes and mathematical concepts: receiving information, generating candidate answers, organizing the answer into clusters, calculating a calibration score, comparing that score to a threshold, and deciding whether to output an answer or ask the user to rephrase. These are acts of collecting, analyzing, classifying, scoring, and making a rule-based decision. The “calibration score”, “clusters,” “N sample responses,” and “threshold” are mathematical or logical concepts. The processor and LLM merely automate those abstract steps and do not improve how a computer or LLM operates. Under Alice/Mayo Step 2B, the claims do not add significantly more than the abstract idea. The claims use only generic computer components, namely “at least one processor” and an LLM, to perform routine data processing steps. It does not recite a new LLM architecture, a new calibration algorithm, or any technical change to the model. Repeating the LLM input N times, grouping responses, calculating a score, and comparing that score to a threshold are ordinary uses of computer logic and mathematics. Thus, the claims merely applies the abstract idea on generic computer technology and lacks an inventive concept. The claims remain ineligible even under the newer Ex parte Desjardins guidance. Desjardins supports eligibility where the claims improve the machine learning model or computer system itself, such as by improving model operation, reducing storage, reducing complexity, or preserving performance. The claims do not do that. They do not modify the LLM, improve its internal operation, reduce computing resources, or solve a technical problem in machine learning. Instead, they use an existing LLM as a tool to generate text and applies a mental/mathematical confidence value to decide what message to output. Therefore, claims still stand rejected. Applicant's arguments with respect to 35 U.S.C. 103 in regards to claims 1, 11 and 20 have been considered, however are not found to be persuasive due to the following reasons. Wang in view of Mirkovic teach all the newly added limitations. See detailed rejection below. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 19 recites the limitation "The apparatus of claim 1". There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-2, 4-9, 11-12 and 14-20 are rejected under 35 U.S.C. 101. Claims 1, 11 and 20 are directed to an abstract idea. These claims are an abstract idea because it is essentially: (1) collect information (get a query and generate multiple outputs), (2) analyze/organize the information (cluster the outputs), (3) compute a score (calibration score), and (4) make a decision using a rule/threshold (answer vs. ask user to rephrase). Those are “information processing” steps that fall within the USPTO’s recognized groupings of abstract ideas such as mental processes and mathematical concepts (e.g., scoring, comparing, and selecting). The claims does not tie the abstract idea to a practical technological application. It does not claim a specific improvement to computer functionality. It broadly uses an LLM as a tool and then applies a generic decision rule to decide whether to output a response or ask the user to rephrase, more managing the content of information presented to a user than improving a technical system. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device. Dependent claims 2, 4-9, 12 and 14-19 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known in information processing. Claims 2 and 12: organizing information (a mental/process-of-sorting concept) with no technical improvement. Claims 4 and 14: a mathematical/statistical evaluation of grouped data. Claims 5 and 15: a straightforward mathematical calculation (normalization) applied to the abstract scoring idea. Claims 6 and 16: just measuring and evaluating data group sizes—an abstract information-analysis step. Claims 7 and 17: a basic math operation applied to the abstract scoring concept. Claims 8 and 18: a mathematical formula applied to data group counts, without adding a concrete technological improvement. Claims 9 and 19: just a field-of-use limitation. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2, 4-9, 11-12 and 14-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (“Self-Consistency Boosts Calibration for Math Reasoning” Mar. 14, 2024, hereinafter Wang24) in view Mirkovic et al. (US 2006/0271364). Claims 1, 11 and 20, Wang24 teaches a method performed by at least one processor, the method comprising: receiving a natural language input query ([1.] Wang states: “Self-consistency performs clustering over multiple LLM samples before picking one from the largest cluster as the response to each input query”; Wang also teaches “input questions”); generating N sample responses based on the natural language input query by inputting the natural language input query into a large language model (LLM) in a pre-calibrated state N number of times, N being an integer greater than zero ([1.] Wang states: “LLMs lack adequate calibration out of the box.”; [2.] “Wang et al. (2022) initially sample various reasoning paths r1…rN from the LLM given input x with Chain-of-Thought (CoT) prompting”; [4.1] Wang further states: “We use nucleus sampling to obtain N = 16 samples by default for each instance”); organizing the N sample responses into one or more clusters ([1.] Wang states: “Self-consistency performs clustering over multiple LLM samples”; [3.] Wang further states: “After performing self-consistency on input x using LLMθ, we obtain a set of clusters C = {c1, …, c|C|} with each cluster ci comprising ni samples responses with the same answers.”); performing a calibration process on the one or more clusters in which the calibration process determines a calibration score ([3.] Wang states: “We design the following strategies, tailored to the characteristics of these clusters, to estimate the confidence of LLMθ”; [3.] [Eq. 3] Wang states: “Cluster Size: the number of samples (e.g. ni) within a specific cluster (e.g. ci). Again, we compute its proportion relative to the total sample size to normalize the score into the range [0, 1]”; FCS(x,θ) = ni/N); and outputting a response to the natural language input query based on the calibration process ([1.] Wang states: “Self-consistency performs clustering over multiple LLM samples before picking one from the largest cluster as the response to each input query”), the outputted response is one of the N sample responses ([2.] [Eq. 1] Wang teaches “Then, the answers a1 …, aN are extracted from the paths, and the most consistent answer … is selected as final answer a”; see eq. 1; [1.] Wang also states: “picking one from the largest cluster as the response to each input query”), that is entered into the LLM another N number of times to calibrate the LLM ([4.1] Wang teaches “We use nucleus sampling to obtain N = 16 samples by default for each instance”; “Out methods are founded on the principle of self-consistency, which relies on sampling multiple times for prediction”; [1.] Wang also teaches handling low confidence by “keep resampling until a confident response is produced”). The difference between the prior art and the claimed invention is that Wang24 does not explicitly teach wherein based on a determination the calibration score is greater than or equal to a threshold, and wherein based on a determination the calibration score is less than the threshold, the outputted response is an output requesting a user to rephrase the natural language input query. Mirkovic teaches wherein based on a determination the calibration score is greater than or equal to a threshold ([0142] Mirkovic states: “Confidence thresholds (upper and lower bounds) set by the dialogue designer specify the levels at which a candidate move is rejected, requires explicit confirmation by the user, or is accepted”; Mirkovic further states: “In 1408 it is determined whether the highest candidate move score is above the high threshold, T1. If it is, then the move candidate move can simply be accepted”; see Fig. 14), and wherein based on a determination the calibration score is less than the threshold ([0142] Mirkovic teaches “If the highest score is below T2, then the candidate move or moves are taken as a failure of interpretation, and the user is asked for clarification”; see Fig. 14), the outputted response is an output requesting a user to rephrase the natural language input query ([0121] Mirkovic states: “If the confidence is low, either there is no match from that move and the system returns a general ‘did-not-understand’ type response”; [0122] Mirkovic further states: “This help feature produces a node specific help message or hint for the user like ‘if you want to play something, try saying something like: play a Beatles song”; Wang also states: “in this case it can help to give the user a specific hint how to rephrase his request rather than provide a general response to the user”). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Wang24 with teachings of Mirkovic by modifying the self-consistency boosts calibration for math reasonings as taught by Wang24 to include wherein based on a determination the calibration score is greater than or equal to a threshold, and wherein based on a determination the calibration score is less than the threshold, the outputted response is an output requesting a user to rephrase the natural language input query as taught by Mirkovic for the benefit of resolution preventing noun-phrases from being properly resolved until the appropriate device has been determined ([0006] Mirkovic). Claims 2 and 12, Wang24 further teaches the method according to claim 1, wherein the organizing the N sample responses into the one or more clusters comprises organizing each sample response having a same answer into a same cluster ([1] [3] obtaining a set of cluster with each c comprising n sampled responses with the same answer). Claims 4 and 14, Wang24 further teaches the method according to claim 1, wherein the calibration process comprises determining a calibration score based on a number of clusters ([3] obtain a set of cluster C1-c with each cluster Ci comprising ni sampled responses with the same answers; the characteristics of these clusters are to estimate the confidence of LLM). Claims 5 and 15, Wang24 further teaches the method according to claim 4, wherein the calibration score is normalized based on dividing the number of clusters by N ([3] divide the cluster number by the sample size N to normalize the score into the range [0,1]). Claims 6 and 16, Wang24 further teaches the method according to claim 1, wherein the calibration process comprises determining a calibration score based on a cluster size of each of the one or more clusters ([3] three different ways to calibrate a set of cluster; cluster number, cluster size (the number of samples within a specific cluster; computing a proportion relative to the total sample size to normalize the score range), pairwise comparison)). Claims 7 and 17, Wang24 further teaches the method according to claim 6, wherein the cluster size of each of the one or more clusters is normalized by dividing each of the one or more clusters by N ([3] divide the cluster number by the sample size N to normalize the score into the range [0,1]). Claims 8 and 18, Wang24 further teaches the method according to claim 1, wherein the calibration process comprises, for each cluster: determining a cluster size of each cluster from the one or more clusters ([3] cluster size), and determining, for each cluster, the calibration score based on a product of (i) the cluster size of a respective cluster divided by a sum of the cluster size of the respective cluster and the cluster size of a first cluster other than the respective cluster with (ii) the cluster size of the respective cluster divided by a sum of the cluster size of the respective cluster and the cluster size of a second cluster other than the respective cluster ([3] [eq. 3] ni is the number of samples; N is the normalized score; Fcs(x,0)=ni/N). Claim 9, Wang24 the method of claim 1, wherein the input query is a word math problem ([Abstract] math reasoning tasks). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Weissenberger et al. (US 2024/0144192) – Some implementations process structured calendar data of an electronic calendar for a first user, to generate a natural language representation of the structured calendar data. Versions of those implementations further, in response to receiving a query determined to be relevant to the electronic calendar, prime a large language model (LLM) using a priming input (e.g., process the priming input using the LLM), where the priming input is based on the natural language representation of the structured calendar data. Following priming of the LLM using the priming input, some of those versions process, using the LLM, query input that is based on the query, to generate a LLM output and determine, based on the LLM output, a response to the query. The response can include a natural language response that can be rendered. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHREYANS A. PATEL Primary Examiner Art Unit 2653 /SHREYANS A PATEL/Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Show 4 earlier events
Feb 13, 2026
Response Filed
Feb 26, 2026
Final Rejection mailed — §101, §103, §112
Apr 27, 2026
Response after Non-Final Action
May 26, 2026
Request for Continued Examination
May 28, 2026
Response after Non-Final Action
Jun 05, 2026
Non-Final Rejection mailed — §101, §103, §112
Jul 27, 2026
Applicant Interview (Telephonic)
Jul 27, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12659658
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
2y 10m to grant Granted Jun 16, 2026
Patent 12646496
METHODS AND SYSTEMS OF TEXT-CONDITIONED AUDIO-VISUAL SPEECH GENERATION WITH MULTI-MODAL LATENT DIFFUSION MODELS
2y 2m to grant Granted Jun 02, 2026
Patent 12608559
METHOD AND SYSTEM FOR ENHANCING A MUTIMODAL INPUT CONTENT
3y 0m to grant Granted Apr 21, 2026
Patent 12609128
METHOD FOR IMPROVING FAR-FIELD SPEECH INTERACTION PERFORMANCE, AND FAR-FIELD SPEECH INTERACTION SYSTEM
2y 0m to grant Granted Apr 21, 2026
Patent 12586597
ENHANCED AUDIO FILE GENERATOR
3y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
89%
Grant Probability
97%
With Interview (+8.5%)
2y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 411 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month