Prosecution Insights
Last updated: October 02, 2026
Application No. 19/176,506

SYSTEM AND METHOD FOR PREVENTING HALLUCINATIONS

Non-Final OA §101§103§112
Filed
Apr 11, 2025
Priority
Apr 12, 2024 — provisional 63/633,608
Examiner
LEVENTAL, YUVAL HAIM
Art Unit
Tech Center
Assignee
Sri International
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
9 currently pending
Career history
5
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending and have been examined. Claims 1, 8, and 15 are independent. This Application was published as U.S. 2025/0322178 A1. Apparent priority: April 12, 2024. Information Disclosure Statement The information disclosure statement (IDS) submitted on January 15, 2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The disclosure is objected to because of the following informalities: a. In paragraph [0004], "Such training, however can be time consuming" should apparently read "Such training, however, can be time consuming." b. In paragraph [0014], the sentence "FIG. 4 depicts a computing device suitable for use with embodiments of a reasoning system in accordance with the present principles" should apparently end with a period. c. In paragraph [0026], "of the determined next word next word," should apparently read "of the determined next word,". d. In paragraph [0027], "such as an measure of inconsistency" should apparently read "such as a measure of inconsistency." Appropriate correction is required. Claim Objections Claims 1, 2, 8, 9, 15, and 16 are objected to because of the following informalities: a. In claims 1, 8, and 15, the clause "determining a measure of uncertainty for the generated token:" ("determine a measure of uncertainty for the generated token:" in claims 8 and 15) ends with a colon and should apparently end with a semicolon, consistent with the other clauses of the claim. b. In claims 2, 9, and 16, "to cause the language to perform at least one other additional computation" should apparently read "to cause the language model to perform at least one other additional computation." Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 (Statutory Category). Claims 1-7 are directed to a method (a process), claims 8-14 are directed to an apparatus (a machine), and claims 15-20 are directed to a system (a machine). Claims 1-20 are therefore directed to a statutory category of invention. Step 2A, Prong One (Recitation of a Judicial Exception). Claim 1 recites: "A method for preventing hallucinations in a language model, comprising: monitoring a generation of a token by the language model; determining a measure of uncertainty for the generated token: comparing the determined measure of uncertainty with an expected measure of uncertainty; generating at least one think token if the determined measure of uncertainty does not comply with the expected measure of uncertainty; and communicating the at least one generated think token to the language model to cause the language model to perform at least one additional computation for determining the token." The limitations "monitoring a generation of a token by the language model," "determining a measure of uncertainty for the generated token," "comparing the determined measure of uncertainty with an expected measure of uncertainty," and "generating at least one think token if the determined measure of uncertainty does not comply with the expected measure of uncertainty," as drafted, are a process that, under its broadest reasonable interpretation, covers performance in the human mind, or by a human using pen and paper, but for the recitation of a language model. Restated in ordinary words, the claim describes a person who watches a language model produce its next token, works out from the model's output how uncertain the model is about that token (the specification gives the entropy of the model's output probability as one such measure, paragraph [0026]), compares that uncertainty with the level the person expects, and, if the uncertainty is too high, decides that more thought is needed before the token is accepted and writes down a note to that effect. Observing an output, evaluating it against a standard, and forming a judgment on the result are mental processes, and the determination of a measure of uncertainty such as entropy from output probabilities is a mathematical calculation. Claim 1 therefore recites an abstract idea in the mental processes and mathematical concepts groupings. Claims 8 and 15 recite the same operations, performed by an apparatus and a system respectively, and recite the same abstract idea. Step 2A, Prong Two (Integration into a Practical Application). The additional elements recited in claim 1 are the language model and the step of communicating the at least one generated think token to the language model to cause the language model to perform at least one additional computation for determining the token. Claims 8 and 15 additionally recite a processor and a memory having stored therein programs or instructions executable by the processor. The language model is recited at the highest level of generality, as any language model (paragraph [0024]), and the processor and memory are the processors 410 and system memory 420 of the computing device 400 of FIG. 4, which the specification describes as any suitable general purpose or embedded processors and any suitable memory technology in a general purpose computer (paragraphs [0055]-[0057] and [0064]). Communicating the think token to the language model to cause an additional computation amounts to instructing the model to compute again when the evaluation calls for it; the specification describes the think token as a generated token that causes the language model to pause from its normal routine of generating tokens and perform at least one additional computation before generating a response (paragraphs [0020] and [0028]), and describes no change to the structure or operation of the language model itself. This is a recitation of the words "apply it" with a generic computer component: the language model and the processor and memory are used as tools on which the abstract idea is carried out, and the claim does not improve the functioning of the language model or of the computer. The additional elements, considered individually and in combination, do not impose any meaningful limits on practicing the abstract idea and do not integrate it into a practical application. Claims 1, 8, and 15 are directed to the abstract idea. Step 2B (Inventive Concept). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are a generic language model and generic computer components used to apply the abstract idea, and mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Considered individually and as an ordered combination, the additional elements do not add significantly more than the abstract idea, and claims 1, 8, and 15 are not patent eligible. The dependent claims have been considered and do not cure the deficiencies of the independent claims. Claims 2, 9, and 16 recite performing the monitoring, determining, comparing, and generating operations again on the token produced by the additional computation and communicating another think token. This is a repetition of the same mental evaluation and of the same instruction to the language model to compute again; no additional element beyond those addressed above is recited. Claims 3, 10, and 17 recite that the measure of uncertainty is at least one of a measure of entropy or a measure of inconsistency. This further characterizes the mathematical calculation and the mental evaluation and adds no additional element. Claims 4 and 11 recite that the expected measure of uncertainty comprises a predetermined threshold value of uncertainty. Comparing a value with a threshold is a mental comparison; no additional element is added. Claims 5, 12, and 18 recite that the additional computation comprises a tokenization computation using the just previously determined token and a just previously implemented hidden state. This describes the ordinary next token computation of a language model, which is the generic operation of the model itself, and does not change the manner in which the abstract idea is applied. Claims 6 and 13 recite the kinds of token that may be monitored. Limiting the data on which the mental evaluation is performed to a portion of a word, a word, a phrase, or a portion of an image or video is a field of use limitation and adds no additional element. Claims 7, 14, and 19 recite that the language model is trained to perform the additional computation every time the token is being generated based on a respective think token. Training a model is a generic use of a computer, recited at a high level of generality, and does not integrate the abstract idea into a practical application. Claim 20 recites that the language model comprises a large language model, which is a generic recitation of the computer component on which the abstract idea is applied. None of the dependent claims recites an additional element that integrates the abstract idea into a practical application or that amounts to significantly more than the abstract idea. Claims 2-7, 9-14, and 16-20 are therefore not patent eligible. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.-The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 9-14 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Each of claims 9-14 recites "The apparatus of claim 1." Claim 1 is a method claim and recites no apparatus, so there is insufficient antecedent basis for "the apparatus" in each of claims 9-14. Because the only apparatus recited in the claims is the apparatus of independent claim 8, and each of claims 9-14 further limits the apparatus in a manner paralleling claims 2-7, claims 9-14 are examined as depending from claim 8. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 4, 6-9, 11, 13-16, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Schuster (U.S. 2024/0020516 A1), hereinafter Schuster, in view of Zelikman et al., "Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking," arXiv:2403.09629v1 [cs.CL], March 14, 2024, https://arxiv.org/abs/2403.09629v1, hereinafter Zelikman. Regarding claim 1, Schuster discloses: 1. A method for preventing hallucinations in a language model, comprising: "Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks. These gains come with a drastic increase in the models' size, potentially leading to slow and costly use at inference time. In practice, however, the series of generations made by LLMs when generating an output sequence is composed of varying levels of difficulty." (Schuster, para. [0009].) monitoring a generation of a token by the language model; "In other words, the decoder neural network 110 is referred to as an auto-regressive neural network because the neural network 110 auto-regressively generates an output sequence of tokens by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular text token in the output sequence" (Schuster, para. [0045].) "When the layer is in the subset, the system generates a confidence score for the layer from at least the updated respective input hidden state for the last input in the current input sequence generated by the layer (step 306)." (Schuster, para. [0107].) Generating a confidence score for the token being generated, at each layer while the token is being generated, is the recited monitoring of the generation of the token. determining a measure of uncertainty for the generated token: The recited measure of uncertainty is interpreted as any quantity indicating how certain or uncertain the language model is about the token being generated. Under the broadest reasonable interpretation of the claim, a confidence score computed from the probability distribution over the token being generated is such a quantity: the less confident the score, the more uncertain the model is about the token. "The system can then determine the confidence score based on a difference between the highest probability in the probability distribution and the second highest probability in the probability distribution. For example, the confidence score can be equal to the difference or equal to the output of an increasing function applied to the difference." (Schuster, para. [0110].) comparing the determined measure of uncertainty with an expected measure of uncertainty; The recited expected measure of uncertainty is interpreted as a value of the measure that the token is expected to satisfy; claim 4 recites that it comprises a predetermined threshold value. A threshold value to which the confidence score is compared is such a value. "The system determines that the termination criterion is satisfied when the confidence score for the layer is greater than or equal to the threshold value for the layer for the output time step (step 308)" (Schuster, para. [0116].) generating at least one think token if the determined measure of uncertainty does not comply with the expected measure of uncertainty; and "and determines that the termination criterion is not satisfied when the confidence score for the layer is less than the threshold value for the layer for the output time step." (Schuster, para. [0116].) communicating the at least one generated think token to the language model to cause the language model to perform at least one additional computation for determining the token. "The system processes the respective hidden states for the inputs in the current input sequence through the layers in the sequence of layers until a termination criterion is satisfied (step 206)." (Schuster, para. [0079].) When the confidence score for a layer is below the threshold, the processing of the next layer in the sequence for the same token is the additional computation for determining the token. Schuster does not teach the think-token aspects of the claim. Zelikman discloses the think-token elements, mapped limitation by limitation: The recited think token is interpreted, consistent with the specification (paragraph [0022]), as a token that is provided to the language model to cause the language model to perform further computation before it determines the output token. A learned start-of-thought token that puts the model into a thinking mode is such a token. generating at least one think token if the determined measure of uncertainty does not comply with the expected measure of uncertainty ("We insert learned <|startofthought|> and <|endofthought|> tokens to mark each rationale’s start and end." (Zelikman, Sec. 4.1, p. 5.) "The <|startofthought|> and <|endofthought|> tokens serve as learned meta-tokens that control the model’s rationale generation ... Intuitively, the start thought tokens can be understood as putting the model into a “thinking mode” and the end thought token can be understood as telling the model when it’s done thinking." (Zelikman, Sec. 4.4.1, p. 6.) The inserted start-of-thought token is the generated think token.) communicating the at least one generated think token to the language model to cause the language model to perform at least one additional computation for determining the token ("Broadly, Quiet-STaR proceeds by generating rationales after every token to explain future text (think), mixing the future-text predictions with and without rationales (talk), and then learning to generate better rationales using REINFORCE (learn)." (Zelikman, Sec. 1, p. 2.) "Given the end-of-thought token’s hidden state and the hidden state of the original text token, the mixing head outputs a weight that determines the extent to which the post-thought prediction logits will be used." (Zelikman, Sec. 4.3, p. 6.) The rationale generated by the language model from the inserted start-of-thought token, and the post-thought prediction of the next token from it, are the additional computation performed by the language model for determining the token.) Schuster and Zelikman pertain to the generation of tokens by auto-regressive Transformer language models and to controlling the computation the model performs for each token. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Schuster so that, when the confidence score for a token is below the threshold value, the system inserts the learned start-of-thought token of Zelikman and has the language model generate a rationale and a post-thought prediction before the token is determined, in place of or in addition to processing the further layers of the decoder. One of ordinary skill in the art would have been motivated to make this modification in order to improve the prediction of the tokens the model finds difficult, which are exactly the tokens for which Schuster's confidence score falls below the threshold, as Zelikman expressly teaches: "Encouragingly, generated rationales disproportionately help model difficult-to-predict tokens and improve the LM’s ability to directly answer difficult questions." (Zelikman, Abstract, p. 1.) Regarding claim 2, Schuster discloses: 2. The method of claim 1, further comprising: monitoring the token generated by the at least one additional computation; determining a measure of uncertainty for the token generated by the at least one additional computation; comparing the measure of uncertainty determined for the token generated by the at least one additional computation with an expected measure of uncertainty; generating at least one other think token if the measure of uncertainty determined for the token generated by the at least one additional computation does not comply with the expected measure of uncertainty; and communicating the at least one generated other think token to the language model to cause the language to perform at least one other additional computation for determining the token. The recited token generated by the at least one additional computation is interpreted as the token the language model predicts as the result of the additional computation, whether or not that prediction is the token finally output. In Schuster, each layer in the subset has its own output subnetwork that generates a probability distribution over the vocabulary from the hidden state the layer produced, and the confidence score for the layer is determined from that distribution; the confidence score is generated and compared against the threshold again at each successive layer processed for the same token, until the termination criterion is satisfied: "The system processes the respective hidden states for the inputs in the current input sequence through the layers in the sequence of layers until a termination criterion is satisfied (step 206)." (Schuster, para. [0079].) "When the layer is in the subset, the system generates a confidence score for the layer from at least the updated respective input hidden state for the last input in the current input sequence generated by the layer (step 306)." (Schuster, para. [0107].) "The respective output subnetwork for each of the layers 120 is configured to, at any given output time step, process the updated hidden state for the last token in the input sequence for the given output time step after being updated by the layer 120 to generate a probability distribution over the tokens in the vocabulary." (Schuster, para. [0058].) The token predicted from the probability distribution generated by the output subnetwork of the next layer processed is the token generated by the additional computation, the confidence score determined from that distribution is the measure of uncertainty determined for it, its comparison against the threshold at step 308 is the recited comparing, and the processing of the layer after it when the score is again below the threshold is the other additional computation. Schuster does not teach the other think token. Zelikman teaches: generating at least one other think token if the measure of uncertainty determined for the token generated by the at least one additional computation does not comply with the expected measure of uncertainty; and communicating the at least one generated other think token to the language model to cause the language to perform at least one other additional computation for determining the token ("Moreover, this parallelized next-sampling token procedure can be repeated arbitrarily many times (or at least, until one runs out of memory)." (Zelikman, Sec. 4.2, p. 6.) Each repetition inserts a further start-of-thought token that the language model then processes for the prediction, which is the recited other think token.) The rationale for the combination is similar to the one provided for claim 1. In the modified system of Schuster, the confidence score is generated and compared with the threshold at each layer processed for the token, so the check is repeated after the additional computation, and the start-of-thought token of Zelikman is inserted at every check at which the score is below the threshold. The further start-of-thought token inserted at the check that follows the additional computation is the other think token, and the rationale and post-thought prediction the language model generates from it are the other additional computation for determining the token. Regarding claim 4, Schuster discloses: 4. The method of claim 1, wherein the expected measure of uncertainty comprises a predetermined threshold value of uncertainty. "In order to employ early exiting, the system 100 maintains a respective threshold value for each of the output time steps for each of a subset of the layers 120." (Schuster, para. [0054].) "FIG. 4 is a flow diagram of an example process 400 for determining a shared threshold value." (Schuster, para. [0117].) The threshold value maintained for each output time step, determined before generation by the process of FIG. 4, is the predetermined threshold value of the claim. Regarding claim 6, Schuster discloses: 6. The method of claim 1, wherein the monitored, generated token comprises at least one of a portion of a word, a word, a phrase, a portion of an image, an image, a portion of a video, or a video. "The tokens in the vocabulary can be any appropriate text tokens, e.g., words, word pieces, characters, punctuation marks, and so on, that represent elements of text in one or more natural languages" (Schuster, para. [0047].) The word pieces and characters are portions of a word, and the words are words, of the claim. Regarding claim 7, Schuster does not teach training the language model on think tokens. Zelikman teaches: 7. The method of claim 1, wherein the language model is trained to perform at least one additional computation every time the token is being generated based on at least one respective, generated think token. Claim 7 is read as follows. Claim 1 conditions the generation of a think token on the measure of uncertainty; claim 7 adds that the language model has been trained so that, each time it generates a token for which a think token has been generated, it performs the additional computation from that think token. Claim 7 is therefore directed to how the language model has been trained to respond to a think token. It does not require that a think token be generated for every token; the condition of claim 1 still governs when one is generated. "Broadly, Quiet-STaR proceeds by generating rationales after every token to explain future text (think), mixing the future-text predictions with and without rationales (talk), and then learning to generate better rationales using REINFORCE (learn)." (Zelikman, Sec. 1, p. 2.) "The gradients from this loss are used to update both the LM parameters and the start-of-thought and end-of-thought token embeddings" (Zelikman, Sec. 4.4.3, p. 7.) Zelikman trains the language model, by updating its parameters and the embeddings of the start-of-thought and end-of-thought tokens, to generate a rationale after every token from the start-of-thought token inserted for that token, so that the trained model performs the additional computation from the think token every time a token is being generated. That is the recited training. In the combination set forth for claim 1, the start-of-thought token is inserted whenever the confidence score is below the threshold, and the language model so trained performs the additional computation from it every time it is inserted. The rationale for the combination is similar to the one provided for claim 1. Claim 8 is an apparatus claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally: Regarding claim 8, Schuster discloses: 8. An apparatus for preventing hallucinations in a language model, comprising: a processor; and a memory coupled to the processor, the memory having stored therein at least one of programs or instructions executable by the processor to configure the apparatus to: monitor a generation of a token by the language model; determine a measure of uncertainty for the generated token: compare the determined measure of uncertainty with an expected measure of uncertainty; generate at least one think token if the determined measure of uncertainty does not comply with the expected measure of uncertainty; and communicate the at least one generated think token to the language model to cause the language model to perform at least one additional computation for determining the token. "The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry." (Schuster, para. [0151].) The central processing unit is the processor of the claim, and the memory devices storing the instructions it executes are the memory coupled to the processor having stored therein instructions executable by the processor. The operations the apparatus is configured to perform are set forth in the claim 1 mapping above. Claim 9 is an apparatus claim with limitations corresponding to the limitations of Claim 2 and is rejected under a similar rationale. Claim 11 is an apparatus claim with limitations corresponding to the limitations of Claim 4 and is rejected under a similar rationale. Claim 13 is an apparatus claim with limitations corresponding to the limitations of Claim 6 and is rejected under a similar rationale. Claim 14 is an apparatus claim with limitations corresponding to the limitations of Claim 7 and is rejected under a similar rationale. Claim 15 is a system claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Claim 15 recites a predetermined threshold where claim 1 recites an expected measure of uncertainty; the threshold value of Schuster to which the confidence score is compared, set forth in the claim 1 mapping above, is the predetermined threshold. Additionally: Regarding claim 15, Schuster discloses: 15. A system for preventing hallucinations in a language model, comprising: a language model; and an apparatus comprising a processor and a memory coupled to the processor, the memory having stored therein at least one of programs or instructions executable by the processor to configure the system to: monitor a generation of a token by the language model; determine a measure of uncertainty for the generated token: compare the determined measure of uncertainty with a predetermined threshold; generate a think token if the determined measure of uncertainty does not comply with the predetermined threshold; and communicate the generated think token to the language model to cause the language model to perform at least one additional computation for determining the token. "Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks." (Schuster, para. [0009].) "The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data." (Schuster, para. [0151].) The large language model is the language model of the claim, and the central processing unit and the memory devices storing the instructions it executes are the apparatus comprising a processor and a memory coupled to the processor. The operations the system is configured to perform are set forth in the claim 1 mapping above. Claim 16 is a system claim with limitations corresponding to the limitations of Claim 2 and is rejected under a similar rationale. Claim 19 is a system claim with limitations corresponding to the limitations of Claim 7 and is rejected under a similar rationale. Regarding claim 20, Schuster discloses: 20. The system of claim 15, wherein the language model comprises a large language model. "Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks." (Schuster, para. [0009].) Claims 3, 10, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Schuster (U.S. 2024/0020516 A1) in view of Zelikman (arXiv:2403.09629v1), and further in view of Sahu et al., "Unpacking Large Language Models with Conceptual Consistency," arXiv:2209.15093v1 [cs.CL], September 29, 2022, https://arxiv.org/abs/2209.15093v1, hereinafter Sahu. Regarding claim 3, Schuster and Zelikman do not teach a measure of entropy or a measure of inconsistency as the measure of uncertainty. Sahu teaches: 3. The method of claim 1, wherein the measure of uncertainty is at least one of a measure of entropy or a measure of inconsistency. Sahu measures how consistent a language model's responses are across conceptually related queries and scores the model from that measure: "This novel metric measures how well a model can be characterized by finding out how consistent its responses to queries about conceptually relevant background knowledge are." (Sahu, Abstract, p. 1.) "At a given threshold t ∈ [0, 1] our predicted task score is" the indicator of whether the background score meets the threshold; "This score predicts the model will be correct when it is 1 or incorrect when it is 0." (Sahu, Sec. 3.3, p. 6.) "In absolute terms average precision is on the lower end, showing significant conceptual inconsistency." (Sahu, Sec. 5, p. 7.) A score that measures the consistency of the model's responses, and whose low values are conceptual inconsistency, is a measure of inconsistency of the claim, and comparing it against the threshold t to predict whether the model's answer will be correct is using it as the measure of uncertainty compared against the expected measure. Schuster, Zelikman, and Sahu pertain to the evaluation and improvement of the outputs of large language models. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use, as the measure of uncertainty compared against the threshold in the combination of Schuster and Zelikman, a measure of the inconsistency of the language model's responses as taught by Sahu. One of ordinary skill in the art would have been motivated to make this modification in order to predict from the model's consistency whether its answer will be correct, as Sahu expressly teaches: "A model’s knowledge of background information is somewhat predictive of its question answering correctness." (Sahu, Sec. 5, p. 7.) Claim 10 is an apparatus claim with limitations corresponding to the limitations of Claim 3 and is rejected under a similar rationale. Claim 17 is a system claim with limitations corresponding to the limitations of Claim 3 and is rejected under a similar rationale. Claims 5, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Schuster (U.S. 2024/0020516 A1) in view of Zelikman (arXiv:2403.09629v1), and further in view of Dehghani et al. (U.S. 2019/0354567 A1), hereinafter Dehghani. Regarding claim 5: 5. The method of claim 1, wherein the at least one additional computation comprises a tokenization computation using the just previously determined token and a just previously implemented hidden state. The recited tokenization computation using the just previously determined token and a just previously implemented hidden state is read, consistent with the FIG. 2 embodiment of the specification (paragraphs [0029]-[0030]), as a computation that determines the token and that takes as its inputs a token the language model determined before the current computation and the hidden state the language model produced in the computation immediately preceding. The just previously determined token is read as the last token the language model output before the token being determined; the claim does not require it to be a tentative value of the token being determined. Schuster and Zelikman teach the additional computation for determining the token, set forth in the claim 1 mapping above, but do not expressly teach that the additional computation is a tokenization computation using the just previously determined token and a just previously implemented hidden state. Dehghani teaches: "The system revises a representation of the next predicted element in the target sequence using two-stage self-attention and a transition function (320)." "The system can autoregressively determine each next symbol in the sequence, which means that each output in the target sequence is conditioned on all of the previously generated outputs in the target sequence." (Dehghani, para. [0046].) "For example, at each revision step t from 1 to T, the system can compute an updated representation Ht. To do so, the system can apply a multihead dot product self-attention mechanism followed by a recurrent transition function. . . . In some implementations, the system computes the updated representation of Ht according to: Ht=LayerNorm(Ai-1+Transition(At)) where At=LayerNorm(Ht-1+MultiHeadSelfAttention(Ht-1+Pt))," (Dehghani, para. [0035].) "If the stop-decoding condition is not met, the system again performs another step of revisions for the next element in the target sequence (branch to 320)." (Dehghani, para. [0049].) "After the N steps of decoding have completed, the system applies a final softmax layer 430 to generate final output probabilities 440." (Dehghani, para. [0062].) The revision steps of Dehghani's decoder are the computation by which the next symbol is determined: the representation of the next predicted element is revised for N steps and the final softmax generates the output probabilities from the final representation. Each step after the first takes as its input the representation Ht-1 that the immediately preceding step produced for the same element, which is the just previously implemented hidden state, and each step conditions on all of the previously generated outputs in the target sequence, the last of which is the just previously determined token. The further step of revisions performed when the stop-decoding condition is not met is therefore a tokenization computation using the just previously determined token and a just previously implemented hidden state. Schuster, Zelikman, and Dehghani pertain to the allocation of additional computation to the generation of individual tokens by a transformer language model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to perform the additional computation of the combination of Schuster and Zelikman, which the think token causes the language model to perform for the token being determined, as a further revision step of the kind taught by Dehghani, taking as its inputs the representation produced by the preceding computation for that token and the previously generated tokens. One of ordinary skill in the art would have been motivated to make this modification because Dehghani teaches that re-applying the same series of operations to the representation from the preceding step is how a transformer devotes more processing resources to a symbol that is more ambiguous than others, which is the purpose for which the think token of the combination causes additional computation: "certain symbols, e.g. some words or phonemes, are usually more ambiguous than others. Therefore, the system can dynamically determine to allocate more processing resources to these more ambiguous symbols." (Dehghani, para. [0031].) The modification applies a known technique to a known system ready for improvement to yield the predictable result of an additional computation for the token that starts from the representation the preceding computation produced rather than discarding it. Claim 12 is an apparatus claim with limitations corresponding to the limitations of Claim 5 and is rejected under a similar rationale. Claim 18 is a system claim with limitations corresponding to the limitations of Claim 5 and is rejected under a similar rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: a. Goyal et al., "Think before you speak: Training language models with pause tokens," arXiv:2310.02226, https://arxiv.org/abs/2310.02226: training and inference of a decoder-only language model with a learnable pause token, a sequence of which is appended to the input prefix, with extraction of the model's output delayed until the last pause token is seen so that the model performs extra computation before committing to an answer. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUVAL H. LEVENTAL whose telephone number is (571) 270-3130. The examiner can normally be reached Monday-Friday, 8:00 AM - 5:00 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, PIERRE-LOUIS DESIR, can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YUVAL HAIM LEVENTAL/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Apr 11, 2025
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month