DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Applicant is advised that should claim 19 be found allowable, claim 20 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 6, 8-11, 13, 15-17, 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Saurav Kadavath et al. (hereinafter Kadavath) (“Language Models (Mostly) Know What They Know,” 11/21/2022) in view of Jiaxin Zhang et al. (hereinafter Zhang) (US 20250077777 A1, 03/06/2023), further in view of Chi-Min Chan et al. (hereinafter Chan) (“CHATEVAL: TOWARDS BETTER LLM-BASED EVALUATORS THROUGH MULTI-AGENT DEBATE,” 08/14/2023).
Regarding claim 1, Kadavath teaches;
A method for predicting probability of hallucination ([pg. 5] P(IK) – The probability a model assigns to "I know", i.e. the proposition that it will answer a given question correctly…) before generation for a query (predict "P(IK)", …, without reference to any particular proposed answer) imposed to a Large Language Model (LLM) ([Abstract] language models can evaluate predict which questions they will be able to answer correctly... [pg. 14] examples of P(IK) scores from a 52 billion parameter model)
NOTE: From the applicant’s specification; “[00108] factuality hallucinations refer to outputs which directly contradict or fabricate the ground truth.” Kadavath determines correctness of model answers by comparing to the ground truth; “[pg. 14] correctness on TriviaQA questions is judged based on a string comparison. Our dataset for training and evaluating the classifier then contains datapoints of the form (Few-Shot Prompt + Question, Ground Truth Label).” From this, Kadavath’s answers which are incorrect relative to the ground truth constitute factuality hallucinations in light of the applicant’s spec, and Kadavath’s P(IK) constitutes a predicted probability that the model will produce a correct, non-hallucinatory output for a given query/question, where the equivalent complementary value 1 – P(IK) represents the predicted probability that the model will produce an incorrect, hallucinatory output for the given query.
implementing a generative model; ([Abstract] We study whether language models can ... predict which questions they will be able to answer correctly.) training the generative model via a method of leveraging a simulation algorithm ([pg. 5] We train models with a value head to predict … P(IK) ... [pg. 14] we sample 30 answers (at temperature = 1) from each TriviaQA question … [pg. 15] We set the ground truth P(IK) as the fraction of samples at T=1 that the model gets correct) for construction of an encoder for hallucination; ([pg. 14] During P(IK) training, we finetune the entire model along with the value head)
NOTE: Kadavath teaches sampling 30 generated answers at temperature 1 for a given question, grades each answer as correct or incorrect, then determines the ground-truth P(IK) as a fraction of the sampled answers that are correct. This sampling process reasonably constitutes a sampling-based simulation because repeated stochastic generations are used as trials to empirically estimate the probability that the model will correctly answer the question. Kadavath then uses the empirically estimated ground-truth P(IK) to train the generative language model and its value head to predict P(IK).
Additionally, Kadavath teaches adding a value head on top of the language model to predict P(IK), wherein the equivalent complementary value 1-P(IK) indicates the predicted hallucination probability, as previously taught. Kadavath teaches fine-tuning the entire language model together with the value head using the empirically estimated ground-truth P(IK). Accordingly, the fine-tuned language model and associated value head collectively form a trained model that encodes an input question into a predicted probability of a correct answer, and, equivalently, the complementary probability of hallucination. The fine-tuned language model and value head constructed by the P(IK) training therefore correspond to the claimed encoder for hallucination.
applying the simulation algorithm on the sampled outputs; ([pg. 14] we sample 30 answers (at temperature = 1) from each TriviaQA question … [pg. 15] We set the ground truth P(IK) as the fraction of samples at T=1 that the model gets correct)
deriving an empirical estimate into an expected rate of hallucination for the original received query as a ground truth ([pg. 14] For a given question Q, … approximate a soft label for the ground-truth P(IK) by using many hard labels) for the encoder; ([pg. 14] ground-truth P(IK) scores that we use for training … During P(IK) training, we finetune the entire model along with the value head)
and outputting a probability of hallucination value for the query received by the generative model ([pg. 5] P(IK) – The probability a model assigns to "I know", i.e. the proposition that it will answer a given question correctly) before the LLM [pg. 14] examples of P(IK) scores from a 52 billion parameter model (note: large language model)) generates an output in response to the received query ([Abstract] predict "P(IK)” … to a question … without reference to any particular proposed answer...)
NOTE: P(IK) equivalently indicates the complementary probability of hallucination, 1 – P(IK), as previously taught.
Kadavath fails to explicitly teach but Zhang teaches;
by utilizing one or more processors along with allocated memory, the method comprising: ([0012] system may include one or more processors and a memory storing instructions for execution by the one or more processors.)
receiving a query by the generative model from a user via a user interface operatively connected to the generative model; ([0039] If the system 300 is local to a user ..., the interface 310 may include ... suitable input or output elements that allow interfacing with the user (such as to provide a prompt to the generative AI models 340 …)
perturbing the received query n times into unique variations that retain the original semantic meaning of the received query yet diverge lexically; ([0012] receiving a first prompt for submission to the first LLM, generating, using the first LLM, a plurality of semantically equivalent prompts to the first prompt)
sample an output from each query including the original received query; ([0012] generating, using the first LLM, a first response to the first prompt, generating a plurality of second responses using the first LLM, each second response generated in response to a semantically equivalent prompt of the plurality of semantically equivalent prompts)
OBVIOUSNESS TO COMBINE ZHANG:
Zhang is analogous art to the present disclosure as it pertains to detecting hallucinations in responses generated by LLMs (see [Abstract]).
Kadavath already teaches a method of applying a simulation algorithm on sampled outputs for a query to train a generative model to predict the complementary probability of hallucination for the query. Zhang teaches receiving a query from a user, generating ‘n’ semantically similar variations of the query, and sampling an output for each of the ‘n’ generated queries and the original query.
Zhang also explains that their method of generating semantically similar prompts to sample outputs from improves hallucination detection;
([0028] if the prompt 101 is not clearly constructed, then inconsistent or inaccurate responses may be generated… The example implementations improve hallucination detection relating to [the limitation] by first generating a plurality of prompts which are semantically equivalent to the user's initial prompt)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Kadavath’s sampling procedure to generate and evaluate outputs for the original query and ‘n’ semantically equivalent variations as taught by Zhang, predictably providing Kadavath’s simulation process with a more robust ground-truth probability reflecting the model’s ability to answer a question across different semantically similar formulations, thereby improving the reliability of Kadavath’s complementary hallucination probability prediction.
Kadavath and Zhang fail to teach but Chan teaches;
implementing
([pg. 3] We treat each individual LLM as an agent and ask them to generate their response from the given prompt)
OBVIOUSNESS TO COMBINE CHAN:
Chan is analogous art to the present disclosure as it pertains to a multi-agent framework.
Zhang teaches using LLMs to sample an output from an original received query and ‘n’ different but semantically similar queries, while Chan teaches implementing a plurality of independent LLM agents that each sample an output from a respective query.
Additionally, Zhang already identifies a problem with relying on one LLM to sample outputs; ([0028] if an LLM is consistently inaccurate in its responses, then a self-check will not catch such persistent hallucinations)
Zhang addresses this problem by using a second LLM different from the first and incorporating both models’ responses into a cross-model consistency score; ([0026] The example implementations improve hallucination detection relating to [this limitation] … by generating response to the user's initial prompt and the plurality of semantically equivalent prompts using an additional LLM… [0046] responses to the first prompt and the semantically equivalent prompts may be generated by the first LLM 220 and the second LLM 230.)
Chan provides a known way to expand that two-model concept into multiple LLM agents. Chan treats each individual LLM as an agent, gives the agents respectively configured prompts, and has each LLM generate its own response.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to assign Zhang’s original query and semantically equivalent query variations to respective Chan agents, thereby reducing dependence on the systematic errors of any one LLM, and providing Kadavath’s simulation process with a more robust set of sampled outputs for estimating complementary hallucination probability.
From this reasoning, Zhang and Chan reasonably teach;
implementing n+1 independent agents (Chan) to sample an output from each query including the original received query; (Zhang)
Regarding claim 2, Kadavath teaches;
wherein the simulation algorithm is a ([pg. 5] We train models with a value head to predict … P(IK) ... [pg. 14] we sample 30 answers (at temperature = 1) from each TriviaQA question … [pg. 15] We set the ground truth P(IK) as the fraction of samples at T=1 that the model gets correct)
NOTE: Kadavath teaches sampling 30 generated answers from the generative language model at temperature 1 for a given question, grades each answer as correct or incorrect, then determines the ground-truth P(IK) as a fraction of the sampled answers that are correct. This sampling process reasonably constitutes a Monte Carlo simulation because repeated stochastic generations are used as trials to empirically estimate the probability that the model will correctly answer the question.
Kadavath and Zhang fail to explicitly teach but Chan teaches;
Multi-agent
([pg. 3] We treat each individual LLM as an agent and ask them to generate their response from the given prompt)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to assign Zhang’s original query and semantically equivalent query variations to respective Chan agents, thereby reducing dependence on the systematic errors of any one LLM, and providing Kadavath’s Monte Carlo simulation process with a more robust set of sampled outputs for estimating complementary hallucination probability.
From this, Kadavath and Chan reasonably teach;
wherein the simulation algorithm is a Multi-Agent (Chan) Monte Carlo Simulation algorithm (Kadavath)
Regarding claim 3, Kadavath teaches;
wherein the empirical estimate (approximate ground truth P(IK) complement) that is provided through the (using the same reasoning from claim 2) is proportional to an approximation of hallucination rate. ([pg. 14] For each question, we generated 30 answer samples at T = 1 … approximate a soft label for the ground-truth P(IK) by using many hard labels)
NOTE: Using the same reasoning from claim 1, the approximate ground truth P(IK) equivalently indicates the approximate complementary probability/rate of hallucination for a given question/query, 1 – approximate ground truth P(IK).
Kadavath fails to explicitly teach but Chan teaches;
Multi-Agent
([pg. 3] We treat each individual LLM as an agent and ask them to generate their response from the given prompt)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to assign Zhang’s original query and semantically equivalent query variations to respective Chan agents, thereby reducing dependence on the systematic errors of any one LLM, and providing Kadavath’s Monte Carlo simulation process with a more robust set of sampled outputs for estimating complementary hallucination probability.
Regarding claim 4, Kadavath teaches;
estimating, by the trained generative model ([Abstract] models can be trained to predict "P(IK)"), a binary classification of the received query’s propensity to hallucinate ([pg. 5] P(IK) is usually computed using a binary classification head on top of a language model) before generation ([Abstract] predict "P(IK)", …, without reference to any particular proposed answer...)
NOTE: Using the same reasoning from claim 1, P(IK) equivalently indicates the complementary probability/propensity of hallucination for a given question/query, 1 – P(IK).
Regarding claim 6, Kadavath teaches;
Training ([Abstract] models can be trained to predict "P(IK)") a binary model ([pg. 5] P(IK) is usually computed using a binary classification head on top of a language model) to estimate propensity the received query can hallucinate. ([pg. 5] P(IK)– The probability a model assigns to "I know", i.e. the proposition that it will answer a given question correctly when samples are generated at unit temperature)
NOTE: Using the same reasoning from claim 1, P(IK) equivalently indicates the complementary probability/propensity of hallucination for a given question/query, 1 – P(IK).
Regarding claim 8,
Claim 8 is a system claim that is substantially similar to method claim 1, with one added limitation, which Kadavath fails to explicitly teach but Zhang teaches;
the system comprising: a processor; and a memory operatively connected to the processor via a communication interface,
PNG
media_image1.png
431
671
media_image1.png
Greyscale
the memory storing computer readable instructions, when executed, causes the processor to: ([0012] system may include one or more processors and a memory storing instructions for execution by the one or more processors.)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Kadavath’s sampling procedure to generate and evaluate outputs for the original query and ‘n’ semantically equivalent variations as taught by Zhang, predictably providing Kadavath’s simulation process with a more robust ground-truth probability reflecting the model’s ability to answer a question across different semantically similar formulations, thereby improving the reliability of Kadavath’s complementary hallucination probability prediction.
Zhang further teaches hardware system capable of implementing an LLM based hallucination detection framework; ([0038] FIG. 3 shows an example system 300 for detecting hallucinations in large language models)
Thus, it further would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement the hallucination detection framework taught by Kadavath as modified by Zhang on the hardware system of Zhang, predictably allowing the hallucination detection framework to be utilized and run.
The remaining limitations are taught using the same reasoning as in claim 1.
Regarding claims 9-11, 13,
Claims 9-11, 13 are system claims directly corresponding to method claims 2-4, 6 respectively, and are rejected using the same reasoning.
Regarding claim 15,
Claim 15 is a non-transitory computer readable medium claim that is substantially similar to method claim 1, with one added limitation, which Kadavath fails to explicitly teach but Zhang teaches;
A non-transitory computer readable medium configured to store instructions for ([0042] The memory 335, … such as … non-transitory memory) may store any number of software programs, executable instructions)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Kadavath’s sampling procedure to generate and evaluate outputs for the original query and ‘n’ semantically equivalent variations as taught by Zhang, predictably providing Kadavath’s simulation process with a more robust ground-truth probability reflecting the model’s ability to answer a question across different semantically similar formulations, thereby improving the reliability of Kadavath’s complementary hallucination probability prediction.
Zhang further teaches hardware system capable of implementing an LLM based hallucination detection framework; ([0038] FIG. 3 shows an example system 300 for detecting hallucinations in large language models, according to some implementations.)
Thus, it further would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement the hallucination detection framework taught by Kadavath as modified by Zhang on the hardware system of Zhang, predictably allowing the hallucination detection framework to be utilized and run.
The remaining limitations are taught using the same reasoning as in claim 1.
Regarding claim 16, Kadavath teaches;
wherein the simulation algorithm is a
([pg. 5] We train models with a value head to predict … P(IK) ... [pg. 14] we sample 30 answers (at temperature = 1) from each TriviaQA question … [pg. 15] We set the ground truth P(IK) as the fraction of samples at T=1 that the model gets correct)
NOTE: Kadavath teaches sampling 30 generated answers from the generative language model at temperature 1 for a given question, grades each answer as correct or incorrect, then determines the ground-truth P(IK) as a fraction of the sampled answers that are correct. This sampling process reasonably constitutes a Monte Carlo simulation because repeated stochastic generations are used as trials to empirically estimate the probability that the model will correctly answer the question.
wherein the empirical estimate (approximate ground truth P(IK) complement) that is provided through the is proportional to an approximation of hallucination rate. ([pg. 14] For each question, we generated 30 answer samples at T = 1 … approximate a soft label for the ground-truth P(IK) by using many hard labels)
NOTE: Using the same reasoning from claim 1, the approximate ground truth P(IK) equivalently indicates the approximate complementary probability/rate of hallucination for a given question/query, 1 – approximate ground truth P(IK).
Kadavath and Zhang fail to explicitly teach but Chan teaches;
Multi-agent
([pg. 3] We treat each individual LLM as an agent and ask them to generate their response from the given prompt)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to assign Zhang’s original query and semantically equivalent query variations to respective Chan agents, thereby reducing dependence on the systematic errors of any one LLM, and providing Kadavath’s Monte Carlo simulation process with a more robust set of sampled outputs for estimating complementary hallucination probability.
Regarding claims 17 and 19,
Claims 17 and 19 are non-transitory computer readable medium claims directly corresponding to method claims 4 and 6 respectively and are rejected using the same reasoning.
Regarding claim 20,
Claim 20 is a non-transitory computer readable medium claim directly corresponding to method claim 6 and is rejected using the same reasoning.
Claim(s) 5, 7, 12, 14, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kadavath (“Language Models (Mostly) Know What They Know,” 11/21/2022) in view of Zhang (US 20250077777 A1, 03/06/2023), further in view of Chan (“CHATEVAL: TOWARDS BETTER LLM-BASED EVALUATORS THROUGH MULTI-AGENT DEBATE,” 08/14/2023) as applied to claims 1, 8, and 15 above, and further in view of Yuyan Chen et al. (hereinafter Chen) (“Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models,” 10/21/2023).
Regarding claim 5, Kadavath teaches;
estimating a ([pg. 5] P(IK)– The probability a model assigns to "I know", i.e. the proposition that it will answer a given question correctly when samples are generated at unit temperature) estimating an expected value of hallucination via sampling ([pg. 5] Ground Truth P(IK)– The fraction of unit temperature samples to a question that are correct) before generation ([Abstract] predict "P(IK)", … to a question, without reference to any particular proposed answer)
NOTE: Using the same reasoning from claim 1, the estimated P(IK) equivalently corresponds to the complementary hallucination probability 1 – P(IK). Kadavath trains the model to estimate P(IK) using the sampled ground truth P(IK), without reference to a particular answer (i.e., before the model generates a response to the question).
Kadavath, Zhang, and Chan fail to explicitly teach but Chen teaches;
estimating a multi-class hallucination rate
([Abstract] we propose a robust discriminator named RelD to effectively detect hallucination in LLMs’ generated answers… [pg. 5] Initially, we employ a regression approach to train the discriminator RelD in order to fit the final score … we normalize the final score into different numbers of classes, such as four, six, eight, and ten, for multi-class classification)
OBVIOUSNESS TO COMBINE CHEN:
Chen is analogous art to the present disclosure as it pertains to detecting hallucinations in LLMs.
Kadavath already teaches estimating a complementary hallucination rate estimating an expected value of hallucination via sampling by using a generative model with a binary head, while Chen teaches converting a hallucination related score into a multi-class classification.
Additionally, Chen states;
([pg. 5] One potential advantage of this method is that the classification task, which focuses on distinguishing different categories, may facilitate capturing subtle differences among the final scores)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the pre-generation hallucination detection system of Kadavath as modified by Zhang and Chan to employ the multiclass score-prediction architecture taught by Chen, to facilitate capturing subtle differences among hallucination related reliability scores.
Regarding claim 7, Kadavath teaches;
training a ([pg. 5] We train models with a value head to predict … P(IK)) expected value of hallucinations when sampled ([pg. 14] For each question, we generated 30 answer samples at T = 1 … estimate a soft label for the ground-truth P(IK) by using many hard labels)
Kadavath fails to explicitly teach but Zhang teaches;
sampling n + 1 times
(using the same reasoning from claim 1; sampling outputs from n + 1 queries including an original query)
OBVIOUSNESS:
Using the same reasoning from claim 1, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Kadavath’s sampling procedure to generate and evaluate outputs for the original query and ‘n’ semantically equivalent variations as taught by Zhang, predictably providing Kadavath’s simulation process with a more robust ground-truth probability reflecting the model’s ability to answer a question across different semantically similar formulations, thereby improving the reliability of Kadavath’s complementary hallucination probability prediction.
Kadavath, Zhang, and Chan fail to explicitly teach but Chen teaches;
multi-class model
([Abstract] we propose a robust discriminator named RelD to effectively detect hallucination in LLMs’ generated answers… [pg. 5] Initially, we employ a regression approach to train the discriminator RelD in order to fit the final score … we normalize the final score into different numbers of classes, such as four, six, eight, and ten, for multi-class classification)
OBVIOUSNESS:
Using the same reasoning from claim 5, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the pre-generation hallucination detection system of Kadavath as modified by Zhang and Chan to employ the multiclass score-prediction architecture taught by Chen, to facilitate capturing subtle differences among hallucination related reliability scores.
Regarding claims 12 and 14,
Claims 12 and 14 are system readable medium claims directly corresponding to method claims 5 and 7 respectively and are rejected using the same reasoning.
Regarding claim 18,
Claim 18 is a non-transitory computer readable medium claim directly corresponding to method claim 5 and is rejected using the same reasoning.
CONCLUSION
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW ALAN CADY/ Examiner, Art Unit 2145
/CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145