DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 – 24, 26 – 37 are rejected under 35 U.S.C. 103 as being unpatentable over Krishna et al. (US PAP 2024/0395404) in view of Imani et al. (US PAP 2024/0362417).
As per claims 1, 26, Krishna et al. teach a method for evaluating an artificial intelligence (AI) large language model (LLM) generated response, the method comprising:
generating a LLM answer for a user query by a LLM (“the generative artificial intelligence module 170 may provide the input to the one or more large language models (and/or other machine learning models) which can interpret the query and determine an answer using the context provided by the large language models and the underlying features of the AI platform 110.”; paragraphs 110 – 112);
performing a domain-specific evaluation of the LLM answer (“domain specific model… determine an answer using the context provided by the large language models and the underlying features of the AI platform 110”; paragraphs 95 – 99; 110 – 112);
wherein the user query is in a form of patient-related question in a medical treatment domain (“the physician may recommend a treatment for a patient (e.g., a patient associated with and/or represented by a particular digital entity) based on successful treatments associated with other digital entities who are most similar (e.g., according to one or more factors, such as demographics, medical conditions, disease stages, and the like) to the particular digital entity.”; paragraphs 5 – 11).
However, Krishna et al. do not specifically teach performing a domain-agnostic evaluation of the LLM answer.
Imani et al disclose that a readability model 108 parses the text of the input 12 and the text of the generated LLM output 14 and evaluates readability metrics of the text included in the input 12 and the text included in the LLM output 14. The readability model 108 creates a feature vector 16 based on the evaluation of the readability metrics… a high value for the feature 18 indicates that the text is complex and more difficult to read, and a low value for the feature 18 indicates that the text is easier to read. In some implementations, a low value for the feature 18 indicates that the text is complex and more difficult to read, and a high value for the feature 18 indicates that the text is easier to read (paragraph 20).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective date of the claimed invention to determine the readability score of the LLM output as taught by Imani et al. in Krishna et al., because that would help improve the quality and interpretability of the generated text by the LLM and enables users to understand the underlying factors that contribute to the output's quality (paragraph 16).
As per claims 2, 27, Krishna et al. in view of Imani et al. further disclose that the process of generating the LLM answer comprises: receiving the user query; and analyzing the user query using the LLM, wherein the LLM is learned with open-source data (Krishna et al. paragraphs 110 – 112, and 175).
As per claims 3, 28, Krishna et al. in view of Imani et al. further disclose the processes of performing the domain-specific evaluation of the LLM answer and the domain-agnostic evaluation of the LLM answer are automated (“use of machine automation of data collection and aggregation may limit or reduce human errors made by physicians and increase an overall quality of data relied on by medical professionals” Krishna et al., paragraphs 28, 76).
As per claims 4, 29, Krishna et al. in view of Imani et al. further disclose generating at least one metric for the LLM answer from the domain-specific evaluation and the domain-agnostic evaluation of the LLM answer; and evaluating the quality of the LLM answer based on the at least one metric (“the computing system may identify digital entities (and/or corresponding information) for comparison based on specific criteria (e.g., disease subgroups), leveraging embeddings that capture similarities between treatments and between drugs. The use of the numeric representations makes it possible for a physician to receive the information in a limited timeframe, which means the information can be used to produce better patient treatment outcomes.”; Krishna et al., paragraphs 28 – 33).
As per claims 5, 30, Krishna et al. in view of Imani et al. further disclose initiating a human-in-the-loop process to evaluate the at least one metric generated from the domain-specific evaluation (“the user of the application (e.g., insurance analyst or physician) can choose which data categories are to be used in computing similarity”; Krishna et al., paragraphs 28 – 33, 80).
As per claims 6, 31, Krishna et al. in view of Imani et al. further disclose the medical treatment domain comprises medical oncology or radiation oncology (Krishna et al., paragraphs 4, 40).
As per claims 7, 32, Krishna et al. in view of Imani et al. further disclose the domain-specific evaluation performs a fact-checking of the LLM answer (“providing the first data as an input to a first domain specific model to generate the first value.”; Krishna et al., paragraphs 34 – 37).
As per claims 8, 33, Krishna et al. in view of Imani et al. further disclose the domain-agnostic evaluation checks relevance of the LLM answer to general public's understanding of the LLM answer (“providing the first data as an input to a first domain specific model to generate the first value…the user of the application (e.g., insurance analyst or physician) can choose which data categories are to be used in computing similarity”; Krishna et al., paragraphs 28 – 37, 80).
As per claims 9, 34, Krishna et al. in view of Imani et al. further disclose the domain-specific evaluation comprises comparing the LLM answer with a reference answer identified by a group of human domain-specific experts (Krishna et al., paragraphs 28 – 37, 80).
As per claims 10, 35, Krishna et al. in view of Imani et al. further disclose a similarity metric is computed to compare the similarity between the LLM answer and the reference answer (“similarity values”; Abstract, Krishna et al., paragraphs 28 – 33, 69 - 82).
As per claims 11, 36, Krishna et al. in view of Imani et al. further disclose a set of domain-specific evaluation metrics is used in the comparison of the LLM answer and the reference answer (“similarity values”; Krishna et al., Abstract, paragraphs 28 – 33, 69 - 82).
As per claims 12, 37, Krishna et al. in view of Imani et al. further disclose the set of domain-specific evaluation metrics comprises potential harm, factual correctness, completeness, and conciseness (“similarity values”; Krishna et al., Abstract, paragraphs 28 – 33, 69 - 82).
As per claim 13, Krishna et al. in view of Imani et al. further disclose each of the domain-specific evaluation metrics is graded with a numeric reference between 1 to 5 in the comparison of the LLM answer and the reference answer(“similarity values”; Krishna et al., Abstract, paragraphs 28 – 33, 69 - 82).
As per claim 14, Krishna et al. in view of Imani et al. further disclose the grade of each of the domain-specific evaluation metrics and the similarity metric between the LLM answer and reference answer are stored in a lookup table(“similarity values”; Krishna et al., Abstract, paragraphs 28 – 33, 69 - 82).
As per claim 15, Krishna et al. in view of Imani et al. further disclose determining at least one readability metric; wherein the readability metric comprises an estimated school grade level required to understand the LLM answer (“the features 18 include different values that quantify the complexity of the text in the input 12 and the text of the LLM output 14. In some implementations, a high value for the feature 18 indicates that the text is complex and more difficult to read, and a low value for the feature 18 indicates that the text is easier to read. In some implementations, a low value for the feature 18 indicates that the text is complex and more difficult to read, and a high value for the feature 18 indicates that the text is easier to read… , the features 18 include readability metrics that evaluate human readability features of the text included in the input 12 and the text of the LLM output 14. Example human readability metrics include the Gunning Fog Index, the Coleman-Liau Index, and the Automated Readability Index. Example human readability features include sentence length, word length, and syllable count”; Imani et al; paragraphs 20, 21).
As per claims 16, 17 , Krishna et al. in view of Imani et al. further disclose determining a sentence complexity for the LLM answer; determining a readability score for the LLM answer (“the features 18 include different values that quantify the complexity of the text in the input 12 and the text of the LLM output 14. In some implementations, a high value for the feature 18 indicates that the text is complex and more difficult to read, and a low value for the feature 18 indicates that the text is easier to read. In some implementations, a low value for the feature 18 indicates that the text is complex and more difficult to read, and a high value for the feature 18 indicates that the text is easier to read… , the features 18 include readability metrics that evaluate human readability features of the text included in the input 12 and the text of the LLM output 14. Example human readability metrics include the Gunning Fog Index, the Coleman-Liau Index, and the Automated Readability Index. Example human readability features include sentence length, word length, and syllable count”; Imani et al; paragraphs 20, 21).
As per claim 18, Krishna et al. in view of Imani et al. further disclose the readability score detects a literacy bias of the LLM answer against a patient whose literacy is below an average literacy of the general public (“a high value for the feature 18 indicates that the text is complex and more difficult to read, and a low value for the feature 18 indicates that the text is easier to read. In some implementations, a low value for the feature 18 indicates that the text is complex and more difficult to read, and a high value for the feature 18 indicates that the text is easier to read.”; Imani et al., paragraphs 20 – 26).
As per claim 19, Krishna et al. in view of Imani et al. further disclose mitigating the literacy bias against the patient's literacy by inputting into the LLM the readability score associated with the LLM answer; and requesting the LLM to generate a modified LLM answer with a modified readability score lower than the readability score (“a high value for the feature 18 indicates that the text is complex and more difficult to read, and a low value for the feature 18 indicates that the text is easier to read. In some implementations, a low value for the feature 18 indicates that the text is complex and more difficult to read, and a high value for the feature 18 indicates that the text is easier to read…The confidence score 20 and/or the insights 22 may also be used to identify areas where the LLM 106 needs improvement and may be used to provide feedback to developers of the LLM 1”; Imani et al., paragraphs 20 – 30).
As per claim 20, Krishna et al. in view of Imani et al. further disclose determining a syllable score; wherein the syllable score comprises a syllable count in the LLM answer (“the features 18 include readability metrics that evaluate human readability features of the text included in the input 12 and the text of the LLM output 14. Example human readability metrics include the Gunning Fog Index, the Coleman-Liau Index, and the Automated Readability Index. Example human readability features include sentence length, word length, and syllable count”; Imani et al. paragraphs 20 - 22).
As per claim 21, Krishna et al. in view of Imani et al. further disclose the syllable score detects a syllable bias of the LLM answer against a patient with reading disorders (Imani et al., paragraphs 20 - 22).
As per claim 22, Krishna et al. in view of Imani et al. further disclose mitigating the syllable bias of the LLM answer against the patient with reading disorders by inputting the syllable score for the LLM answer into the LLM and requesting the LLM to generate a modified LLM answer with a lower syllable count (Imani et al., paragraphs 20 – 30).
As per claims 23, 24, Krishna et al. in view of Imani et al. further disclose determining a score for word count or number of sentences; wherein the score for word count or number of sentences determines an ease level of comprehension and interpretation of the LLM answer; determining a lexicon score; wherein the lexicon score determines a subjectivity and contextual impact of the LLM answer (“the features 18 include readability metrics that evaluate human readability features of the text included in the input 12 and the text of the LLM output 14. Example human readability metrics include the Gunning Fog Index, the Coleman-Liau Index, and the Automated Readability Index. Example human readability features include sentence length, word length, and syllable count “; Imani et al., paragraphs 20 - 22).
Allowable Subject Matter
Claim 25 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter:
As to claim 25, the prior art made of record does not teach or suggest converting the user query and the LLM answer into a vectorized user query and a vectorized LLM answer, respectively; comparing the vectorized user query with a database of reference queries comprising more than one reference queries; computing a vectorized query similarity metric between the vectorized user query and each of the reference queries; selecting a surrogate user query from the reference queries wherein the surrogate user query has highest vectorized query similarity metric; returning the reference answer that corresponds to the surrogate user query, wherein the reference answer is a surrogate reference answer to the vectorized user query; and computing a vectorized answer similarity metric between the vectorized LLM answer with the surrogate reference answer.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Liu et al. teach NATURAL LANGUAGE TRAINING AND/OR AUGMENTATION WITH LARGE LANGUAGE MODELS. Bell et al. teach Methods For Generating Task-Specific Agent Modules Based On User Requests.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEONARD SAINT-CYR whose telephone number is (571)272-4247. The examiner can normally be reached Monday- Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LEONARD SAINT-CYR/Primary Examiner, Art Unit 2658