Prosecution Insights
Last updated: August 17, 2026
Application No. 18/634,394

Framework for Trustworthy Generative Artificial Intelligence

Non-Final OA §101§102§103
Filed
Apr 12, 2024
Examiner
LEE, WILLIAM MICHAEL
Art Unit
Tech Center
Assignee
ServiceNow Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
17 currently pending
Career history
15
Total Applications
across all art units

Statute-Specific Performance

§101
29.2%
-10.8% vs TC avg
§103
48.6%
+8.6% vs TC avg
§102
2.8%
-37.2% vs TC avg
§112
19.4%
-20.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION The action is in response to the original filing on April 12, 2024. Claims 1-20 are pending and have been considered below. Claims 1, 15, and 18 are dependent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1 – The claim is directed to a method: a method comprising… Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mental processes (see MPEP 2106.04(a)(2)(III)): determining that the metric satisfies a fault threshold… a human can reasonably determine whether a metric satisfies a threshold within the human mind. in response to determining that the metric satisfies the fault threshold, labeling the output as untrustworthy… a human can reasonably label an output in response to determining whether a metric satisfies a threshold within the human mind or with the aid of a pen and paper. Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the judicial exception into a practical application: obtaining a prompt for a large-language model (LLM)… obtaining a prompt is data gathering (see MPEP 2106.05(g)). generating, using the LLM, an output of an artificial intelligence system… generating an output using a model is data outputting (see MPEP 2106.05(g)). obtaining a validation model configured to detect a property in the output, the property indicating a fault in the output… obtaining a model is data gathering (see MPEP 2106.05(g)). generating, using the validation model on the output, a metric indicating likelihood of the property in the output… generating a metric using a model is data outputting (see MPEP 2106.05(g)). Step 2B – These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they only amount to data gathering and outputting (MPEP 2106.05(g)) without significantly more. These limitations, taken either alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible. Claims 2-14 recite limitations which further narrow the abstract idea of claim 1 by specifying more details of the mental processes that occur: Regarding claim 2, this claim further limits the abstract idea of claim 1 to be based on a mental process: wherein after obtaining the prompt for the LLM, evaluating the prompt, wherein evaluating the prompt comprises at least one of: determining whether an answer to the prompt exists within a database, appending predetermined inputs related to operational guidelines for the LLM to the prompt, or providing the prompt to a use-case filter configured to reject prompts unrelated to predetermined categories… a human can reasonably perform “evaluating the prompt, wherein evaluating the prompt comprises at least one of: determining whether an answer to the prompt exists… appending predetermined inputs… or providing the prompt to a use-case filter” within the human mind or with the aid of a pen and paper. Regarding claim 3, describing further comprising: in response to determining that the metric satisfies the fault threshold, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce likelihood of the property in further output from the LLM amounts to data gathering (see MPEP 2106.05(g)). Regarding claim 4, describing further comprising: ased upon the second prompt, generating, using the LLM, a second output of the artificial intelligence system and generating, using the validation model on the second output, a second metric indicating the likelihood of the property in the second output amounts to data gathering and outputting (see MPEP 2106.05(g)). Regarding claim 5, this claim further limits the abstract idea of claim 4 to be based on a mental process: further comprising: determining that the second metric does not satisfy the fault threshold and in response to determining that the second metric does not satisfy the fault threshold, labeling the second output as trustworthy… a human can reasonably perform determining that a metric doesn’t satisfy a threshold and labeling an output in response to that determination within the human mind or with the aid of a pen and paper. Regarding claim 6, this claim further limits the abstract idea of claim 4 to be based on a mental process: further comprising: in response to labeling the second output as trustworthy, modifying the validation model based on the second output being trustworthy… for example, given a simple enough pen and paper “validation model,” a human can reasonably perform “modifying the validation model” in response to labeling an output. Regarding claim 7, specifying wherein the fault in the output relates to one or more of bias, hallucination, toxic behavior, threat, or readability does not overcome the rejection of claim 1 as modifying “the fault” in this manner does not make “determining” and “labeling” in response to determining to not be mental processes. Regarding claim 8, specifying wherein the metric indicating the likelihood of the property in the output comprises one of a Boolean value or a degree of confidence in this manner does not overcome the rejection of claim 1 as modifying “the metric” does not make “determining” and “labeling” in response to determining to not be mental processes. Regarding claim 9, this claim further limits the abstract idea of claim 1 to be based on a mental process: determining one or more validation metrics related to presence of the property within the output… a human can reasonably determine metrics related to presence of the property in the output within the human mind or with the aid of a pen and paper. Furthermore, creating the validation model that outputs a presence indicator of the property based upon the validation metrics and training the validation model based on datasets containing prior examples of the validation metrics amount to data gathering and outputting (see MPEP 2106.05(g)). Regarding claim 10, this claim further limits the abstract idea of claim 9 to be based on a mental process: wherein training the validation model based on datasets containing prior examples of the validation metrics comprises determining that acceleration hardware is present in a computing system, and, in response to determining that acceleration hardware is present, utilizing parallelization capabilities of the acceleration hardware during training of the validation model… a human can reasonably determine whether acceleration hardware is present in a computing system within the human mind. Regarding claim 11, describing in response to labeling the output as untrustworthy, outputting related factors to the likelihood of the property in the output, wherein the related factors comprise a portion of the outputs relevant to the metric or reasoning for why the metric was provided amounts to data outputting (see MPEP 2106.05(g)). Regarding claim 12, this claim further limits the abstract idea of claim 1 to be based on a mental process: and based on the metric and the second metric, ranking the prompt and the second prompt… a human can reasonably perform ranking prompts based on metrics within the human mind or with the aid of a pen and paper. Furthermore, describing in response to determining that the metric satisfies the fault threshold, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce presence of the property indicating the fault in the output… based upon the second prompt, generating, using the LLM, a second output of the artificial intelligence system… and generating, using the validation model on the second output, a second metric indicating likelihood of the property in the second output amounts to data gathering and outputting (see MPEP 2106.05(g)). Regarding claim 13, this claim further limits the abstract idea of claim 1 to be based on a mental process: wherein generating, using the validation model on the output, the metric indicating the likelihood of the property in the output comprises: computing, by one or more pre-processing modules, a validation metric related to presence of the property within the output… a human can reasonably compute “a validation metric related to presence of the property within the output” within the human mind or with the aid of a pen and paper. Furthermore, describing and propagating the computed metric to the validation model amounts to data gathering and outputting (see MPEP 2106.05(g)). Regarding claim 14, specifying wherein the validation metric comprises one or more of: semantic similarity with a reference dataset, conformance to a pre-determined principle, a sentiment analysis score, or a determination that pre-determined numeric patterns exist in the output in this manner does not overcome the rejection of claim 13 as modifying “the validation metric” does not make “computing” to not be a mental process. Regarding claim 15: Step 1 – The claim is directed to a system: a computing system comprising… Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mental processes (see MPEP 2106.04(a)(2)(III)): determining that the metric satisfies a fault threshold… a human can reasonably determine whether a metric satisfies a threshold within the human mind. in response to determining that the metric satisfies the fault threshold, labeling the output as untrustworthy… a human can reasonably label an output in response to determining whether a metric satisfies a threshold within the human mind or with the aid of a pen and paper. Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the judicial exception into a practical application: one or more processors; memory; and program instructions, stored in the memory, that upon execution by the one or more processors cause the computing system to perform operations comprising… one or more processors and a memory used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). Claims 15-17 are system claims that contain similar limitations to the methods of claims 1, 3-5, and 6, respectively. Therefore, claims 15-17 are rejected under substantially the same rationale as claims 1, 3-5, and 6, respectively. Regarding claim 18: Step 1 – The claim is directed to a product: a non-transitory computer-readable medium… Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mental processes (see MPEP 2106.04(a)(2)(III)): determining that the metric satisfies a fault threshold… a human can reasonably determine whether a metric satisfies a threshold within the human mind. in response to determining that the metric satisfies the fault threshold, labeling the output as untrustworthy… a human can reasonably label an output in response to determining whether a metric satisfies a threshold within the human mind or with the aid of a pen and paper. Step 2A, Prong 2 – The following limitations are additional elements that fail to integrate the judicial exception into a practical application: non-transitory computer-readable medium storing program instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising… a computer-readable medium storing program instructions executable by a processor to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)). Claims 18-20 are computer-readable medium claims that contain similar limitations to the methods of claims 1, 3-5, and 6, respectively. Therefore, claims 18-20 are rejected under substantially the same rationale as claims 1, 3-5, and 6, respectively. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1 and 7-9 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Inan et al. (“Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations,” 2023, hereinafter Inan). Regarding claim 1: Inan teaches a method comprising: obtaining a prompt for a large-language model (LLM) (Page 2, ¶2 “we publicly release an input-output safeguard tool for classifying safety risks in prompts and responses for conversational AI agent use cases”). Inan further teaches generating, using the LLM, an output of an artificial intelligence system (Page 2, ¶2 “we publicly release an input-output safeguard tool for classifying safety risks in prompts and responses for conversational AI agent use cases… We provide different instructions for classifying human prompts (input to the LLM) vs AI model responses (output of the LLM)”). Inan further teaches obtaining a validation model configured to detect a property in the output, the property indicating a fault in the output (Page 2, ¶2 “We introduce Llama Guard, an LLM-based input-output safeguard model, fine-tuned on data labeled according to our taxonomy,” wherein “Llama Guard” encompasses a validation model, Section 2, ¶1 “Building automated input-output safeguards relies on classifiers to make decisions about content in real time. A prerequisite to building these systems is to have the following components: 1. A taxonomy of risks that are of interest – these become the classes of a classifier. 2. Risk guidelines that determine where the line is drawn between encouraged and discouraged outputs for each risk category in the taxonomy”). Inan further teaches generating, using the validation model on the output, a metric indicating likelihood of the property in the output (Pages 4-5, Section 3.3, ¶1 “we use one of our internal Llama checkpoints to generate a mix of cooperating and refusing responses for these prompts. We employ our expert, in-house red team to label the prompt and response pairs for the corresponding category based on the taxonomy… The red-teamers annotate the dataset for 4 labels: prompt-category, response-category, prompt-label (safe or unsafe), and response-label (safe or unsafe) … The final dataset comprises of 13,997 prompts and responses, with their respective annotations,” wherein “responses” encompass the output, Page 5, Section 4, ¶¶1-2 “The absence of standardized taxonomies makes comparing different models challenging, as they were trained against different taxonomies (for example, Llama Guard recognizes Guns and Illegal Weapons as a category, while Perspective API focuses on toxicity and does not have this particular category)… we evaluate Llama Guard on two axes: 1. In-domain performance on its own datasets (and taxonomy) to gauge absolute performance; 2. Adaptability to other taxonomies. Since Llama Guard is an LLM, we use zero-shot and few-shot prompting and fine-tuning using the taxonomy applicable to the dataset for evaluating it,” Pages 5-6, Section 4.1, ¶1 “Given that we are interested in evaluating different methods on several datasets, each with distinct taxonomies, we need to decide how to evaluate the methods in different settings. Evaluating a model, especially in an off-policy setup (i.e., to a test set that uses foreign taxonomy and guidelines), makes fair comparisons challenging and requires trade-offs… we take a different approach… for obtaining scores in the off-policy setup. We list the three techniques we employ for evaluating different methods in on- and off- policy settings,” Page 6, ¶1 “Overall binary classification for APIs that provide per-category output. Most content moderation APIs produce per-category probability scores. Given the probability scores from a classifier, the probability score for binary classification across all categories is computed as: PNG media_image1.png 46 606 media_image1.png Greyscale where ŷi is the predicted score for the i-th example, c1, c2, …, cn are the classes (from the classifier’s taxonomy), with c0 being the benign class, ŷc,i are the predicted scores for each of the positive categories c1, c2, …, cn for the ith example… we consider that a classifier assigns a positive label if it predicts a positive label due to any of its own categories,” wherein “the probability score” or “predicted score” encompasses a metric indicating a likelihood of the property or “the positive categories” or “a positive label” in the output). Inan further teaches determining that the metric satisfies a fault threshold (Page 2, Section 2, ¶1 “A prerequisite to building these systems is to have… Risk guidelines that determine where the line is drawn between encouraged and discouraged outputs for each risk category in the taxonomy,” Page 7, Section 4.3.3, ¶1 “For all experiments, we use the area under the precision-recall curve (AUPRC) as our evaluation metrics… AUPRC focuses on the trade-off between precision and recall, highlight the model’s performance… on the positive (“unsafe”) class, and is useful for selecting the classification threshold that balances precision and recall based on the specific requirements of use cases,” Section 4.4, ¶1 “Table 2 contains the comparison between Llama Guard against the probability-score-based baseline APIs on various benchmarks,” Page 14, Section B, ¶1 “we could not compute AUPRC for baselines that did not offer output probabilities… we compare them here using metrics that do not require access to probabilities. We set every threshold to 0.5 and compute Precision, Recall and F1 Score,” Page 8, Table 2 depicts “Evaluation results” against baselines “OpenAI API” and “Perspective API” or “the probability-score-based baseline APIs” which uses the metric or “the probability score,” implying determining that the metric satisfies a fault threshold or “the classification threshold” of “AUPRC” was required to compute the results of Table 2). Inan further teaches and in response to determining that the metric satisfies the fault threshold, labeling the output as untrustworthy (Page 3, Section 3.1, ¶1 “In our work, we… fine-tune LLMs with tasks that ask to classify content as being safe or unsafe,” ¶5 “In Llama Guard, the output contains two elements. First, the model should output “safe” or “unsafe”, both of which are single tokens… If the model assessment is “unsafe,” then the output should contain a new line, listing the taxonomy categories that are violated in the given piece of content. We train Llama Guard to use a format for the taxonomy categories that consists of a letter (e.g. ’O’) followed by the 1-based category index. With this output format, Llama Guard accommodates binary and multi-label classification, where the classifier score can be read off from the probability of the first token,” Page 5, ¶1 “The final dataset comprises of 13,997 prompts and responses, with their respective annotations,” Page 7, Section 4.3.3, ¶1 “For all experiments, we use the area under the precision-recall curve (AUPRC)… AUPRC focuses on the trade-off between precision and recall, highlight the model’s performance… on the positive (“unsafe”) class, and is useful for selecting the classification threshold that balances precision and recall,” Page 8, Table 2, wherein it is implicit during “all experiments” that output or “responses” from the “dataset” are labeled as untrustworthy or the positive “unsafe” class in response to determining that the metric satisfies the fault threshold or whether “the probability score” satisfies the “AUPRC” threshold). Regarding claim 7, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Inan further teaches wherein the fault in the output relates to one or more of bias, hallucination, toxic behavior, threat, or readability (Pages 2-3, Section 2.1, ¶1 “Below, we provide both the content types themselves and also examples of the specific kinds of content that we consider inappropriate for this purpose under each category: Violence & Hate… Sexual Content… Guns & Illegal Weapons… Regulated or Controlled Substances… Suicide & Self Harm… Criminal Planning”). Regarding claim 8, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Inan further teaches wherein the metric indicating the likelihood of the property in the output comprises one of a Boolean value or a degree of confidence (Pages 3-4, Section 3.1, ¶5 “the model should output “safe” or “unsafe”, both of which are single tokens… If the model assessment is “unsafe”, then the output should contain a new line, listing the taxonomy categories that are violated in the given piece of content. We train Llama Guard to use a format for the taxonomy categories that consists of a letter (e.g. ’O’) followed by the 1-based category index. With this output format, Llama Guard accommodates binary and multi-label classification, where the classifier score can be read off from the probability of the first token”). Regarding claim 9, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Inan further teaches wherein obtaining the validation model comprises: determining one or more validation metrics related to presence of the property within the output (Section 2, ¶1 “Building automated input-output safeguards relies on classifiers to make decisions about content in real time. A prerequisite to building these systems is to have the following components: 1. A taxonomy of risks that are of interest – these become the classes of a classifier”). Inan further teaches creating the validation model that outputs a presence indicator of the property based upon the validation metrics (Page 2, ¶2 “We introduce Llama Guard, an LLM-based input-output safeguard model, fine-tuned on data labeled according to our taxonomy,” Page 3, Section 3.1, ¶1 “we… fine-tune LLMs with tasks that ask to classify content as being safe or unsafe”). Inan further teaches and training the validation model based on datasets containing prior examples of the validation metrics (Page 5, ¶1 “The red-teamers annotate the dataset for 4 labels: prompt-category, response-category, prompt-label (safe or unsafe), and response-label (safe or unsafe) … The final dataset comprises of 13,997 prompts and responses, with their respective annotations,” Section 3.4, ¶1 “We train for 500 steps, which corresponds to ∼1 epoch over our training set”). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Koide et al. (“ChatSpamDetector: Leveraging Large Language Models for Effective Phishing Email Detection,” 2024, hereinafter Koide). Regarding claim 2, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation wherein after obtaining the prompt for the LLM, evaluating the prompt, wherein evaluating the prompt comprises at least one of: determining whether an answer to the prompt exists within a database, appending predetermined inputs related to operational guidelines for the LLM to the prompt, or providing the prompt to a use-case filter configured to reject prompts unrelated to predetermined categories, Inan teaches wherein after obtaining a prompt for the validation model, evaluating the prompt, wherein evaluating the prompt comprises at least one of… appending predetermined inputs related to operational guidelines for the LLM to the prompt (Page 3, Section 3.1, ¶¶1-2 “In our work, we… fine-tune LLMs with tasks that ask to classify content as being safe or unsafe. For input-output safeguarding tasks… Each task takes a set of guidelines as input, which consist of numbered categories of violation, as well as plain text descriptions as to what is safe and unsafe within that category,” please note that the items in this list are interpreted disjunctively; for example, evaluating the prompt comprises at least one of: determining… appending… or providing has been interpreted to mean “evaluating the prompt comprises determining… OR appending… OR providing”). However, Inan fails to teach after obtaining the prompt for the LLM, evaluating the prompt… Koide, in the same field of endeavor, teaches after obtaining the prompt for the LLM, evaluating the prompt… wherein evaluating the prompt comprises appending predetermined inputs related to operational guidelines for the LLM to the prompt (Page 3, Section 3, ¶1 “we present ChatSpamDetector, a system that uses LLMs to detect phishing emails. Our system converts emails into appropriate prompts for LLMs to analyze the entire email content,” Page 4, Section 3.2, ¶1 “Emails can be long or contain complex HTML structures. Given the token limit constraints of LLMs, it is essential to adjust the input data to not exceed these limits,” Page 6, ¶1 “we created Prompt Template 1 to analyze an email assigned to the template,” Page 5, “Prompt Template 1: Normal Prompt” depicts appending predetermined inputs related to operational guidelines for the LLM along with the “email text data”). Inan and Koide are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the evaluating of the prompt of Koide with the methodology of Inan. The motivation to do so is to “to conduct more advanced analyses by using LLMs” (Page 16, Section 6, ¶2). Regarding claim 11, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation further comprising: in response to labeling the output as untrustworthy, outputting related factors to the likelihood of the property in the output, wherein the related factors comprise a portion of the outputs relevant to the metric or reasoning for why the metric was provided, Inan teaches labeling the output as untrustworthy (Page 3, Section 3.1, ¶1 “In our work, we… fine-tune LLMs with tasks that ask to classify content as being safe or unsafe,” ¶3 “Each task indicates whether the model needs to classify the user messages (dubbed “prompts”) or the agent messages (dubbed “responses”),” ¶5 “Each task specifies the desired output format, which dictates the nature of the classification problem… the model should output ‘safe’ or ‘unsafe’ … If the model assessment is “unsafe”, then the output should contain a new line, listing the taxonomy categories that are violated in the given piece of content”), the likelihood of the property in the output (Page 6, ¶2 “Given the probability scores from a classifier, the probability score for binary classification across all categories is computed”), and the metric (Page 6, ¶2 “probability score for binary classification across all categories”). However, Inan fails to teach further comprising: in response to labeling the output as untrustworthy, outputting related factors to the likelihood of the property in the output, wherein the related factors comprise a portion of the outputs relevant to the metric or reasoning for why the metric was provided. Koide teaches further comprising: in response to labeling a text as untrustworthy, outputting related factors to untrustworthiness, wherein the related factors comprise a portion of the outputs relevant to the metric or reasoning for why the label was provided (Page 14, ¶1 “The body of an email often contains various SE techniques aimed at encouraging users to click on links. Phishing emails typically exploit the trust and recognition of well-known brands to lower users’ caution and encourage them to follow the instructions… In these cases, LLMs can accurately identify the brand impersonation, not merely through Named Entity Recognition (NER) but by analyzing the context to determine which brand is being impersonated,” ¶2 “LLMs can identify evidence leading to the determination of phishing emails by combining clues extracted from both headers and the body. Equipped with knowledge of brands and their associated legitimate domain names, LLMs can distinguish mismatches between these and the sender’s address or URLs in the email body. Additionally, LLMs accurately recognize the use of domain names similar to legitimate ones or URL shortening services… to redirect users to phishing sites,” ¶3 “Our system enables LLMs to output both detailed rationales and simplified explanations, allowing users to understand why an analyzed email is suspected of being phishing or is considered legitimate,” Page 5, “Prompt Template 1: Normal Prompt” – “Provide a comprehensive evaluation of the email… Include a detailed explanation of any phishing or legitimacy indicators… Summarize your findings and provide your final verdict on the legitimacy of the email, supported by the evidence you gathered,” please note that the items a portion and reasoning in the list are interpreted to be read disjunctively). Inan and Koide are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the outputting of related factors comprising reasoning for untrustworthiness of Koide with the methodology of Inan. The motivation to do so is to “to conduct more advanced analyses by using LLMs” (Page 16, Section 6, ¶2). Claims 3-5 are rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Li et al. (“SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models,” 2024, hereinafter Li). Regarding claim 3, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation further comprising: in response to determining that the metric satisfies the fault threshold, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce likelihood of the property in further output from the LLM, Inan teaches determining that the metric satisfies the fault threshold (Page 2, Section 2, ¶1, Page 7, Section 4.3.3, ¶1 and Section 4.4, ¶1, Page 14, Section B, ¶1, Page 8, Table 2 all as explained above with respect to claim 1). However, Inan fails to teach further comprising: in response to determining that the metric satisfies the fault threshold, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce likelihood of the property in further output from the LLM. Li, in the same field of endeavor, teaches further comprising: in response to an unsafe prompt, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce likelihood of the property (Page 3, Col. 1, Fig. 2 and Section 2.1, ¶1 “we propose a hierarchical three-level safety taxonomy for LLMs… Generally, SALAD-Bench includes six domain-level harmfulness areas,” wherein “taxonomy” or “harmfulness areas” encompasses the property) in further output from the LLM (Page 4, Col. 2, Section 3.2, ¶1 “To extensively measure the effectiveness of various attack methods, we also construct corresponding defense-enhanced subset QD. Contrary to the attack-enhanced subset, this subset comprises questions that are less likely to elicit harmful responses from LLMs,” Page 5, Col. 1, ¶2 “For each unsafe question qD from QD, we pick the most effective defense prompt… to enhance qD and collect all enhanced questions as Q̄D,” Page 6, Col. 1, Section 5.1, ¶2 “During experiments, we also incorporate different paraphrasing-based methods… perturbation-based methods… and prompting-based methods as defense methods,” Page 16, Col. 1, ¶1 “for prompting-based methods, we utilize the recently proposed Safe/XSafe prompts… and Self-Reminder prompt… in our experiments, which have shown effective defense abilities in small-scale experiments”). Inan and Li are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the obtaining of a second prompt of Li with the methodology of Inan. The motivation to do so is to design a safety benchmark that “goes beyond mere safety assessment of LLMs, providing a robust source for evaluating both attack and defense algorithms notably tailored for these models” (Li, Page 8, Col. 2, Section 6, ¶1). Regarding claim 4, Inan in view of Li teaches the method of claim 3 (and thus the rejection of claim 3 is incorporated). Li further teaches further comprising: based upon the second prompt, generating, using the LLM, a second output of the artificial intelligence system (Page 7, Col. 2, Footnote 2: “Given a question, we count an attack success if at least one prompt triggers a harmful response,” Page 8, Col. 1, Table 6 and ¶2 “We evaluate the performance of the defense methods on the attack-enhanced subset- with different LLMs, as shown in Table 6… GPT-paraphrasing method… and Self-Reminder prompt… obtain the best defense ability against unsafe instructions and attack methods,” wherein the “LLMs” depicted in Table 6 are implied to produce “responses” or a second output based upon the second prompt or “the defense methods” applied to the “unsafe instructions and attack methods”). Li further teaches and generating, using the validation model (Page 2, Col. 2, ¶2 “MD-Judge, is an LLM-based evaluator tailored for question-answer pairs,” Page 5, Col. 2, ¶¶1-2 “our task involves evaluating not only plain question answer pairs but also attack-enhanced question answer pairs. Our evaluator is named MD-Judge. To make our MD-Judge capable of both plain and attack-enhanced questions, we collect plain QA pairs… and construct both safe and unsafe answers to enhanced questions… During fine-tuning, we propose a safety evaluation template to reformat question-answer pairs for MD-Judge predictions… to enhance MD-Judge’s capabilities,” wherein MD-Judge encompasses the validation model) on the second output, a second metric indicating the likelihood of the property in the second output (Page 5, Col. 1, Section 5.1, ¶3 “The effectiveness of attack and defense strategies is evaluated using the Attack Success Rate (ASR) based on our MD-Judge. Note that ASR equals 1 minus the corresponding safety rate for each LLM,” Page 8, Col. 1, Table 6, Caption: “Attack success rate (ASR) comparison of different defense methods on attack-enhanced subset among multiple LLMs” and ¶2 “after introducing GPT-paraphrasing as the defense method, the ASR of Mistral7B… largely drops from 93.6% to 24.98%... after using self-reminder prompts, the ASR of Llama-2-13B even largely drops to 12.68%”). Inan and Li are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the methodologies of Li and Inan. The motivation to do so is to design a safety benchmark that “goes beyond mere safety assessment of LLMs, providing a robust source for evaluating both attack and defense algorithms notably tailored for these models” (Li, Page 8, Col. 2, Section 6, ¶1). Regarding claim 5, Inan in view of Li teaches the method of claim 4 (and thus the rejection of claim 4 is incorporated). Regarding the limitation further comprising: determining that the second metric does not satisfy the fault threshold, Inan teaches further comprising: determining that a metric does not satisfy the fault threshold (Page 3, Section 3.1, ¶1 “In our work, we adopt this paradigm as well, and fine-tune LLMs with tasks that ask to classify content as being safe or unsafe,” ¶5 “In Llama Guard… the model should output ‘safe’ or ‘unsafe,’” Page 7, Section 4.3.3, ¶1 “AUPRC focuses on the trade-off between precision and recall, highlight the model’s performance… on the positive (“unsafe”) class, and is useful for selecting the classification threshold that balances precision and recall,” wherein “selecting the classification threshold that balances precision and recall” implies determining whether metrics do or do not satisfy the fault threshold to find the right “trade-off between precision and recall”). However, Inan fails to teach the second metric. Li teaches the second metric (Page 5, Col. 1, Section 5.1, ¶3 “The effectiveness of attack and defense strategies is evaluated using the Attack Success Rate (ASR),” Page 7, Col. 2, Footnote 2: “Given a question, we count an attack success if at least one prompt triggers a harmful response”). Regarding the limitation and in response to determining that the second metric does not satisfy the fault threshold, labeling the second output as trustworthy, Inan teaches and in response to determining that a metric does not satisfy the fault threshold, labeling an output as trustworthy (Page 3, Section 3.1, ¶1 “In our work, we adopt this paradigm as well, and fine-tune LLMs with tasks that ask to classify content as being safe or unsafe,” ¶5 “In Llama Guard… the model should output ‘safe’ or ‘unsafe,’” Page 6, ¶2 “Given the probability scores from a classifier, the probability score for binary classification across all categories is computed… we consider that a classifier assigns a positive label if it predicts a positive label due any of its own categories,” Page 7, Section 4.3.3, ¶1 “AUPRC focuses on the trade-off between precision and recall, highlight the model’s performance… on the positive (“unsafe”) class,” wherein it is implicit that metrics which do not satisfy the fault threshold are labeled as belonging to a negative “safe” class, or a trustworthy class, Page 14, Section B, ¶1 “We set every threshold to 0.5”). However, Inan fails to teach the second metric and the second output. Li teaches the second metric (Page 5, Col. 1, Section 5.1, ¶3, Page 7, Col. 2, Footnote 2, Page 8, Col. 1, Table 6 and ¶2 “after introducing GPT-paraphrasing as the defense method, the ASR of Mistral7B… largely drops from 93.6% to 24.98%... after using self-reminder prompts, the ASR of Llama-2-13B even largely drops to 12.68%,” wherein the fault threshold of Inan that was set to “0.5” can be equivalent to, for example, 50%) and the second output (Page 7, Col. 2, Footnote 2, Page 8, Col. 1, Table 6 and ¶2 as explained above with respect to claim 4). Inan and Li are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the second output and metric of Li with the threshold comparison and labeling of Inan. The motivation to do so is to design a safety benchmark that “goes beyond mere safety assessment of LLMs, providing a robust source for evaluating both attack and defense algorithms notably tailored for these models” (Li, Page 8, Col. 2, Section 6, ¶1). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Li and further in view of Markov et al. (“A Holistic Approach to Undesired Content Detection in the Real World,” 2023, hereinafter Markov). Inan in view of Li teaches the method of claim 5 (and thus the rejection of claim 5 is incorporated). Regarding the limitation further comprising: in response to labeling the second output as trustworthy, modifying the validation model based on the second output being trustworthy, Inan teaches labeling an output as trustworthy (Page 3, Section 3.1, ¶1 “In our work, we… fine-tune LLMs with tasks that ask to classify content as being safe or unsafe”), the validation model (Page 2, ¶2 “We introduce Llama Guard, an LLM-based input-output safeguard model, fine-tuned on data labeled according to our taxonomy”), and an output being trustworthy (Page 3, Section 3.1, ¶3 “Each task indicates whether the model needs to classify the user messages (dubbed “prompts”) or the agent messages (dubbed “responses”),” ¶5 “Each task specifies the desired output format, which dictates the nature of the classification problem… the model should output ‘safe’ or ‘unsafe’”). However, Inan fails to teach further comprising: in response to labeling the second output as trustworthy, modifying the validation model based on the second output being trustworthy. Li teaches the second output (Page 7, Col. 2, Footnote 2, Page 8, Col. 1, Table 6 and ¶2 as explained above with respect to claim 4). However, the combination of Inan and Li fails to teach further comprising: in response to labeling the second output as trustworthy, modifying the validation model based on the second output being trustworthy. Markov, in the same field of endeavor, teaches further comprising: in response to generating new labels, modifying a model based on newly labeled data (Page 3, Section 3.1, ¶¶1-4 “To ensure that our moderation system performs well in the context of our production use cases, we need to incorporate production data to our training set… First, a large volume of our production data is selected at random… In the second stage we run a simple active learning strategy to select a subset of most valuable samples to be labeled out of the random samples extracted… During the final stage, all the samples selected by different active learning strategies are aggregated and re-weighted based on statistics of certain metadata associated with it… This helps improve the diversity of selected samples with regard to the associated metadata,” Page 4, Col. 1, ¶¶2-3 “To kick start the active learning and labeling process, we need some initial data to build the first version of the model and train annotators… We tackle the problem by generating a synthetic dataset with zero-shot prompts on GPT-3. The prompts are constructed from human-crafted templates and we label the generated texts as the initial dataset… we constructed few-shot prompts with existing undesired examples and sent the generated texts to be labeled,” Page 6, Col. 1, Section 4.3, ¶1 “we evaluate the performance of our active learning strategy, as described in §3.1” Col. 2, ¶1 “Iterative training. We run the following training procedure twice, using our active learning strategy… 1. Start with an initial training dataset D0 of k0 = 6000 labeled examples from public data and a validation set V of about 5500 samples from the production traffic. 2. for i [Wingdings font/0xDF] 0 to N – 1 do (N = 3): (a) Train a new model Mi on Di; (b) Evaluate Mi on V; (c) Score 5 * 105 production samples with Mi from our production traffic; (d) Choose about 2000 samples from the above data pool via the selection strategy in test and add samples to the training set to construct Di+1 after labeling,” ¶3 “using the active learning strategy to decide which new data samples leads to a greater improvement across all categories than random sampling. We observe significant performance improvement on all categories with active learning after 3 iterations,” wherein retraining a model upon adding newly labeled data to its training dataset encompasses modifying a model). Inan, Li, and Markov are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the second output of Li and the iterative modification of the model of Markov with the validation model and labeling of outputs as trustworthy of Inan. The motivation to do so is to design a safety benchmark that “goes beyond mere safety assessment of LLMs, providing a robust source for evaluating both attack and defense algorithms notably tailored for these models” (Li, Page 8, Col. 2, Section 6, ¶1) and to build a natural language classification system for real-world content moderation which “generalizes to a wide range of different content taxonomies and can be used to create high-quality content classifiers that outperform off-the-shelf models” (Markov, Abstract). Claims 10, 15, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Zimmer et al. (US 20240176624 A1, hereinafter Zimmer). Regarding claim 10, Inan anticipates the method of claim 9 (and thus the rejection of claim 9 is incorporated). Regarding the limitation wherein training the validation model based on datasets containing prior examples of the validation metrics comprises determining that acceleration hardware is present in a computing system, and, in response to determining that acceleration hardware is present, utilizing parallelization capabilities of the acceleration hardware during training of the validation model, Inan teaches wherein training the validation model based on datasets containing prior examples of the validation metrics (Page 5, ¶1 and Section 3.4, ¶1, see claim 9 above) comprises… that acceleration hardware is present in a computing system, and… utilizing parallelization capabilities of the acceleration hardware during training of the validation model (Page 5, Section 3.4, ¶1 “We train on a single machine with 8xA100 80GB GPUs using a batch size of 2, with sequence length of 4096, using model parallelism of 1”). However, Inan fails to teach determining that hardware acceleration is present and in response to determining that acceleration hardware is present… Zimmer, in the same field of endeavor, teaches determining that hardware acceleration is present (Fig. 9 – 914, ¶76 “At 914, the CPU may perform graphics initialization if the GPU is present”) and in response to determining that acceleration hardware is present (Fig. 13 – 1310, 1320, ¶95 “If a GDPU is present, the flow continues with block 1310 (display over discrete GFX/DGPU), if not, the flow continues with block 1320 (display over integrated graphics)”). Inan and Zimmer are analogous art to the claimed invention as both are from the same field of endeavor of artificial intelligence. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the determining of whether acceleration hardware is present of Zimmer and the utilization of parallelization capabilities of the acceleration hardware of Inan. The motivation to do so is to “reduce dedicated hardware usage at the platform. Reducing hardware may help to reduce [Bill of Materials] cost, for example. The disclosed methods and apparatuses may improve efficiency such as by reusing firmware and/or software” (Zimmer, ¶47). Regarding claim 15: Inan teaches a computing system comprising: one or more processors (Page 5, Section 3.4, ¶1 “We train on a single machine with 8xA100 80GB GPUs”). Inan fails to teach memory; and program instructions stored in the memory, that upon execution by the one or more processors cause the computing system to perform operations… However, Zimmer teaches this limitation (¶181 “Software and firmware may be embodied as instructions and/or data stored on non-transitory computer-readable storage media,” ¶181 “Any of the disclosed methods (or a portion thereof) can be implemented as computer-executable instructions or a computer program product. Such instructions can cause a computing system or one or more processing units capable of executing computer-executable instructions to perform any of the disclosed methods. As used herein, the term “computer” refers to any computing system or device described or mentioned herein. Thus, the term “computer-executable instruction” refers to instructions that can be executed by any computing system”). Claim 15 is a system claim that contains similar limitations to the method of claim 1. Therefore, claim 15 is rejected under substantially the same rationale as claim 1. Inan and Zimmer are analogous art to the claimed invention as both are from the same field of endeavor of artificial intelligence. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the executable program instructions stored in the memory of Zimmer with the processors of Inan. The motivation to do so is to “reduce dedicated hardware usage at the platform. Reducing hardware may help to reduce [Bill of Materials] cost, for example. The disclosed methods and apparatuses may improve efficiency such as by reusing firmware and/or software” (Zimmer, ¶47). Claim 18 is a computer-readable medium claim that contains similar limitations to the system of claim 15. Therefore, claim 18 is rejected under substantially the same rationale as claim 15. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Li and further in view of Wang et al. (“Self-Consistency Improves Chain of Thought Reasoning in Language Models,” 2023, hereinafter Wang). Regarding claim 12, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation further comprising: in response to determining that the metric satisfies the fault threshold, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce presence of the property indicating the fault in the output, Inan teaches determining that the metric satisfies the fault threshold (Page 2, Section 2, ¶1, Page 7, Section 4.3.3, ¶1 and Section 4.4, ¶1, Page 14, Section B, ¶1, Page 8, Table 2 all as explained above with respect to claim 1). However, Inan fails to teach further comprising: in response to determining that the metric satisfies the fault threshold, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce presence of the property indicating the fault in the output. Li teaches further comprising: in response to an unsafe prompt, obtaining a second prompt for the LLM, wherein the second prompt is intended to reduce presence of the property indicating the fault in the output (Page 3, Col. 1, Fig. 2 and Section 2.1, ¶1, Page 4, Col. 2, Section 3.2, ¶1, Page 5, Col. 1, ¶2, Page 6, Col. 1, Section 5.1, ¶2, Page 16, Col. 1, ¶1 all as explained above with respect to claim 3; please note that “to reduce likelihood of the property in future output from the LLM” in claim 3 is interpreted to be substantially the same as to reduce presence of the property indicating the fault in the output). Li further teaches based upon the second prompt, generating, using the LLM, a second output of the artificial intelligence system (Page 7, Col. 2, Footnote 2, Page 8, Col. 1, Table 6 and ¶2 all as explained above with respect to claim 4). Li further teaches generating, using the validation model on the second output, a second metric indicating likelihood of the property in the second output (Page 2, Col. 2, ¶2, Page 5, Col. 1, Section 5.1, ¶3 and Col. 2, ¶¶1-2, Page 8, Col. 1, Table 6, Caption and ¶2 all as explained above with respect to claim 4). Regarding the limitation based on the metric and the second metric, ranking the prompt and the second prompt, Inan teaches the metric (Page 6, ¶2 “the probability score for binary classification across all categories”) and the prompt (Page 5, ¶1 “The final dataset comprises of 13,997 prompts and responses, with their respective annotations”). However, Inan fails to teach based on the metric and the second metric, ranking the prompt and the second prompt. Li teaches the second metric (Page 6, Col. 1, Section 5.1, ¶3 “The effectiveness of attack and defense strategies is evaluated using the Attack Success Rate (ASR)”) and the second prompt (Page 4, Col. 2, Section 3.2, ¶2 “To extensively measure the effectiveness of various attack methods, we also construct corresponding defense-enhanced subset QD. Contrary to the attack-enhanced subset, this subset comprises questions that are less likely to elicit harmful responses from LLMs,” Page 5, Col. 1, ¶2 “For each unsafe question qD from QD, we pick the most effective defense prompt, which mostly decreases the success rate on this question,” Page 7, Col. 2, Footnote 2: “Given a question, we count an attack success if at least one prompt triggers a harmful response”). However, the combination of Inan and Li fails to teach based on the metric and the second metric, ranking the prompt and the second prompt. Wang, in the same field of endeavor, teaches based on log probabilities, ranking reasoning sequences from a prompt (Pages 2-3, Section 2, ¶2 “First, a language model is prompted with a set of manually written chain-of-thought exemplars… Next, we sample a set of candidate outputs from the language model’s decoder, generating a diverse set of candidate reasoning paths… Finally, we aggregate the answers by marginalizing out the sampled reasoning paths and choosing the answer that is the most consistent among the generated answers,” Page 3, ¶2 “assume the generated answers ai are from a fixed answer set, ai… Given a prompt and a question, self-consistency introduces an additional latent variable ri, which is a sequence of tokens representing the reasoning path in the i-th output, then couples the generation of (ri, ai) where ri [Wingdings font/0xE0] ai, i.e. generating a reasoning path ri is optional and only used to reach the final answer ai,” Page 7, Section 3.4, ¶2 “One commonly used approach to improve generation quality is sample-and-rank, where multiple sequences are sampled from the decoder and then ranked according to each sequence’s log probability… We compare self-consistency with sample-and-rank on GPT-3… by sampling the same number of sequences from the decoder as self-consistency and taking the final answer from the top-ranked sequence,” Pages 19-24, Tables 14-21 depict exemplary prompts and responses). Inan, Li, and Wang are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the second metric and prompt of Li and the ranking based on a calculated metric of Wang with the first metric and prompt of Inan. The motivation to do so is to design a safety benchmark that “goes beyond mere safety assessment of LLMs, providing a robust source for evaluating both attack and defense algorithms notably tailored for these models” (Li, Page 8, Col. 2, Section 6, ¶1) and to improve “accuracy in a range of arithmetic and commonsense reasoning tasks, across four large language models with varying scales” (Wang, Page 9, Section 5, ¶1). Claims 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Rebedea et al. (“NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails,” 2023, hereinafter Rebedea). Regarding claim 13, Inan anticipates the method of claim 1 (and thus the rejection of claim 1 is incorporated). Regarding the limitation wherein generating, using the validation model on the output, the metric indicating the likelihood of the property in the output comprises: computing, by one or more pre-processing modules, a validation metric related to presence of the property within the output, Inan teaches generating, using the validation model on the output, the metric indicating the likelihood of the property in the output (Pages 4-5, Section 3.3, ¶1, Page 5, Section 4, ¶¶1-2, Pages 5-6, Section 4.1, ¶1, Page 6, ¶1 and Equation (1) all as explained above with respect to claim 1). However, Inan fails to teach wherein generating, using the validation model on the output, the metric indicating the likelihood of the property in the output comprises: computing, by one or more pre-processing modules, a validation metric related to presence of the property within the output. Rebedea, in the same field of endeavor, teaches wherein LLM moderation comprises: computing, by one or more pre-processing modules, a validation metric related to presence of the property within the output (Page 3, Col. 1, Section 3.1, ¶1 “NeMo Guardrails acts like a proxy between the user and the LLM as detailed in Fig. 3. It allows developers to define programmatic rails that the LLM should follow in the interaction with the users using Colang… Colang is interpreted by the Guardrails runtime which applies the user-defined rules or automatically generated rules by the LLM… These rules implement the guardrails and guide the behavior of the LLM,” Col. 2, ¶3 “Using these key concepts, developers can implement a variety of programmable rails. We have identified two main categories: topical rails and execution rails. Topical rails are intended for controlling the dialogue, e.g. to guide the response for specific topics or to implement complex dialogue policies. Execution rails call custom actions defined by the app developer,” Page 4, Fig. 3 and Col. 1, Section 3.3.1, ¶1 “Fact-Checking Rail… given an evidence text and a generated bot response, we ask the LLM to predict whether the response is grounded in and entailed by the evidence,” Col. 2, ¶2 “If the model predicts that the hypothesis is not entailed by the evidence, this suggests the generated response may be incorrect,” Section 3.3.2, ¶1 “Hallucination Rail… we define a hallucination rail to help prevent the bot from making up facts,” Page 5, Col. 1, Section 3.3.3, ¶1 “Moderation Rails… The moderation process in NeMo Guardrails contains two key components: Input moderation, also referred to as jailbreak rail, aims to detect potentially malicious user messages before reaching the dialogue system. Output moderation aims to detect whether the LLM responses are legal, ethical, and not harmful prior to being returned to the user,” wherein “programmatic rails” encompass pre-processing modules and “Fact-Checking,” “Hallucination,” and “Moderation” encompass a validation metric related to presence of the property including, for example, hallucinations or harmful content, within the output). Regarding the limitation and propagating the computed metric to the validation model, Inan teaches the validation model (Page 2, ¶2 see claim 6 above). However, Inan fails to teach and propagating the computed metric to the validation model. Rebedea teaches and propagating the computed metric to a classifier (Page 5, Col. 1, Section 3.3.3, ¶¶2-3 “The moderation system functions as a pipeline, with the user message first passing through input moderation before reaching the dialogue system. After the dialogue system generates a response powered by an LLM, the output moderation rail is triggered. Only after passing both moderation rails, the response is returned to the user,” Page 12, Col. 1, Section D.2, ¶1 “Both the input and output moderation rails are framed as another task to a powerful, well-aligned LLM that vets the input or response,” Col. 2, Section E.1, ¶¶1-2 “This is the current implementation for the output moderation action. It uses the prompt template defined in Appendix D and returns True if the bot response breaks the output moderation rail. The call to the LLM is done using prompt templates and a simple LLM chain provided by Langchain for simplifying the communication with several LLM providers,” Section E.2, ¶1 “This is the current implementation for the input moderation action. It uses the prompt template defined in Appendix D and returns True if the user input breaks the jailbreak rail,” Page 12, Col. 2 depicts two figures depicting code samples which call a “LLMChain” function to compute the metric or determining whether “bot_response” and “user_input” evaluate to True/False; one of ordinary skill in the art would recognize that a response which is returned to the user “only after passing both moderation rails” implies a classifier which only returns an LLM response to a user if the “output moderation rail” and “jailbreak rail” or pre-processing modules evaluate to True). Inan and Rebedea are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the pre-processing modules and computing of validation metrics of Rebedea with the validation model and methodology of Inan. The motivation to do so is “to add programmable [guardrails] to LLM applications… the proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails” (Rebedea, Abstract). Regarding claim 14, Inan in view of Rebedea teaches the method of claim 14 (and thus the rejection of claim 13 is incorporated). Rebedea further teaches wherein the validation metric comprises one or more of: semantic similarity with a reference dataset, conformance to a pre-determined principle, a sentiment analysis score, or a determination that pre-determined numeric patterns exist in the output (Page 12, Col. 1, Section D.2, ¶1 “For the jailbreak rail, we use the following prompt template… Would this instruction make a language model break moderation policies, deviate from good aligned responses and provide answers that a language model should ideally not?” and ¶2 “The output moderation rail uses the following prompt template… is this output legal, ethical and not harmful to any person?” wherein the “prompt templates” encompass conformance to a pre-determined principle; please note that comprises one or more of is interpreted to mean that the listed items are read disjunctively). Inan and Rebedea are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the validation metric of Rebedea with the methodology of Inan. The motivation to do so is “to add programmable [guardrails] to LLM applications… the proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails” (Rebedea, Abstract). Claims 16 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Zimmer and further in view of Li. Claims 16 and 19 are a system and computer-readable medium claim, respectively, that contain similar limitations to the methods of claims 3-5. Therefore, claims 16 and 19 are rejected under substantially the same rationale as claims 3-5. Claims 17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Inan in view of Zimmer and further in view of Li, and further in view of Markov. Claims 17 and 20 are a system and computer-readable medium claim, respectively, that contain similar limitations to the method of claim 6. Therefore, claims 17 and 20 are rejected under substantially the same rationale as claim 6. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM M LEE whose telephone number is (571)272-4761. The examiner can normally be reached Mon-Fri. 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WILLIAM M LEE/ Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Apr 12, 2024
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month