Prosecution Insights
Last updated: August 15, 2026
Application No. 18/617,994

SYSTEM AND METHOD FOR EVALUATING GENERATIVE ARTIFICIAL INTELLIGENCE MODELS

Non-Final OA §101§103
Filed
Mar 27, 2024
Priority
Mar 28, 2023 — provisional 63/492,674 +1 more
Examiner
KAPOOR, DEVAN
Art Unit
Tech Center
Assignee
Mastercard Technologies Canada Ulc
OA Round
1 (Non-Final)
7%
Grant Probability
At Risk
1-2
OA Rounds
1y 11m
Est. Remaining
18%
With Interview

Examiner Intelligence

Grants only 7% of cases
7%
Career Allowance Rate
1 granted / 14 resolved
-52.9% vs TC avg
Moderate +11% lift
Without
With
+11.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
24 currently pending
Career history
45
Total Applications
across all art units

Statute-Specific Performance

§101
39.2%
-0.8% vs TC avg
§103
50.2%
+10.2% vs TC avg
§102
7.4%
-32.6% vs TC avg
§112
2.5%
-37.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§101 §103
DETAILED ACTION This action is responsive to the application filed on 03/27/2024. Claims 1-20 are pending and have been examined. This action is Non-final. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1: This claim is directed to a system, which is a machine, one of the four statutory categories of invention. Therefore, the claim satisfies Step 1. Step 2A Prong 1: (a) “organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity” - This limitation is directed to a mathematical concept, as organizing data points according to their proximity/distance relationships to one another is a mathematical relationship and/or mathematical calculation. (b) “selecting labeled features in the feature space within a radius of the test feature” - This limitation is directed to a mathematical concept (comparing numerical distance values against a defined radius) and/or a mental process, as a person could evaluate which data points fall within a specified distance of another point using pen and paper or in their mind, given a small enough data set. (c) “computing a first metric corresponding to a count of the selected labeled features” - This limitation is directed to a mathematical concept, as counting a number of selected features is a mathematical calculation that could also be performed mentally or with pen and paper. (d) “computing a second metric corresponding to distances between the selected labeled features and the test feature” - This limitation is directed to a mathematical concept, as computing distances between data points is a mathematical calculation (e.g., Euclidean or Manhattan distance), and thus the limitation is directed to math. (e) “computing a risk score based on the first metric and the second metric” - This limitation is directed to a mathematical concept, as combining two numerical metrics into a single score is a mathematical formula/calculation. (f) “in response to the risk score being above a threshold, assigning a first label to the test output...and in response to the risk score not being above the threshold, (i) assigning a second label to the test output” - This limitation is directed to a mental process, as comparing a computed value to a threshold and making an evaluation/judgment/observation based on that comparison is a task that can practically be performed in the human mind or with the aid of pen and paper. Step 2A Prong 2 and Step 2B: (a) “A system comprising: memory hardware storing instructions; and one or more electronic processors configured to execute the instructions” - This limitation is directed to generic computer components recited at a high level of generality and amounts to no more than mere instructions to apply the exception using generic computer components, which cannot integrate the judicial exception into a practical application, nor can it provide significantly more than the judicial exception (see MPEP 2106.05(f)). (b) “providing a plurality of training inputs to a first artificial intelligence model to generate a plurality of training outputs…providing a test input to the first artificial intelligence model to generate a test output” - These limitations amount to necessary data-gathering steps and insignificant extra-solution activity performed by a generically-recited “first artificial intelligence model,” used according to its known, conventional function of generating an output from an input. This cannot integrate the abstract idea into a practical application, nor provide significantly more under Step 2B (see MPEP 2106.05(g) and 2106.05(d)). (c) “preprocessing the plurality of training inputs and/or training outputs…preprocessing the test output” - These limitations amount to insignificant extra-solution data-gathering/data-manipulation activity performed at a high level of generality, and are also well-understood, routine, and conventional data preprocessing activity. This cannot integrate the abstract idea into a practical application, nor provide significantly more under Step 2B (see MPEP 2106.05(g) and 2106.05(d)). (d) “labeling one or more of the plurality of preprocessed training inputs and/or training outputs” - This limitation is directed to labelling preprocessed training inputs/outputs. The limitation is recited in a high level of generality, and thus it does not integrate to a practical application, nor provide significantly more than (e) “adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model” - This limitation invokes a generically-recited “second artificial intelligence model” performing its known, conventional function of embedding data into a feature space. This amounts to no more than applying a generic, conventional tool to implement the abstract idea, and cannot integrate the abstract idea into a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(f)). (f) “transmitting the test output to a user device” - This limitation amounts to insignificant post-solution activity of merely transmitting data, and cannot integrate the abstract idea into a practical application, nor provide significantly more than the judicial exception under 2B (see MPEP 2106.05(g) and (2106.05(d)). Thus, claim 1 is non-patent eligible. Claim 11 is analogous to claim 1, aside from claim type, and thus the same rejection applies as above. Regarding claim 2: Step 1: The claim depends from claim 1, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. There are no additional elements to be evaluated under Step 2A Prong 1. Step 2A Prong 2 and Step 2B: (a) “the first artificial intelligence model includes a large language model” - This limitation merely narrows the generic “first artificial intelligence model” of claim 1 to a particular, but still generically-recited and conventionally-used, type of AI model. This amounts to limiting the abstract idea to a particular technological environment/field of use, which cannot integrate the abstract idea into a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(h)). (b) “the plurality of training inputs includes one or more input text strings; and the plurality of training outputs includes one or more output text strings” - These limitations merely narrow the data type acted upon by the abstract idea to text strings, without adding any additional technical element. This amounts to insignificant application of the abstract idea to a particular data type/field of use, which cannot integrate the abstract idea into a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(h)). Thus, claim 2 is non-patent eligible. Claim 12 is analogous to claim 2, aside from claim type, and thus the same rejection applies as above. Regarding claim 3: Step 1: The claim depends from claim 2, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. There are no additional elements to be evaluated under Step 2A Prong 1. Step 2A Prong 2 and Step 2B: “wherein preprocessing the plurality of training inputs and/or training outputs includes applying at least one of stemming, lemmatization, stop word removal, part-of-speech tagging, and tokenization operations to the plurality of training inputs and/or training outputs” - The limitation recites the application of these known, generic preprocessing techniques, for which amounts to insignificant extra-solution activity performed using well-understood, routine, and conventional technology stemming, lemmatization, stop word removal, part-of-speech tagging, and tokenization are well-understood, routine, and conventional natural language processing techniques, as acknowledged in the specification's background discussion of standard NLP processing steps. The limitation thus cannot integrate into a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(g) and 2106.05(d)). Thus, claim 3 is non-patent eligible. Claim 13 is analogous to claim 3, aside from claim type, and thus the same rejection applies as above. Regarding claim 4: Step 1: The claim depends from claim 3, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. Step 2A Prong 1: “The system of claim 3, wherein labeling one or more of the plurality of preprocessed training inputs and/or training outputs includes at least one of: assigning a third label to each preprocessed training input and/or training output that contains a term present in a list; assigning a fourth label to each preprocessed training input and/or training output marked by a user; and assigning a fifth label to each preprocessed training input and/or training output conforming to a criteria.” -- The limitation is directed as determining whether a text string contains a term from a list, and labeling it accordingly, and evaluating whether data conforms to a criterion, which is an evaluation that could practically be performed in the human mind or with pen and paper, and thus the limitation is directed to a mental process. There are no elements to be evaluated under Step 2A Prong 2 and Step 2B. Thus, claim 4 is non-patent eligible. Claim 14 is analogous to claim 4, aside from claim type, and thus the same rejection applies as above. Regarding claim 5: Step 1: The claim depends from claim 4, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. Step 2A Prong 1: (a) “generating feature vectors corresponding to each of the preprocessed training inputs and/or preprocessed training outputs” - This limitation is directed to a mathematical concept, as transforming preprocessed data into feature vectors (numerical representations) is a mathematical operation/transformation. (b) “generating edges between the nodes based on distances between feature vectors corresponding to the nodes” - This limitation is directed to a mathematical concept, as computing distances between feature vectors, and defining relationships (edges) based on those computed distances, is a mathematical calculation. Step 2A Prong 2 and Step 2B: (a) “generating a graph with each feature vector as a node” - This limitation amounts to organizing data into a particular, but generically-recited, data structure (a graph). This is insignificant extra-solution activity of mere data organization/gathering, as well as a well-understood, routine and conventional activity (WURC) that cannot provide significantly more than the judicial exception that does not integrate the abstract idea into a practical application, nor does it provide significantly more than the judicial exception (see MPEP 2106.05(g) and 2106.05(d)). (b) “applying a graph neural network to the graph to allow each node to aggregate information from its neighbors” - This limitation invokes a generically-recited “graph neural network” performing its own known, conventional function of neighbor information aggregation. The specification (~[0021-0022, 0069] itself describes graph neural networks as an existing, known machine-learning tool applied to implement the recited feature-space organization, rather than disclosing any improvement to the internal architecture or technical operation of the graph neural network itself. This amounts to no more than applying a generic, well-understood, routine, and conventional machine-learning tool to implement the abstract idea, and cannot integrate the abstract idea into a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(g/d) and 2106.05(f)). (c) “training the graph neural network using the labeled features” - This limitation recites training a graph neural network using labeled features. The limitation amounts to an insignificant extra-solution activity/mere instructions to apply a generic training process and cannot integrate the abstract idea into a practical application, as well as applying a well-understood, routine, and conventional machine-learning model-training technique, recited at a high level of generality without any specific, unconventional training methodology, which furthermore does not provide significantly more than the judicial exception (see MPEP 2106.05(g) and 2106.05(d)). Thus, claim 5 is non-patent eligible. Claim 15 is analogous to claim 5, aside from claim type, and thus the same rejection applies as above. Regarding claim 6: Step 1: The claim depends from claim 5, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. Step 2A Prong 1: “generating feature vectors corresponding to each of the preprocessed training inputs and/or preprocessed training outputs includes transforming text of each of the preprocessed training inputs and/or preprocessed training outputs into a numerical vector in a high-dimensional space” -- The limitation is directed to generating vectors corresponding to training inputs/outputs that includes transforming text of the inputs/outputs into a numerical vector in high dimensional space. Transforming text data into numerical vector representations (such as TF-IDF vectors, word embeddings, or similar mathematical transformations) constitutes a mathematical operation and mathematical formula. The transformation of abstract data into numerical vectors using mathematical operations, and thus is directed to math. There are no elements to be evaluated under Step 2A Prong 2 and Step 2B. Thus, claim 6 is non-patent eligible. Claim 16 is analogous to claim 6, aside from claim type, and thus the same rejection applies as above. Regarding claim 7: Step 1: The claim depends from claim 4, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. Step 2A Prong 1: (a) “The system of claim 4, wherein organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model includes:” - The limitation recites that the organizing of the preprocessed training inputs in feature space based on the proximity using the second AI model includes generating feature vectors corresponding to each of the preprocessed training inputs and/or training outputs” – This limitation is directed to a mathematical concept. Transforming preprocessed data into feature vectors is a mathematical operation/transformation, as discussed with respect to claims 5 and 6, and thus the limitation is directed to math. (b) “computing a distance matrix representing pairwise distances between feature vectors” - This limitation is directed to computing pairwise distances between vectors and organizing them into a matrix representation. Hence the limitation is directed to a mathematical concept. (c) “assigning each node to a cluster corresponding to a nearest cluster center” - This limitation is directed to determining which cluster center is nearest to a given node and assigning membership accordingly involves comparing numerical values and making an evaluation, which is a task that could be performed in the human mind, using observation and judgement, with aid of pen and paper, and thus it is directed to a mental process. (d) “applying a clustering algorithm to the graph to determine cluster centers of clusters of nodes” - This limitation invokes a generic clustering algorithm performing its known, conventional function. Applying a standard clustering algorithm (such as k-means or other well-known clustering techniques) to determine cluster centers is a process that can be performed in the human mind using evaluation, observation, and judgement, with aid of pen and paper, and thus the limitation is directed to a mental process. Step 2A Prong 2 and Step 2B: (a) “constructing a graph with each node representing a feature vector, wherein edges are drawn between nodes that are proximate in the feature space” - This limitation is directed to organizing data into a data structure (a graph) based on mathematical proximity relationships. This is insignificant extra-solution activity of data organization that cannot integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(g)). Thus, claim 7 is non-patent eligible. Claim 17 is analogous to claim 7, aside from claim type, and thus the same rejection applies as above. Regarding claim 8: Step 1: The claim depends from claim 4, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. There are no elements to be evaluated under Step 2A Prong 1. Step 2A Prong 2 and Step 2B: “The system of claim 4, wherein: the test input includes one or more input text strings, and the test output includes one or more output text strings.” -- The limitation recites that the test inputs will include one or more input text strings, and same for the test output. The limitation amounts to no more than mere further limiting to a field of use/environment, and thus does not integrate to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(h)). Thus, claim 8 is non-patent eligible. Claim 18 is analogous to claim 8, aside from claim type, and thus the same rejection applies as above. Regarding claim 9: Step 1: The claim depends from claim 8, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. There are no elements to be evaluated under Step 2A Prong 1. Step 2A Prong 2 and Step 2B: “The system of claim 8, wherein preprocessing the test output includes applying at least one of stemming, lemmatization, stop word removal, part-of-speech tagging, and tokenization operations to the test output.” -- The limitation recited stemming, lemmatization, stop word removal, part-of-speech tagging, and tokenization, and applying these preprocessing techniques. The limitation is directed to insignificant extra-solution activity that does not integrate the judicial exception into a practical application and are well-understood, routine, and conventional data preprocessing operations in the field of natural language processing, and thus nor does it provide significantly more than the judicial exception under Step 2B (see MPEP 2106.05(g) and 2106.05(d)). Thus, claim 9 is non-patent eligible. Claim 19 is analogous to claim 9, aside from claim type, and thus the same rejection applies as above. Regarding claim 10: Step 1: The claim depends from claim 8, which is directed to a system, a machine, one of the four statutory categories. Therefore, the claim satisfies Step 1. Step 2A Prong 1: “The system of claim 9, wherein the second metric represents a sum of distances between the test feature and each selected label feature.” -- The limitation is directed to computing a sum of distance values between the test and selected label features. The limitation is directed to the use of mathematical concept/calculation/operation, and thus the limitation is directed to math. There are no elements to be evaluated under Step 2A Prong 2 and Step 2B. Thus, claim 10 is non-patent eligible. Claim 20 is analogous to claim 101 aside from claim type, and thus the same rejection applies as above. Top of Form Bottom of Form Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over US20220366280A1, by Rowe et. al. (referred herein by Rowe) in view of NPL reference “To Trust Or Not To Trust A Classifier”, by Jiang et. al. (referred herein as Jiang) in view of NPL reference “LUNAR:Unifying Local Outlier Detection Methods via Graph Neural Networks”, by Goodge et. al. (referred herein as Goodge). Regarding claim 1, Rowe teaches: A system comprising: memory hardware storing instructions; and one or more electronic processors configured to execute the instructions, wherein the instructions include: providing a plurality of training inputs to a first artificial intelligence model to generate a plurality of training outputs, ([Rowe, [0040]] “The data analysis platform 110 includes a machine learning engine 111 that trains a machine learning model 112 based on a machine learning model training subset of data 122. The machine learning model 112 is trained to receive a data point including a set of values as an input and to generate a label for the data point. The label may be referred to as a prediction.”, AND [Rowe, 0043] “The data application layer 130 includes system attributes 131 and system data 132. System attributes 131 may include, for example, attributes of one or more processors, data storage devices, network interfaces, operating systems, and applications.”, wherein the examiner interprets Rowe’s system attributes for processors and its operations and the model being “trained to receive a data point including a set of values as an input and to generate a label for the data point” to be the same as the claimed instructions including the subsequently recited operations and providing the training inputs to a model for the purpose of obtaining an output because they are both directed to stored executable instructions defining the operations performed by the processor-based system.) providing a test input to the first artificial intelligence model to generate a test output,preprocessing the test output, ([Rowe, [0055]] “A system obtains a dataset and trains a machine learning model using a first subset of data from the dataset (Operation 202). The system trains the machine learning model to generate a particular prediction, of a particular label, associated with a particular input data point. The first subset of data includes a plurality of data points input to the model to train the model.”, AND [Rowe, [0043]] “The data analysis platform 110 may convert the received system attributes 131 into a data point to provide to the trained machine learning model 112”, wherein the examiner interprets “first artificial intelligence model” to be the same as “machine learning model using a first subset of data from the dataset”, because they are both feeding a test/target instance to the first model to obtain the test output (predicted label/test sample) that will subsequently be risk-scored. Also they are both transforming the test instance into the same feature representation used for the training data (converting it “into a data point”/a chosen “representation of the data”) before it is scored.) selecting labeled features in the feature space within a radius of the test feature, ([Rowe, [0026]] “the number of data points within a threshold distance from the data point for which the confidence score is being calculated.”, wherein the examiner interprets “within threshold distance” to the be same as “feature space within a radius of the test feature”, because they are both selecting the reference (labeled) points that fall within a bounded distance, a radius/threshold distance/ball, around the test point.) computing a first metric corresponding to a count of the selected labeled features, ([Rowe, [0035]] “the confidence score takes into account (1) a number of nearest neighbors in a machine learning model training data set, (2) an accuracy of the nearest neighbors, and (3) a distance to the nearest neighbors.”, wherein the examiner interprets metric corresponding to a count” to be the same as “a score…takes into account .. .number of nearest neighbors” to be the same as because they are both computing a count of the selected neighbouring labeled points as one component of the score.) computing a second metric corresponding to distances between the selected labeled features and the test feature, ([Rowe, [0025]] “the distance of each of the k data points from the target data point are used to compute a confidence score.”, wherein the examiner interprets “distance of k data points… from the target data point” to be the same as “distance between the selected labeled features and test feature”, because they are both computing a count of the selected neighbouring labeled points as one component of the score.) computing a risk score based on the first metric and the second metric, ([Rowe, [0036]] “factoring into the confidence value: (1) the number of nearest neighbors in a machine learning model training data set, (2) the accuracy of the nearest neighbors, and (3) the distance to the nearest neighbors.” AND [Rowe, 0080]] “For example, in a system in which taking an action based on a prediction carries little risk, the system may be configured to take action when the confidence score is greater than 50%. Alternatively, in an environment in which taking action incorrectly is determined to be a greater risk than not taking action, the system may be configured to take action only when the confidence score is greater than 90%.” wherein the examiner interprets “confidence value” and the fact the system will be aware that prediction will be at risk to the confidence score and will be configured to take action “when the score is greater than 50%” to be the same as “risk score” because they are both combining the neighbour-count metric and the neighbour-distance metric into a single scalar score that expresses how anomalous /(un)trustworthy/hallucinated (i.e., how risky) the test output is.) in response to the risk score not being above the threshold, (i) assigning a second label to the test output and ([Rowe, [0031]] “A set of confidence scores may be used to determine whether to re-train a machine learning model. For example, if the system determines that the confidence scores for 80% of the last one hundred predicted labels did not meet a threshold confidence score, the system may trigger a re-training of the machine learning model to update the model based on more recently-received data”, wherein the examiner interprets “risk score not .. above the threshold…assigning a second label…”, because they are both applying the opposite branch of the same threshold test, when the score signals low risk (high confidence / normal / grounded in valid information), the output is treated as valid /non-erroneous.) (ii) transmitting the test output to a user device, wherein the second label is indicative of a non-erroneous output from the first artificial intelligence model. ([Rowe, [0044]] “the data application layer 130 may display on a graphical user interface (GUI) of the user interface 136 the prediction and the confidence value.” wherein the examiner interprets “output to a user device … indicative of a non-erroneous output” to be the same as “display a GUI of the … prediction and confidence value”, because they are both delivering the model’s output to the user, displaying it on a GUI/showing it to users, for the user to see and act on.) Rowe does not teach preprocessing the plurality of training inputs and/or training outputs, labeling one or more of the plurality of preprocessed training inputs and/or training outputs, … organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model, .. adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model. Jiang teaches: preprocessing the plurality of training inputs and/or training outputs, [Jiang, page 3, sec 3] “we first pre-process the training data, as described in Algorithm 1, to find the α-high-density-set of each class”, wherein the examiner interprets “pre-process the training data” to be the same as “preprocessing the plurality of training inputs” because they are both applying pre-processing steps (density filtering / n-gram tokenization / TF-IDF and probability features) to the training data before it is placed into the reference feature space. ) labeling one or more of the plurality of preprocessed training inputs and/or training outputs, ([Jiang, page 2, sec 1] “we use a set of labeled examples (e.g. training data or validation data) to help determine a classifier’s trustworthiness for a particular testing example.”, wherein the examiner interprets “labeled examples” to be the same as “labeling one or more of the plurality of preprocessed training inputs”, because they are both attaching labels (including a valid vs. hallucination/erroneous label) to the training data points that form the labeled reference set used for scoring.) Jiang does not teach organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model,...adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model. Goodge teaches: organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model, ([Goodge, page 6740] “LUNAR Methodology … we represent a set of data as a graph, with a node corresponding to each data sample and directed edges connecting a target node to a set of source nodes, which are the nearest neighbours of the samples”, and [Goodge, page 1] “we propose LUNAR, a novel, graph neural network-based anomaly detection method”, wherein the examiner interprets “we represent a set of data as a graph” and “graph neural network-based anomaly detection” to be the same as “feature space based on proximity” , because they are both placing the data points as features in a common feature space organized by proximity (nearest neighbours) and operated on by a second model (a GNN/confidence estimator), including both labeled reference points and unlabeled/unseen test points.) adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model, ([Goodge, 6739-6740] “each data sample corresponds to one node in a graph and node i is connected to each of its k nearest neighbours, j ∈ Ni, via a directed edge (j, i) … For a given target node, the network utilises information from its nearest neighbouring nodes to learn its anomaly score.”, wherein the examiner interprets “adding the preprocessed test output to the feature space as a test feature using the second artificial intelligence model” to be the same as “the network utilises information from its nearest neighbouring nodes to learn its anomaly score”, because they are both inserting the test instance as a new node/point (the “target node”) into the proximity graph/feature space, where the second model (the GNN) then processes it via its neighbours.) Rowe, Jiang, Goodge, and the instant application are analogous art because they are all directed to evaluating the reliability of machine-learning model outputs using training data represented in a feature space and proximity relationships. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the machine learning platform disclosed by Rowe to include the use of data labeling technique disclosed by Jiang. One would have been motivated to do so to effectively improve the identification of model outputs that are likely correct or incorrect using labeled training examples and distances calculated in a selected feature representation, as suggested by Jiang ([Jiang, page 1] “We show empirically that high (low) trust scores produce surprisingly high precision at identifying correctly (incorrectly) classified examples, consistently outperforming the classifier’s confidence score as well as many other baselines.”). It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include a GNN method disclosed by Goodge. One would have been motivated to do so to effectively provide a trainable and adaptable mechanism for combining proximity and nearest-neighbor distance information into an anomaly score for a test example, as suggested by Goodge ([Goodge, page 1] “We show that our method performs significantly better than existing local outlier methods, as well as state-of-the-art deep baselines.”). Claim 11 is analogous to claim 1, aside from claim type and minute differences, and thus the same rejection applies as above. Claim(s) 2-3, and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Rowe in view Jiang in view of Goodge in view of NPL reference “SELFCHECKGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models”, by Manakul et. al. (referred herein as Manakul). Regarding claim 2, Rowe, Jiang, Goodge teach The system of claim 1, (see rejection of claim 1). Rowe, Jiang, Goodge teach do not teach wherein: the first artificial intelligence model includes a large language model; the plurality of training inputs includes one or more input text strings; and the plurality of training outputs includes one or more output text strings. Manakul teaches: wherein: the first artificial intelligence model includes a large language model; ([Manakul, page 1] “Large Language Models (LLMs) such as GPT-3 (Brown et al., 2020) and PaLM (Chowdhery et al., 2022) are capable of generating fluent and realistic responses to a variety of user prompts.”, wherein the examiner interprets the disclosed GPT-3 and PaLM large language models to be the same as the first artificial intelligence model including a large language model because they are both artificial-intelligence models configured to receive language-based inputs and generate language-based responses.) the plurality of training inputs includes one or more input text strings; ([Manakul, page 3] “draws a further N stochastic LLM response samples {S1, S2, ..., Sn, ..., SN} using the same query”, and “generating fluent and realistic responses to a variety of user prompts”, wherein the examiner interprets “user prompts” and “user query” to be the same as one or more input text strings because they are both textual inputs provided to a large language model for generating corresponding language responses. Rowe supplies the previously mapped plurality of training inputs.) and the plurality of training outputs includes one or more output text strings ([Manakul, page 1] “using GPT-3 to generate passages about individuals from the WikiBio dataset, and manually annotate the factuality of the generated passages.”, wherein the examiner interprets “passages” and “generated passages” to be the same as one or more output text strings because they are both sequences of textual content generated as outputs by a large language model.) Rowe, Jiang, Goodge, Manakul, and the instant application are analogous art because they are both directed to evaluating outputs generated by large language models using textual inputs and textual outputs. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 1 disclosed by Rowe, Jiang, and Goodge to include the process of generating fluent and realistic responses to a variety of user prompts disclosed by Manakul. One would be motivated to do so to effectively enable the confidence evaluation system of Rowe to evaluate predictions generated by large language models using textual prompts and textual responses, as suggested by Manakul ([Manakul, page 1] “Large Language Models (LLMs) such as GPT-3 (Brown et al., 2020) and PaLM (Chowdhery et al., 2022) are capable of generating fluent and realistic responses to a variety of user prompts.”). Claim 12 is analogous to claim 2, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 3, Rowe, Jiang, Goodge, Manakul teach The system of claim 2, (see rejection of claim 1) Manakul further teaches wherein preprocessing the plurality of training inputs and/or training outputs includes applying at least one of stemming, lemmatization, stop word removal, part-of-speech tagging, and tokenization operations to the plurality of training inputs and/or training outputs ([Manakul, page 3] “ Given the LLM’s response R, let i denote the i-th sentence in R, j denote the j-th token in the i-th sentence, J is the number of tokens in the sentence, and pij be the probability of the word generated by the LLM at the j-th token of the i-th sentence”, wherein the examiner interprets “the j-th token in the i-th sentence” to be the same as tokenization operations because they are both directed to representing text as individual tokens for subsequent processing by a machine-learning model.) Rowe, Jiang, Goodge, Manakul, and the instant application are analogous art because they are all directed to preprocessing textual inputs and outputs for processing by a large language model. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system claim 2 disclosed by Rowe, Jiang, Goodge, and Manakul to include the tokenization process disclosed by Manakul. One would have been motivated to do so to effectively prepare textual inputs and outputs for subsequent machine-learning processing by representing the text as individual tokens, as suggested by Manakul ([Manakul, page 3] “let i denote the i-th sentence in R, j denote the j-th token in the i-th sentence, J is the number of tokens in the sentence, and pij be the probability of the word generated by the LLM at the j-th token of the i-th sentence.”). Claim 13 is analogous to claim 3, aside from claim type and minute differences, and thus the same rejection applies as above. Claim(s) 4-10, and 14-20 are rejected under 35 U.S.C. 103 as being unpatentable over Rowe in view Jiang in view of Goodge in view of Goodge in view of Manakul further NPL reference “A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation”, by Liu et. al. (referred herein as Liu). Regarding claim 4, Rowe, Jiang, Goodge, Manakul teach The system of claim 3, (see rejection of claim 3). Manakul further assigning a fifth label to each preprocessed training input and/or training output conforming to a criteria. ([Manakul, page 3] “We design SelfCheckGPT to predict the hallucination score of the i-th sentence, S(i), such that S(i) ∈ [0.0, 1.0] where S(i) → 0.0 if the i-th sentence is grounded in valid information and S(i) → 1.0 if the i-th sentence is hallucinated.3 The following subsections will describe each of the SelfCheckGPT variants”, wherein the examiner interprets a sentence having a hallucination score approaching 0.0 or 1.0 to be the same as a preprocessed training input and/or training output conforming to a criterion because they are both directed to evaluating textual data according to a defined condition that distinguishes text grounded in valid information from hallucinated text.) Rowe, Jiang, Goodge, Manakul do not teach wherein labeling one or more of the plurality of preprocessed training inputs and/or training outputs includes at least one of: assigning a third label to each preprocessed training input and/or training output that contains a term present in a list; assigning a fourth label to each preprocessed training input and/or training output marked by a user; and assigning a fifth label to each preprocessed training input and/or training output conforming to a criteria. Liu teaches: wherein labeling one or more of the plurality of preprocessed training inputs and/or training outputs includes at least one of: ([Liu, page 6724] “We formulate our hallucination detection task as a binary classification task. As shown in Fig 1, our goal is to assign either a “hallucination” (abbreviated as “H”) or a “not hallucination” (abbreviated as “N “) label to the highlighted spans.”, wherein the examiner interprets “assign either a ‘hallucination’ (abbreviated as ‘H’) or a ‘not hallucination’ (abbreviated as ‘N’) label to the highlighted spans” to be the same as labeling one or more of the plurality of preprocessed training inputs and/or training outputs because they are both directed to assigning classification labels to processed textual data according to whether the textual data contains hallucinated or non-hallucinated information.) assigning a third label to each preprocessed training input and/or training output that contains a term present in a list; ([Liu, page 6723] “To mitigate label imbalance during annotation, we utilize an iterative model-in-loop strategy. We conduct comprehensive data analyses and create multiple baseline models.” AND [Liu, page 6725] “we do not mask stop words or punctuation identified by NLTK (Bird, 2006).”, wherein the examiner interprets “stop words or punctuation identified by NLTK” and “To mitigate label imbalance during annotation, we utilize an iterative model-in-loop strategy” to be the same as a term present in a list because they are both directed to identifying particular textual terms through a predefined lexical resource containing recognized stop words or punctuation.) assigning a fourth label to each preprocessed training input and/or training output marked by a user; and ([Liu, page 6724] “We then ask human annotators to assess whether the perturbed text spans are hallucinations given the original text”, and [Liu, page 6724] “To this end, we propose a reference-free, token level hallucination detection task and introduce an annotated training and benchmark testing dataset that we call HADES (Hallucination Detection dataSet). The reference-free property of this task yields greater flexibility in a broad range of generation applications. We expect the token-level property of this task to foster the development of models that can detect fine-grained signals of potential hallucination.… assign either a “hallucination” … or a “not hallucination” … label to the highlighted spans”, wherein the examiner interprets the disclosed human annotators to be the same as the claimed user because they are both human individuals who review textual data and provide an assessment, and further interprets “assign either a ‘hallucination’ … or a ‘not hallucination’ … label to the highlighted spans” to be the same as assigning a fourth label to each preprocessed training input and/or training output marked by a user because they are both directed to a human reviewer assigning a classification label to identify textual content based on the reviewer’s assessment of that content.) Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are all directed to evaluating and labeling textual outputs generated by artificial-intelligence models. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 3 disclosed by Rowe, Jiang, Goodge, and Manakul to include the human “binary classification task” of labeling ‘hallucination’ (abbreviated as ‘H’) or a ‘not hallucination’ (abbreviated as ‘N’) as disclosed by Liu. One would have been motivated to do so to effectively produce reliable classification labels for processed textual examples and thereby distinguish erroneous model outputs from non-erroneous model outputs using human-reviewed reference data, as suggested by Liu ([Liu, page 6724] “We then ask human annotators to assess whether the perturbed text spans are hallucinations given the original text”). Claim 14 is analogous to claim 4, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 5, Rowe, Jiang, Goodge, Manakul, and Liu teaches The system of claim 4, (see rejection of claim 4). Goodge further teaches: wherein organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model includes ([Goodge, page 6740] “We represent a set of data as a graph, with a node corresponding to each data sample and directed edges connecting a target node to a set of source nodes, which are the nearest neighbours of the samples.”), wherein the examiner interprets “We represent a set of data as a graph, with a node corresponding to each data sample and directed edges connecting a target node to a set of source nodes, which are the nearest neighbours of the samples” to be the same as organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model because they are both arranging data samples as nodes within a graph-based feature space according to nearest-neighbor proximity relationships, using a graph neural network model.) generating feature vectors corresponding to each of the preprocessed training inputs and/or preprocessed training outputs; ([Goodge, page 6740] “We construct the k-NN graph of any feature-based, tabular dataset, rather than being restricted to graph datasets”, wherein the examiner interprets “feature-based, tabular dataset, rather than being restricted to graph datasets.” to be the same as generating vectors that correspond to preprocessed training inputs/outputs, because they are both representing each preprocessed example as a numerical feature vector (a point in a multi-dimensional/feature space). generating a graph with each feature vector as a node; ([Goodge, page 6740] “We represent a set of data as a graph, with a node corresponding to each data sample and directed edges connecting a target node to a set of source nodes, which are the nearest neighbours of the samples.” AND [Goodge, page 6739] “each data sample corresponds to one node in a graph”, wherein the examiner interprets “We represent a set of data as a graph, with a node corresponding to each data sample” to be the same as generating a graph with each vector as a node, because they are both building a graph in which each data sample / feature vector is represented as a node.) generating edges between the nodes based on distances between feature vectors corresponding to the nodes; ([Goodge, page 6739 “Recall that KNN computes the anomaly score based on the distance to the kth nearest neighbour of a point. In the context of message passing, each data sample corresponds to one node in a graph and node i is connected to each of its k nearest neighbours, j 2 Ni, via a directed edge (j; i), with edge feature ej;i equal to the distance between them (k-NN graph):”] AND [Goodge, page 6740] “we define a target node i and edge (j; i) connecting it to a source node j for all j where xj is in the set of k nearest neighbours to xi. The edge feature vector is equal to the Euclidean distance between the two points”, wherein the examiner interprets “LUNAR learns to use information from the nearest neighbours of each node in a trainable way to find anomalies…we define a target node i and edge (j; i) connecting it to a source node j for all j where xj is in the set of k nearest neighbours…The edge feature vector is equal to the Euclidean distance between the two points” to be the same as generating edges between the nodes based on distances in relation to the nodes, because they are both drawing edges from each node to its nearest neighbours, where the edge is defined by the distance between the corresponding feature vectors.) applying a graph neural network to the graph to allow each node to aggregate information from its neighbors; and ([Goodge, page 6737] “specifically, we propose LUNAR, a novel, graph neural network-based anomaly detection method. LUNAR learns to use information from the nearest neighbours of each node in a trainable way to find anomalies.” AND [Goodge, page 6739] “This relies on a message passing scheme, made up of message, aggregation and update steps. The message function determines the information to be sent to the node in question from each neighbour. The aggregation function summarises these incoming messages into one message,:”, wherein the examiner interprets “learns to use information from the nearest neighbours of each node…This relies on a message passing scheme, made up of message, aggregation and update steps. The message function determines the information to be sent to the node in question from each neighbour” to be the same as applying a GNN to the graph to allow nodes to aggregate information from its neighbours, because they are both applying a graph neural network whose message-passing / aggregation lets each node combine information from its neighbouring nodes.) training the graph neural network using the labeled features. ([Goodge, page 6741] “We use a loss function which trains the GNN to output a score of 0 for normal nodes and 1 for anomalous nodes.”, wherein the examiner interprets “output a score of 0 for normal nodes and 1 for anomalous nodes” to be the same as training a graph using labeled features, because they are both training the graph neural network in a supervised manner using the labeled nodes/features (normal vs. anomalous)). Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are all directed to evaluating the reliability or anomalousness of machine-learning outputs using labeled sample data organized in a feature space and proximity or distance relationships between a target sample and neighboring samples. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system claim 4 disclosed by Rowe, Jiang, Goodge, Manakul, and Liu to include the graph neural network implementation disclosed by Goodge. One would have been motivated to do so to effectively provide a trainable mechanism for aggregating nearest-neighbor information and detecting anomalous outputs based on relationships between neighboring samples, as suggested by Goodge ([Goodge, page 6737] “LUNAR learns to use information from the nearest neighbours of each node in a trainable way to find anomalies.”). Claim 15 is analogous to claim 5, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 6, Rowe, Jiang, Goodge, Manakul, and Liu teach The system of claim 5, (see rejection of claim 5). Manakul further teaches wherein generating feature vectors corresponding to each of the preprocessed training inputs and/or preprocessed training outputs includes transforming text of each of the preprocessed training inputs and/or preprocessed training outputs into a numerical vector in a high-dimensional space. ([Manakul, page 4] “Let B(., .) denote the BERTScore between two sentences. SelfCheckGPT with BERTScore finds the average BERTScore of the i-th sentence with the most similar sentence from each drawn sample”, wherein the examiner interprets the disclosed BERTScore between two sentences to be the same as using numerical vectors generated from text in a high-dimensional space because they are both directed to converting textual units into numerical feature representations that permit sentence-level similarity comparisons.) Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are both directed to representing textual data as numerical feature representations that can be compared in a high-dimensional feature space to evaluate outputs generated by a large language model. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system claim 5 disclosed by Rowe, Jiang, Goodge, Manakul, and Liu to include “Let B(., .) denote the BERTScore between two sentences. SelfCheckGPT with BERTScore finds the average BERTScore of the i-th sentence with the most similar sentence from each drawn sample” disclosed by Manakul. One would have been motivated to do so to effectively represent textual inputs and outputs as numerical feature vectors in a high-dimensional embedding space so that similarity and distance computations may be performed during reliability evaluation, as suggested by Manakul ([Manakul, page 4,] “SelfCheckGPT with BERTScore finds the average BERTScore of the i-th sentence with the most similar sentence from each drawn sample.”). Claim 16 is analogous to claim 6 aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 7, Rowe, Jiang, Goodge, Manakul, and Liu teaches The system of claim 4, (see rejection of claim 4). Goodge further teaches: wherein organizing the preprocessed training inputs and/or training outputs as labeled and unlabeled features in a feature space based on proximity using a second artificial intelligence model includes: generating feature vectors corresponding to each of the preprocessed training inputs and/or training outputs; ([Goodge, page 6740] “We construct the k-NN graph of any feature-based, tabular dataset, rather than being restricted to graph datasets. We use a node’s distances to its k nearest neighbours as input, which is more generalizable than using its feature vector.”, wherein the examiner interprets “its feature vector” to be the same as feature vectors corresponding to each of the preprocessed training inputs and/or training outputs because they are both directed to representing each data sample as a numerical feature vector that defines the sample’s location in a multidimensional feature space.) computing a distance matrix representing pairwise distances between feature vectors; ([Goodge, page 6740] “The edge feature vector is equal to the Euclidean distance between the two points: [Eq 11] , … we use a learnable aggregation, which is suitable for our setting as we are dealing with node neighbourhoods of a fixed size (k). Our message aggregation involves concatenating them to give a k-dimensional vector, e(i), where each entry represents the distance of xi to its corresponding neighbour [Eq 13] “, wherein the examiner interprets the disclosed collection of Euclidean-distance values indexed by pairs of data points and the k-dimensional vector corresponding to a neighbour to be the same as a distance matrix representing pairwise distances between feature vectors because they are both directed to numerically organizing distances between pairs of feature vectors, with each value identifying the distance between a particular feature vector and another feature vector in its neighborhood.) constructing a graph with each node representing a feature vector, wherein edges are drawn between nodes that are proximate in the feature space; ([Goodge, page 6740] “LUNAR Methodology … we represent a set of data as a graph, with a node corresponding to each data sample and directed edges connecting a target node to a set of source nodes, which are the nearest neighbours of the samples”, [Goodge, page 6739] “each data sample corresponds to one node in a graph and node i is connected to each of its k nearest neighbours, j ∈ Ni , via a directed edge (j, i), with edge feature ej,i equal to the distance between them (k-NN graph): [Eq.6]”, wherein the examiner interprets “a node corresponding to each data sample” to be the same as each node representing a feature vector because each data sample is represented by its corresponding feature vector in the feature space. The examiner further interprets “node i is connected to each of its k nearest neighbours … via a directed edge” to be the same as edges are drawn between nodes that are proximate in the feature space because nearest-neighbor nodes are nodes having the smallest distances from one another in the feature space.) applying a clustering algorithm to the graph to determine cluster centers of clusters of nodes; and ([Goodge, page 6737] “We show that many popular local outlier methods, such as KNN, LOF and DBSCAN, can be unified under a single framework based on graph neural networks”, AND [Goodge, page 6738] “DBSCAN (Ester et al. 1996), which simultaneously learns to cluster normal data while also detecting outliers, uses the number of points within a pre-defined distance... Ruff et al. (2018) learn a normality-encoding hypersphere in the latent space and the anomaly score is the distance from the centre.”, wherein the examiner interprets “DBSCAN … simultaneously learns to cluster normal data” to be the same as applying a clustering algorithm to the graph because DBSCAN is a clustering algorithm that identifies groups of data points based on their proximity and density, and further interprets “learn a normality-encoding hypersphere in the latent space” and “the centre” to be the same as determine cluster centers of clusters of nodes because they are both directed to identifying a central reference location for a group of feature-space data points from which the positions or distances of individual points may be evaluated.) assigning each node to a cluster corresponding to a nearest cluster center. ([Goodge, 6738] “They rely on the assumption that anomalies are in sparse regions of the data space, far away from highly dense clusters of normal points. Points that are close to their neighbours are more likely to be normal themselves.”, wherein the examiner interprets determining that a point is associated with a highly dense cluster of normal points based on its closeness to neighboring points to be the same as assigning each node to a cluster corresponding to a nearest cluster center because they are both directed to associating a feature-space point with the closest dense grouping of related points according to proximity.) Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are all directed to evaluating the reliability or anomalousness of artificial-intelligence model outputs using feature-space representations, distance relationships, and proximity-based groupings of data samples. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 4 disclosed by Rowe, Jiang, Goodge, Manakul, and Liu to include the “k-NN graph of any feature-based, tabular dataset” disclosed by Goodge. One would have been motivated to do so to effectively organize feature-space samples according to pairwise proximity and facilitate clustering and anomaly detection within the resulting graph, as suggested by Goodge ([Goodge, page 6738] “simultaneously learns to cluster normal data while also detecting outliers”). Claim 17 is analogous to claim 7 aside from claim type and minute differences, and thus the same rejection applies as above. Regarding Claim 8, Rowe, Jiang, Goodge, Manakul, and Liu teaches The system of claim 4, (see rejection of claim 4). Manakul further teaches: wherein: the test input includes one or more input text strings ([Manakul, page 3] “Let R refer to an LLM response drawn from a given user query. SelfCheckGPT draws a further N stochastic LLM response samples {S1, S2, ..., Sn, ..., SN} using the same query,” AND [Manakul, page 1] “Generative Large Language Models (LLMs) such as GPT-3 are capable of generating highly fluent responses to a wide variety of user prompts.”, wherein the examiner interprets “a given user query,” “the same query,” and “user prompts” to be the same as one or more input text strings because they are all directed to textual language inputs provided to a large language model to cause the model to generate corresponding textual responses, and further interprets the query used to draw the LLM response samples to be the same as the test input because they are both directed to the input supplied to the model to generate the output being evaluated.) and the test output includes one or more output text strings ([Manakul, page 1, Fig. 1] “LLM’s passage to be evaluated at sentence-level”, [Manakul, page 3] “Given the LLM’s response R, let i denote the i-th sentence in R, j denote the j-th token in the i-th sentence, J is the number of tokens in the sentence”, and [Manakul, page 6] “To obtain the main response, we set the temperature to 0.0 and use standard beam search decoding.”, wherein the examiner interprets “LLM’s passage…the LLM’s response R,” and “the main response” to be the same as one or more output text strings because they are all directed to natural-language text generated by the large language model, including a passage or response composed of one or more sentences and tokens, and further interprets the LLM’s passage or main response selected for sentence-level evaluation to be the same as the test output because they are both directed to the model-generated output that is subsequently evaluated.) Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are all directed to evaluating artificial-intelligence model outputs, including natural-language outputs generated in response to textual inputs/strings. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 4 disclosed by Rowe, Jiang, Goodge, Manakul, and Liu to include “an LLM response drawn from a given user query” disclosed by Manakul. One would have been motivated to do so to effectively extend the evaluation of artificial-intelligence model outputs to natural-language generation systems and permit textual user prompts and the corresponding generated passages to be evaluated, as suggested by Manakul ([Manakul, page 1] “capable of generating highly fluent responses to a wide variety of user prompts.”). Claim 18 is analogous to claim 8 aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 9, Rowe, Jiang, Goodge, Manakul, and Liu teach The system of claim 8, (see rejection of claim 8). Manakul further teaches: wherein preprocessing the test output includes applying at least one of stemming, lemmatization, stop word removal, part-of-speech tagging, and tokenization operations to the test output. [Manakul, page 4] “we train a simple n-gram model using the samples {S1, ..., SN} as well as the main response R (which is assessed), where we note that including R can be considered as a smoothing method where the count of each token in R is increased by 1”, AND [Manakul, page 3] “Given the LLM’s response R, let i denote the i-th sentence in R, j denote the j-th token in the i-th sentence, J is the number of tokens in the sentence”, wherein the examiner interprets “the main response R (which is assessed)” and “the LLM’s response R” to be the same as the test output because they are both directed to the model-generated response that is subjected to subsequent evaluation, and further interprets “each token in R” and “the j-th token in the i-th sentence” to be the same as tokenization operations because they are both directed to representing the test output as individual tokens for subsequent n-gram modeling and evaluation.) Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are all directed to evaluating outputs generated by artificial-intelligence models, including processing the generated outputs to determine their reliability. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 8 disclosed by Rowe, Jiang, Goodge, Manakul, and Liu to include the use of a smoothing technique disclosed by Manakul. One would have been motivated to do so to effectively improve the robustness of the statistical evaluation of the generated test output by representing the response as individual tokens and smoothing the token counts used by the n-gram model, as suggested by Manakul ([Manakul, page 4] “including R can be considered as a smoothing method where the count of each token in R is increased by 1”). Claim 19 is analogous to claim 9, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 10, Rowe, Jiang, Goodge, Manakul, and Liu teach The system of claim 9, (see rejection of claim 9). Goodge further teaches wherein the second metric represents a sum of distances between the test feature and each selected label feature. ([Goodge, page 6738] “lrdk (A) := [ Σ j∈Ni reachk (xi , xj ) / |Ni | ]−1 …where Ni is the set of k nearest neighbours of xi”, AND [Goodge, 6739] “Table 1: Local outlier methods as they relate to the message passing framework defined in (5)” AND [Goodge, page 6740] “Our message aggregation involves concatenating them to give a k-dimensional vector, e(i) , where each entry represents the distance of xi to its corresponding neighbour … This vector is mapped to a single, scalar value representing the anomalousness of node i”, wherein the examiner interprets “xi” to be the same as the test feature because they are both directed to the target feature whose relationship to neighboring features is being evaluated; “xj” for each member of “the set of k nearest neighbours of xi” to be the same as each selected label feature because they are both directed to the neighboring reference features selected based on proximity to the target feature; and the disclosed “Σ”, and aggregation of entries in which “each entry represents the distance of xi to its corresponding neighbour” to be the same as a sum of distances between the test feature and each selected label feature because they are both directed to aggregating, by summation, the respective distance values between the target feature and its selected neighboring features to obtain a metric used in determining a scalar anomalousness score.). Rowe, Jiang, Goodge, Manakul, Liu, and the instant application are analogous art because they are all directed to evaluating the reliability or anomaly capability of a machine-learning model output using distances between a target feature and selected neighboring features. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 9 disclosed by Rowe, Jiang, Goodge, Manakul, and Liu to include the distance calculation technique as disclosed by Goodge. One would have been motivated to do so to efficiently consolidate multiple neighboring-feature distance measurements into a single metric suitable for determining the reliability or anomalousness of the target feature, as suggested by Goodge ([Goodge, page 6740] “This vector is mapped to a single, scalar value representing the anomalousness of node i”). Claim 20 is analogous to claim 10, aside from claim type and minute differences, and thus the same rejection applies as above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVAN KAPOOR/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Mar 27, 2024
Application Filed
Aug 03, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
7%
Grant Probability
18%
With Interview (+11.1%)
4y 4m (~1y 11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month