Prosecution Insights
Last updated: August 17, 2026
Application No. 18/190,937

SYSTEMS AND METHODS FOR GENERATING RECOMMENDATIONS FOR CAUSES OF LABELING DETERMINATIONS THAT ARE GENERATED BY NON-DIFFERENTIABLE ARTIFICIAL INTELLIGENCE MODELS

Final Rejection §103
Filed
Mar 27, 2023
Examiner
MORALES, PEDRO JESUS
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Capital One Services LLC
OA Round
2 (Final)
62%
Grant Probability
Moderate
3-4
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
8 granted / 13 resolved
+6.5% vs TC avg
Strong +56% interview lift
Without
With
+55.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
20 currently pending
Career history
34
Total Applications
across all art units

Statute-Specific Performance

§101
24.4%
-15.6% vs TC avg
§103
47.5%
+7.5% vs TC avg
§102
11.9%
-28.1% vs TC avg
§112
13.1%
-26.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 13 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to Applicant’s reply filed May 21st 2026. This action is made final. Status of the Claims Claims 1, 2 and 15 are amended. Claim status is currently pending and under examination for Claims 1-20 of which independent claims are 1, 2 and 15. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Applicant’s amendments to the Claims have overcome each and every 101 rejections previously set forth in the Non-Final Office Action mailed February 26th 2026. Applicant’s arguments regarding the art rejections are moot in view of the new grounds of rejection necessitated by Applicant’s amendment. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The following are the references relied upon in the rejections below: Merrill (US 20190378210 A1) Mishra (US 20220121744 A1) J. Ables et al., "Creating an Explainable Intrusion Detection System Using Self Organizing Maps," 2022 IEEE Symposium Series on Computational Intelligence (SSCI), Singapore, Singapore, 2022, pp. 404-412, doi: 10.1109/SSCI51031.2022.10022255. Dugger (US 20240005150 A1) Tahir, Ghalib Ahmed, and Chu Kiong Loo. "Progressive kernel extreme learning machine for food image analysis via optimal features from quality resilient CNN." Applied Sciences 11.20 (2021): 9562. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Ables / Dugger / Tahir. With respect to claim 1, Merrill teaches: A system for generating recommendations for causes of … labels that are generated by non-differentiable artificial intelligence models …, comprising ([0029] “explanation system (e.g., 120 of FIGS. 1A and 1B) uses a non-differentiable model decomposition module (e.g., 121) to explain models by reference by transforming SHAP attributions (which are model-based attributions) to reference-based attributions (that are computed with respect to a reference population of data sets)” Merrill discloses explanation information (‘causes’) can be generated to explain how a score (‘label’) was generated for a credit applicant, “S271 can include functions generating explanation information for a test data point (e.g., a test data point representing a credit applicant). The explanation information can include score explanation information that provides information that can be used to explain how a score was generated for the test data point” [0199]. Merrill discloses adverse action information (‘recommendations’) can be generated from score explanation information (‘causes’), “the score explanation information is used to generate Adverse Action information” [0033]. Merrill discloses adverse action information, “when generating a decision to deny a consumer credit application, lenders are required to provide to each consumer the reasons why the credit application was denied, in terms of factors the model actually used, that the consumer can take practical steps to improve. These adverse action reasons and notices …” [0018].): one or more processors; and a non-transitory, computer-readable medium comprising instructions recorded thereon that when executed by the one or more processors cause operations comprising ([0235] “the processing unit includes one or more processors communicatively coupled to one or more of a RAM, ROM, and machine-readable storage medium; the one or more processors of the processing unit receive instructions stored by the one or more of a RAM, ROM, and machine-readable storage medium via a bus; and the one or more processors execute the received instructions”): receiving a first feature input corresponding to a dataset … (Merrill discloses a test data point (‘first feature input’), “selecting a single test data point (e.g., an input data set whose score/output generated by at least the non-differentiable model is to be explained). In some embodiments, the test data point represents a credit applicant” [0084].), wherein the first feature input comprises a plurality of values (Merrill discloses a test data point (‘first feature input’) is comprised of multiple features (‘values’), “for each feature of a test data point, generating a difference value, the difference value for the test data point relative to a corresponding reference data point,” [0030].), inputting the first feature input into an artificial intelligence model, wherein the artificial intelligence model is non-differentiable ([0084] “selecting a single test data point (e.g., an input data set whose score/output generated by at least the non-differentiable model is to be explained”), receiving a first prediction from the artificial intelligence model … (See [0084] describing how an output for a test data point is generated by a non-differentiable model.), receiving a second prediction for the artificial intelligence model (Merrill discloses “a novel method for the decomposition of ensembles (combinations) of tree and neural network models that combines the SHAP and Integrated Gradients methods to produce a new and useful result (decomposition of ensemble models) which is used to perform new and useful analysis that results in tangible outputs that support human decision-making, e.g.: feature importance, adverse action, and disparate impact analysis, as described herein” [0028]. [0168] “both Shapley and Integrated Gradient methods enforce nullity, e.g., features that do not contribute to the score will receive attribution values of zero” SHAP and integrated gradients can be used to calculate each feature’s contribution towards a model’s score (output), therefore each calculated feature’s importance is a second prediction.), determining an effect of each value of the first feature input on the first prediction … ([0048] “a Shapley value decomposition (e.g., generated by the non-differentiable model decomposition module) is a linear combination of feature attribution values ϕi (Shapley value). … Shapley value decompositions are SHAP (SHapley Additive exPlanation) values … SHAP values explain the output of a model ƒ as a sum of the effects ϕi of each feature being introduced into a conditional expectation” [0168] “both Shapley and Integrated Gradient methods enforce nullity, e.g., features that do not contribute to the score will receive attribution values of zero” [0031] “the non-differentiable model decomposition module computes a score decomposition for a test data point relative to a reference data point for the non-differentiable model (as described herein), and the differentiable model decomposition module computes a score decomposition for the test data point relative to the reference data point for the differentiable model, and combines the decomposition of the non-differentiable model with the decomposition for the differentiable model by using an ensembling function of the ensemble model, to generate a decomposition for an ensemble model score for the test data point relative to the reference data point” SHAP and integrated gradients are used to determine feature importance towards a model score/output (‘first prediction’). Both SHAP and integrated gradients determine feature importance (‘an effect of each value’) according to the same test data point (‘first feature input’) and reference data point (training data). The results of SHAP and integrated gradients are combined into a single decomposition (therefore a decomposition represents ‘effect of each value of the first feature input’)).); generating for display, on a user interface, a recommendation for a cause of the known label in the dataset based on the effect of each value of the first feature input on the first prediction (Merrill discloses “the model evaluation system uses a decomposition generated for a model score to generate adverse action information (e.g., at S271) (as described herein) and provide the generated adverse action information to the operator device 171” [0131-0132]. See Figure 2 depicting the decomposition at S250 is a combined decomposition using SHAP and integrated gradients. [0033] “the score explanation information is used to generate Adverse Action information” Adverse action information (‘first recommendation’) can be generated from score explanation information (‘cause of the known label’). See [0018, 0033] describing how adverse action information explains why a credit application was denied and steps a consumer can take to improve. Merrill discloses a test data point (‘first feature input’) is comprised of features of a denied credit applicant (‘known label in the dataset’), “the model evaluation system 120 evaluates a specific denied credit applicant. In this embodiment the specific denied credit applicant comprises the test set (test data point) (selected at S220)” [0153]. [0190-0191] “S272 can include identifying features having decomposition values (in the generated decompositions) above a threshold. In some embodiments, the method 200 includes providing the identified features to an operator device (e.g., 171) via a network. In some embodiments, the method 200 includes displaying the identified features on a display device of an operator device (e.g., 171). In other embodiments, the method 20 includes displaying natural language explanations generated based on the decomposition described above. In some embodiments the method 200 includes displaying the identified features and their decomposition in a form similar to the table presented in FIG. 3”); determining a … response based on the cause (Merrill discloses above adverse action information can be generated from score explanation information (‘cause’). The adverse action information consists of reasons why a credit application was denied (‘response’) and steps a consumer can take to improve, see [0018].); and generating for display a second recommendation for executing the … response ([0018] “when generating a decision to deny a consumer credit application, lenders are required to provide to each consumer the reasons why the credit application was denied, in terms of factors the model actually used, that the consumer can take practical steps to improve. These adverse action reasons and notices are easily provided” Adverse action information can be generated from score explanation information (‘cause’), see [0033]. The adverse action information consists of reasons why a credit application was denied (‘response’) and steps a consumer can take to improve (‘second recommendation’). See [0131-0132, 0190-0191] describing how an operator device can be used to display natural language explanations and adverse action information.). However, Merrill does not teach a first feature input with an unknown label and detecting a known label for the first feature input, which is taught by Mishra: receiving a first feature input corresponding to a dataset with an unknown label (Mishra discloses “hardware processor is configured to execute a software application suspected of being malware; monitor behavior of the software application at run-time; and acquire an input time sequence of data records based on a trace analysis of the software application, wherein the input time sequence comprises a plurality of features of the software application. The hardware processor is further configured to classify the software application as being a malicious software application based on the plurality of features of the software application; and output a ranking of a subset plurality of features by their respective contributions towards the classification of the software application as being malicious software” [Abstract]. Mishra discloses “A classic structure of RNN is shown in FIG. 4. In the figure, A represents the neural network architecture, where x0, x1, x2, . . . represents the time series inputs and his represent the outputs of hidden layers” [0039]. A time series (sequence) comprised of a plurality of features is input into a RNN to classify a software application (therefore a time series that is being input into a RNN for classification has an unknown label).), wherein the first feature input comprises a plurality of values (See [Abstract, 0039] describing how a time series is comprised of a plurality of features.); inputting the first feature input into an artificial intelligence model (See [0039] describing how a time series is input into a RNN model.), and wherein the artificial intelligence model is trained to detect a known label based on a set of training data comprising labeled feature inputs corresponding to the known label ([0038] “an exemplary machine learning model for the present disclosure should satisfy the following two properties: (1) Ability to accept time series type data as input; and (2) Ability to make decisions utilizing potential information concealed in consecutive adjacent inputs. Accordingly, exemplary embodiments utilize a Recurrent Neural Network (RNN) training to satisfy these properties, since RNN is powerful in handling sequential input data.” [0058] “To do so, both malicious and benign software are executed on an exemplary hardware platform, in which a total of 367 programs (including both malicious and benign ones) are executed. All the traced data were mixed up and further split into training (80%) and test (20%) sets after labeling. Total training epochs are 200 for every model and test accuracy was plotted every 10 epochs” [0039] “the RNN accepts sequential inputs. … For trace-data-based malware detection, each column of a trace table can be set as inputs, and the hidden state of a final stage, i.e. ht, can be set as the final output” [0053] “exemplary RNN-classifier that utilizes the structure outlined in FIG. 4. … After passing through RNN units, the outputs are fed into a fully connected layer to achieve dimension reduction. Finally, a Softmax layer takes the reduced outputs from the fully connected layer to produce classification labels”); receiving a first prediction from the artificial intelligence model (See [0053] describing how a classification label is obtained from an RNN-classifier.), wherein the first prediction indicates whether the first feature input corresponds to the known label (Mishra discloses a classification result is either labeled as benign or malicious (‘known label’), “hardware-assisted malware detection provides transparency in malware detection by providing interpretable explanations for classification results of benign and malicious programs. In one embodiment, an exemplary system/method interprets the outputs of a machine learning model with a ranking of contribution factors, which explicitly provides a detailed feature importance map and explains the internal mechanism of each individual prediction” [0025].); Mishra teaches classifying inputs by using a RNN trained with labeled training data is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to combine the model explainability method of Merrill with the RNN of Mishra to classify unlabeled inputs. By classifying unlabeled inputs, a trained machine learning model can label incoming inputs based on what it has learned, thereby allowing incoming data to be classified without human supervision or manual labeling. Furthermore, the combination of Merrill / Mishra does not teach detecting cyber incidents, which is taught by Ables: generating … causes of computer security labels that are generated by … artificial intelligence models processing datasets built by monitoring network activity ((P. 409, Sec. V-A, ¶2) “The results for the NSL-KDD dataset can be found in Figures 2a and 2b. The local explanation example shows that the most important features for its prediction were ‘Duration’, ‘Destination (dst) bytes’, and ‘Source (src) bytes’. The remaining features, ‘Service (srv) count’, ‘Count’, and ‘Destination (dst) host count’ are considered less significant because of their distance from the BMU” (P. 408, Sec. IV-C, ¶3) “we can see the features with the largest impact on a prediction: duration, dst bytes, and src bytes. These features were the closest to the BMU, and they played a large role in computing the predicted value. Seeing the specific features that influence predictions provides insight about samples labeled as malicious or benign and can further help operators determine the reason of incorrect predictions” (P. 406, Sec. III-A, ¶1) “Self Organizing Maps (SOMs), sometimes referred to as Kohonen Maps [8], [38], Kohonen Self Organizing Maps [39], or Kohonen Networks [40], are a class of unsupervised machine learning algorithms. SOMs are comprised of a network of individual units, each of which has a feature vector of the same size as the dimension of training data” The NSL-KDD dataset is comprised of features collected from network activity (duration, destination/source bytes). Local explanations explain which features were most important in determining a label prediction (therefore local explanations are the causes of computer security labels).), and wherein the plurality of values indicates networking activity of a user ((P. 409, Sec. V-A, ¶2) “The results for the NSL-KDD dataset can be found in Figures 2a and 2b. The local explanation example shows that the most important features for its prediction were ‘Duration’, ‘Destination (dst) bytes’, and ‘Source (src) bytes’. The remaining features, ‘Service (srv) count’, ‘Count’, and ‘Destination (dst) host count’ are considered less significant because of their distance from the BMU”); and wherein the known label comprises a detected cyber incident ((P. 404, Sec. I, ¶2) “systems monitor networks and automate attack detection by comparing network activity to the signature of known attacks or by detecting behavior that is anomalous to benign network patterns [2]. Through these methods, a security analyst can use an IDS to detect improper use, unauthorized access, or the abuse of a network” (P. 408, Sec. IV-C, ¶3) “Figure 2a shows the local explanations for a prediction, where each feature on the y-axis has a value representing distance from its respective BMU value (See Section III). In this example, we can see the features with the largest impact on a prediction: duration, dst bytes, and src bytes. These features were the closest to the BMU, and they played a large role in computing the predicted value. Seeing the specific features that influence predictions provides insight about samples labeled as malicious or benign and can further help operators determine the reason of incorrect predictions”), and determining a cyber incident response based on the cause ((P. 408, Sec. IV-C, ¶3) “we can see the features with the largest impact on a prediction: duration, dst bytes, and src bytes. These features were the closest to the BMU, and they played a large role in computing the predicted value. Seeing the specific features that influence predictions provides insight about samples labeled as malicious or benign and can further help operators determine the reason of incorrect predictions. These features can also be further investigated with feature value heat maps” (P. 406, Sec. II-C, ¶1) “the users need to be confident in the predictions or recommendations computed by an IDS. Understandable explanations allow users to perform their tasks correctly. The stakeholders of an IDS (e.g. CSoC operators, developers, and investors) are individuals who will be dependent on the performance of the system. CSoC operators will be performing defensive actions based on prediction and explanation results. Developers can use explanations to fortify the model in areas where it is weak. Investors may need explanations to help them in making budgeting decisions for their company”); Ables teaches an intrusion detection system that generates explanations to explain which features influenced a sample being labeled as malicious or benign is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra with the intrusion detection system disclosed by Ables to classify network activity. By classifying network activity, network activity that is classified as malicious can be further investigated and preventative defensive actions can be taken to mitigate threats and prevent system compromises. However, the combination of Merrill / Mishra / Ables does not teach a second prediction that indicates an approximated integrated gradient for a non-differentiable artificial intelligence model, which is taught by Dugger: inputting the first feature input into an artificial intelligence model, wherein the artificial intelligence model is non-differentiable ([0236] “the process 3000 involves evaluating an integrated gradients calculation numerically along a chosen path. In some examples, the chosen path could be a straight line path, in attribute space from the baseline values x′ to the given input values x, taking account of the partial derivatives of any derived functions of the time series inputs that enter the output function.” [0236] “For a non-differentiable model, such as a tree based model, it may be necessary to estimate the gradient numerically.”), receiving a second prediction for the artificial intelligence model ([0244] “applying integrated gradients to generate model explanations for models with time series inputs includes selecting a representative baseline set of attribute values x′ including time series inputs, that represents either an optimal or average set of values.”), wherein the second prediction indicates an approximated integrated gradient for the artificial intelligence model (Integrated gradients are calculated to generate model explanations (‘second prediction’) for a non-differentiable model (‘artificial intelligence model’). For a non-differentiable model, gradients can be estimated numerically when calculating integrated gradients (see [0236]), therefore the estimated gradients calculated result in an approximated integrated gradient.), wherein the approximated integrated gradient comprises respective rates of change of one value with respect to another (The Examiner interprets “respective rates of change” according to its broadest reasonable interpretation (BRI) in view of the Applicant’s specification as encompassing a calculated gradient of a path. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [020], (see excerpt below). Applicant’s written description at [020]: “the gradient of any line or curve (e.g., a line or curve based on the model and/or function) indicates the rate of change of one variable with respect to another. This rate of change for one variable can then be used to determine the conditional expectation of another.” [0236] “the formulation of the integrated gradients approach depends upon a path …in an attribute space. The default choice may be the straight line path … The integral can be evaluated numerically, simply by calculating the gradient ∇ƒ at equally spaced points along the path … For a non-differentiable model, such as a tree based model, it may be necessary to estimate the gradient numerically.” Dugger discloses Equation 37 (reproduced below) in [0236] describing an equation for calculating integrated gradients. A gradient ∇ƒ is calculated for a straight line path, therefore the gradient of the straight line path indicates the rate of change of one variable (input variable x) with respect to another (baseline value x’). Multiple gradients are calculated and then summed, therefore multiple respective rates of change are calculated for each input and baseline value. PNG media_image1.png 515 1540 media_image1.png Greyscale ); determining an effect of each value of the first feature input on the first prediction based on the approximated integrated gradient … ([0233] “A reason code or other explanatory data may be generated using an integrated gradients approach.” [0212] “example of explanatory data is a reason code, adverse action code, or other data indicating an impact of a given variable on a predictive output. For instance, explanatory reason codes may indicate why an entity received a particular predicted output (e.g. an adverse event prediction in a timing-prediction model). The explanatory reason codes can be generated from a wavelet based model to satisfy suitable requirements.” A reason code represents an impact (‘effect’) each variable (‘value of the first feature input’) has on a predictive output (‘first prediction’).); Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Ables with the integrated gradients technique disclosed by Dugger to estimate integrated gradients for a non-differentiable model. By estimating integrated gradients for a non-differentiable model, integrated gradients can be calculated for a non-differentiable model to explain the impact a variable has on a predictive output, thereby increasing model interpretability for a non-differentiable model. However, the combination of Merrill / Mishra / Ables / Dugger does not teach determining an effect of each value of a first feature input on a first prediction based on determining a respective conditional expectation based on the respective rates of change, which is taught by Tahir: determining an effect of each value of the first feature input on the first prediction based on … by determining a respective conditional expectation based on the respective rates of change ((P. 2, Sec. 1, Bullet 2) “We analyzed the visualizations from the Gradient Explainer, which shows that all the pixels in an area of interest do not contribute equally or positively to the desired output. Based on our findings, we introduced a feature selection technique that uses SHAP scores from the gradient explainer to select the optimal subset of features from the high dimension space by filtering out the irrelevant ones;” (P. 6, Sec. 3.3, ¶1) “The Gradient Explainer combines the ideas from Integrated Gradients, SHAP, and Smooth Grad into a single expected equation. It approximates the model as a linear function between each background data sample and input feature vector. The attributes are assumed to be independent of each other and then expected gradients compute approximate SHAP values.” (P. 6, Sec. 3.3, ¶3) “SHAP values of both techniques integrate the conditional expectations with game theory and classic Shapley values to assign the ϕ i values to the attributes of the feature vector.” (P. 6, Sec. 3.3, ¶4) “Here, f x ( S ) = f ( h x ( z ' ) ) = E [ f ( x ) | x s ] , S represents the set of non-zero indexes in 𝑧′ …. E f x x s   is the expected value of the function conditioned on a subset S of the input features, M indicates the number of features and N refers to the set of all input features. Tahir discloses Equation 4 (reproduced below) on P. 6 describing an equation for calculating a SHAP value ϕ i for a feature i. E [ f ( x ) | x s ] represents the expected value of function f conditioned on a subset of input features, therefore the expected value is a conditional expectation. To approximate SHAP values, expected gradients are used, therefore expected values (conditional expectations) are based on gradients (‘respective rates of change’). The SHAP values are used to determine how much each pixel (‘each value of the first feature input’) contributes to a desired output (‘first prediction’), therefore determining an effect of each value of the first feature input on the first prediction (the first feature input being pixels in an area of interest). PNG media_image2.png 416 1229 media_image2.png Greyscale ); Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Ables / Dugger with the technique disclosed by Tahir to use conditional probabilities to calculate feature contribution. By using conditional probabilities to calculate feature contribution, each feature’s contribution to a model prediction can be measured based on whether a feature is included or excluded in a subset of features, thereby creating accurate and consistent feature contribution scores. Claims 2-4, 9-10, 13 and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir. With respect to claim 2, Merrill teaches: A method for generating recommendations for causes of labeling determinations that are generated by non-differentiable artificial intelligence models, comprising ([0029] “explanation system (e.g., 120 of FIGS. 1A and 1B) uses a non-differentiable model decomposition module (e.g., 121) to explain models by reference by transforming SHAP attributions (which are model-based attributions) to reference-based attributions (that are computed with respect to a reference population of data sets)” Merrill discloses explanation information (‘causes’) can be generated to explain how a score (‘labeling determination’) was generated for a credit applicant, “S271 can include functions generating explanation information for a test data point (e.g., a test data point representing a credit applicant). The explanation information can include score explanation information that provides information that can be used to explain how a score was generated for the test data point” [0199]. Merrill discloses adverse action information (‘recommendations’) can be generated from score explanation information (‘causes’), “the score explanation information is used to generate Adverse Action information” [0033]. Merrill discloses adverse action information, “when generating a decision to deny a consumer credit application, lenders are required to provide to each consumer the reasons why the credit application was denied, in terms of factors the model actually used, that the consumer can take practical steps to improve. These adverse action reasons and notices …” [0018].): receiving a first feature input corresponding to a dataset … (Merrill discloses a test data point (‘first feature input’), “selecting a single test data point (e.g., an input data set whose score/output generated by at least the non-differentiable model is to be explained). In some embodiments, the test data point represents a credit applicant” [0084].), wherein the first feature input comprises a plurality of values (Merrill discloses a test data point (‘first feature input’) is comprised of multiple features (‘values’), “for each feature of a test data point, generating a difference value, the difference value for the test data point relative to a corresponding reference data point,” [0030].), inputting the first feature input into an artificial intelligence model, wherein the artificial intelligence model is non-differentiable ([0084] “selecting a single test data point (e.g., an input data set whose score/output generated by at least the non-differentiable model is to be explained”), receiving a first prediction from the artificial intelligence model … (See [0084] describing how an output for a test data point is generated by a non-differentiable model.), receiving a second prediction for the artificial intelligence model (Merrill discloses “a novel method for the decomposition of ensembles (combinations) of tree and neural network models that combines the SHAP and Integrated Gradients methods to produce a new and useful result (decomposition of ensemble models) which is used to perform new and useful analysis that results in tangible outputs that support human decision-making, e.g.: feature importance, adverse action, and disparate impact analysis, as described herein” [0028]. [0168] “both Shapley and Integrated Gradient methods enforce nullity, e.g., features that do not contribute to the score will receive attribution values of zero” SHAP and integrated gradients can be used to calculate each feature’s contribution towards a model’s score (output), therefore each calculated feature’s importance is a second prediction.), determining an effect of each value of the first feature input on the first prediction … ([0048] “a Shapley value decomposition (e.g., generated by the non-differentiable model decomposition module) is a linear combination of feature attribution values ϕi (Shapley value). … Shapley value decompositions are SHAP (SHapley Additive exPlanation) values … SHAP values explain the output of a model ƒ as a sum of the effects ϕi of each feature being introduced into a conditional expectation” [0168] “both Shapley and Integrated Gradient methods enforce nullity, e.g., features that do not contribute to the score will receive attribution values of zero” [0031] “the non-differentiable model decomposition module computes a score decomposition for a test data point relative to a reference data point for the non-differentiable model (as described herein), and the differentiable model decomposition module computes a score decomposition for the test data point relative to the reference data point for the differentiable model, and combines the decomposition of the non-differentiable model with the decomposition for the differentiable model by using an ensembling function of the ensemble model, to generate a decomposition for an ensemble model score for the test data point relative to the reference data point” SHAP and integrated gradients are used to determine feature importance towards a model score/output (‘first prediction’). Both SHAP and integrated gradients determine feature importance (‘an effect of each value’) according to the same test data point (‘first feature input’) and reference data point (training data). The results of SHAP and integrated gradients are combined into a single decomposition (therefore a decomposition represents ‘effect of each value of the first feature input’)).); and generating for display, on a user interface, a first recommendation for a cause of the known label in the dataset based on the effect of each value of the first feature input on the first prediction (Merrill discloses “the model evaluation system uses a decomposition generated for a model score to generate adverse action information (e.g., at S271) (as described herein) and provide the generated adverse action information to the operator device 171” [0131-0132]. See Figure 2 depicting the decomposition at S250 is a combined decomposition using SHAP and integrated gradients. [0033] “the score explanation information is used to generate Adverse Action information” Adverse action information (‘first recommendation’) can be generated from score explanation information (‘cause of the known label’). See [0018, 0033] describing how adverse action information explains why a credit application was denied and steps a consumer can take to improve. Merrill discloses a test data point (‘first feature input’) is comprised of features of a denied credit applicant (‘known label in the dataset’), “the model evaluation system 120 evaluates a specific denied credit applicant. In this embodiment the specific denied credit applicant comprises the test set (test data point) (selected at S220)” [0153]. [0190-0191] “S272 can include identifying features having decomposition values (in the generated decompositions) above a threshold. In some embodiments, the method 200 includes providing the identified features to an operator device (e.g., 171) via a network. In some embodiments, the method 200 includes displaying the identified features on a display device of an operator device (e.g., 171). In other embodiments, the method 20 includes displaying natural language explanations generated based on the decomposition described above. In some embodiments the method 200 includes displaying the identified features and their decomposition in a form similar to the table presented in FIG. 3”). However, Merrill does not teach a first feature input with an unknown label and detecting a known label for the first feature input, which is taught by Mishra: receiving a first feature input corresponding to a dataset with an unknown label (Mishra discloses “hardware processor is configured to execute a software application suspected of being malware; monitor behavior of the software application at run-time; and acquire an input time sequence of data records based on a trace analysis of the software application, wherein the input time sequence comprises a plurality of features of the software application. The hardware processor is further configured to classify the software application as being a malicious software application based on the plurality of features of the software application; and output a ranking of a subset plurality of features by their respective contributions towards the classification of the software application as being malicious software” [Abstract]. Mishra discloses “A classic structure of RNN is shown in FIG. 4. In the figure, A represents the neural network architecture, where x0, x1, x2, . . . represents the time series inputs and his represent the outputs of hidden layers” [0039]. A time series (sequence) comprised of a plurality of features is input into a RNN to classify a software application (therefore a time series that is being input into a RNN for classification has an unknown label).), wherein the first feature input comprises a plurality of values (See [Abstract, 0039] describing how a time series is comprised of a plurality of features.); inputting the first feature input into an artificial intelligence model (See [0039] describing how a time series is input into a RNN model.), and wherein the artificial intelligence model is trained to detect a known label based on a set of training data comprising labeled feature inputs corresponding to the known label ([0038] “an exemplary machine learning model for the present disclosure should satisfy the following two properties: (1) Ability to accept time series type data as input; and (2) Ability to make decisions utilizing potential information concealed in consecutive adjacent inputs. Accordingly, exemplary embodiments utilize a Recurrent Neural Network (RNN) training to satisfy these properties, since RNN is powerful in handling sequential input data.” [0058] “To do so, both malicious and benign software are executed on an exemplary hardware platform, in which a total of 367 programs (including both malicious and benign ones) are executed. All the traced data were mixed up and further split into training (80%) and test (20%) sets after labeling. Total training epochs are 200 for every model and test accuracy was plotted every 10 epochs” [0039] “the RNN accepts sequential inputs. … For trace-data-based malware detection, each column of a trace table can be set as inputs, and the hidden state of a final stage, i.e. ht, can be set as the final output” [0053] “exemplary RNN-classifier that utilizes the structure outlined in FIG. 4. … After passing through RNN units, the outputs are fed into a fully connected layer to achieve dimension reduction. Finally, a Softmax layer takes the reduced outputs from the fully connected layer to produce classification labels”); receiving a first prediction from the artificial intelligence model (See [0053] describing how a classification label is obtained from an RNN-classifier.), wherein the first prediction indicates whether the first feature input corresponds to the known label (Mishra discloses a classification result is either labeled as benign or malicious (‘known label’), “hardware-assisted malware detection provides transparency in malware detection by providing interpretable explanations for classification results of benign and malicious programs. In one embodiment, an exemplary system/method interprets the outputs of a machine learning model with a ranking of contribution factors, which explicitly provides a detailed feature importance map and explains the internal mechanism of each individual prediction” [0025].); Mishra teaches classifying inputs by using a RNN trained with labeled training data is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to combine the model explainability method of Merrill with the RNN of Mishra to classify unlabeled inputs. By classifying unlabeled inputs, a trained machine learning model can label incoming inputs based on what it has learned, thereby allowing incoming data to be classified without human supervision or manual labeling. However, the combination of Merrill / Mishra does not teach a second prediction that indicates an approximated integrated gradient for a non-differentiable artificial intelligence model, which is taught by Dugger: inputting the first feature input into an artificial intelligence model, wherein the artificial intelligence model is non-differentiable ([0236] “the process 3000 involves evaluating an integrated gradients calculation numerically along a chosen path. In some examples, the chosen path could be a straight line path, in attribute space from the baseline values x′ to the given input values x, taking account of the partial derivatives of any derived functions of the time series inputs that enter the output function.” [0236] “For a non-differentiable model, such as a tree based model, it may be necessary to estimate the gradient numerically.”), receiving a second prediction for the artificial intelligence model ([0244] “applying integrated gradients to generate model explanations for models with time series inputs includes selecting a representative baseline set of attribute values x′ including time series inputs, that represents either an optimal or average set of values.”), wherein the second prediction indicates an approximated integrated gradient for the artificial intelligence model (Integrated gradients are calculated to generate model explanations (‘second prediction’) for a non-differentiable model (‘artificial intelligence model’). For a non-differentiable model, gradients can be estimated numerically when calculating integrated gradients (see [0236]), therefore the estimated gradients calculated result in an approximated integrated gradient.), wherein the approximated integrated gradient comprises respective rates of change of one value with respect to another (The Examiner interprets “respective rates of change” according to its broadest reasonable interpretation (BRI) in view of the Applicant’s specification as encompassing a calculated gradient of a path. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [020], (see excerpt below). Applicant’s written description at [020]: “the gradient of any line or curve (e.g., a line or curve based on the model and/or function) indicates the rate of change of one variable with respect to another. This rate of change for one variable can then be used to determine the conditional expectation of another.” [0236] “the formulation of the integrated gradients approach depends upon a path …in an attribute space. The default choice may be the straight line path … The integral can be evaluated numerically, simply by calculating the gradient ∇ƒ at equally spaced points along the path … For a non-differentiable model, such as a tree based model, it may be necessary to estimate the gradient numerically.” Dugger discloses Equation 37 (reproduced below) in [0236] describing an equation for calculating integrated gradients. A gradient ∇ƒ is calculated for a straight line path, therefore the gradient of the straight line path indicates the rate of change of one variable (input variable x) with respect to another (baseline value x’). Multiple gradients are calculated and then summed, therefore multiple respective rates of change are calculated for each input and baseline value. PNG media_image1.png 515 1540 media_image1.png Greyscale ); determining an effect of each value of the first feature input on the first prediction based on the approximated integrated gradient … ([0233] “A reason code or other explanatory data may be generated using an integrated gradients approach.” [0212] “example of explanatory data is a reason code, adverse action code, or other data indicating an impact of a given variable on a predictive output. For instance, explanatory reason codes may indicate why an entity received a particular predicted output (e.g. an adverse event prediction in a timing-prediction model). The explanatory reason codes can be generated from a wavelet based model to satisfy suitable requirements.” A reason code represents an impact (‘effect’) each variable (‘value of the first feature input’) has on a predictive output (‘first prediction’).); Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra with the integrated gradients technique disclosed by Dugger to estimate integrated gradients for a non-differentiable model. By estimating integrated gradients for a non-differentiable model, integrated gradients can be calculated for a non-differentiable model to explain the impact a variable has on a predictive output, thereby increasing model interpretability for a non-differentiable model. However, the combination of Merrill / Mishra / Dugger does not teach determining an effect of each value of a first feature input on a first prediction based on determining a respective conditional expectation based on the respective rates of change, which is taught by Tahir: determining an effect of each value of the first feature input on the first prediction based on … by determining a respective conditional expectation based on the respective rates of change ((P. 2, Sec. 1, Bullet 2) “We analyzed the visualizations from the Gradient Explainer, which shows that all the pixels in an area of interest do not contribute equally or positively to the desired output. Based on our findings, we introduced a feature selection technique that uses SHAP scores from the gradient explainer to select the optimal subset of features from the high dimension space by filtering out the irrelevant ones;” (P. 6, Sec. 3.3, ¶1) “The Gradient Explainer combines the ideas from Integrated Gradients, SHAP, and Smooth Grad into a single expected equation. It approximates the model as a linear function between each background data sample and input feature vector. The attributes are assumed to be independent of each other and then expected gradients compute approximate SHAP values.” (P. 6, Sec. 3.3, ¶3) “SHAP values of both techniques integrate the conditional expectations with game theory and classic Shapley values to assign the ϕ i values to the attributes of the feature vector.” (P. 6, Sec. 3.3, ¶4) “Here, f x ( S ) = f ( h x ( z ' ) ) = E [ f ( x ) | x s ] , S represents the set of non-zero indexes in 𝑧′ …. E f x x s   is the expected value of the function conditioned on a subset S of the input features, M indicates the number of features and N refers to the set of all input features. Tahir discloses Equation 4 (reproduced below) on P. 6 describing an equation for calculating a SHAP value ϕ i for a feature i. E [ f ( x ) | x s ] represents the expected value of function f conditioned on a subset of input features, therefore the expected value is a conditional expectation. To approximate SHAP values, expected gradients are used, therefore expected values (conditional expectations) are based on gradients (‘respective rates of change’). The SHAP values are used to determine how much each pixel (‘each value of the first feature input’) contributes to a desired output (‘first prediction’), therefore determining an effect of each value of the first feature input on the first prediction (the first feature input being pixels in an area of interest). PNG media_image2.png 416 1229 media_image2.png Greyscale ); Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger with the technique disclosed by Tahir to use conditional probabilities to calculate feature contribution. By using conditional probabilities to calculate feature contribution, each feature’s contribution to a model prediction can be measured based on whether a feature is included or excluded in a subset of features, thereby creating accurate and consistent feature contribution scores. With respect to claims 3 and 16, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: The method of claim 2, further comprising: receiving a test feature input, wherein the test feature input represents test values corresponding to datasets that correspond to the known label (Mishra discloses traced data is split into a test set with labels (‘known labels’), “All the traced data were mixed up and further split into training (80%) and test (20%) sets after labeling. Total training epochs are 200 for every model and test accuracy was plotted every 10 epochs. Accordingly, FIGS. 8A-8C compares the prediction accuracy of an exemplary hardware-assisted malware detection approach with PREEMPT RF and PREEMPT DT. As we can see, the exemplary method (referred to as “proposed” in the figures) provides the best malware detection accuracy” [0058-0059].); labeling the test feature input with the known label (Mishra discloses test accuracy is plotted using test data, see [0058-0059]. To plot accuracy, a test data input must be labeled with a classification generated by a model. Accurate results indicate that a generated classification matches with a correct, known test set label (therefore test data inputs are labeled with correct, known labels).); and training the artificial intelligence model to detect the known label based on the test feature input (Mishra discloses “exemplary embodiments utilize a Recurrent Neural Network (RNN) training to satisfy these properties, since RNN is powerful in handling sequential input data.” [0038]. Mishra discloses “All the traced data were mixed up and further split into training (80%) and test (20%) sets after labeling. Total training epochs are 200 for every model and test accuracy was plotted every 10 epochs” [0058]. Test accuracy is plotted during training epochs, therefore training of a RNN to classify input data occurs alongside determining test data labeling correctness (detecting known labels) using the RNN.). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the technique disclosed by Mishra to test a machine learning model’s classification accuracy during training. By testing a machine learning model’s classification accuracy during training, machine learning engineers can retrain or fine tune a model based on classification performance, thereby developing an optimal, accurate model. With respect to claims 4 and 17, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, wherein receiving the second prediction for the artificial intelligence model further comprises: determining a numerical approximation of gradients and integrals for the artificial intelligence model ([0236] “the formulation of the integrated gradients approach depends upon a path … The integral can be evaluated numerically, simply by calculating the gradient ∇ƒ at equally spaced points along the path … For a non-differentiable model, such as a tree based model, it may be necessary to estimate the gradient numerically.” Dugger discloses Equation 37 (reproduced below) in [0236] describing an equation for calculating integrated gradients. A gradient ∇ƒ is estimated for a point on a path, and an integral can be evaluated by calculating and summing gradients along the path, therefore determining a numerical approximation of gradients (sum of estimated gradients) and integrals for a model. PNG media_image1.png 515 1540 media_image1.png Greyscale ); and determining the approximated integrated gradient based on the numerical approximation of gradients and integrals (Integrated gradients are determined by evaluating integrals calculated by summing estimated gradients along a path.). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the integrated gradients technique disclosed by Dugger to estimate integrated gradients for a non-differentiable model. By estimating integrated gradients for a non-differentiable model, integrated gradients can be calculated for a non-differentiable model to explain the impact a variable has on a predictive output, thereby increasing model interpretability for a non-differentiable model. With respect to claim 9, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, wherein determining the effect of each value of the first feature input on the first prediction comprises determining a SHAP (SHapley Additive exPlanations) value for each value of the first feature input (Merrill discloses “a Shapley value decomposition (e.g., generated by the non-differentiable model decomposition module) is a linear combination of feature attribution values ϕi (Shapley value). … Shapley value decompositions are SHAP (SHapley Additive exPlanation) values … SHAP values explain the output of a model ƒ as a sum of the effects ϕi of each feature being introduced into a conditional expectation” [0048].). With respect to claim 10, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, wherein determining the effect of each value of the first feature input on the first prediction based on the approximated integrated gradient further comprises: determining a respective contribution of each value to a difference between an actual prediction and a mean prediction (Merrill discloses “M is the number of input features, N is the set of input features, S is the set features constructed from superset N. The function ƒ(hx(z′)) defines a manner to remove features so that an expected value of f(x) can be computed which is conditioned on the subset of a feature space xS. The missingingness is defined by z′, each zi′ variable represents a feature being observed (zi′=1) or unknown (zi′=0)” [0042]. Merrill discloses Equations 1 and 2 at [0042] (reproduced below) depicting equations for calculating SHAP values. The contribution of feature i is calculated by the difference between a prediction when feature i is known and an average model prediction. PNG media_image3.png 533 1410 media_image3.png Greyscale ); determining a respective SHAP value based on the respective contribution (See Equation 1 above depicting how a SHAP value ϕi is calculated based on a feature i’s contribution to the difference between a prediction when feature i is known and an average model prediction.); and determining the effect of each value based on the respective contribution (Merrill discloses “a Shapley value decomposition (e.g., generated by the non-differentiable model decomposition module) is a linear combination of feature attribution values ϕi (Shapley value). … Shapley value decompositions are SHAP (SHapley Additive exPlanation) values … SHAP values explain the output of a model ƒ as a sum of the effects ϕi of each feature being introduced into a conditional expectation” [0048].). With respect to claim 13, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, wherein the known label comprises a refusal of a credit application (Merrill discloses a test data point is comprised of features of a denied credit applicant (‘known label in the dataset’), “the model evaluation system 120 evaluates a specific denied credit applicant. In this embodiment the specific denied credit applicant comprises the test set (test data point) (selected at S220)” [0153].), wherein the plurality of values indicates a credit history of a user (Merrill discloses Figure 3 (reproduced below) depicting variables used in the decomposition of a model, along with their feature description and calculated importance. PNG media_image4.png 371 614 media_image4.png Greyscale ), and wherein the method further comprises: determining a response based on the cause (Merrill discloses adverse action information can be generated, “the model evaluation system evaluates and explains the model (or ensemble) by generating score explanation information for a specific score generated by the ensemble model for a particular input data set. In some embodiments, the score explanation information is used to generate Adverse Action information” [0033]. Merrill discloses “when generating a decision to deny a consumer credit application, lenders are required to provide to each consumer the reasons why the credit application was denied, in terms of factors the model actually used, that the consumer can take practical steps to improve. These adverse action reasons and notices …” [0018]. Merrill discloses adverse action information can be generated from score explanation information (‘cause’). The adverse action information consists of reasons why a credit application was denied (‘response’) and steps a consumer can take to improve (‘second recommendation’).); and generating for display a second recommendation for executing the response (Merrill discloses “the model evaluation system uses a decomposition generated for a model score (e.g., at one or more of S230, S240, S250) to generate feature importance information and provide the generated feature importance information to the operator device 171 … the model evaluation system … provide the generated adverse action information to the operator device 171” [0131-0132]. See [0190-0191] describing how an operator device 171 can be used to display natural language explanations.). With respect to claim 15, the rejection of claim 2 is incorporated. The difference in scope being: A non-transitory, computer-readable medium comprising instructions that, when executed by one or more processors, cause operations comprising ([0235] “one or more processors of the processing unit receive instructions stored by the one or more of a RAM, ROM, and machine-readable storage medium via a bus; and the one or more processors execute the received instructions”). The following are the references relied upon in the rejections below: Baydin, Atilim Gunes, et al. "Automatic differentiation in machine learning: a survey." Journal of machine learning research 18.153 (2018): 1-43. Claims 5 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Baydin. With respect to claims 5 and 18, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches the method of claim 2, however, the combination does not teach approximating a derivative using finite differences, which is taught by Baydin: wherein receiving the second prediction for the artificial intelligence model further comprises: approximating a derivative for the artificial intelligence model using finite differences by solving differential equations (Baydin discloses “Numerical differentiation is the finite difference approximation of derivatives using values of the original function evaluated at some sample points (Burden and Faires, 2001) (Figure 2, lower right). In its simplest form, it is based on the limit definition of a derivative. For example, for a multivariate function f :   R n → R , one can approximate the gradient ∇ f   = ( ∂ f ∂ x 1 ,   … ,   ∂ f ∂ x n ) using PNG media_image5.png 174 902 media_image5.png Greyscale where ei is the i-th unit vector and h > 0 is a small step size. This has the advantage of being uncomplicated to implement” (P. 4, Sec. 2.1, ¶1).); and determining numerical approximations of gradients for the artificial intelligence model based on the derivative (Baydin discloses Equation 1 on P. 4 (reproduced above) describing how a gradient ∇ f can be approximated using numerical differentiation.). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the finite differences disclosed by Baydin to approximate a gradient using finite differences. By approximating a gradient using finite differences, finite differences are uncomplicated to implement, thereby enabling a more straightforward and interpretable way to approximate gradients. The following are the references relied upon in the rejections below: Kirsch, U., Bogomolni, M. & Sheinman, I. Efficient structural optimization using reanalysis and sensitivity reanalysis. Engineering with Computers 23, 229–239 (2007). https://doi.org/10.1007/s00366-007-0062-1 Claims 6 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Baydin / Kirsch. With respect to claims 6 and 19, the combined model explainability method of Merrill / Mishra / Dugger / Tahir / Baydin teaches the method of claim 5, however, the combination does not teach approximating a derivative using a predetermined step-size, which is taught by Kirsch: the method of claim 5, wherein approximating the derivative for the artificial intelligence model using finite differences by solving differential equations further comprises: receiving a predetermined step-size for a first application (Kirsch discloses “we use efficient finite-differences for purposes of illustration. Assuming forward-differences, the response derivatives are approximated from the displacements at the original design point X0 and at the perturbed point X0 + δX by PNG media_image6.png 123 811 media_image6.png Greyscale where δX is predetermined step-size” (P. 231, Sec. 2.1.1, ¶2).); and using the predetermined step-size for approximating the derivative (Kirsch discloses Equation 2 on P. 231 (reproduced above) describing derivates are approximated using finite differences with a predetermined step-size δX.). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir / Baydin with the predetermined step-size disclosed by Kirsch to approximate derivatives with a predetermined step size. By approximating derivatives with a predetermined step-size, an optimal step-size known to balance the trade-off between truncation error and round-off error can be used to consistently maximize the accuracy of a derivative approximation. The following are the references relied upon in the rejections below: Mishra et al. (US 20230281047 A1), hereinafter Mishra ‘047 Claims 7 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Mishra ‘047. With respect to claims 7 and 20, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches the method of claim 2, however, the combination does not teach approximating an integral by approximating a region under a graph, which is taught by Mishra ‘047: wherein receiving the second prediction for the artificial intelligence model further comprises: approximating an integral for the artificial intelligence model by approximating a region under a graph of a function that defines the artificial intelligence model (Mishra ‘047 discloses “the transformation of one or more IG computations can be used to generate an XAI model (XAI model 620) and/or can be utilized for feature attribution to provide explainable ML in accordance with one or more embodiments of the present disclosure. As described herein, the computation of IG is straightforward using Equation 3. In many circumstances, the output function F (e.g., as provided in Equation 3) is too complicated to have an analytically solvable integral. However, this challenge can be mitigated using two strategies. In the first strategy, numerical integration with polynomial interpolation can be applied to approximate the integral” [0085]. Mishra ‘047 discloses “Integrated Gradients (IG) serves as another technique to explain ML models. The equation to compute the IG attribution for an input record x and a baseline record x′ is as follows … where F :   R n → [ 0 ,   1 ] represents the ML model” [0045-0046]. See [0045] describing integrated gradients can be computed using equation 3, where function F represents a machine learning model. Mishra ‘047 discloses “The numerical integration is computed through the trapezoidal rule. Formally, the trapezoidal rule works by approximating the region under the graph of the function F(x) as a trapezoid and calculating its area to approach the definite integral, which is actually the result obtained by averaging the left and right Riemann sums. The interpolation improves approximation by partitioning the integration interval, applying the trapezoidal rule to each sub-interval, and summing the results” [0086].); and determining numerical approximations of integrals for the artificial intelligence model based on the integral (Mishra ‘047 discloses “the output function F (e.g., as provided in Equation 3) is too complicated to have an analytically solvable integral. However, this challenge can be mitigated using two strategies. In the first strategy, numerical integration with polynomial interpolation can be applied to approximate the integral” [0085]. See [0086] describing how approximated areas are summed to approximate an integral.). Mishra ‘047 teaches approximating an integral by approximating a region under a graph is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the approximation technique disclosed by Mishra ‘047 to approximate an integral by approximating a region under a graph. By approximating an integral using region under a graph, complicated functions that represent machine learning models can have solvable integrals, thereby enabling machine learning models to use integrals to generate predictions or provide feature attribution. The following are the references relied upon in the rejections below: Holmes, Mark H. "Numerical Integration." Introduction to Scientific Computing and Data Analysis. Cham: Springer International Publishing, 2016. 231-274. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Holmes. With respect to claim 8, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches the method of claim 2, wherein receiving the second prediction for the artificial intelligence model further comprises: … and determining numerical approximations of integrals for the artificial intelligence model based on the integral ([0236] “the formulation of the integrated gradients approach depends upon a path … The integral can be evaluated numerically, simply by calculating the gradient ∇ƒ at equally spaced points along the path … For a non-differentiable model, such as a tree based model, it may be necessary to estimate the gradient numerically.”). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the integrated gradients technique disclosed by Dugger to estimate integrated gradients for a non-differentiable model. By estimating integrated gradients for a non-differentiable model, integrated gradients can be calculated for a non-differentiable model to explain the impact a variable has on a predictive output, thereby increasing model interpretability for a non-differentiable model. However, the combination does not teach approximating an integral by approximating an integrand f(x) by a quadratic interpolant P(x), which is taught by Holmes: approximating an integral for the artificial intelligence model by approximating an integrand f (x) by a quadratic interpolant P(x) of a function that defines the artificial intelligence model (Holmes discloses “The objective of this chapter is to derive and then test methods that can be used to evaluate the definite integral ∫ a b f x d x ” (P. 231, Sec. 6.1, ¶1). Holmes discloses the integrand f(x) is approximated by using quadratic approximation, “The next step is to try a quadratic approximation for f(x), and for this we need three data points. One option is to use xi, xi+1, and some point within the subinterval. Another option is to pair up the subintervals and use an approximation over x1 ≤ x ≤ x3, another one over x3 ≤ x ≤ x5, etc. We will use the latter option although this will require n to be even. From (5.4), the quadratic that interpolates f(x) over the interval xi−1 ≤ x ≤ xi+1 is PNG media_image7.png 104 1106 media_image7.png Greyscale An example of the resulting approximation is shown in Figure 6.7 in the case when n = 4. There are two quadratics in this case, one used for 0 ≤ x ≤ 1 2   and another for 1 2   ≤ x ≤ 1. If you look closely, you will notice that over each subinterval the quadratic is above the function on one half, and below the function on the other half. This is also what happened for the midpoint rule, as seen in Figure 6.3, and it will result in the integration rule being more accurate than might be expected” (P. 241, Sec. 6.3.2, ¶1-2). Holmes disclose Figure 6.7 on P. 240 (reproduced below) depicting piecewise quadratic approximation p2(x) used to approximate f(x). PNG media_image8.png 356 1028 media_image8.png Greyscale ); Holmes teaches approximating an integrand f(x) by using quadratic approximation is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the quadratic approximation technique disclosed by Holmes to approximate an integral by using quadratic approximation. By approximating an integral using quadratic approximation, the curvature of f(x) can be approximated more accurately, thereby resulting in a more precise approximation of the area under a curve for an integral. The following are the references relied upon in the rejections below: Gonzales (US 20230245128 A1) Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Gonzales. With respect to claim 11, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, … and wherein the method further comprises: determining a … response based on the cause; and generating for display a second recommendation for executing the … response (Merrill discloses adverse action information can be generated from score explanation information (‘cause’), see [0033]. The adverse action information consists of reasons why a credit application was denied (‘response’) and steps a consumer can take to improve (‘second recommendation’), see [0018]. See [0131-0132] describing how adverse action information is sent to an operator device.). However, the combination does not teach detecting fraudulent transactions, which is taught by Gonzales: wherein the known label comprises a detected fraudulent transaction (Gonzales discloses “the data harvesting detection system 106 can determine that the account number is involved in fraudulent activity (e.g., via the account number being flagged, location of transaction, amount of transaction, time of transaction)” [0061].), wherein the plurality of values indicates a transaction history of a user (Gonzales discloses “the data harvesting detection system receives the network transaction requests, within a time period, having an account number (e.g., a credit card number) and a transaction amount and determines transaction request response codes for the requests” [0020]. Gonzales discloses “the client applications 114a-114n (via the client devices 112a-112n) can provide user data activity (e.g., network transaction requests) to the data harvesting detection system 106 (via the transaction facilitator network device(s) 110 to the server device(s) 102) to detect account number harvesting” [0044].), and wherein the method further comprises: determining a fraudulent transaction response based on the cause (Gonzales discloses “As an example, the data harvesting detection system 106 can utilize transaction request response codes for fraud alert decline responses. For instance, a transaction request response code for a fraud alert decline response can include a code to indicate a response for (detected) fraudulent activity corresponding to the account number. In particular, the data harvesting detection system 106 can determine that the account number is involved in fraudulent activity (e.g., via the account number being flagged, location of transaction, amount of transaction, time of transaction). In some instances, the data harvesting detection system 106 can determine a transaction request response code for a fraud alert decline response that indicates irregular transaction activity for the account number (e.g., increased usage, usage with an irregular transaction facilitator, irregular time or location of transaction)” [0061]. Gonzales discloses “the term “transaction request response code” refers to a label, descriptor, or other text or numeric that describes or indicates a response to a network transaction request. In particular, a transaction request response code includes a label or descriptor for a verification or rejection performed on an account number that corresponds to a network transaction request” [0030]. A transaction request response code (‘fraudulent transaction response’) can be generated to explain which irregular transaction activity (‘cause’) indicated an account number as being involved in fraudulent activity.); Gonzales teaches a detection system that labels transactions as fraudulent and generates transaction request response codes is known in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the detection system disclosed by Gonzales to detect fraudulent transactions. By detecting fraudulent transactions, financial businesses and customers can protect financial assets, thus preventing financial losses and ensuring trust in financial institutions. The following are the references relied upon in the rejections below: J. Ables et al., "Creating an Explainable Intrusion Detection System Using Self Organizing Maps," 2022 IEEE Symposium Series on Computational Intelligence (SSCI), Singapore, Singapore, 2022, pp. 404-412, doi: 10.1109/SSCI51031.2022.10022255. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Ables. With respect to claim 12, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, … and wherein the method further comprises: determining a … response based on the cause; and generating for display a second recommendation for executing the … response (Merrill discloses adverse action information can be generated from score explanation information (‘cause’), see [0033]. The adverse action information consists of reasons why a credit application was denied (‘response’) and steps a consumer can take to improve (‘second recommendation’), see [0018]. See [0131-0132] describing how adverse action information is sent to an operator device.). However, the combination does not teach detecting cyber incidents, which is taught by Ables: wherein the known label comprises a detected cyber incident (Ables discloses “systems monitor networks and automate attack detection by comparing network activity to the signature of known attacks or by detecting behavior that is anomalous to benign network patterns [2]. Through these methods, a security analyst can use an IDS to detect improper use, unauthorized access, or the abuse of a network” (P. 404, Sec. I, ¶2). Ables discloses “we can see the features with the largest impact on a prediction: duration, dst bytes, and src bytes. These features were the closest to the BMU, and they played a large role in computing the predicted value. Seeing the specific features that influence predictions provides insight about samples labeled as malicious or benign and can further help operators determine the reason of incorrect predictions” (P. 408, Sec. IV-C, ¶3).), wherein the plurality of values indicates networking activity of a user (Ables discloses “The results for the NSL-KDD dataset can be found in Figures 2a and 2b. The local explanation example shows that the most important features for its prediction were ‘Duration’, ‘Destination (dst) bytes’, and ‘Source (src) bytes’. The remaining features, ‘Service (srv) count’, ‘Count’, and ‘Destination (dst) host count’ are considered less significant because of their distance from the BMU” (P. 409, Sec. V-A, ¶2).), and wherein the method further comprises: determining a cyber incident response based on the cause (Ables discloses “Seeing the specific features that influence predictions provides insight about samples labeled as malicious or benign and can further help operators determine the reason of incorrect predictions. These features can also be further investigated with feature value heat maps” (P. 408, Sec. IV-C, ¶3). Ables discloses “the users need to be confident in the predictions or recommendations computed by an IDS. Understandable explanations allow users to perform their tasks correctly. The stakeholders of an IDS (e.g. CSoC operators, developers, and investors) are individuals who will be dependent on the performance of the system. CSoC operators will be performing defensive actions based on prediction and explanation results. Developers can use explanations to fortify the model in areas where it is weak. Investors may need explanations to help them in making budgeting decisions for their company” (P. 406, Sec. II-C, ¶1).); Ables teaches an intrusion detection system that generates explanations to explain which features influenced a sample being labeled as malicious or benign is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the intrusion detection system disclosed by Ables to classify network activity. By classifying network activity, network activity that is classified as malicious can be further investigated and preventative defensive actions can be taken to mitigate threats and prevent system compromises. The following are the references relied upon in the rejections below: João Bento, et al. 2021. TimeSHAP: Explaining Recurrent Models through Sequence Perturbations. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD '21). Association for Computing Machinery, New York, NY, USA, 2565–2573. https://doi.org/10.1145/3447548.3467166 Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Merrill / Mishra / Dugger / Tahir / Bento. With respect to claim 14, the combined model explainability method of Merrill / Mishra / Dugger / Tahir teaches: the method of claim 2, … and wherein the method further comprises: determining a … response based on the cause; and generating for display a second recommendation for executing the … response (Merrill discloses adverse action information can be generated from score explanation information (‘cause’), see [0033]. The adverse action information consists of reasons why a credit application was denied (‘response’) and steps a consumer can take to improve (‘second recommendation’), see [0018]. See [0131-0132] describing how adverse action information is sent to an operator device.). However, the combination does not teach detecting identity theft, which is taught by Bento: wherein the known label comprises a detected identity theft (Bento discloses “We use TimeSHAP to explain the predictions of a real-world bank account takeover fraud detection RNN model, and draw key insights from its explanations: i) the model identifies important features and events aligned with what fraud analysts consider cues for account takeover” (P. 2565, Abstract).), wherein the plurality of values indicates a user transaction history (Bento discloses “Each instance, dubbed from here on as an event, represents one of three behaviours, or event types: transaction, where a client performs a monetary transaction; login, representing a client login on the banking application or website; or enrollment, representing account settings behaviours, such as logging into a new device or changing the password” (P. 2569, Sec. 4, ¶1-2). Bento discloses “Figure 4 shows a plot of the global feature-wise explanations. We observe that some features have predominantly positive contributions, meaning that the model routinely relies on these features to predict account takeover fraud. These include the Transaction (0.29) and Event (0.092) types, the clients’ age (0.090), and features related to the IP and location of the events (0.08 to 0.03)” (P. 2570-2571, Sec. 4.2, Last Paragraph).), and wherein the method further comprises: determining an identity theft response based on the cause (Bento discloses “Regarding feature importances, the most relevant features are related to the transaction type, event type, the clients’ age, and the location. When inspecting the raw feature data, we observe that the client is in the elderly age range, which, as previously mentioned, may indicate a more susceptible demographic. When analyzing the location features Location feature A and Location feature D, we observe a discrepancy between the location of the enrollment, login and transactions from the account’s history. This discrepancy in physical location is highly suspicious and indicates that there was an enrollment on the account from a previously unused location” (P. 2571-2572, Sec. 4.3, Last Paragraph). Bento discloses “Local explanations explain the model’s rationale regarding one specific instance. These explanations can be used in several use cases, for example, for bias auditing or model debugging. However, these explanations can mostly be used by end-users, the fraud analysts, to aid their decision-making tasks” (P. 2571, Sec. 4.3, ¶1). Local explanations that explain which features contribute the most towards a model’s account takeover prediction (therefore local explanations are a cause), can be used for bias auditing, model debugging, and decision-making tasks (therefore these actions are an identity theft response).); Bento teaches using local explanations to explain which features contribute the most towards predicting identity theft is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined model explainability method of Merrill / Mishra / Dugger / Tahir with the technique disclosed by Bento to use a machine learning model to predict identity theft. By using a machine learning model to predict identity theft, a model can automatically classify transactions as fraudulent, thereby enabling real-time fraud detection and preventing account takeover. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Dhurandhar et al. (US 20200193243 A1) teaches estimating gradients for a non-differential model to generate explanations for a decision made by a classifier. Abeyagunasekera et al. ("LISA: Enhance the explainability of medical images unifying current XAI techniques.") teaches using expected gradients to recreate integrals as expectations to approximate SHAP values. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PEDRO J MORALES/Examiner, Art Unit 2124 /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Mar 27, 2023
Application Filed
Feb 26, 2026
Non-Final Rejection mailed — §103
May 13, 2026
Interview Requested
May 21, 2026
Examiner Interview Summary
May 21, 2026
Applicant Interview (Telephonic)
May 21, 2026
Response Filed
Jul 30, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664036
ANOMALY DETECTION IN COMPUTER SYSTEMS
4y 0m to grant Granted Jun 23, 2026
Patent 12651179
REINFORCEMENT LEARNING MODEL FOR BALANCED UNIT RECOMMENDATION
4y 5m to grant Granted Jun 09, 2026
Patent 12639625
BIAS ADJUSTMENT DEVICE, INFORMATION PROCESSING DEVICE, INFORMATION PROCESSING METHOD, AND INFORMATION PROCESSING PROGRAM
4y 1m to grant Granted May 26, 2026
Patent 12591803
SYSTEMS AND METHODS FOR APPLYING MACHINE LEARNING BASED ANOMALY DETECTION IN A CONSTRAINED NETWORK
3y 11m to grant Granted Mar 31, 2026
Patent 12530412
SEARCH-QUERY SUGGESTIONS USING REINFORCEMENT LEARNING
4y 2m to grant Granted Jan 20, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+55.6%)
3y 8m (~3m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 13 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month