Prosecution Insights
Last updated: August 17, 2026
Application No. 18/427,467

SYSTEMS AND METHODS FOR APPROXIMATION OF SHAPLEY VALUES IN MEMORY-CONSTRAINED ENVIRONMENTS

Non-Final OA §101§103
Filed
Jan 30, 2024
Examiner
ROHD, BENJAMIN MATTHEW
Art Unit
Tech Center
Assignee
JPMorgan Chase Bank, N.A.
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
1y 8m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 2 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
21 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
24.9%
-15.1% vs TC avg
§103
49.2%
+9.2% vs TC avg
§102
9.6%
-30.4% vs TC avg
§112
15.8%
-24.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 2 resolved cases

Office Action

§101 §103
DETAILED ACTION This office action is in response to submission of application on 01/30/2024. Claims 1-12 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1: Step 1: The claim is directed to a method, which falls within the statutory category of a process. Step 2A Prong 1: The claim is directed to an abstract idea. Specifically, the claim recites: estimating, [by a model explanation computer program], a Shapley value for the prediction. (Abstract idea – mental process. Estimating a Shapley value is a judgement/evaluation which can practically be performed in the human mind or with the aid of pen and paper. The courts have recognized that claims can recite a mental process even if they are claimed as being performed on a computer. See MPEP 2106.04(a)(2)(III).) Step 2A Prong 2: The additional elements recited in the claim do not integrate the abstract idea into a practical application, individually or in combination. Specifically, the claim recites the additional elements: receiving, by a training computer program executed on a training electronic device, a training dataset; (Receiving a training dataset amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).) training, by the training computer program, a machine learning model on the training dataset; (Training a generic machine learning model on training data is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) deploying, by the training computer program, the machine learning model to a deployment electronic device; (Deploying the model to an edge device amounts to adding insignificant extra-solution activity to the judicial exception – see MPEP2106.05(g).) receiving, by the deployment electronic device, an incoming data query; (Receiving a data query amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).) generating, using the machine learning model, a prediction for the incoming data query; (Generating a prediction using a generic machine learning model is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) estimating the Shapley value by a model explanation computer program (Estimating the Shapley value using a computer program amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Specifically, the claim recites the additional elements: receiving, by a training computer program executed on a training electronic device, a training dataset; (Receiving a training dataset amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) training, by the training computer program, a machine learning model on the training dataset; (Training a generic machine learning model on training data is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) deploying, by the training computer program, the machine learning model to a deployment electronic device; (Deploying the model to an edge device amounts to adding insignificant extra-solution activity to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) receiving, by the deployment electronic device, an incoming data query; (Receiving a data query amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) generating, using the machine learning model, a prediction for the incoming data query; (Generating a prediction using a generic machine learning model is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) estimating the Shapley value by a model explanation computer program (Estimating the Shapley value using a computer program amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) Claims 2-8: Claim 2 recites The method of claim 1, further comprising: selecting, by the training computer program, a subset of the training dataset; and communicating, by the training computer program, the subset to the model explanation computer program, wherein the model explanation computer program estimates the Shapley value using kernel density estimation. Selecting a subset of training data is a judgement/evaluation which can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process). Communicating the subset to the model explanation computer program amounts to adding insignificant extra-solution activity to the judicial exception, and is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts. Estimating the Shapley value using kernel density estimation is a mathematical concept – see MPEP 2106.04(a)(2)(I). Therefore, the claim merges with the abstract idea recited in claim 1, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claim 3 recites The method of claim 1, further comprising: clustering, by the training computer program, the training dataset into a plurality of clusters; and communicating, by the training computer program, centroids for the plurality of clusters to the deployment electronic device; wherein the model explanation computer program estimates the Shapley value by aggregating data from the plurality of clusters and weighing the plurality of clusters by similarity to the incoming query data. Clustering the training data and estimating the Shapley value by aggregating the cluster data weighted by similarity to the query data can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally grouping similar training instances, mentally determining a similarity of the query data to each group, and then summing the average Shapley values of each group weighted by their similarity to the query data using pen and paper. Communicating cluster centroids to the deployment electronic device amounts to adding insignificant extra-solution activity to the judicial exception, and is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts. Therefore, the claim merges with the abstract idea recited in claim 1, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claim 4 recites The method of claim 1, further comprising: hashing, by the training computer program, the training dataset into a plurality of arrays using locality sensitive hashing; and communicating, by the training computer program, the plurality of arrays to the deployment electronic device; wherein the model explanation computer program estimates the Shapley value by using kernel density estimation. Hashing training data using locality sensitive hashing and estimating the Shapley value using kernel density estimation are mathematical concepts. Communicating the hashed arrays to the deployment electronic device amounts to adding insignificant extra-solution activity to the judicial exception, and is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts. Therefore, the claim merges with the abstract idea recited in claim 1, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claims 5-8 are system claims containing substantially the same elements as method claims 1-4, respectively, and are rejected on the same grounds under 35 U.S.C. 101 as claims 1-4, respectively, mutatis mutandis. The additional components of a training electronic device comprising a first computer processor and a deployment electronic device comprising a second computer processor are interpreted as generic computer components, and thus amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea. Therefore, the claims do not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claim 9: Step 1: The claim is directed to a method, which falls within the statutory category of a process. Step 2A Prong 1: The claim is directed to an abstract idea. Specifically, the claim recites: estimating, [by a model explanation computer program], a Shapley value for the prediction. (Abstract idea – mental process. Estimating a Shapley value is a judgement/evaluation which can practically be performed in the human mind or with the aid of pen and paper. The courts have recognized that claims can recite a mental process even if they are claimed as being performed on a computer. See MPEP 2106.04(a)(2)(III).) Step 2A Prong 2: The additional elements recited in the claim do not integrate the abstract idea into a practical application, individually or in combination. Specifically, the claim recites the additional elements: receiving, by a training computer program executed by a deployment electronic device, a trainable machine learning model; (Receiving a machine learning model amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).) reserving, by the training computer program, an amount of memory on the deployment electronic device; (Reserving memory is a standard function of a generic computing device, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) receiving, by the training computer program, a plurality of training examples in a stream; (Receiving a stream of training data amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).) training, by the training computer program, the trainable machine learning model with the training examples; (Training a generic machine learning model on training data is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) storing, by the training computer program, some of the training examples in the amount of memory; (Storing training data in memory amounts to adding insignificant extra-solution activity to the judicial exception – see MPEP2106.05(g).) receiving, by the trainable machine learning model, an incoming data query; (Receiving a data query amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).) generating, by the trainable machine learning model, a prediction for the incoming data query; (Generating a prediction using a generic machine learning model is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) estimating the Shapley value by a model explanation computer program (Estimating the Shapley value using a computer program amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Specifically, the claim recites the additional elements: receiving, by a training computer program executed by a deployment electronic device, a trainable machine learning model; (Receiving a machine learning model amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) reserving, by the training computer program, an amount of memory on the deployment electronic device; (Reserving memory is a standard function of a generic computing device, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) receiving, by the training computer program, a plurality of training examples in a stream; (Receiving a stream of training data amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) training, by the training computer program, the trainable machine learning model with the training examples; (Training a generic machine learning model on training data is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) storing, by the training computer program, some of the training examples in the amount of memory; (Storing training data in memory amounts to adding insignificant extra-solution activity to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to storing and retrieving information in memory, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) receiving, by the trainable machine learning model, an incoming data query; (Receiving a data query amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).) generating, by the trainable machine learning model, a prediction for the incoming data query; (Generating a prediction using a generic machine learning model is standard in the field of machine learning, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) estimating the Shapley value by a model explanation computer program (Estimating the Shapley value using a computer program amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).) Claims 10-12: Claim 10 recites The method of claim 9, wherein the model explanation computer program estimates the Shapley value based on a similarity of the incoming data query to one of the plurality of training examples. Estimating the Shapley value based on the similarity of the data query to a training example can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally identifying a training example with similar features to the data query and estimating the data query’s Shapley value to be equal to the Shapley value of the identified training example. Therefore, the claim merges with the abstract idea recited in claim 9, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claim 11 recites The method of claim 9, further comprising: initializing, by the training computer program, a plurality of centroid vectors in the amount of memory; and identifying, by the training computer program and for each training example, one of the plurality of centroid vectors that is closest to the training example and updating the closest centroid vector with the training example; wherein the model explanation computer program estimates the Shapley value by identifying one of the plurality of centroid vectors that is closest to the incoming data query. Initializing centroid vectors, identifying the closest centroid vector for a training example, and updating the centroid vector with the training example can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally selecting training data points to use as centroids, writing the centroid vectors out on paper, mentally determining which centroid is most similar to a training example, and adjusting the centroid vector’s values by hand such that it is closer to the training example. Storing the centroid vectors in memory amounts to adding insignificant extra-solution activity to the judicial exception, and storing and retrieving information in memory is well-understood, routine, and conventional in the computer arts. Estimating the Shapley value by identifying the closest centroid vector can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally determining the centroid vector most similar to the data query and estimating the data query’s Shapley value to be equal to the Shapley value of the centroid. Therefore, the claim merges with the abstract idea recited in claim 9, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claim 12 recites The method of claim 9, further comprising: initializing, by the training computer program, a random number of vectors in the amount of memory as centroids; calculating, by the training computer program, a hash string for each centroid using locality sensitive hashing; and hashing, by the training computer program, each of the plurality of training examples using locality sensitive hashing and identifying one of the plurality of centroids corresponding to the hash; wherein the model explanation computer program estimates the Shapley value by identifying one of the plurality of centroids that is closest to the incoming data query. Initializing a random number of centroid vectors is an evaluation/judgement that can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process). Hashing centroids and training examples using locality sensitive hashing is a mathematical concept. Estimating the Shapley value by identifying the closest centroid can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally determining the centroid vector most similar to the data query and estimating the data query’s Shapley value to be equal to the Shapley value of the centroid. Therefore, the claim merges with the abstract idea recited in claim 9, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 5, and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Sharma et al. (hereinafter Sharma), U.S. Patent Application Publication US-20200327371-A1 (filed 04/09/2019), in view of Yang et al. (hereinafter Yang), “Efficient Shapley Values Estimation by Amortization for Text Classification” (published 05/31/2023). Regarding Claim 1, Sharma teaches A method, comprising: receiving, by a training computer program executed on a training electronic device, a training dataset; (0113: “Further, the data publisher can be used to retrieve aggregated data stored in the local time-series database and to transfer the data to the remote cloud storage, for example to facilitate developing machine learning models for deployment…” Aggregated data (i.e. a training dataset) is received by the remote network (i.e. a training electronic device) to facilitate developing machine learning models (i.e. a training computer program).) training, by the training computer program, a machine learning model on the training dataset; (0012: “A machine learning model is created and trained in the remote network using aggregated sensor data and deployed to the edge platform.” A machine learning model is trained using the aggregated sensor data (i.e. training dataset) in the remote network (i.e. by the training computer program).) deploying, by the training computer program, the machine learning model to a deployment electronic device; (See the portion of 0012 cited above. The machine learning model is deployed to an edge platform (i.e. a deployment electronic device).) receiving, by the deployment electronic device, an incoming data query; (0013: “The edge computing platform receives a first sensor data stream from a first sensor of the plurality of sensors.” The edge computing platform (i.e. deployment electronic device) receives sensor data (i.e. an incoming data query).) generating, using the machine learning model, a prediction for the incoming data query; and (0012: “The ‘edge-ified’ model is adapted to operate on continuous streams of sensor data in real-time and produce inferences.” The machine learning model produces inferences (i.e. generates predictions) for the sensor data (i.e. incoming data query).) Sharma does not appear to explicitly disclose estimating, by a model explanation computer program, a Shapley value for the prediction. However, Yang teaches estimating, by a model explanation computer program, a Shapley value for the prediction. (Pg. 1, Abstract: “To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature’s Shapley Value…” Pg. 2, section 1: “At inference time, our amortized model directly outputs explanation scores for new instances.” The amortized model (i.e. model explanation computer program) outputs an estimated Shapley Value for the instance.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma and Yang. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. One of ordinary skill would have motivation to combine Sharma and Yang in order to increase the interpretability of the model taught by Sharma while adhering to the resource constraints of the edge platform. According to Yang, Shapley values are a popular way to introduce explainability to a model, but “computing them is prohibitive for large pretrained models due to a large number of model evaluations… our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods” (Yang, pg. 1, Abstract). Claim 5 is a system claim containing substantially the same elements as method claim 1. Sharma and Yang teach the elements of claim 1, as shown above. Sharma also teaches a training electronic device comprising a first computer processor (0178-0180: “ML models are typically created in a ‘development’ computing environment, such as a cloud computing environment… Computing resources may suitably include various known processors.”) a deployment electronic device comprising a second computer processor (0117: “The edge infrastructure comprises certain hardware components, such as network connections and a processor…”) Regarding Claim 9, Sharma teaches A method, comprising: receiving, by a training computer program executed by a deployment electronic device, a trainable machine learning model; (0012: “A machine learning model is created and trained in the remote network using aggregated sensor data and deployed to the edge platform.” A trainable machine learning model is deployed to an edge platform (i.e. received by a deployment electronic device).) reserving, by the training computer program, an amount of memory on the deployment electronic device; (0079: “The edge software stack services also may include a local time-series database in which sensor and other data may be aggregated locally and from which applications can make time-based sensor data queries.” A local time-series database (i.e. an amount of memory) is reserved on the edge platform (i.e. deployment electronic device).) receiving, by the training computer program, a plurality of training examples in a stream; (See the portion of 0079 cited above. Time-series data (i.e. training examples in a stream) is aggregated (i.e. received).) training, by the training computer program, the trainable machine learning model with the training examples; (0104: “[T]he software development kit can access aggregated time-series sensor data stored locally in the time-series database to facilitate developing and training machine learning models on the edge platform…” The machine learning model is trained using the time-series data (i.e. training examples).) storing, by the training computer program, some of the training examples in the amount of memory; (See the portions of 0079 and 0104 cited above. The time series data (i.e. training examples) is stored in the local time-series database (i.e. amount of memory).) receiving, by the trainable machine learning model, an incoming data query; (0013: “The edge computing platform receives a first sensor data stream from a first sensor of the plurality of sensors.” The edge computing platform (i.e. deployment electronic device) receives sensor data (i.e. an incoming data query).) generating, by the trainable machine learning model, a prediction for the incoming data query; (0012: “The ‘edge-ified’ model is adapted to operate on continuous streams of sensor data in real-time and produce inferences.” The machine learning model produces inferences (i.e. generates predictions) for the sensor data (i.e. incoming data query).) Sharma does not appear to explicitly disclose estimating, by a model explanation computer program, a Shapley value for the prediction. However, Yang teaches estimating, by a model explanation computer program, a Shapley value for the prediction. (Pg. 1, Abstract: “To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature’s Shapley Value…” Pg. 2, section 1: “At inference time, our amortized model directly outputs explanation scores for new instances.” The amortized model (i.e. model explanation computer program) outputs an estimated Shapley Value for the instance.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma and Yang. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. One of ordinary skill would have motivation to combine Sharma and Yang in order to increase the interpretability of the model taught by Sharma while adhering to the resource constraints of the edge platform. According to Yang, Shapley values are a popular way to introduce explainability to a model, but “computing them is prohibitive for large pretrained models due to a large number of model evaluations… our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods” (Yang, pg. 1, Abstract). Claims 2-3 and 6-7 are rejected under 35 U.S.C. 103 as being unpatentable over Sharma in view of Yang, and further in view of Wang et al. (hereinafter Wang), “A Simple Nadaraya-Watson Head can offer Explainable and Calibrated Classification” (published 12/07/2022). Regarding Claim 2, Sharma and Yang teach The method of claim 1, as shown above. Sharma and Yang do not appear to explicitly disclose the remaining features of claim 2. However, Wang teaches further comprising: selecting, by the training computer program, a subset of the training dataset; and (Pg. 4, section 3.3: “To characterize the effect of the support set on inference, in our experiments we implement the following ‘inference modes’… Random: Sample uniformly at random over the dataset, such that each class is represented k times: S ∼ D and   | S | = k | C | .” A support set (i.e. subset) S is selected from the training dataset D .) communicating, by the training computer program, the subset to the model explanation computer program, (See the portions of section 3.3 cited above and section 3.1 cited below. The subset is used as the support set for inferencing and is therefore necessarily communicated to the prediction model (i.e. the model explanation computer program which performs the estimation).) wherein the model explanation computer program estimates the Shapley value using kernel density estimation. (Examiner notes that per specification paragraph 0054 of the instant application, “the model explanation computer program may estimate a Shapley value for incoming data query prediction by using a kernel density estimator, such as the Nadaraya-Watson estimator…” Wang teaches prediction using the Nadaraya-Watson estimator, e.g. pg. 3, section 3.1: “Consider a ‘support set’ of N s examples and their associated labels, S = { z i : ( x i , y i ) } i = 1 N s . This support set can be a subset, or the whole, of the training dataset. In this paper, we will focus on the image classification setting… We note that the NW [Nadaraya-Watson] head can be used with any input data type, not just images, and can be readily extended to the regression setting. For a query image x , the Nadaraya-Watson prediction is computed as the weighted sum of support labels y → i , where the weights quantify the similarity of the query image with each support image x i …” Prediction is performed using the Nadaraya-Watson estimator (i.e. kernel density estimation), which is applicable to regression settings (e.g. estimating Shapley values).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, and Wang. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. Wang teaches using a Nadaraya-Watson estimator for prediction based on a support set sampled from the training dataset. One of ordinary skill would have motivation to combine Sharma, Yang, and Wang because replacing Yang’s prediction model with Wang’s Nadaraya-Watson prediction model amounts to a simple substitution of known alternatives, and Wang’s nonparametric model “can be more interpretable, since the dependence on other datapoints gives an indication about what is driving the prediction” (Wang, pg. 1, section 1). Regarding Claim 3, Sharma and Yang teach The method of claim 1, as shown above. Sharma and Yang do not appear to explicitly disclose the remaining features of claim 3. However, Wang teaches further comprising: clustering, by the training computer program, the training dataset into a plurality of clusters; and (Pg. 4, section 3.3: “To characterize the effect of the support set on inference, in our experiments we implement the following ‘inference modes’… Cluster: Given the trained model, we first compute the features for all the training datapoints. Then, we perform k -means clustering on the features of the training datapoints for each class. These k cluster centroids are then used as the support features for each class.”) communicating, by the training computer program, centroids for the plurality of clusters to the deployment electronic device; (See the portions of section 3.3 cited above and section 3.1 cited below. The cluster centroids are used as the support set for inferencing and are therefore necessarily communicated to the prediction model (i.e. the deployment electronic device which performs the estimation).) wherein the model explanation computer program estimates the Shapley value by aggregating data from the plurality of clusters and weighing the plurality of clusters by similarity to the incoming query data. (Pg. 3, section 3.1: “Consider a ‘support set’ of N s examples and their associated labels, S = { z i : ( x i , y i ) } i = 1 N s . This support set can be a subset, or the whole, of the training dataset. In this paper, we will focus on the image classification setting… We note that the NW [Nadaraya-Watson] head can be used with any input data type, not just images, and can be readily extended to the regression setting. For a query image x , the Nadaraya-Watson prediction is computed as the weighted sum of support labels y → i , where the weights quantify the similarity of the query image with each support image x i …” Prediction is performed by a weighted sum (i.e. aggregation) of the support set labels (i.e. data from the plurality of clusters), weighted by similarity to the query data. The estimator is applicable to regression settings (e.g. estimating Shapley values).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, and Wang. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. Wang teaches using a Nadaraya-Watson estimator for prediction based on a support set sampled from the training dataset. One of ordinary skill would have motivation to combine Sharma, Yang, and Wang because replacing Yang’s prediction model with Wang’s Nadaraya-Watson prediction model amounts to a simple substitution of known alternatives, and Wang’s nonparametric model “can be more interpretable, since the dependence on other datapoints gives an indication about what is driving the prediction” (Wang, pg. 1, section 1). Claims 6-7 are system claims containing substantially the same elements as method claims 2-3, respectively. Sharma, Yang, and Wang teach the elements of claims 2-3, as shown above. Claims 4 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Sharma in view of Yang, and further in view of Daghaghi et al. (hereinafter Daghaghi), “Adaptive Sampling for Deep Learning via Efficient Nonparametric Proxies” (published 11/22/2023). Regarding Claim 4, Sharma and Yang teach The method of claim 1, as shown above. Sharma and Yang do not appear to explicitly disclose the remaining features of claim 4. However, Daghaghi teaches further comprising: hashing, by the training computer program, the training dataset into a plurality of arrays using locality sensitive hashing; and (Pg. 2-3, section 1.2-1.3: “We will estimate the numerator and denominator of the Nadaraya-Watson kernel estimator using recent techniques from randomized algorithms for kernel density estimation. These techniques rely on a particular kind of hash function known as a locality-sensitive hash (LSH)… We begin by constructing a sketch S ∈ Z R × W , a 2D array of integers… To construct the sketch, we create R independent hash functions { h 1 , . . h R } – one for each row. For each element x i ∈ D , we increment the corresponding bucket of the sketch.” As can be seen in algorithm 1, construction of the NWS [Nadaraya-Watson Sketch] includes applying each locality sensitive hash function h 1 , . . h R to each dataset element x i (i.e. hashing the training dataset) to obtain sketch arrays S t and S b (i.e. a plurality of arrays).) communicating, by the training computer program, the plurality of arrays to the deployment electronic device; (Pg. 6, section 3.1: “After the warm-up phase, we query both sketches with the incoming batch of data, and compute scores for both arrays…” The sketch arrays are queried during inferencing and are therefore necessarily communicated to the prediction model (i.e. the deployment electronic device which performs the estimation).) wherein the model explanation computer program estimates the Shapley value by using kernel density estimation. (Examiner notes that per specification paragraph 0054 of the instant application, “the model explanation computer program may estimate a Shapley value for incoming data query prediction by using a kernel density estimator, such as the Nadaraya-Watson estimator…” Daghaghi teaches prediction using the Nadaraya-Watson estimator, e.g. pg. 5, section 2.2.2: “To demonstrate that the NWS sketch [Nadaraya-Watson sketch] is a useful model, we apply NWS to standard regression datasets.”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, and Daghaghi. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. Daghaghi teaches using a Nadaraya-Watson sketch estimator for prediction based on a support set hashed using locality sensitive hashing. One of ordinary skill would have motivation to combine Sharma, Yang, and Daghaghi because replacing Yang’s prediction model with Daghaghi’s Nadaraya-Watson prediction model amounts to a simple substitution of known alternatives, and Daghaghi’s model “provably approximates the kernel regression model with O ( N d ) training and O ( 1 ) inference complexity” (Daghaghi, pg. 2, section 1). Claim 8 is a system claim containing substantially the same elements as method claim 4. Sharma, Yang, and Daghaghi teach the elements of claim 4, as shown above. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Sharma in view of Yang, and further in view of Cover, “Estimation by the Nearest Neighbor Rule” (published 01/31/1968). Regarding Claim 10, Sharma and Yang teach The method of claim 9, as shown above. Sharma and Yang do not appear to explicitly disclose the remaining features of claim 10. However, Cover teaches wherein the model explanation computer program estimates the Shapley value based on a similarity of the incoming data query to one of the plurality of training examples. (Pg. 50, section I: “The nearest neighbor (NN) estimate of θ on the basis of the knowledge of x and the representative samples x 1 , θ 1 , x 2 , θ 2 , … , x n , θ n is defined to be θ n ' , the parameter associated with x n ' , the nearest neighbor to x .” The estimate of parameter value θ (i.e. the Shapley value) for an input x (i.e. data query) is based on the parameter value of the nearest neighbor to the input x (i.e. similarity of the data query to a training example).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, and Cover. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. Cover teaches using a nearest neighbor estimator for prediction based on a collection of labeled samples. One of ordinary skill would have motivation to combine Sharma, Yang, and Cover because replacing Yang’s prediction model with Cover’s nearest neighbor prediction model amounts to a simple substitution of known alternatives, and “it may be concluded that at least half the decision information in an infinite set of classified samples is contained in the nearest neighbor” (Cover, pg. 55, section VI). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Sharma in view of Yang, and further in view of MacQueen, “Some methods for classification and analysis of multivariate observations” (published 1967). Regarding Claim 11, Sharma and Yang teach The method of claim 9, as shown above. Sharma and Yang do not appear to explicitly disclose the remaining features of claim 11. However, MacQueen teaches further comprising: initializing, by the training computer program, a plurality of centroid vectors in the amount of memory; and (Pg. 283, section 2: “Stated informally, the k -means procedure consists of simply starting with k groups each of which consists of a single random point…” k random points are initialized as centroid vectors representing k groups.) identifying, by the training computer program and for each training example, one of the plurality of centroid vectors that is closest to the training example and updating the closest centroid vector with the training example; (Pg. 283, section 2: “…and thereafter adding each new point to the group whose mean the new point is nearest. After a point is added to a group, the mean of that group is adjusted in order to take account of the new point.” For each new point (i.e. training example), the new point’s nearest group mean (i.e. a centroid vector that is closest to the training example) is identified and adjusted (i.e. updated) based on the new point.) wherein the model explanation computer program estimates the Shapley value by identifying one of the plurality of centroid vectors that is closest to the incoming data query. (Pg. 291, section 3.2: “…a prediction, A or B, was made for each of a new sample of 250 points on the basis of whether or not each point was nearest to an A mean or a B mean.” A prediction (e.g. an estimated Shapley value) is made by identifying the closest group mean (i.e. centroid vector) to a new sample (i.e. incoming data query).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, and MacQueen. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. MacQueen teaches using a k -means inference model for prediction based on a collection of samples. One of ordinary skill would have motivation to combine Sharma, Yang, and MacQueen because replacing Yang’s prediction model with MacQueen’s k -means prediction model amounts to a simple substitution of known alternatives, and in the experiment conducted by MacQueen, “[t]hese predictions turned out to be 87% correct. As this example shows, the method is potentially capable of taking advantage of a highly nonlinear relationship. Also, the method has something to recommend it from the point of view of simplicity, and can easily be applied in many dimensions and to more than two-valued dependent variables” (MacQueen, pg. 291, section 3.2). Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Sharma in view of Yang, and further in view of MacQueen and Malewicz, U.S. Patent Application Publication US-20150213112-A1 (filed 01/24/2014). Regarding Claim 12, Sharma and Yang teach The method of claim 9, as shown above. Sharma and Yang do not appear to explicitly disclose the remaining features of claim 12. However, MacQueen teaches further comprising: initializing, by the training computer program, a random number of vectors in the amount of memory as centroids; (Pg. 283, section 2: “Stated informally, the k -means procedure consists of simply starting with k groups each of which consists of a single random point…” k random points are initialized as centroids representing k groups.) wherein the model explanation computer program estimates the Shapley value by identifying one of the plurality of centroids that is closest to the incoming data query. (Pg. 291, section 3.2: “…a prediction, A or B, was made for each of a new sample of 250 points on the basis of whether or not each point was nearest to an A mean or a B mean.” A prediction (e.g. an estimated Shapley value) is made by identifying the closest group mean (i.e. centroid vector) to a new sample (i.e. incoming data query).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, and MacQueen. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. MacQueen teaches using a k -means inference model for prediction based on a collection of samples. One of ordinary skill would have motivation to combine Sharma, Yang, and MacQueen because replacing Yang’s prediction model with MacQueen’s k -means prediction model amounts to a simple substitution of known alternatives, and in the experiment conducted by MacQueen, “[t]hese predictions turned out to be 87% correct. As this example shows, the method is potentially capable of taking advantage of a highly nonlinear relationship. Also, the method has something to recommend it from the point of view of simplicity, and can easily be applied in many dimensions and to more than two-valued dependent variables” (MacQueen, pg. 291, section 3.2). Sharma, Yang, and MacQueen do not appear to explicitly disclose the remaining features of claim 12. However, Malewicz teaches calculating, by the training computer program, a hash string for each centroid using locality sensitive hashing; and (0032: “At block 508, the system hashes each of the K centroids using the LSH [locality sensitive hashing] operation with the re-computed parameters.”) hashing, by the training computer program, each of the plurality of training examples using locality sensitive hashing and identifying one of the plurality of centroids corresponding to the hash; (0033: “At blocks 510-518, the system considers each of the query vectors for reassignment into a different cluster. At block 512, it hashes the current query vector using the LSH operation with the re-computed parameters. At block 514, it identifies the union of buckets to which the current query vector is hashed, computes the distance between the current query vector and each of the centroids in the union of the buckets, and determines the centroid corresponding to the shortest distance.” Each query vector in the set of query vectors (i.e. plurality of training examples) is hashed using locality sensitive hashing, and the closest centroid is identified.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sharma, Yang, MacQueen, and Malewicz. Sharma teaches training a machine learning model on a remote network and then deploying it on a resource constrained edge device. Yang teaches estimating Shapley values for ML model explainability using a model trained on instance-Shapley value pairs. MacQueen teaches using a k -means inference model for prediction based on a collection of samples. Malewicz teaches an implementation of k-means which utilizes locality sensitive hashing. One of ordinary skill would have motivation to combine Sharma, Yang, MacQueen, and Malewicz because “K-means clustering computations are complex, an LSH operation can be added to find data points that are proximate (e.g., close) to one another and thereby simultaneously reduce computational complexity and increase performance” (Malewicz, 0016). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN M ROHD whose telephone number is (571)272-6445. The examiner can normally be reached Mon-Thurs 8:00-6:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.M.R./Examiner, Art Unit 2147 /VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147
Read full office action

Prosecution Timeline

Jan 30, 2024
Application Filed
Jul 13, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
4y 3m (~1y 8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 2 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month