Prosecution Insights
Last updated: October 02, 2026
Application No. 18/059,277

ADJUSTMENT OF TRAINING DATA SETS FOR FAIRNESS-AWARE ARTIFICIAL INTELLIGENCE MODELS

Non-Final OA §101§103§112
Filed
Nov 28, 2022
Examiner
SHALU, ZELALEM W
Art Unit
2145
Tech Center
2100 — Computer Architecture & Software
Assignee
PayPal Inc.
OA Round
3 (Non-Final)
32%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
52%
With Interview

Examiner Intelligence

Grants only 32% of cases
32%
Career Allowance Rate
37 granted / 117 resolved
-23.4% vs TC avg
Strong +20% interview lift
Without
With
+20.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
24 currently pending
Career history
154
Total Applications
across all art units

Statute-Specific Performance

§101
12.6%
-27.4% vs TC avg
§103
66.9%
+26.9% vs TC avg
§102
7.1%
-32.9% vs TC avg
§112
11.3%
-28.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 117 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the Application filed on 08/17/2026. Claims 1-20 are pending in the case. Applicant Response 3. In Applicant’s response dated 08/17/2026, Applicant amended Claims 1, 5-7, 10, 12-17, and 20 and argued against all objections and rejections previously set forth in the Office Action dated 05/15/2026. Continued Examination under 37 CFR 1.114 4. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 08/17/2026 has been entered. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 1 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites “… wherein training the ML model includes reducing a bias in the ML model by including, in the training, a subset of the plurality of data records corresponding to uncertain predictions by the ML model”. The recitation of “uncertain predictions” lacks antecedent bases in the claim. Claim 1 comprises an entropy of predictive output of ML model and model attribution score that is increased based on an increase in the entropy, however, the claim does not introduce what “uncertain predictions” includes. It is unclear whether uncertain predictions refer to previously recited predictive output with increased entropy or some other output in the ML model. The claim fails to recite a threshold, range or other criterion for determining when the predictive output comprise “uncertain predictions”. Therefore, the lack of antecedent bases results in an unclear definition of “uncertain predictions” and makes the scope of the claimed subject indefinite under 35 U.S.C. 112(b). Independent claim 10 and 17 recite similar/same limitations as Claim 1 and are rejected under the same rationale. Dependent claims 2-9, 11-16 and 18-20 are rejected under 35 U.S.C. 112(b) for the same reason as the independent claim they depend. Examiner Comments 5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claim Rejections - 35 USC § 103 6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 7. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Anand (Pub. No. US 20220351055 B1, Pub. Date 2022-11-03) in view of KOBAYASHI (Pub. No. US 20130097107 A1, Pub. Date 2013-04-18.) in further view of Calapodescu (Pub. No. US 0160307113 A1, Pub. Date 2016-10-20.) in further view of Yang (Pub. No. US 20100217732 A1, Pub. Date 2010-08-26.) Anand teaches a system comprising: a non-transitory memory; and one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform operations (see Anand: Fig.17, [0169], “computer 1702, the computer 1702 including a processing unit 1704, a system memory 1706 and a system bus 1708. The system bus 1708 couples’ system components including, but not limited to, the system memory 1706 to the processing unit 1704.”), comprising: receiving training data for a machine learning (ML) model comprising a plurality of data records (see Anand: Fig.16, [0152], “act 1602 can include accessing, by a device (e.g., 114) operatively coupled to a processor, a first set of data candidates (e.g., 106) and a second set of data candidates (e.g., 108), wherein a machine learning model (e.g., 104) is trained on the first set of data candidates.”) determining a plurality of observations in a feature space for the plurality of data records (see Anand: Fig.16, [0154], “act 1606 can include generating, by the device (e.g., 118), a first set of compressed data points (e.g., 502) by applying a dimensionality reduction technique to the first set of latent activations, and generating, by the device (e.g., 118), a second set of compressed data points (e.g., 504) by applying the dimensionality reduction technique to the second set of latent activations.”) wherein the plurality of observations are associated with data points in the feature space for corresponding ones of the plurality of data records (see Anand: Fig.16, [0153], “a first set of latent activations (e.g., 202) generated by the machine learning model based on the first set of data candidates, and obtaining, by the device (e.g., 116), a second set of latent activations (e.g., 204) generated by the machine learning model based on the second set of data candidates”) determining distances between each of the plurality of observations in the feature space based on the data points (see Anand: Fig.14, [0138], “act 1410 can include, for each compressed training data point that belongs to the selected cluster, computing, by the device (e.g., 120), the Euclidean distance between the compressed training data point and the center of the selected cluster. When this is performed for each compressed training data point that belongs to the selected cluster, this is can result in a set of Euclidean distances that are associated with the selected cluster.”) estimating a distribution of the plurality of observations based on the distances (see Anand: Fig.14, [0140], “act 1414 can include computing, by the device (e.g., 120), a standard deviation distance value for the selected cluster, which standard deviation distance value can be denoted as σ, based on the set of Euclidean distances associated with the selected cluster. In various aspects, the computer-implemented method 1400 can proceed back to act 1404.”) calculating, for each of the plurality of data records a diversity score of each of the plurality of data records (see Anand: Fig.16, [0155], “act 1608 can include computing, by the device (e.g., 120), a diversity score (e.g., 702) based on the first set of compressed data points and the second set of compressed data points.”) based on the distances and distribution (see Anand: Fig.14, [0141], “the computer-implemented method 1400 can iterate through acts 1404-1414, until a μ and a σ are computed for each cluster of compressed training data points.”), […] Anand does not teach system wherein: the diversity score is further calculated based, at least in part, on a density estimate of each of the plurality of data records within the training data using a density estimation technique calculating, for each of the plurality of data records, a model attribution score of each of the plurality of data records based on an entropy of a predictive output of the ML model for each of the plurality of data records, wherein the model attribution score is increased based on an increase in the entropy. sampling, based on the diversity scores and the model attribution scores, the plurality of data records from the training data based on the plurality of observations, wherein the sampling includes selecting a portion of the plurality of data records based on the distances between each of corresponding ones of the plurality of observations being greater than a minimum distance between a set of the plurality of observations. generating, based on the sampling, a sampled training data set. training the ML model using the sampled training data set, wherein training the ML model includes reducing a bias in the ML model by including, in the training, a subset of the plurality of data records corresponding to uncertain predictions by the ML model. However, KOBAYASHI teaches the system wherein: the diversity score is further calculated based, at least in part, on a density estimate of each of the plurality of data records within the training data using a density estimation technique (see KOBAYASHI: Fig.32, [0229], “The information processing apparatus 10 first models the density of the feature amount coordinates (S261). To model the density, as one example a density estimating method such as GMM (Gaussian Mixture Model) is used. Next, the information processing apparatus 10 calculates the density of the respective feature amount coordinates based on the constructed model (S262). After this, the information processing apparatus 10 randomly selects, out of the feature amount coordinates that are yet to be selected, the feature amount coordinates with a probability that is proportional to the reciprocal of the density (S263).”). see also Fig.33, [0232], “calculates the density of the feature amount coordinates based on the constructed model (S272). After this, the information processing apparatus 10 sets the reciprocals of the calculated densities as the weightings and ends the series of processes”.) Examiner notes that Anand teaches the claimed diversity score and KOBAYASHI teaches the known technique of making a selection based on density estimate of each of data points. Because both Anand and KOBAYASHI are analogues art because both are directed to in machine learning training systems that involve management of training data, accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the teaching of Anand to include the a system that calculate diversity score based, at least in part, on a density estimate of each of the plurality of data records within the training data using a density estimation technique as taught by KOBAYASHI. One would have been motivated to make such a combination in order to reduce redundant data representation of data and create cleanser and more divers datasets to generate quality machine learning training datasets. Calapodescu teaches the system wherein: calculating, for each of the plurality of data records, a model attribution score of each of the plurality of data records based on an entropy of a predictive output of the ML model for each of the plurality of data records (see Calapodescu: Fig63, [0078], At S300, using the current classifier model 56, the entropy H(d) for all documents remaining in the pool is computed, e.g., by the entropy computation component 36. The entropy H(d) of each document in the pool set relative to the classifier model 56 can be computed.”, examiner notes that entropy H(d) is a model based score associated with outputs of the ML models), wherein the model attribution score is increased based on an increase in the entropy (see Calapodescu: Fig.6, [0084], “At S304, a first document d.sub.1 is drawn from the pool 14, such as the first document in the Q (i.e., the one with the highest entropy). S304 is used to initialize the selected documents in the batch (the first document in the batch is always the one with the highest entropy in the exemplary embodiment). This step is also used as back-off when reaching the threshold, as described for S312.”) sampling, based on the diversity scores and the model attribution scores, the plurality of data records from the training data based on the plurality of observations (see Calapodescu: Fig.1, [0007], “quantifying to what extent a candidate sample is new with respect to the samples already selected in a batch during its construction. In practice, a hybrid criterion is often used, which aggregates the uncertainty value (or expected added value) with the diversity measure. The MMR (Maximum Marginal Relevance) principle is an example of such a hybrid criterion”), wherein the sampling includes selecting a portion of the plurality of data records based on the distances between each of corresponding ones of the plurality of observations being greater than a minimum distance between a set of the plurality of observations (see Calapodescu: Fig.1, [0062], “In selecting the family of hash functions to be used (e.g., based on a training set of objects), the (d.sub.1,d.sub.2,p.sub.1,p.sub.2)—sensitive criteria a) considers only those objects with a high probability of collision (low distance/high similarity between them) and requires selection of a family of hash functions which provide a high probability that these will be assigned to the same bucket, while the (d.sub.1, d.sub.2, p.sub.1,p.sub.2)—sensitive criteria b) considers only those objects with a low probability of collision (high distance/low similarity between them) and requires a family of hash functions which provide a low probability that these will be assigned to the same bucket. Both criteria are met in the family of hash functions which are selected for use in the method.”). generating, based on the sampling, a sampled training data (see Calapodescu: Fig.1, [0106], “The classifier model is then retrained using all (or at least some) of the labeled training objects in the set 70 (S116). Specifically, a classification function is learned which best fits the labels and representations of 50 of the objects in the training set. As will be appreciated, rather than using the same representations that are used for generation of the signatures, another type of multidimensional vectorial representation of the objects can be used.”) Because Anand, KOBAYASHI and Calapodescu are in the same/similar field of endeavor of selecting machine learning training data, accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the teaching of Anand to include the a system that calculate a model attribution(uncertainty) score and diversity score of the plurality of data records to generate a sampled training data set that enables the ML model to be trained as taught by Calapodescu. One would have been motivated to make such a combination in provide smaller, more efficient training datasets and a lag-free and efficient machine learning model training system. (see Calapodescu [0005]) Anand, KOBAYASHI and Calapodescu does not teach the system wherein: training the ML model using the sampled training data set, wherein training the ML model includes reducing a bias in the ML model by including, in the training, a subset of the plurality of data records corresponding to uncertain predictions by the ML model. However, Yang teaches the system wherein: training the ML model using the sampled training data set (see Yang: Fig.2, [0033], “Once operation 208 estimates the sample selection bias, operation 210 trains or learns the classifier or model over the training set L, whose samples are re-weighted by the bias factor.”), wherein training the ML model includes reducing a bias in the ML model by including, in the training, a subset of the plurality of data records ( see Yang: Fig.2, [0028], “named unbiased active learning, is motivated to introduce the sample selection bias so that the distribution difference is treated explicitly. The algorithm iteratively estimates the bias, learns the classifier (or model) f, and then selects samples to label for the next round learning, until the stop condition is met. At each round, the sample selection bias is not only considered as a weight factor in the learning of classifier, but also in the sample selection strategy.”), corresponding to uncertain predictions by the ML model (see Yang: Fig.2, [00037], “select sample elements with the most uncertainty, whose prediction loss will be reduced to zero after labeling. However, it is the risk of the whole unlabeled set which will be reduced by training a better model from the labeled sample (instead of the labeled sample's own expected loss) that should be examined and used in selecting sample elements.”) it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the teaching of Anand to include the system that train ML model using the sampled training data set, wherein training the ML model includes reducing a bias in the ML model by including, in the training, a subset of the plurality of data records corresponding to uncertain predictions by the ML model as taught by Yang. One would have been motivated to make such a combination in provide smaller, more efficient training datasets and a lag-free and efficient machine learning model training system. Regarding Claim 2, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 1. Anand further teaches the system wherein: the distance comprises a vector distance between different ones of the plurality of observations in the feature space from data features utilized by the ML model with the training data (see Anand: Fig.15, [0145], “ act 1506 can include computing, by the device (e.g., 120), the Euclidean distance between the selected compressed test data point and the nearest cluster of compressed training data points (e.g., the cluster whose center is closest and/or nearest in terms of Euclidean distance to the selected compressed test data point).”) Regarding Claim 3, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 2. Anand further teaches the system wherein: the data point comprises a coordinate placement for the corresponding one of the plurality of data records based on feature data associated with the ML model for each of the plurality of data records (see Anand: Fig.9, [0109], “the graph 900 shows a non-limiting example where the set of compressed test data points 504 nicely fit into and/or otherwise conform to the clusters of the set of compressed training data points 502. In other words, the green dots in FIG. 9 are all located very close to the red dot clusters. Accordingly, the diversity score 702 that corresponds to the graph 900 can be small in magnitude. Based on the graph 900, the operator can determine that the machine learning model 104 can be accurately deployed on the set of test data candidates 108, that the machine learning model 104 cannot be trained on the set of test data candidates 108 without overfitting, and/or that automatic annotation can be accurately applied to the set of test data candidates 108.”), and wherein the diversity score is increased when the vector distance between coordinate placements of the different ones of the plurality of data records is increased (see Anand: Fig.15, [0107], “The set of compressed training data points 502 and the set of compressed test data points 504 were then generated as described herein, using t-SNE as the dimensionality reduction technique, and where each compressed data point was a two-element vector. The graphs 900-1100 were then plotted.”) Regarding Claim 4, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 3. Anand further teaches the system wherein: the coordinate placement is determined based on one of a kernel density, a gaussian mixture, or a clustering algorithm (see Anand: Fig.11, [0111], “the set of compressed training data points 502 is illustrated in red, and the set of compressed test data points 504 is illustrated in green. As shown, the set of compressed training data points 502 are grouped into five distinct clusters 902-910, which correspond to the five distinct classes (e.g., Tibia, Femoral, Sagittal Relevant, Coronal Relevant, Irrelevant). As can be easily seen, the graph 1100 shows a non-limiting example where the set of compressed test data points 504 do not nicely fit into and/or otherwise conform to the clusters of the set of compressed training data points 502. In other words, the green dots in FIG. 11 are mostly located away from the red dot clusters.”) Regarding Claim 5, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 1. Anand further teaches the system wherein: calculating the model attribution score is further based on a level of confidence of an accuracy of the predictive outputs associated with each of the plurality of data records by the ML model (see Anand: Fig.4B, [0083], “a decoder may reconstruct the sequence of layers from the last hidden state representation. Upon comparison of the reconstructed layer surface data with a layer surface data of the layer being printed, the encoder-decoder based model may identify a deviation between the reconstructed layer surface data and the layer surface data of the layer being printed. Such a deviation is indicated as a predicted anomaly score.”) Regarding Claim 6, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 1. Anand, KOBAYASHI, Calapodescu and Yang further teaches the system wherein: determining a sampling score for each of the plurality of data records based on the diversity score (see Anand: Fig.16, [0155], “act 1608 can include computing, by the device (e.g., 120), a diversity score (e.g., 702) based on the first set of compressed data points and the second set of compressed data points.”) and the model attribution score (see Calapodescu: Fig63, [0078], At S300, using the current classifier model 56, the entropy H(d) for all documents remaining in the pool is computed, e.g., by the entropy computation component 36. The entropy H(d) of each document in the pool set relative to the classifier model 56 can be computed.”, examiner notes that entropy H(d) is a model based score associated with outputs of the ML models),; and comparing the sampling score to a threshold, wherein the generating the sampled training data set is further based on the comparing (see Anand: Fig.16, [0152], “act 1602 can include accessing, by a device (e.g., 114) operatively coupled to a processor, a first set of data candidates (e.g., 106) and a second set of data candidates (e.g., 108), wherein a machine learning model (e.g., 104) is trained on the first set of data candidates.”) See motivation to combine Anand, KOBAYASHI, Calapodescu and Yang in claim 1. Regarding Claim 7, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 1. Anand further teaches the system wherein: prior to the training the ML model was trained from a previous ML model configuration (see Anand: Fig.12, [0121], “iterate through acts 1206-1214 until every training data candidate has been analyzed (e.g., until a hidden activation map has been inserted into the set of training activation maps for each training data candidate). At this point, the computer-implemented method 1200 can proceed to act 1216.”) Regarding Claim 8, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 1. Erenrich further teaches the system wherein: determining a first weight to apply to the diversity score and a second weight to apply to the model attribution score (see Calapodescu: Fig.2, [0008], “where H(d) is the entropy score derived from the current classifier model estimated probabilities P(c|d), the weight β can be learned on a calibration set, and sim(d.sub.i,d.sub.i) can be the cosine distance calculated on a bag-of-words representation of the documents d.sub.i and d.sub.j. At each iteration, all documents in the pool have their MMR score computed and the document with highest score is added to the batch.”); and applying the first weight to the diversity score and the second weight the model attribution score prior to the sampling (see Calapodescu: Fig.2, [0008], “where H(d) is the entropy score derived from the current classifier model estimated probabilities P(c|d), the weight β can be learned on a calibration set, and sim(d.sub.i,d.sub.i) can be the cosine distance calculated on a bag-of-words representation of the documents d.sub.i and d.sub.j. At each iteration, all documents in the pool have their MMR score computed and the document with highest score is added to the batch.”) It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the teaching of Anand to include a system that determining and applying prior to sampling a first weight to apply to the diversity score and a second weight to apply to the model attribution as taught by Calapodescu. One would have been motivated to make such a combination in provide smaller, more efficient training datasets and a lag-free and efficient machine learning model training system. Regarding Claim 9, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 1. Anand further teaches the system wherein: the sampled training data set enables the ML model to determine one of a policy selection determination, a risk and fraud analysis, or a marketing model (see Calapodescu: Fig.2, [0049], “the trained classifier model 56 may be used, by the classification component 42, to label a new object 60, such as some or all the remaining objects in the pool, or a new object not initially in the pool. At S126, the label is output.). See motivation to combine Anand, KOBAYASHI, Calapodescu and Yang Claim 1 above. Regarding Claim independent 10, Claim 10 is directed to a method claim and has similar/same claim limitation as claim 1 and is rejected under same rationale. Regarding Claim 11, As shown above, Anand, KOBAYASHI, Calapodescu and Yang and teaches all the limitations of claim 10. Anand further teaches the system wherein: determining the diversity scores based on the estimating (see Anand: Fig.16, [0155], “act 1608 can include computing, by the device (e.g., 120), a diversity score (e.g., 702) based on the first set of compressed data points and the second set of compressed data points.”) Regarding Claim 12, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 10. Anand further teaches the system wherein: determining the distances is further based on a vector computation between vector d corresponding to the data points (see Anand: Fig.14, [140], “act 1414 can include computing, by the device (e.g., 120), a standard deviation distance value for the selected cluster, which standard deviation distance value can be denoted as σ, based on the set of Euclidean distances associated with the selected cluster. In various aspects, the computer-implemented method 1400 can proceed back to act 1404.”) Regarding Claim 13, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 10. Anand further teaches the system wherein: calculating a likelihood of one of the plurality of data records to be observed during training of the ML model based on the distribution (see Anand: Fig.14, [0137], “act 1408 can include computing, by the device (e.g., 120), the center of the selected cluster. For example, the center of a given cluster of compressed training data points can be equal to the average of all the compressed training data points that belong to that given cluster.”), wherein the diversity scores are further based on the calculated likelihood (see Anand: Fig.16, [act 1608 can include computing, by the device (e.g., 120), a diversity score (e.g., 702) based on the first set of compressed data points and the second set of compressed data points.”) Regarding Claim 14, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 11. Anand further teaches the system wherein: iterating the calculating of the likelihood over the plurality of data records using the distribution (see Anand: Fig.14, [0139], “act 1412 can include computing, by the device (e.g., 120), an average distance value for the selected cluster, which average distance value can be denoted as μ, based on the set of Euclidean distances associated with the selected cluster.”) Regarding Claim 15, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 11. Erenrich further teaches the system wherein: calculating the ML model for the plurality of data records based on certainties that the plurality of data records are correctly classified by the ML model (see Erenrich: Fig.8, [0079], “The P(c|d) vales may be retrieved by the classification component 42. There may be any number of classes c, such as 2, 3 or more, depending on the type of classifier model. For a binary classifier that is uncertain as to which of two classes to assign to an object, the probability for each class may be about 0.5, resulting in an entropy close to 1. Where the classifier is more certain, the entropy will be less than 1.”) It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the teaching of Anand to include a system that calculating the confidences in the output of the ML model are correctly classified by the ML model as taught by Calapodescu. One would have been motivated to make such a combination in provide smaller, more efficient training datasets and a lag-free and efficient machine learning model training system. Regarding Claim 16, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 11. Erenrich further teaches the system wherein: the calculating the entropy is performed using a previous iteration of the ML model (see Calapodescu: Fig63, [0078], At S300, using the current classifier model 56, the entropy H(d) for all documents remaining in the pool is computed, e.g., by the entropy computation component 36. The entropy H(d) of each document in the pool set relative to the classifier model 56 can be computed.”, examiner notes that entropy H(d) is a model based score associated with outputs of the ML models), Regarding independent Claim 17, Claim 17 is directed to a non-transitory machine-readable medium claim and has similar/same claim limitation as claim 1 and is rejected under same rationale. Regarding Claim 18, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 11. Anand further teaches the system wherein: the ML model is previously trained using a set of sampled data from at least a portion of the plurality of data records (see Anand: Fig.12, [0116], “act 1206 can include determining, by the device (e.g., 116), whether each training data candidate in the set of training data candidates has been analyzed by the device. If not, the computer-implemented method 1200 can proceed to act 1208. If so, the computer-implemented method 1200 can proceed to act 1216.”) Regarding Claim 19, As shown above, Anand and Calapodescu teaches all the limitations of claim 17. Anand further teaches the non-transitory machine-readable medium wherein: the retraining comprises reconfiguring at least one of a weight or a value of one or more nodes of the ML model based on the sampled training data set (see Anand: Fig.13, [0131], “, the computer-implemented method 1300 can iterate through acts 1306-1314 until every test data candidate has been analyzed (e.g., until a hidden activation map has been inserted into the set of test activation maps for each test data candidate). At this point, the computer-implemented method 1300 can proceed to act 1316.”) Regarding Claim 20, As shown above, Anand, KOBAYASHI, Calapodescu and Yang teaches all the limitations of claim 17. Calapodescu further teaches the non-transitory machine-readable medium wherein: the model attribution scores are based on a confidence of a predictive output for each of the plurality of data records by the ML model (see Calapodescu: Fig.2, [0078], “using the current classifier model 56, the entropy H(d) for all documents remaining in the pool is computed, e.g., by the entropy computation component 36. The entropy H(d) of each document in the pool set relative to the classifier model 56 can be computed.”) Because both Anand and Calapodescu are in the same/similar field of endeavor of selecting machine learning training data, accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the teaching of Anand to include model attribution scores are based on a confidence of a predictive output for each of the plurality of data records by the ML model as taught by Calapodescu. One would have been motivated to make such a combination in provide smaller, more efficient training datasets and a lag-free and efficient machine learning model training system. Response to Arguments Claim Rejections - 35 U.S.C. § 101, Regarding the 35 U.S.C. 101 rejection for being directed non-statutory subject matter has been updated and withdrawn based on applicant amendments and. Therefore, the 35 U.S.C. 101 rejection has been withdrawn. Claim Rejections - 35 U.S.C. § 103, Applicant’s arguments with respect to claim amendments have been considered but are moot considering the new combination of references being used in the current rejection. The new combination of references was necessitated by Applicant’s claim amendments. Therefore, the claims are rejected under the new combination of references as indicated above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. PGPUB NUMBER: INVENTOR-INFORMATION: TITLE / DESCRIPTION US 20200372406 A1 Wick; Michael Louis Title: Enforcing Fairness On Unlabeled Data To Improve Modeling Performance Description: his disclosure relates to generating data sets to improve classification performance in machine learning. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZELALEM W SHALU whose telephone number is (571)272-3003. The examiner can normally be reached M- F 0800am- 0500pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Zelalem Shalu/Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Show 5 earlier events
Jan 24, 2026
Examiner Interview Summary
May 15, 2026
Final Rejection mailed — §101, §103, §112
Jun 25, 2026
Interview Requested
Jul 21, 2026
Examiner Interview Summary
Jul 21, 2026
Applicant Interview (Telephonic)
Aug 17, 2026
Request for Continued Examination
Aug 18, 2026
Response after Non-Final Action
Sep 24, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737103
GENERATING INTERACTIVE, DIGITAL DATA NARRATIVE ANIMATIONS BY DYNAMICALLY ANALYZING UNDERLYING LINKED DATASETS
7y 10m to grant Granted Sep 15, 2026
Patent 12670381
FINDING K EXTREME VALUES IN CONSTANT PROCESSING TIME
5y 4m to grant Granted Jun 30, 2026
Patent 12619879
TRAINING NEURAL NETWORKS USING LEARNED OPTIMIZERS
4y 7m to grant Granted May 05, 2026
Patent 12477016
AUTOMATION OF VISUAL INDICATORS FOR DISTINGUISHING ACTIVE SPEAKERS OF USERS DISPLAYED AS THREE-DIMENSIONAL REPRESENTATIONS
3y 5m to grant Granted Nov 18, 2025
Patent 12468969
METHODS FOR CORRELATED HISTOGRAM CLUSTERING FOR MACHINE LEARNING
3y 4m to grant Granted Nov 11, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
32%
Grant Probability
52%
With Interview (+20.4%)
3y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 117 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month