Prosecution Insights
Last updated: October 01, 2026
Application No. 18/624,245

AUTOMATED TRAINING DATASET MODIFICATIONS TO BALANCE DATA VARIATION

Non-Final OA §101§102§103
Filed
Apr 02, 2024
Examiner
ASEGDEW, NATNAEL AREGA
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
13 currently pending
Career history
9
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the instant application filed on 04/02/2024. Claims 1-20 are pending. Information Disclosure Statement The information disclosure statement (IDS) submitted on 04/02/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1: The claim recites a method which falls into the statutory category of process. Step 2A Prong 1: The claim recites multiple abstract ideas: detecting that a confidence score associated with a machine learning prediction is below a configured threshold, wherein the machine learning prediction is based on input data applied to a machine learning model; inserting, in response to detecting that the confidence score is below the configured threshold, a data point for the input data in a review dataset; detecting, based on a pattern of data point attributes, a cluster of data points among a plurality of data points in the review dataset; and modifying, based on identifying that the cluster includes a threshold number of data points, a training dataset for the machine learning model. Given a human being can reasonably detect whether the confidence score is low and then add points to a dataset all within the mind or with the aid of a generic computer. Further, detecting groups of data points and altering datasets are all things that a human being can do in the mind. Step 2A Prong 2: Claim 1 does not integrate the abstract idea into a practical application since the additional element of a machine learning model merely links the abstract idea to a technological environment. Step 2B: Claim 1 does not integrate the abstract idea into a practical application since the machine learning model falls under generally linking the abstract idea to a technological environment (2106.05(h)). Claim 1 is not patent eligible. Regarding claim 2, the rejection of claim 1 is incorporated, further the claim recites: wherein the machine learning prediction is a classification of the input data. This limitation amounts to generally linking the abstract idea to a field of use: classification (MPEP 2106.05(h)). Claim 2 is not patent eligible. Regarding claim 3, the rejection of claim 1 is incorporated, further the claim recites: wherein the confidence score is a function of a probability associated with the machine learning prediction and a minimum probability threshold. This limitation amounts to more specifics of the abstract idea of detecting that a confidence score is below a threshold given it merely describes the confidence score. Claim 3 is not patent eligible. Regarding claim 4, the rejection of claim 1 is incorporated, further the claim recites: wherein the cluster of data points indicates a degree of similarity in the input data associated with each data point. This limitation amounts to more specifics of the abstract idea of detecting a cluster given it merely describes the relationship between the clusters and the input data. Claim 4 is not patent eligible. Regarding claim 5, the rejection of claim 1 is incorporated, further the claim recites: retraining the machine learning model based on the modified training dataset. This limitation amounts to mere instructions to apply the abstract idea by a generic computer (machine learning model) (MPEP 2106.05(f)). Claim 5 is not patent eligible. Regarding claim 6, the rejection of claim 1 is incorporated, further the claim recites: wherein modifying, based on identifying that the cluster includes a threshold number of data points, a training dataset for the machine learning model includes: adding data samples, based on input data associated with the cluster of data points, to the training dataset for the machine learning model. This limitation amounts to more specifics of the abstract idea of modifying the dataset given it merely describes how to modify the dataset while it remains a mental process. Claim 6 is not patent eligible. Regarding claim 7, the rejection of claim 1 is incorporated, further the claim recites: wherein modifying, based on identifying that the cluster includes a threshold number of data points, a training dataset for the machine learning model includes: mapping the cluster of data points to one or more samples in the training dataset; and increasing a weight of the one or more samples in the training dataset. This limitation amounts to more specifics of the abstract idea of modifying a dataset given it merely describes how by mapping clusters to samples and increasing weights, which are both mental processes. Claim 7 is not patent eligible. Regarding claim 8, the rejection of claim 7 is incorporated, further the claim recites: training a second learning model based on a reduced dataset that includes the one or more Samples, which amounts to generally linking the abstract idea to a technological environment; providing input data associated with the cluster of data points to the second learning model, is considered well-understood, conventional, routine activity (MPEP 2106.05(d)(II)(I)), wherein increasing the weight of the one or more samples in the training dataset includes: increasing the weight when the second learning model meets a performance goal, a mental process given a human being can increase a weight in their mind or with the aid of a generic computer. Claim 8 is not patent eligible. Regarding claim 9, the rejection of claim 8 is incorporated, further the claim recites: wherein the input data associated with the cluster of data points is added to the training dataset when the second learning model does not meet the performance goal. This limitation amounts to a mental process given a human being can look at the performance of a model and decide whether or not to add data to the dataset. Claim 9 is not patent eligible. Regarding claims 10-20, the inventive concept is essentially the same as claims 1-9, as such the rejections above are incorporated. Claims 10-20 are not patent eligible. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-6 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Mousavi (COLLABORATIVE LEARNING OF SEMI-SUPERVISED CLUSTERING AND CLASSIFICATION FOR LABELING UNCURATED DATA). Regarding claim 1, Mousavi teaches detecting that a confidence score associated with a machine learning prediction is below a configured threshold, wherein the machine learning prediction is based on input data applied to a machine learning model (Fig 1, The iteration repeats for images with low confidence predictions. High confidence predictions are fed back to the classifier to enable self-learning); inserting, in response to detecting that the confidence score is below the configured threshold, a data point for the input data in a review dataset (Fig 1, The iteration repeats for images with low confidence prediction, review set is the set of images that are reiterated through); detecting, based on a pattern of data point attributes, a cluster of data points among a plurality of data points in the review dataset (Fig 1, Unlabeled data are mapped to numerical feature embeddings and then clustered together, Pg. 3, High-confidence predictions are used to expand the training data using the predicted labels and low confidence predictions are assigned for clustering and manual labeling, low confidence predictions are assigned to a review dataset and then clustered); and modifying, based on identifying that the cluster includes a threshold number of data points, a training dataset for the machine learning model (Pg.3, High-confidence predictions are used to expand the training data using the predicted labels and low confidence predictions are assigned for clustering and manual labeling. The threshold for considering high or low confidence values is determined by manually exploring the predictions, Fig 1, The iteration repeats for images with low confidence predictions, the training data set is expanded if the cluster has even 1 data point, which is the threshold). Regarding claim 2, Mousavi teaches wherein the machine learning prediction is a classification of the input data (Abs, Plud is an iterative sequence of unsupervised clustering, human assistance, and supervised classification). Regarding claim 3, Mousavi teaches wherein the confidence score is a function of a probability associated with the machine learning prediction and a minimum probability threshold (Section 2.3, the predictions are ordered by the confidence level the classifier assigns to each prediction. High confidence predictions are used to expand the training data using the predicted labels, implies probability given high confidence means more likely to be right but not guaranteed and a minimum threshold which is the line between high and low confidence). Regarding claim 4, Mousavi teaches wherein the cluster of data points indicates a degree of similarity in the input data associated with each data point (Section 1, Using unsupervised methods one can cluster image data using their numerical embeddings in groups that share similar characteristics). Regarding claim 5, Mousavi teaches retraining the machine learning model based on the modified training dataset (Fig 1, The iteration repeats for images with low confidence predictions, retraining is any iteration after the first). Regarding claim 6, Mousavi teaches wherein modifying, based on identifying that the cluster includes a threshold number of data points, a training dataset for the machine learning model includes: adding data samples, based on input data associated with the cluster of data points, to the training dataset for the machine learning model (Section 2.3, Adding images with high confidence predicted labels to the classifier’s training data improve its accuracy on those type of images, the added data samples are input data samples that are not in the cluster of data points (since they have high confidence values)). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 7-20 are rejected under 35 U.S.C. 103 as being unpatentable over Mousavi as applied to claims 1-6 above, and further in view of Shu (CMW-Net: Learning a Class-Aware Sample Weighting Mapping for Robust Deep Learning). Regarding claim 7, Mousavi teaches mapping the cluster of data points to one or more samples in the training dataset (Fig 1, Unlabeled data are mapped to numerical feature embeddings and then clustered together, mapping between data samples and clusters). Mousavi fails to teach increasing a weight of the one or more samples in the training dataset. Shu teaches increasing a weight of the one or more samples in the training dataset (Fig 2, data is broken up into clusters, Fig. 3, some clusters have their weights increased and others decreased). Mousavi and Shu are analogous to the claimed invention because they are in the field of solving classification problems using clustering. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the reweighting system described in Shu along with the existing retraining system in Mousavi to “achieve proper weighting schemes in various data bias cases, like class imbalance, feature-independent and dependent label noises, and more complicated bias scenarios beyond conventional cases” (Shu Abs). Regarding claim 8, Mousavi in view of Shu teaches the method of claim 7 further, Shu teaches, which Mousavi fails to teach, training a second learning model based on a reduced dataset that includes the one or more samples (Alg 1, the second learning model is the CMW-Net which is trained on batches of data/reduced data); and providing input data associated with the cluster of data points to the second learning model (Fig 2, clustering done to data), wherein increasing the weight of the one or more samples in the training dataset includes: increasing the weight when the second learning model meets a performance goal (Section 3C, Now, the objective function of CWM-Net can be written as the following bi-level optimization problem…., Section 2, There are mainly two manners to design such weighting function. One is to make it monotonically increasing, which is specifically effective in class imbalance case, Alg 1, weights are increased to optimize loss (in the monotonically increasing case), therefore the performance goal is lower loss which is done through the updates that follow gradient descent). Regarding claim 9, Mousavi in view of Shu teaches the method of claim 8, further Mousavi teaches wherein the input data associated with the cluster of data points is added to the training dataset when the second learning model does not meet the performance goal (Fig 1, data associated with the clusters is added back into the training set, and given the second model in Shu is used before training the classifier, then the method of Mousavi in view of Shu would involve adding data to training set after the second model is used, whether the performance goal is met or not, since the broadest reasonable interpretation of “when” does not necessarily imply strict causality). Regarding claims 10-20, the inventive concept is essentially the same as claims 1-9 with the addition of a computing device and a computer program product, which are implied by Shu (Code for reproducing our experiments is available at https://github.com/xjtushujun/CMW-Net). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NATNAEL A ASEGDEW whose telephone number is (571)270-0407. The examiner can normally be reached 7:30-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NATNAEL A ASEGDEW/Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Apr 02, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month