DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8-12 and 15-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Desmond et al. (US 20210174196 A1) in view of Chalamalasetti et al. (US 20220121885 A1).
Regarding claim 1.
Desmond teaches a system for mitigating biases during training of a machine learning model, comprising: a memory configured to store a machine learning model and a training dataset, wherein the training dataset comprises a set of datapoints; and a processor, operably coupled to the memory (see ¶ 33-35, “processing system 100 includes the processing device 102, the memory 104, a model generation engine 106, a vector representation generation engine 108, a vector representation clustering engine 110 and a ground truth analysis engine 112. According to some embodiments, processing system 100 can be any suitable computing device or collection of computing devices that are sufficient to perform the functionalities described herein. The processing system 100 can be configured to communicate with a user device 120, which can display notifications and/or data to and receive user inputs from a user 121.”), and configured to:
train the machine learning model using the training dataset, wherein training the machine learning model using the training dataset (see ¶ 38, “block 204, the method includes training (e.g., via model generation engine 106) a model based on the plurality of data inputs…the model generation engine 106 can build and train one or more predictive models based on a plurality of labeled data inputs (i.e., ground truth data), such as for example, the cat/dog images shown in FIG. 3.”, also see ¶ 39) comprises:
inputting a first datapoint from among the set of datapoints to the machine learning model (see ¶ 42-43, “The input layer 410 is made up of a plurality of inputs 412, 414, 416, the hidden layer(s) are made up of a plurality of hidden layer neurons 422, 424, 426 and 428, and the output layer is made up of a plurality of output neurons 432, 434 (which may be referred to as the “final layer” or “output layer” of the neural network)… as data is input to the neural network via the input layer 410, the data propagated along the paths shown by multiplying the value of the data by the dot product of the weight of the path and then adding the bias of the destination neuron and then passed through an activation function to convert the input signal to an output signal. As will be understood by those of skill in the art, weights can be applied to the inputs and then the activation function can be applied over an aggregate of the weighted inputs. The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect… each data input of the ground truth data can be received by the input layer 410 (e.g., each pixel of an image is received as an input value to an input node 412, 414, 416, etc.) and the values are propagated through the paths of the neural network 400 by applying the activation functions, weights and biases and updating the weights and biases as described above to train the model. ”);
receiving a first output from the machine learning model, wherein the first output is a prediction of the machine learning model with respect to a first label associated with the first datapoint (see ¶ 42-43, “The input layer 410 is made up of a plurality of inputs 412, 414, 416, the hidden layer(s) are made up of a plurality of hidden layer neurons 422, 424, 426 and 428, and the output layer is made up of a plurality of output neurons 432, 434 (which may be referred to as the “final layer” or “output layer” of the neural network)… as data is input to the neural network via the input layer 410, the data propagated along the paths shown by multiplying the value of the data by the dot product of the weight of the path and then adding the bias of the destination neuron and then passed through an activation function to convert the input signal to an output signal. As will be understood by those of skill in the art, weights can be applied to the inputs and then the activation function can be applied over an aggregate of the weighted inputs. The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. … each data input of the ground truth data can be received by the input layer 410 (e.g., each pixel of an image is received as an input value to an input node 412, 414, 416, etc.) and the values are propagated through the paths of the neural network 400 by applying the activation functions, weights and biases and updating the weights and biases as described above to train the model… if the model is intended to identify images of cats and dogs, following training of the model with some portion of the ground truth data, new images can be input into the neural network 400 and the neural network can output an identification of the image (e.g., either “cat” or “dog”) via the output layer 430 (which may also be referred to as the softmax output layer). According to some embodiments, the values of the final hidden layer 430 (i.e., if there are multiple connected hidden layers 420, the final hidden layer 430 is the layer connected to the output layer 430) can be considered to be an n-dimensional vector representation a given data input (e.g., the image of a cat).”);
inputting a second datapoint from among the set of datapoints to the machine learning model (see ¶ 37-43, “The input layer 410 is made up of a plurality of inputs 412, 414, 416, the hidden layer(s) are made up of a plurality of hidden layer neurons 422, 424, 426 and 428, and the output layer is made up of a plurality of output neurons 432, 434 (which may be referred to as the “final layer” or “output layer” of the neural network)… as data is input to the neural network via the input layer 410, the data propagated along the paths shown by multiplying the value of the data by the dot product of the weight of the path and then adding the bias of the destination neuron and then passed through an activation function to convert the input signal to an output signal. As will be understood by those of skill in the art, weights can be applied to the inputs and then the activation function can be applied over an aggregate of the weighted inputs. The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. … each data input of the ground truth data can be received by the input layer 410 (e.g., each pixel of an image is received as an input value to an input node 412, 414, 416, etc.) and the values are propagated through the paths of the neural network 400 by applying the activation functions, weights and biases and updating the weights and biases as described above to train the model… if the model is intended to identify images of cats and dogs, following training of the model with some portion of the ground truth data, new images can be input into the neural network 400 and the neural network can output an identification of the image (e.g., either “cat” or “dog”) via the output layer 430 (which may also be referred to as the softmax output layer). According to some embodiments, the values of the final hidden layer 430 (i.e., if there are multiple connected hidden layers 420, the final hidden layer 430 is the layer connected to the output layer 430) can be considered to be an n-dimensional vector representation a given data input (e.g., the image of a cat).”);
and receiving a second output from the machine learning model, wherein the second output is a prediction of the machine learning model with respect to the first label associated with the second datapoint (see ¶ 37-43, “The input layer 410 is made up of a plurality of inputs 412, 414, 416, the hidden layer(s) are made up of a plurality of hidden layer neurons 422, 424, 426 and 428, and the output layer is made up of a plurality of output neurons 432, 434 (which may be referred to as the “final layer” or “output layer” of the neural network)… as data is input to the neural network via the input layer 410, the data propagated along the paths shown by multiplying the value of the data by the dot product of the weight of the path and then adding the bias of the destination neuron and then passed through an activation function to convert the input signal to an output signal. As will be understood by those of skill in the art, weights can be applied to the inputs and then the activation function can be applied over an aggregate of the weighted inputs. The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. … each data input of the ground truth data can be received by the input layer 410 (e.g., each pixel of an image is received as an input value to an input node 412, 414, 416, etc.) and the values are propagated through the paths of the neural network 400 by applying the activation functions, weights and biases and updating the weights and biases as described above to train the model… if the model is intended to identify images of cats and dogs, following training of the model with some portion of the ground truth data, new images can be input into the neural network 400 and the neural network can output an identification of the image (e.g., either “cat” or “dog”) via the output layer 430 (which may also be referred to as the softmax output layer). According to some embodiments, the values of the final hidden layer 430 (i.e., if there are multiple connected hidden layers 420, the final hidden layer 430 is the layer connected to the output layer 430) can be considered to be an n-dimensional vector representation a given data input (e.g., the image of a cat).”, i.e. teaches multiple labeled data inputs being processed by the NN and classifier);
see ¶ 42, “The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. Following this forward propagation through the neural network, the system performs a backward propagation to update the weight parameters of the paths and the biases of the neurons. These steps can be repeated to train the model by updating the weights and biases until a cost value is met or a predefined number of iterations are run.”);
in response to determining that the machine learning model is biased, update the machine learning model by updating one or more parameters of a neural network associated with the machine learning model, wherein the one or more parameters comprise a weight value or a bias value (see ¶ 42, “The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. Following this forward propagation through the neural network, the system performs a backward propagation to update the weight parameters of the paths and the biases of the neurons. These steps can be repeated to train the model by updating the weights and biases until a cost value is met or a predefined number of iterations are run.”);
and output the updated machine learning model (see ¶ 60, “the method 200 can further include forming a new plurality of data inputs by removing the at least one anomalous data input from the plurality of data inputs and automatically retraining the model based on the new plurality of data inputs. In some embodiments, the method can include utilizing the retrained model to classify one or more test data inputs. For example, after being retrained, the processing system 100 can receive new test data inputs (e.g., new images) and can utilize the retrained model to classify the new test data inputs. In this way, the method 200 can be used to perform quality control on the ground truth data and generate an improved model having better accuracy by modifying or removing training data as appropriate.” also see ¶ 67-68, “block 612, the method includes forming (e.g., via processing system 100) a new plurality of data inputs having associated labels by, responsive to identifying an outlier data input, removing the outlier data input from the plurality of data inputs and responsive to identifying a mislabeled data input, relabeling the data input to have an associated label of a classification type of the predominant classification type of the vector representations in the same cluster as the vector representation corresponding to the mislabeled data input. As shown at block 614, the method includes automatically retraining (e.g., via processing system 100) the model based on the new plurality of data inputs. In this way, the processing system 100 can automatically perform a quality control process on the ground truth data and generate a new, more accurate model.”).
Desmond do no teach compare the first output with the second output; determine that the first output does not correspond with the second output; determine that the machine learning model is biased in response to determining that the first output does not correspond with the second output.
Chalamalasetti teaches compare the first output with the second output; determine that the first output does not correspond with the second output; determine that the machine learning model is biased in response to determining that the first output does not correspond with the second output (see ¶ 19, 23, “The bias detection functions performed by the bias detection engine 202 may further include comparing the bias test output of an ML model for the bias test data to expected bias test results to determine whether bias is present in the ML model, and if so, the level of bias present, and comparing the bias level to a threshold bias level to determine whether a bias alert should be generated, and optionally, whether the ML model should be re-trained to mitigate the bias present in the model.”, also see ¶ 36, “block 412, instructions of the bias detection engine 202 may be executed by the hardware processors 402 to compare the bias test results 112 outputted by the ML model 104 in response to the inputted bias test data 110 to expected bias test results for the bias test data 110. In example embodiments, comparing the bias test results 112 to the expected bias test results for the bias test data 110 may include determining a set of statistical metrics based on a deviation between the actual bias test results 112 and the expected bias test results.”).
Both Desmond and Chalamalasetti pertain to the problem of improving machine learning training, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Desmond and Chalamalasetti to teach the above limitations. The motivation for doing so would be “machine learning (ML) model in a manner that is independent of the code/weights deployment path is described. If bias is detected, an alert for bias is generated, and optionally, the ML model can be incrementally re-trained to mitigate the detected bias. Re-training the ML model to mitigate the bias may include enforcing a bias cost function to maintain a level of bias in the ML model below a threshold bias level. One or more statistical metrics representing the level of bias present in the ML model may be determined and compared against one or more threshold values. If one or more metrics exceed corresponding threshold value(s), the level of bias in the ML model may be deemed to exceed a threshold level of bias, and re-training of the ML model to mitigate the bias may be initiated.” (see Chalamalasetti abstract).
Regarding claim 2.
Desmond and Chalamalasetti teach the system of Claim 1,
Desmond further teach wherein determining that the machine learning model is biased is further in response to: comparing the first output with a first expected output, wherein the first expected output is associated with the first label; and determining that the first output does not correspond with the first expected output (see ¶ 42, “The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. Following this forward propagation through the neural network, the system performs a backward propagation to update the weight parameters of the paths and the biases of the neurons. These steps can be repeated to train the model by updating the weights and biases until a cost value is met or a predefined number of iterations are run.”).
Regarding claim 3.
Desmond and Chalamalasetti teach the system of Claim 1,
Desmond further teach wherein the processor is further configured to: input the set of datapoints from the training dataset to the machine learning model; receive a set of outputs from the machine learning model, wherein each of the set of outputs is a prediction of the machine learning model with respect to a label of a respective datapoint from among the set of datapoints; compare each output from among the set of outputs with a counterpart expected output (see ¶ 37-43, “The input layer 410 is made up of a plurality of inputs 412, 414, 416, the hidden layer(s) are made up of a plurality of hidden layer neurons 422, 424, 426 and 428, and the output layer is made up of a plurality of output neurons 432, 434 (which may be referred to as the “final layer” or “output layer” of the neural network)… as data is input to the neural network via the input layer 410, the data propagated along the paths shown by multiplying the value of the data by the dot product of the weight of the path and then adding the bias of the destination neuron and then passed through an activation function to convert the input signal to an output signal. As will be understood by those of skill in the art, weights can be applied to the inputs and then the activation function can be applied over an aggregate of the weighted inputs. The output layer 430 provides a classification of the data input that can be compared to the associated label of the data input to determine if the classification was correct or incorrect. … each data input of the ground truth data can be received by the input layer 410 (e.g., each pixel of an image is received as an input value to an input node 412, 414, 416, etc.) and the values are propagated through the paths of the neural network 400 by applying the activation functions, weights and biases and updating the weights and biases as described above to train the model… if the model is intended to identify images of cats and dogs, following training of the model with some portion of the ground truth data, new images can be input into the neural network 400 and the neural network can output an identification of the image (e.g., either “cat” or “dog”) via the output layer 430 (which may also be referred to as the softmax output layer). According to some embodiments, the values of the final hidden layer 430 (i.e., if there are multiple connected hidden layers 420, the final hidden layer 430 is the layer connected to the output layer 430) can be considered to be an n-dimensional vector representation a given data input (e.g., the image of a cat).”); and determine that more than a threshold number of outputs do not correspond to counterpart expected outputs (see ¶ 54-59, “based on a partitioned vector space, the ground truth analysis engine 112 can examine the partitions for homogeneity and flag clusters that do not have a threshold level of homogeneity. A threshold of homogeneity can be based on the number of vector representations in a cluster. For example, a cluster with a small number (e.g., 1-100) of vector representations can have a lower threshold than a cluster with a larger number (e.g., more than 1000) of vector representations. Although it should be understood that the threshold of homogeneity can be any percentage that is specified by a user, some examples can be 80%, 85%, 90%, 90%, 95% and 98%. Thus, for example, a cluster with only 5 vector representations can have a threshold of homogeneity of 80%, as a single vector representation with a different label can be more indicative of a mislabeled data input than an ambiguous class structure.”).
Regarding claim 4.
Desmond and Chalamalasetti teach the system of Claim 3,
Chalamalasetti further teach wherein determining that the machine learning model is biased is further in response to determining that more than the threshold number of outputs do not correspond to the counterpart expected outputs (see ¶ 36-43, “block 412, instructions of the bias detection engine 202 may be executed by the hardware processors 402 to compare the bias test results 112 outputted by the ML model 104 in response to the inputted bias test data 110 to expected bias test results for the bias test data 110. In example embodiments, comparing the bias test results 112 to the expected bias test results for the bias test data 110 may include determining a set of statistical metrics based on a deviation between the actual bias test results 112 and the expected bias test results.”).
The motivation utilized in the combination of claim 1, super, applies equally as well to claim 4.
Regarding claim 5.
Desmond and Chalamalasetti teach the system of Claim 1,
Desmond further teach wherein the processor is further configured to determine that the training dataset is biased, wherein determining that the training dataset is biased comprises determining that a third datapoint is incompatible with the machine learning model (see ¶ 60, “the method 200 can further include forming a new plurality of data inputs by removing the at least one anomalous data input from the plurality of data inputs and automatically retraining the model based on the new plurality of data inputs. In some embodiments, the method can include utilizing the retrained model to classify one or more test data inputs. For example, after being retrained, the processing system 100 can receive new test data inputs (e.g., new images) and can utilize the retrained model to classify the new test data inputs. In this way, the method 200 can be used to perform quality control on the ground truth data and generate an improved model having better accuracy by modifying or removing training data as appropriate.” also see ¶ 67-68, “block 612, the method includes forming (e.g., via processing system 100) a new plurality of data inputs having associated labels by, responsive to identifying an outlier data input, removing the outlier data input from the plurality of data inputs and responsive to identifying a mislabeled data input, relabeling the data input to have an associated label of a classification type of the predominant classification type of the vector representations in the same cluster as the vector representation corresponding to the mislabeled data input. As shown at block 614, the method includes automatically retraining (e.g., via processing system 100) the model based on the new plurality of data inputs. In this way, the processing system 100 can automatically perform a quality control process on the ground truth data and generate a new, more accurate model.”).
Claims 8-12 recites a method to perform the system recited in claims 1-5. Therefore, the rejection of claims 1-5 above applies equally here.
Claims 15-19 recites a non-transitory computer-readable medium to perform the system recited in claims 1-5. Therefore, the rejection of claims 1-5 above applies equally here.
Allowable Subject Matter
Claims 6-7, 13-14 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Related prior arts:
Sun et al. (“Prediction Consistency Regularization for Learning with Noise Labels Based on Contrastive Clustering”, 2024, 26, 308. https://doi.org/ 10.3390/e26040308) teaches prediction consistency regularization that mitigates the impact of label noise on neural networks by imposing constraints on the prediction consistency of similar samples. However, determining which samples should be similar is a primary challenge. We formalize the similar sample identification as a clustering problem and employ twin contrastive clustering (TCC) to address this issue. To ensure similarity between samples within each cluster, we enhance TCC by adjusting clustering prior to distribution using label information. Based on the adjusted TCC’s clustering results, we first construct the prototype for each cluster and then formulate a prototype-based regularization term to enhance prediction consistency for the prototype within each cluster and counteract the adverse effects of label noise.
Gulamali et al. (“An AI-Guided Data Centric Strategy to Detect and Mitigate Biases in Healthcare Datasets”, 6 Nov 2023, arXiv:2311.03425) teaches adoption of diagnosis and prognostic algorithms in healthcare has led to concerns about the perpetuation of bias against disadvantaged groups of individuals. Deep learning methods to detect and mitigate bias have revolved around modifying models, optimization strategies, and threshold calibration with varying levels of success. Here, we generate a data-centric, model-agnostic, task-agnostic approach to evaluate dataset bias by investigating the relationship between how easily different groups are learned at small sample sizes (AEquity)
Cox et al. (US 20210263949 A1) teaches computerized pipelines can transform input data into data structures compatible with models in some examples. In one such example, a system can obtain a first table that includes first data referencing a set of subjects. The system can then execute a sequence of processing operations on the first data in a particular order defined by a data-processing pipeline to modify an analysis table to include features associated with the set of subjects. Executing each respective processing operation in the sequence to generate the modified analysis table may involve: deriving a respective set of features from the first data by executing a respective feature-extraction operation on the first data; and adding the respective set of features to the analysis table. The system may then execute a predictive model on the modified analysis table for generating a predicted value based on the modified analysis table.
Basu et al. (US 20220300557 A1) teaches updating a classifier model for a set of labels includes accessing a set of datapoints that includes a first datapoint and a second datapoint. A ground-truth label may be assigned to each datapoint of the set of datapoints. The classifier model is enabled to predict a predicted label for each datapoint. The ground-truth label and the predicted label for each datapoint is included in the set of labels. A first data structure that encodes an error matrix (e.g., a confusion matrix) may be generated based on a comparison between the ground-truth label and the predicted label for each datapoint. Components of the error matrix may indicate instances of correct and incorrect label predictions of the classifier model. For each label of the set of labels, a credibility interval may be determined.
DURAND et al. (US 20200302340 A1) teaches training machine learning models for predicting data labels presents challenging technical problems because predicting accurate data labels requires incorporating user specific bias and subjectivity in assigning data labels to data objects, which are incorporated into training data sets. Training a machine learning model with subjective data labels is difficult because model accuracy may be reliant upon the interrelationships not only between the user and the data label or data object, for instance a user's likelihood to assign certain data labels to certain data objects, but also the correlations between the data labels and the data objects themselves.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IMAD M KASSIM whose telephone number is (571)272-2958. The examiner can normally be reached 10:30AM-5:30PM, M-F (E.S.T.).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297 - 4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IMAD KASSIM/Primary Examiner, Art Unit 2129