DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3-7, 9-11, 13-17, 19-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1:
Claims 1-10 are directed to a system, and claims 11-20 are directed to a method. Therefore, each of these claims is directed to one of the four statutory categories of patent eligible subject matter.
Step 2A Prong 1:
Claims 1 and 11 recite:
“determining a training dataset for a first machine learning module based on historical change data by”; determining a dataset is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“detecting and correcting correctable label noises in the historical change data based at least in part on”; detecting and correcting label noises is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“[training a second machine learning module based on at least a sub-dataset of the historical change data in a warm-up stage to] determine a respective predicted label associated with each data point of the historical change data”; determining a predicted label is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“after determining[, by the second machine learning module,] the respective predicated label associated with each data point of the historical change data, determining whether a label noise associated with a candidate data point of the historical change data is correctable based on (a) the candidate data point being in a confident data range of the historical change data and (b) a prediction confidence for the respective predicted label associated with the candidate data point at least as great as a confidence threshold”; determining whether a noise is correctable based on a confident data range and confidence threshold is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“upon determining that the label noise associated with the candidate data point of the historical change data is correctable, correcting the label noise by replacing a current label of the candidate data point with the respective predicted label associated with the candidate data point”; relabeling data based on a determination of correctability is an evaluation that can be carried out by a human with pen and paper, and is thus a mental process
“after detecting and correcting the correctable label noises, determining the training dataset based on a label-error-free portion of the historical change data”; determining an error free portion of data is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“[training the first machine learning module based on the training dataset to] determine a respective risk score associated with a respective change request”; determining a risk score is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“determining[, via the first machine learning module, as trained,] a first risk score associated with a first change request”; determining a risk score is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
“determining a change approval based on the first risk score and a risk threshold”; determining a decision based on a risk score is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
Step 2A Prong 2:
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“one or more processors; and one or more non-transitory computer-readable media storing computing instructions configured to, when run on the one or more processors, cause the one or more processors to perform”; “being implemented via execution of computing instructions configured to run at one or more processors and stored at one or more non-transitory computer-readable media”; these limitations amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f)
“training a second machine learning module based on at least a sub-dataset of the historical change data in a warm-up stage to”; “training the first machine learning module based on the training dataset to”; these limitations, which broadly recite training and execution of machine learning models at a high level of generality, wherein the models are trained to perform abstract ideas such as determining drift and risk scores, amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f)
“transmitting the change approval request”; this limitation amounts to insignificant extra solution activity, mere data gathering and outputting, as per MPEP 2106.05(g)
“in response to the change approval automatically implementing the change request” these limitations, which broadly recite applying a change or simply performing generic operation using a computer without any specific detail step of the performed function which amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f)
Step 2B:
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“one or more processors; and one or more non-transitory computer-readable media storing computing instructions configured to, when run on the one or more processors, cause the one or more processors to perform”; “being implemented via execution of computing instructions configured to run at one or more processors and stored at one or more non-transitory computer-readable media”; these limitations amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f)
“training a second machine learning module based on at least a sub-dataset of the historical change data in a warm-up stage to”; “training the first machine learning module based on the training dataset to”; these limitations, which broadly recite training and execution of machine learning models at a high level of generality, wherein the models are trained to perform abstract ideas such as determining drift and risk scores, amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f)
“transmitting the change approval to cause an implementation of the first change request”; this limitation amounts to insignificant extra solution activity, mere data gathering and outputting, as per MPEP 2106.05(g); furthermore, this amounts to well-understood, routine, and conventional activity as per MPEP 2106.05(d) (“i. Receiving or transmitting data over a network”)
“in response to the change approval automatically implementing the change request” these limitations, which broadly recite applying a change or simply performing generic operation using a computer without any specific detail step of the performed function which amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f)
Dependent Claims
Claims 3-7, 9-10, 13-17, and 19-22 are also rejected under 35 USC 101 for the following reasons:
Claims 3 and 13 recite: “wherein: detecting and correcting the correctable label noises in the historical change data is further based on determining the confident data range of the historical change data based at least in part on a distribution of prediction confidence for predicted labels that are associated with the historical change data”; determining a confident data range based on a distribution is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process under Step 2A Prong 1.
determined by the second machine learning module, as trained in the warm-up stage”; the use of a broadly recited machine learning model at a high level of generality to perform the mental process amounts to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f) under Steps 2A Prong 2 and 2B.
Claims 4 and 14 recite: “wherein determining the training dataset for the first machine learning module further comprises one or more of:(a) imputing missing feature values for the historical change data; (b) encoding one or more categorical features of respective features for each data point of the historical change data to create a respective feature vector for the each data point; or (c) augmenting minority class data points in a minority class of the historical change data”; imputing, encoding, or augmenting data points are evaluations that can be carried out by a human in the mind or with pen and paper, and is thus a mental process.
Claims 5 and 15 recite: “wherein one or more of:(a) imputing the missing feature values for the historical change data further comprises applying linear regression based on existing feature values for the historical change data; (b) encoding the one or more categorical features of the respective features for each data point further comprises applying a first encoding technique to encode a first categorical feature of the one or more categorical features and a second encoding technique to encode a second categorical feature of the one or more categorical features, the first encoding technique being different from the second encoding technique; or (c) augmenting the minority class data points in the minority class further comprises generating new minority class data points of the minority class data points based on existing minority class data points of the minority class data points, wherein the new minority class data points, as generated, are distributed toward a class periphery of the minority class”; applying linear regression, encoding, and generating new data points are evaluations that can potentially be performed by a human with pen and paper, and are thus a mental process
Claims 6 and 16 recite: “wherein the computing instructions are further configured to cause the one or more processors to perform: detecting a concept drift in the training dataset relative to a prior training dataset; and upon determining that the concept drift is above a concept drift threshold, re-training the first machine learning module based on the training dataset associated with the concept drift”; detecting a drift and making a decision to retrain based on the detection of drift is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process under Step 2A Prong 1; the broadly recited retraining of a model at a high level of generality amounts to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f) under Steps 2A Prong 2 and 2B.
Claims 7 and 17 recite: “wherein detecting the concept drift in the training dataset relative to the prior training dataset further comprises comparing a data distribution of the training dataset with a data distribution of the prior training dataset”; comparing data distributions is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process
Claims 9 and 19 recite: “wherein the computing instructions are further configured to cause the one or more processors to perform: after determining, by the first machine learning module, risk scores associated with multiple change requests, transmitting uncertain results, through a computer network, to a computing device for a domain expert”; determining which results are uncertain is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process under Step 2A Prong 1; transmitting them to an expert amounts to insignificant extra solution activity, mere data gathering and outputting, as per MPEP 2106.05(g); furthermore, this amounts to well-understood, routine, and conventional activity as per MPEP 2106.05(d) (“i. Receiving or transmitting data over a network”)
“wherein: the uncertain results comprise one or more change-risk combinations of the multiple change requests and the risk scores associated with the multiple change requests, selected based on a respective predictive uncertainty of each of the risk scores”; selecting uncertain results based on risk scores is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process under Step 2A Prong 1.
“upon receiving, via the computing device through the computer network, feedback from the domain expert, incorporating the feedback and the uncertain results into the training dataset for re-training the first machine learning module”; transmitting results to and from an expert amounts to insignificant extra solution activity, mere data gathering and outputting, as per MPEP 2106.05(g); furthermore, this amounts to well-understood, routine, and conventional activity as per MPEP 2106.05(d) (“i. Receiving or transmitting data over a network”)
Claims 10 and 20 recite: “wherein one or more of: the risk scores associated with the multiple change requests are determined by the first machine learning module in a current feedback cycle; the uncertain results are further determined based on a ranking of the respective predictive uncertainty for each of the risk scores; or the respective predictive uncertainty for each of the risk scores is determined based at least in part on a posterior predictive distribution for the risks scores”; determine risk scores based on rankings and distributions is an evaluation that can be carried out by a human in the mind or with pen and paper, and is thus a mental process under Step 2A Prong 1; the use of a broadly recited machine learning model at a high level of generality to perform the mental process amounts to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f) under Steps 2A Prong 2 and 2B.
Claim 21 recite: “[processor performs automatically] correcting label noise with the respective predicted label associated with the candidate data point when the prediction confidence is above a predetermined threshold” is essentially correcting a label with corrected label which is a mental process of judgement under Step 2A Prong 1 and the additional limitation of “processor performs automatically correcting label noise” as recited broadly recite implementing a processor as a generic tool to perform a function or mental process without any specific detail step of the performed function which amount to nothing more than an instruction to apply the abstract idea using a generic computer as per MPEP 2106.05(f) under Steps 2A Prong 2 and 2B.
Claim 22 recite: “correcting label noise with the respective predicted label associated with the candidate data point when the prediction confidence is above a predetermined threshold” is essentially correcting a label with corrected label which is a mental process of judgement under Step 2A Prong 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 6-7, 11, 13, and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Szanto et al. (US 2020/0286002 A1; hereinafter “Szanto”) in view of Rotta et al. (US 2019/0164100 A1; hereinafter “Rotta”).
As per Claim 1, Szanto teaches a system comprising: one or more processors; and one or more non-transitory computer-readable media storing computing instructions configured to, when run on the one or more processors, cause the one or more processors to perform (Szanto [0054]: “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus.”)
determining a training dataset for a first machine learning module based on historical change data by (Szanto [0004]: “determining that the second data has a difference metric relative to the first data that exceeds a difference threshold; retraining the first machine learning model using at least part of the second data, the retraining producing a second machine learning model.”
Notes: The “second data” has already been collected, and is therefore “historical” and it has a “difference” from the first data, and is thus “change” data. This data is used for training (“retraining”) a first machine learning model (“first machine learning model”)).
detecting and correcting correctable label noises in the historical change data based at least in part on: (correcting mapped below)
training a second machine learning module based on at least a sub- dataset of the historical change data in a warm-up stage to determine a respective predicted label associated with each data point of the historical change data (Szanto [0031]: “A drift detection engine 122 monitors overall distributional statistics of article language and model output. The drift detection engine 122 can identify when the distribution of words that the concept labeling engine 114 is exposed to at prediction time has deviated beyond a threshold from the distribution on which the engine was trained, resulting in degraded model performance.”
Notes: the “drift detection engine” is a “second machine learning model”. This model is trained (“engine was trained”) based on at least a sub-dataset of the historical change data (you cannot train on data future data that does not exist yet) to determine a respective predicted label (“model output”). Also, there is no definition given for “warm-up stage”, which is interpreted by Examiner to merely refer to training.)
after determining, by the second machine learning module, the respective predicated label associated with each data point of the historical change data, determining whether a label noise associated with a candidate data point of the historical change data is correctable based on (a) the candidate data point being in a confident data range of the historical change data and (b) a prediction confidence for the respective predicted label associated with the candidate data point is at least as great as a confidence threshold (Szanto [0031]: “The drift detection engine 122 can identify when the distribution of words that the concept labeling engine 114 is exposed to at prediction time has deviated beyond a threshold from the distribution on which the engine was trained, resulting in degraded model performance. This kind of drift can cause the concept model's predictions to degrade in quality because the language model in the concept labeling engine no longer reflects the data that the classifier currently is processing.” Szanto [0034]: “The system 100 can further include an active learning engine 126. The active learning engine 126 determines 128, for each document in a repository, e.g., the document repository 116, whether it should be given a ground-truth label, e.g., by a human … In one implementation, a concept label confidence score can have a range of from 0 to 1 with a score of 1 reflecting highest confidence and the system can set a threshold metric, e.g., a confidence score of 0.7 as a threshold for attaching a concept label to a document.”
Note: Here, they detect a label noise (“deviated”), and determine if it is correctable (a label that is wrong is “correctable” because one could provide the right label) based on a confident data range (“threshold”) and a prediction confidence (“confidence score”)).
and upon determining that the label noise associated with the candidate data point of the historical change data is correctable, correcting the label noise by replacing a current label of the candidate data point with the respective predicted label associated with the candidate data point (Szanto [0034]: “If so determined, that document is fed back into the annotation process, associating it with an expert label, adding it to labeled data, and possibly retraining the current model.”)
and after detecting and correcting the correctable label noises, determining the training dataset based on a label-error-free portion of the historical change data (Szanto [0034]: “If so determined, that document is fed back into the annotation process, associating it with an expert label, adding it to labeled data, and possibly retraining the current model.”
Here, the corrected “expert label” is label-error-free, and this is used as part of a training dataset.)
training the first machine learning module based on the training dataset [to determine a respective risk score associated with a respective change request] (Szanto [0034]: “If so determined, that document is fed back into the annotation process, associating it with an expert label, adding it to labeled data, and possibly retraining the current model.”)
However, Szanto does not teach to determine a respective risk score associated with a respective change request; determining, via the first machine learning module, as trained, a first risk score associated with a first change request; determining a change approval based on the first risk score and a risk threshold; and transmitting the change approval to cause an implementation of the first change request.
Rotta teaches training the first machine learning module based on the training dataset to determine a respective risk score associated with a respective change request (Rotta, Abstract: “The present invention is a system and method for evaluating an IT change request system based on cognitive and machine learning technologies. The system includes a computing device having a change request evaluator based on a machine learning trained model and in digital communication with a server … A business mapping tool interfaces with the change request evaluator and determines a business impact of the model change record and associates the business impact with the change request.”
Here, a machine learning model evaluates a change request and determines a risk score (“business impact”)).
determining, via the first machine learning module, as trained, a first risk score associated with a first change request (Rotta [0022]: “As will be discussed below, the business impact of prior changes can be categorized by change request type and used to predict the business impact of a new change request through machine learning techniques.”)
determining a change approval based on the first risk score and a risk threshold (Rotta [0028]: “In one embodiment, the determined business impact is compared to a disruption threshold. If the business impact extends beyond the disruption threshold, the change request will not be approved.”)
and transmitting the change approval and in response to the change approval, automatically implementing the first change request (Rotta [0005]: claim 8-9 “The change request evaluator approves the change request at the one or more servers by leveraging the machine learning trained model, change policies, and complexity, and if the business impact of the change request is below a threshold.” Here, an evaluator sends the approval via a server, which will cause the change request to be implemented.)
Rotta is analogous art because it is in the field of endeavor of applying machine learning to change requests. It would have been obvious before the effective filing date of the claimed invention to combine the machine learning model with drift detection of Szanto with the application to change request approvals of Rotta. One of ordinary skill in the art would have been motivated to do so in order to reduce the time and effort required to approve change requests (Rotta [0003-0004]: “Inefficiency in the change approval process leads to significant delays in executing upgrades or changes ordered by the customer. Such delays result in wasted productivity for the customer and ultimately, low customer satisfaction. Therefore, there is a need for a system and method to significantly reduce the time and effort expended in the change approval process.”) and to do so in such a way that changes in business impact over time which cause drift in change risks over time are accounted for (Szanto [0031]: “This kind of drift can cause the concept model's predictions to degrade in quality because the language model in the concept labeling engine no longer reflects the data that the classifier currently is processing.”)
As per Claim 3, the combination of Szanto and Rotta teaches the system in claim 1. Szanto teaches wherein: detecting and correcting the correctable label noises in the historical change data is further based on determining the confident data range of the historical change data based at least in part on a distribution of prediction confidence for predicted labels that are associated with the historical change data and determined by the second machine learning module, as trained in the warm-up stage. (Szanto [0031]: “The drift detection engine 122 can identify when the distribution of words that the concept labeling engine 114 is exposed to at prediction time has deviated beyond a threshold from the distribution on which the engine was trained, resulting in degraded model performance. This kind of drift can cause the concept model's predictions to degrade in quality because the language model in the concept labeling engine no longer reflects the data that the classifier currently is processing.”
Note: Here a “confident data range” is based on a distribution (“deviated … from the distribution”)).
As per Claim 6, the combination of Szanto and Rotta teaches the system in claim 1. Szanto teaches wherein the computing instructions are further configured to cause the one or more processors to perform: detecting a concept drift in the training dataset relative to a prior training dataset; and upon determining that the concept drift is above a concept drift threshold, re-training the first machine learning module based on the training dataset associated with the concept drift. (Szanto [0031]: “The drift detection engine 122 can identify when the distribution of words that the concept labeling engine 114 is exposed to at prediction time has deviated beyond a threshold from the distribution on which the engine was trained, resulting in degraded model performance. This kind of drift can cause the concept model's predictions to degrade in quality because the language model in the concept labeling engine no longer reflects the data that the classifier currently is processing.” Szanto [0034]: “If so determined, that document is fed back into the annotation process, associating it with an expert label, adding it to labeled data, and possibly retraining the current model.”)
Note: Here, they detect drift from the prior training set to the new training set, and upon determining it, retrain the model.)
As per Claim 7, the combination of Szanto and Rotta teaches the system in claim 6. Szanto teaches wherein detecting the concept drift in the training dataset relative to the prior training dataset further comprises comparing a data distribution of the training dataset with a data distribution of the prior training dataset. (Szanto [0031]: “The drift detection engine 122 can identify when the distribution of words that the concept labeling engine 114 is exposed to at prediction time has deviated beyond a threshold from the distribution on which the engine was trained, resulting in degraded model performance. This kind of drift can cause the concept model's predictions to degrade in quality because the language model in the concept labeling engine no longer reflects the data that the classifier currently is processing.”)
As per Claim 11, this is a computer implemented method claim corresponding to system Claim 1, and is rejected for similar reasons.
As per Claim 13, this is a computer implemented method claim corresponding to system Claim 3, and is rejected for similar reasons.
As per Claim 16, this is a computer implemented method claim corresponding to system Claim 6, and is rejected for similar reasons.
As per Claim 17, this is a computer implemented method claim corresponding to system Claim 7, and is rejected for similar reasons.
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Szanto and Rotta in view of Badawy et al. (US 2022/0269990 A1; hereinafter “Badawy”).
As per Claim 4, the combination of Szanto and Rotta teaches the system in claim 1. However, the combination does not teach wherein determining the training dataset for the first machine learning module further comprises one or more of:(a) imputing missing feature values for the historical change data; (b) encoding one or more categorical features of respective features for each data point of the historical change data to create a respective feature vector for the each data point; or (c) augmenting minority class data points in a minority class of the historical change data.
Badawy teaches wherein determining the training dataset for the first machine learning module further comprises one or more of:(a) imputing missing feature values for the historical change data; (b) encoding one or more categorical features of respective features for each data point of the historical change data to create a respective feature vector for the each data point; or (c) augmenting minority class data points in a minority class of the historical change data. (Badawy [0129]: “As mentioned above, however, many drift detection models may be more performant (or simpler to implement) on numerical data. Thus, such drift detection models may not be effectively utilized with categorical data or these types of identity graphs 565. Therefore, in some embodiments, to implement drift detection with respect to identity graph 565, graph embeddings may be utilized. A graph embedding model 592 may be used to transform the nodes, edges or features of identity graph 565 into a (e.g., lower dimension) vector representing the nodes or edges of the graph (or portion thereof) embedded. By utilizing graph embedding models 592 that are trained on identity management graph 565, this graph embedding model 592 can be used on new or different graphs (e.g., when an underlying attribute schema remains the same). These embeddings 596, which are a vector of numerical features, can then be used to detect drifts in the categorical features by applying the drift detection model 588 to comparing a dataset comprising an embedding 596 a of a previous instance of graph 565 (e.g., when the machine learning model 572 was trained) to a dataset comprising an embedding 596 b representing a current instance of the identity graph 565.”
Here, Badawy discloses (b): encoding categorical features as a vector in order to use them as training data.)
Badawy is analogous art because it is in the field of endeavor of detecting machine learning model concept drift. It would have been obvious before the effective filing date of the claimed invention to combine the drift detection of Szanto with the encoding of categorical features of Badawy. One of ordinary skill in the art would have been motivated to do so in order to be able to detect drift in categorical data (Badawy [0129]: “As mentioned above, however, many drift detection models may be more performant (or simpler to implement) on numerical data. Thus, such drift detection models may not be effectively utilized with categorical data or these types of identity graphs 565. Therefore, in some embodiments, to implement drift detection with respect to identity graph 565, graph embeddings may be utilized … These embeddings 596, which are a vector of numerical features, can then be used to detect drifts in the categorical features.”)
As per Claim 14, this is a computer implemented method claim corresponding to system Claim 4, and is rejected for similar reasons.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Szanto, Rotta, and Badawy in view of Chaterjee et al. (US 2020/0097545 A1; hereinafter “Badawy”).
As per Claim 5, the combination of Szanto, Rotta, and Badawy teaches the system in claim 4. However, the combination does not teach wherein one or more of:(a) imputing the missing feature values for the historical change data further comprises applying linear regression based on existing feature values for the historical change data; (b) encoding the one or more categorical features of the respective features for each data point further comprises applying a first encoding technique to encode a first categorical feature of the one or more categorical features and a second encoding technique to encode a second categorical feature of the one or more categorical features, the first encoding technique being different from the second encoding technique; or (c) augmenting the minority class data points in the minority class further comprises generating new minority class data points of the minority class data points based on existing minority class data points of the minority class data points, wherein the new minority class data points, as generated, are distributed toward a class periphery of the minority class.
Chatterjee teaches wherein one or more of: (a) imputing the missing feature values for the historical change data further comprises applying linear regression based on existing feature values for the historical change data; (b) encoding the one or more categorical features of the respective features for each data point further comprises applying a first encoding technique to encode a first categorical feature of the one or more categorical features and a second encoding technique to encode a second categorical feature of the one or more categorical features, the first encoding technique being different from the second encoding technique; or (c) augmenting the minority class data points in the minority class further comprises generating new minority class data points of the minority class data points based on existing minority class data points of the minority class data points, wherein the new minority class data points, as generated, are distributed toward a class periphery of the minority class. (Chatterjee [0071]: “In some implementations, the feature encoding platform, when performing the feature engineering on the numeric features, may convert the numeric features into similarity scores, that are included in the converted features, based on a pre-defined set of words, or convert the numeric features into fixed sized vectors that are included in the converted features. In some implementations, the feature encoding platform, when performing the feature engineering on the categorical features, may convert the categorical features into n-gram sequences that are included in the converted features, convert the categorical features into primary forms that are included in the converted features, convert the categorical features into variable size vectors that are included in the converted features, and/or the like.”
Here, Chatterjee discloses (b), as they disclose the possibility of using a combination of different embedding techniques for different categorical variables (“may convert the categorical features into n-gram sequences … convert the categorical features into primary forms … convert the categorical features into variable size vectors … and/or the like”).
Chatterjee is analogous art because it is in the field of endeavor of encoding categorical features for use in machine learning models. It would have been obvious before the effective filing date of the claimed invention to combine the drift detection with categorical embeddings of Szanto and Badawy with the combination of different encoding techniques for different categorical features of Chatterjee. One of ordinary skill in the art would have been motivated to do so in order to optimally structure the data to maximize efficiency (Chatterjee [0012]: “The feature encoding platform may conserve time and resources (e.g., processing resources, memory resources, and/or the like) associated with training and testing of a machine learning model with unstructured data, and may improve machine learning model training accuracy and testing accuracy when unstructured data is utilized.”)
As per Claim 15, this is a computer implemented method claim corresponding to system Claim 5, and is rejected for similar reasons.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Szanto and Rotta in view of Shumpert (US 2016/0342903 A1).
As per Claim 8, the combination of Szanto and Rotta teaches the system in claim 6. However, the combination does not teach wherein: re-training the first machine learning module based on the training dataset associated with the concept drift further comprises assigning a respective weight to each data point of the training dataset, wherein: the respective weight for a more-recent-in-time data point of the training dataset is higher relative to the respective weight for a less-recent-in-time data point of the training dataset; and re-training the first machine learning module is further based on the respective weight for each data point of the training dataset.
Shumpert teaches wherein: re-training the first machine learning module based on the training dataset associated with the concept drift further comprises assigning a respective weight to each data point of the training dataset, wherein: the respective weight for a more-recent-in-time data point of the training dataset is higher relative to the respective weight for a less-recent-in-time data point of the training dataset; and re-training the first machine learning module is further based on the respective weight for each data point of the training dataset. (Shumpert [0077]: “In order for the models to handle concept drift and adapt to changing conditions over time, the shared learning and prediction component 506 may give stronger weights to newer data instances than older data when training the shared models as described below.”)
Shumpert is analogous art because it is in the field of endeavor of machine learning with concept drift. It would have been obvious before the effective filing date of the claimed invention to combine the concept drift aware machine learning of Szanto with the data weighting of Shumpert. One of ordinary skill in the art would have been motivated to do so in order to adapt to changing conditions over time (Shumpert [0077]: “In order for the models to handle concept drift and adapt to changing conditions over time, the shared learning and prediction component 506 may give stronger weights to newer data instances than older data when training the shared models as described below.”)
As per Claim 18, this is a computer implemented method claim corresponding to system Claim 8, and is rejected for similar reasons.
Claims 9-10 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Szanto and Rotta in view of Wittenbach et al. (US 2022/0067737 A1; hereinafter “Wittenbach”).
As per Claim 9, the combination of Szanto and Rotta teaches the system in claim 1 as well as risk scores associated with multiple change requests (see Rotta in rejection to Claim 1). However, the combination does not teach wherein the computing instructions are further configured to cause the one or more processors to perform: after determining, by the first machine learning module, risk scores associated with multiple change requests, transmitting uncertain results, through a computer network, to a computing device for a domain expert, wherein: the uncertain results comprise one or more change-risk combinations of the multiple change requests and the risk scores associated with the multiple change requests, selected based on a respective predictive uncertainty of each of the risk scores; and upon receiving, via the computing device through the computer network, feedback from the domain expert, incorporating the feedback and the uncertain results into the training dataset for re-training the first machine learning module.
Wittenbach teaches wherein the computing instructions are further configured to cause the one or more processors to perform: after determining, by the first machine learning module, risk scores associated with multiple change requests, transmitting uncertain results, through a computer network, to a computing device for a domain expert, wherein: the uncertain results comprise one or more change-risk combinations of the multiple change requests and the risk scores associated with the multiple change requests, selected based on a respective predictive uncertainty of each of the risk scores; and upon receiving, via the computing device through the computer network, feedback from the domain expert, incorporating the feedback and the uncertain results into the training dataset for re-training the first machine learning module. (Recall above that Rotta teaches risk scores associated with change requests. Wittenbach [0033]: “As noted herein, active learning systems build on machine learning models by augmenting the model with an algorithm that allows the active learning systems to request additional labels from unlabeled data. The use of Bayesian neural networks allows for the decomposition of uncertainty in two different ways to get an aleatoric and an epistemic uncertainty of the predictive model, the combination of which can lead to decisions that can help improve the performance of the model and layer better judgement on the use of model outputs. For example, a predictive model may be trained on an initial labeled data set. As new unlabeled data comes in, epistemic uncertainty scores are computed and collected. Human experts can be assigned to label the data points with the highest epistemic uncertainty scores. Simultaneously, data points with high aleatoric uncertainty in the training data can be reviewed by data scientists and new features or data collection processes can be engineered that increase accuracy on these data points. These data points and features can then be computed and added to the training set, and the model can again be refit.”)
Wittenbach is analogous art because it is in the field of endeavor of machine learning with human intervention for uncertainty. It would have been obvious before the effective filing date of the claimed invention to combine the concept drift aware machine learning of Szanto with the uncertainty intervention of Wittenbach. One of ordinary skill in the art would have been motivated to do so in order to efficiently handle uncertain data by only requiring human intervention for identified uncertain labels (Wittenbach [0033]: “Identification of the source of uncertainty allows for more efficient treatment of the uncertainty by the human-in-the-loop component. For example, aleatoric data will not decrease through the collection of more data, but requires changes to the data collection process itself. Identifying that an uncertainty is of an aleatoric nature allows humans to not only diagnose sources of uncertainty, but to also treat them properly.”)
As per Claim 10, the combination of Szanto, Rotta, and Wittenbach teaches the system in claim 9. Wittenbach teaches wherein one or more of: the risk scores associated with the multiple change requests are determined by the first machine learning module in a current feedback cycle; the uncertain results are further determined based on a ranking of the respective predictive uncertainty for each of the risk scores; or the respective predictive uncertainty for each of the risk scores is determined based at least in part on a posterior predictive distribution for the risks scores. (Recall above that Rotta teaches risk scores. Wittenbach [0033]: “As noted herein, active learning systems build on machine learning models by augmenting the model with an algorithm that allows the active learning systems to request additional labels from unlabeled data. The use of Bayesian neural networks allows for the decomposition of uncertainty in two different ways to get an aleatoric and an epistemic uncertainty of the predictive model, the combination of which can lead to decisions that can help improve the performance of the model and layer better judgement on the use of model outputs. For example, a predictive model may be trained on an initial labeled data set. As new unlabeled data comes in, epistemic uncertainty scores are computed and collected. Human experts can be assigned to label the data points with the highest epistemic uncertainty scores.”
Examiner notes that here, the uncertainty results are based on a ranking of scores, because in order to determine the “highest” uncertainty scores, they must be ordered or ranked.)
As per Claim 19, this is a computer implemented method claim corresponding to system Claim 9, and is rejected for similar reasons.
As per Claim 20, this is a computer implemented method claim corresponding to system Claim 10, and is rejected for similar reasons.
Claims 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Szanto and Rotta in view of Jayne et al. (US 2021/0110439 A1; hereinafter “Jayne”)
As per claim 21, Szanto and Rosso do not specifically disclose wherein the processor performs automatically correcting label noise with the respective predicted label associated with the candidate data point when the prediction confidence is above a predetermined threshold.
Jayne teaches wherein the processor performs automatically correcting label noise with the respective predicted label associated with the candidate data point when the prediction confidence is above a predetermined threshold (par. 94).
As per Claim 22, Jayne teaches further comprising correcting label noise with the respective predicted label associated with the candidate data point when the prediction confidence is above a predetermined threshold (par. 94).
It would have been obvious to a person or ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Szanto and Rotta with the automatic label correction of Jayne. One of ordinary skill in the art would have been motivated to combine the automated label correction with confidence value above threshold to improve the efficiency of label correction without the need of an expert or human intervention for efficient label correction.
Response to Arguments
Applicant's arguments filed 2/25/2026 has been fully considered but they are not persuasive.
Argument:
a) In remarks regarding 101 rejection, applicant argues “training a second machine learning module base on a sub-dataset of the historical change data in a warm-up stage", and "training the first machine learning module based on the training dataset", neither of which can be performed in the human mind as mere mental steps. Additionally, Applicants have amended independent claims 1 and 11 to clarify that a practical application is recited. That is ,as amended, claims 1 and 11 recite implementing the approved change request, thus resulting in a practical real-world application.
b) Regarding 103 rejection applicant argues “Szanto merely describes sending uncertain terms for a human to label. However, Szanto clearly fails to teach or suggest a second machine learning module that is trained on a sub-dataset of historical change data in a warm-up stage to determine a respective predicted label associated with each data point of the historical change data. Szanto's human labeling of uncertain terms fails to teach or suggest automatically correcting erroneous labels that are determined to be correctable, as in embodiments of the present application as claimed.”.
Response to argument:
Regarding argument (a) examiner respectfully disagrees with the applicant. The limitations of training a machine learning model was never treated as a mental process rather has been addressed as generic implementation of the machine learning model as a tool. The claimed limitation does not specifically disclose any details regarding the training of the machine learning model rather discloses the limitation in a high level of generality which appears to be a generic implementation of a model being training in a generic manner as a generic tool to implement the claim abstract. Similarly applicant’s argument regarding “implement the change request” is also not persuasive as the claimed limitation fails to provide any details or real world implementation of the change rather general recitation or implement a change which appears to be a generic implementation of a computer or tool to perform the claimed abstract idea. Therefore, applicant’s argument regarding the amended claims being eligible is not persuasive.
Regarding argument (b) examiner respectfully disagrees with the applicant. In remarks applicants argument regarding the reference failing to teach training of a machine learning model fail to comply with 37 CFR 1.111(b) because they amount to a general allegation that the claims define a patentable invention without specifically pointing out how the language of the claims patentably distinguishes them from the references. Applicant fails to provide any analysis how the cited reference Szanto fails to disclose the claimed limitation or any difference between the cited prior art and the claimed limitation. Moreover the specification of the instant applicant fails to disclose any specific definition or explanation that clearly states training during a warm-up stage rather than simple recitation of the claim language. The issued office action clearly stated the teaching and interpretation of the warm-up stage in the previous office action as stated “ Szanto [0031]: "A drift detection engine 122 monitors overall distributional statistics of article language and model output. The drift detection engine 122 can identify when the distribution of words that the concept labeling engine 114 is exposed to at prediction time has deviated beyond a threshold from the distribution on which the engine was trained, resulting in degraded model performance." Notes: the "drift detection engine" is a "second machine learning model". This model is trained ("engine was trained") based on at least a sub-dataset of the historical change data (you cannot train on data future data that does not exist yet) to determine a respective predicted label ("model output"). Also, there is no definition given for "warm-up stage", which is interpreted by Examiner to merely refer to training.”. The argument presented by applicant fails to address any analysis or interpretation provided by the examiner in the previous office action and fails to provide any reasoning why the cited reference fails to disclose the claimed limitation. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., fails to teach or suggest automatically correcting erroneous labels that are determined to be correctable instead of human labeling) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Therefore applicant’s argument regarding the independent claims are not persuasive and similarly the argument regarding the dependent claims are not persuasive.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDULLAH AL KAWSAR whose telephone number is (571)270-3169. The examiner can normally be reached M-F 7:30am-4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Wiley can be reached at (571) 272-4150. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127