DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to the instant application filed on 03/21/2024. Claims 1-20 are pending and have been examined.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 1, the claim recites: “generating, with the at least one processor, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels, and a plurality of groups of samples of the plurality of samples;”. This claim fails to make clear what is being generated but merely that it is being done based on the predictions, labels, and groups of samples. For the purpose of examination, the claim will be interpreted as “generating, with the at least one processor, based on the plurality of first predictions, the plurality of second predictions, and the plurality of labels, a plurality of groups of samples of the plurality of samples;”.
Regarding claim 9, the claim recites: “generate, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels, and a plurality of groups of samples of the plurality of samples;”. This claim fails to make clear what is being generated but merely that it is being done based on the predictions, labels, and group of samples. For the purpose of examination, the claim will be interpreted as: “generate, based on the plurality of first predictions, the plurality of second predictions, and the plurality of labels, a plurality of groups of samples of the plurality of samples;”.
Regarding claim 17, the claim recites: “generate, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels, and a plurality of groups of samples of the plurality of samples;”. This claim fails to make clear what is being generated but merely that it is being done based on the predictions, labels, and group of samples. For the purpose of examination, the claim will be interpreted as “generate, based on the plurality of first predictions, the plurality of second predictions, and the plurality of labels, a plurality of groups of samples of the plurality of samples;”.
Regarding dependent claims 2-8, 10-16, 18-20, these claims do not cure the deficiencies identified above in the claims they depend on, thus are also rejected for the same reason as noted above.
Claims 6 and 14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 6, the claim recites: “and determining, for each second prediction score, an aligned score aligned to the same scale as the plurality of first predictions scores, the value of the first prediction score assigned to the first bucket of the plurality of first buckets for which the rate of positive first predictions is a same rate as the rate of positive second predictions of the second bucket to which that prediction score is assigned.” The claim fails to make clear what connects the initial determining step and the phrase “the value of the first prediction score assigned to the first bucket of the plurality of first buckets for which the rate of positive first predictions is a same rate as the rate of positive second predictions of the second bucket to which that prediction score is assigned”. It is not clear how the portion after the comma relates to the initial determining step before the comma. Furthermore, the claim recites “that prediction score” given there is a first and second prediction score from a plurality of prediction scores, it is unclear which prediction score is being discussed. For the purpose of examination, the limitation will be interpreted as “and determining, for each second prediction score, an aligned score aligned to the same scale as the plurality of first predictions scores, wherein the value of the first prediction score assigned to the first bucket of the plurality of first buckets for which the rate of positive first predictions is a same rate as the rate of positive second predictions of the second bucket to which that second prediction score is assigned.”
Claim 14 recites essentially the same limitation and as such is deemed indefinite and interpreted the same as claim 6.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1: The claim recites a method which falls into the statutory category of process.
Step 2A Prong 1: The claim recites multiple abstract ideas: generating, with the at least one processor, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels, and a plurality of groups of samples of the plurality of samples, a mental process given a human being can use predictions to group up samples and generate a plurality of groups of samples; determining, with the at least one processor, based on the plurality of groups of samples, a first success rate associated with the first machine learning model and a second success rate associated with the second machine learning model, a mental process given a human being can determine a simple success rate for a machine learning model with the aid of pen and paper; identifying, with the at least one processor, based on the first success rate and the second success rate, a weak point in the second machine learning model associated with a first portion of samples of the plurality of samples including a same first value for a same first feature of the plurality of features and for which the first success rate associated with the first machine learning model is different than the second success rate associated with the second machine learning model, a mental process given a human being can look at the output of models/their success rates and determine where the models are weak.
Step 2A Prong 2: Claim 1 does not integrate the abstract idea into a practical application since
the additional elements of:
obtaining, with at least one processor, a plurality of features associated with a plurality of samples and a plurality of labels for the plurality of samples, is insignificant extra-solution activity: data gathering.
generating, with the at least one processor, a plurality of first predictions for the plurality of samples by providing, as input to a first machine learning model, a first subset of features of the plurality of features, and receiving, as output from the first machine learning model, the plurality of first predictions for the plurality of samples, merely links the abstract idea to a field of use or technological environment.
generating, with the at least one processor, a plurality of second predictions for the plurality of samples by providing, as input to a second machine learning model, a second subset of features of the plurality of features, and receiving, as output from the second machine learning model, the plurality of second predictions for the plurality of samples, merely links the abstract idea to a field of use or technological environment.
Step 2B: Claim 1 does not include additional elements that are sufficient to amount to
significantly more than a judicial exception when taken alone or in combination. As discussed, the additional element of receiving information is well-understood, routine, conventional activity as evidence by MPEP§2106.05(d)(II)(I). Further, the additional elements of generating predictions by using a machine learning model is considered generally linking the abstract idea to a field of use/technological environment (MPEP§2106.05(h)).
Regarding claim 2, the rejection of claim 1 is incorporated, further the claim recites: wherein at least one of: (i) the first subset of features is different than the second subset of features; (ii) a first set of hyperparameters for a machine learning algorithm used to generate the first machine learning model is different than a second set of hyperparameters for a same machine learning algorithm used to generate the second machine learning model; (iii) a first machine learning algorithm used to generate the first machine learning model is different than a second machine learning algorithm used to generate the second machine learning model; and (iv) a first training data set used to train the first machine learning model is different than a second training data set used to train the second machine learning model. This limitation amounts to generally linking the abstract idea to a technological environment: one with two models with differing features, training data sets, hyperparameters, or algorithms (see MPEP 2106.05(h)).
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions given the limitation merely links the judicial exceptions to a technological environment. The claim is not patent eligible.
Regarding claim 3, the rejection of claim 1 is incorporated, further the claim recites: wherein the first subset of features is different than the second subset of features (generally linking the abstract idea to a technological environment), and wherein identifying the weak point in the second machine learning model further includes: determining a difference in features between the first subset of features and the second subset of features (a mental process given a human being can look at two subsets of features and determine a difference mentally); selecting, based on the same first feature included in the first portion of samples and the difference in features, one or more features of the plurality of features (a mental process given a human being can mentally pick features); adjusting the second subset of features based on the selected one or more features (a mental process given a human being can mentally or with the aid of pen and paper change a subset of features); and generating, using the adjusted second subset of features, an updated second machine learning model (mere instructions to apply the abstract idea by a generic computer (a machine learning model) (MPEP 2106.05(f))).
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions given the limitation only recites a single additional element which amounts to mere instructions to apply the judicial exceptions. The claim is not patent eligible.
Regarding claim 4, the rejection of claim 1 is incorporated, further the claim recites: wherein a first set of hyperparameters for a machine learning algorithm used to generate the first machine learning model is different than a second set of hyperparameters for a same machine learning algorithm used to generate the second machine learning model (generally linking the abstract idea to a technological environment see MPEP 2106.05(h)), and wherein identifying the weak point in the second machine learning model further includes: determining a difference in hyperparameters between the first set of hyperparameters and the second set of hyperparameters (a mental process given a human being can find the difference between two sets of hyperparameters); determining, based on the same first feature included in the first portion of samples and the difference in the hyperparameters, one or more hyperparameters (a mental process); adjusting the second set of hyperparameters based on the selected one or more hyperparameters (a mental process); and generating, using the adjusted second set of hyperparameters, an updated second machine learning model (mere instructions to apply the abstract idea by a generic computer (a machine learning model) see MPEP 2106.05(f)).
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions given the additional elements amount to mere instructions to apply the judicial exceptions and merely linking the judicial exceptions to a technological environment. The claim is not patent eligible.
Regarding claim 5, the rejection of claim 1 is incorporated, further the claim recites: wherein the plurality of first predictions include a plurality of first prediction scores, wherein the plurality of second predictions include a plurality of second prediction scores, and wherein generating the plurality of groups of samples of the plurality of samples further includes: aligning the plurality of second prediction scores to a same scale as the plurality of first prediction scores (a mental process given a human being can align scores to a singular scale); applying an operating point to the plurality of first prediction scores to determine a plurality of first positive predictions and a plurality of first negative predictions (a mental process given a human being can decide some operating point that separates positive and negative predictions); applying the operating point to the plurality of aligned second prediction scores to determine a plurality of second positive predictions and a plurality of second negative predictions (a mental process given a human being can mentally decide some operating point that separates positive and negative predictions); and generating, based on the plurality of first positive predictions, the plurality of first negative predictions, the plurality of second positive predictions, the plurality of second negative predictions, the plurality of labels, and the plurality of groups of samples of the plurality of samples (a mental process given a human being can generate groups of samples).
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. The claim is not patent eligible.
Regarding claim 6, the rejection of claim 5 is incorporated, further the claim recites: wherein aligning the plurality of second prediction scores to the same scale as the plurality of first prediction scores includes: assigning each first prediction score of the plurality of prediction scores to a first bucket of a plurality of first buckets according to a value of that first prediction score (a mental process given a human being can mentally assign scores to buckets); determining, for each first bucket, a rate of positive first predictions up to the value of the first prediction score assigned to that first bucket (a mental process given a human being can mentally determine rates for buckets); assigning each second prediction score of the plurality of prediction scores to a second bucket of a plurality of second buckets according to a value of that second prediction score (a mental process given a human being can mentally assign scores to buckets); determining, for each second bucket, a rate of positive second predictions up to the value of the second prediction score assigned to that second bucket (a mental process given a human being can mentally determine rates for buckets); and determining, for each second prediction score, an aligned score aligned to the same scale as the plurality of first predictions scores, the value of the first prediction score assigned to the first bucket of the plurality of first buckets for which the rate of positive first predictions is a same rate as the rate of positive second predictions of the second bucket to which that prediction score is assigned (a mental process given a human being can mentally determine aligned scores).
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. The claim is not patent eligible.
Regarding claim 7, the rejection of claim 1 is incorporated, further the claim recites: wherein generating the plurality of groups of samples of the plurality of samples includes: determining, with the at least one processor, a first group of samples of the plurality of samples for which a first prediction of the plurality of first predictions matches a label of the plurality of labels and a second prediction of the plurality of second predictions matches the label of the plurality of labels (a mental process given a human being can mentally determine groups of samples); determining, with the at least one processor, a second group of samples of the plurality of samples for which the second prediction of the plurality of second predictions matches the label of the plurality of labels and the first prediction of the plurality of first predictions does not match the label of the plurality of labels (a mental process given a human being can mentally determine groups of samples); determining, with the at least one processor, a third group of samples of the plurality of samples for which the first prediction of the plurality of first predictions matches the label of the plurality of labels and the second prediction of the plurality of second predictions does not match the label of the plurality of labels (a mental process given a human being can mentally determine groups of samples); determining, with the at least one processor, a fourth group of samples of the plurality of samples for which the second prediction of the plurality of second predictions does not match the label of the plurality of labels and the first prediction of the plurality of first predictions does not match the label of the plurality of labels (a mental process given a human being can mentally determine groups of samples); determining, with the at least one processor, a fifth group of samples of the plurality of samples for which the first prediction of the plurality of first predictions does not match the label of the plurality of labels and the second prediction of the plurality of second predictions matches the label of the plurality of labels (a mental process given a human being can mentally determine groups of samples); and determining, with the at least one processor, a sixth group of samples of the plurality of samples for which the second prediction of the plurality of second predictions does not match the label of the plurality of labels and the first prediction of the plurality of first predictions matches the label of the plurality of labels (a mental process given a human being can mentally determine groups of samples).
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. The claim is not patent eligible.
Regarding claim 8, the rejection of claim 7 is incorporated, further the claim recites: wherein the first success rate associated with the first machine learning model and the second success rate associated with the second machine learning model are determined according to the following Equations (1) and (2). This limitation amounts to more specifics of the abstract idea of determining success rates given it provides a formula that can be executed in the human mind with the aid of pen and paper.
The claim does not include any additional elements that amount to an integration of the judicial exceptions into a practical application, nor to significantly more than the judicial exceptions. The claim is not patent eligible.
Claim 9 recites features similar to claim 1 and is rejected for at least the same reasons therein. Claim 9 additionally requires analysis for a “A system, comprising: at least one processor programmed and/or configured to….” however these additional elements amount to mere instructions to apply the judicial exception using a generic computer component. Please see MPEP 2106.05(f).
Regarding claims 10-16, they recite features similar to claims 2-8 and are rejected for at least the same reasons therein.
Regarding claim 17, it recites features similar to claim 1 and is rejected for at the least the same reasons therein. Claim 17 additionally requires analysis for “A computer program product comprising at least one non-transitory computer-readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to…” however these additional elements amount to mere instructions to the apply the judicial exception using a generic computer component. Please see MPEP 2106.05(f).
Regarding claim 18-20, they recite features similar to claims 2, 7, 8 and are rejected for at least the same reasons therein.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 5, 7-11, 13, 15, and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Breckenridge (US8250009B1) in view of Zhang (Manifold: A Model-Agnostic Framework for Interpretation and Diagnosis of Machine Learning Models).
Regarding claim 1, Breckenridge teaches obtaining, with at least one processor, a plurality of features associated with a plurality of samples and a plurality of labels for the plurality of samples (Abs, A series of training data sets for predictive modeling can be received, e.g., over a network from a client computing system, Col. 13 lines 23-28, In the above example command, “data” refers to data used in training the models (i.e., training data); “mixture' refers to a combination of text and numeric data, “input' refers to data to be used to update the model (i.e., new training data), “bucket' refers to a location where the models to be updated are stored, “X”, “y” and “Z” refer to other potential data values for a given feature, Col. 16, if the training data is categorized, then when the training data in a particular category included in the new training data reaches a fraction of the initial training data in the particular category, then the first condition can be satisfied…….., categories are labels); generating, with the at least one processor, a plurality of first predictions for the plurality of samples by providing, as input to a first machine learning model, a first subset of features of the plurality of features, and receiving, as output from the first machine learning model, the plurality of first predictions for the plurality of samples; generating, with the at least one processor, a plurality of second predictions for the plurality of samples by providing, as input to a second machine learning model, a second subset of features of the plurality of features, and receiving, as output from the second machine learning model, the plurality of second predictions for the plurality of samples (Col. 8 lines 28-36, each type of predictive model may have multiple training functions and that multiple hyper parameter configurations and selected features may be used for each of the multiple training functions, there are many different trained predictive models that can be generated. Depending on the nature of the input data to be used by the trained predictive model to predict an output, different trained predictive models perform differently, one of these models can be considered a first model and one of the other models a second model).
Breckenridge fails to teach generating, with the at least one processor, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels, identifying, with the at least one processor, based on the first success rate and the second success rate, a weak point in the second machine learning model associated with a first portion of samples of the plurality of samples including a same first value for a same first feature of the plurality of features and for which the first success rate associated with the first machine learning model is different than the second success rate associated with the second machine learning model and generating, with the at least one processor, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels,
Zhang teaches generating, with the at least one processor, based on the plurality of first predictions, the plurality of second predictions, the plurality of labels, (Section 4.1.1, For example, instances in the positive half of the X axis (Q1 and Q4 in Figure 2) indicates that they are predicted by Mi as Ci. In contrast, instances in the negative half of the X axis (Q2 and Q3 in Figure 2) indicates that they are predicted by Mi as another class rather than Ci. Moreover, instances in the fourth quadrant (Q4 in Figure 2) indicates that they are predicted by Mi as Ci, but predicted by Mj as another class instead (since Q4 is within the negative half of the Y axis); determining, with the at least one processor, based on the plurality of groups of samples, a first success rate associated with the first machine learning model and a second success rate associated with the second machine learning model (Table 2, Figure 5, Model performance comparison between model pairs for the bike sharing regression problem, determines model performance over the classes (groups of samples)); identifying, with the at least one processor, based on the first success rate and the second success rate, a weak point in the second machine learning model associated with a first portion of samples of the plurality of samples including a same first value for a same first feature of the plurality of features and for which the first success rate associated with the first machine learning model is different than the second success rate associated with the second machine learning model (Section 3.2, The user filters a subset of instances that have some features in common, which is typically done through a faceted search on the input features or the metadata information…. The user identifies a subset where the results generated by the model are erroneous or suspicious, for example, the instances where the new model has low accuracy while others have high accuracy. We define this type of subset as a symptom set since it is representative of a potential fault within the model. The symptom set is of particular interest to the user during the diagnosis process)
Breckenridge and Zhang are analogous to the claimed invention because they are in the field of predictive models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the manifold system in Zhang along with the method in Breckenridge to improve predictive models by allowing “for comparing feature distributions and generating explanations for the issue” (Zhang Section 4).
Regarding claim 2, Breckenridge in view of Zhang teaches the method of claim 1, further Breckenridge teaches the first subset of features is different than the second subset of features (Col. 8 lines 18-20, Additionally, a predictive model can be trained with different features, again generating different trained models) and a first set of hyperparameters for a machine learning algorithm used to generate the first machine learning model is different than a second set of hyperparameters for a same machine learning algorithm used to generate the second machine learning model (Col. 8 lines 13-16, For a given training function, multiple different hyper parameter configurations can be applied to the training function, again generating multiple different trained predictive models.)
Regarding claim 3, Breckenridge in view of Zhang teaches the method of claim 1, further Breckenridge teaches wherein the first subset of features is different than the second subset of features (Col. 8 lines 18-20, Additionally, a predictive model can be trained with different features, again generating different trained models). Zhang teaches, which Breckenridge fails to teach, and wherein identifying the weak point in the second machine learning model further includes: determining a difference in features between the first subset of features and the second subset of features (Section 3.2, The user filters a subset of instances that have some features in common, which is typically done through a faceted search on the input features or the meta data information); selecting, based on the same first feature included in the first portion of samples and the difference in features, one or more features of the plurality of features (Section 3.2, The user identifies a subset where the results generated by the model are erroneous or suspicious, for example, the instances where the new model has low accuracy while others have high accuracy. We define this type of subset as a symptom set since it is representative of a potential fault within the model…...Comparative analysis is intensively involved in this phase. For example, after selecting a subset of false positive instances (the ground truth class is A while the result predicted by the model is B), the user may want to investigate what features or local structures of the selected subset are more similar to the set which ground truth class is B and less similar to the set which ground truth class is A. These features could be influential to the false positive results and hence are regarded as an explanation of the symptom); adjusting the second subset of features based on the selected one or more features; and generating, using the adjusted second subset of features, an updated second machine learning model (Section 3.2, First, depending on the model type, the user can either apply feature engineering strategies or adjust the internal architecture of the model (typically for deep neural networks)).
Breckenridge and Zhang are analogous to the claimed invention because they are in the field of predictive models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the manifold system in Zhang, particularly the feature engineering, to verify explanations for issues in models and test performance (Zhang Section 3.2, In this phase, the user attempts to verify the explanations generated from the previous phase through encoding the knowledge extracted from the explanation into the model and testing the performance…. First, depending on the model type, the user can either apply feature engineering strategies or adjust the internal architecture of the model...).
Regarding claim 5, Breckenridge in view of Zhang teaches the method of claim 1, further Breckenridge teaches wherein the plurality of first predictions include a plurality of first prediction scores, wherein the plurality of second predictions include a plurality of second prediction scores (Col. 7, lines 43-47, Some examples of training functions that can be used to train a static predictive model include (without limitation): regression (e.g., linear regression, logistic regression), classification and regression tree, multivariate adaptive regression), and wherein generating the plurality of groups of samples of the plurality of samples further includes: aligning the plurality of second prediction scores to a same scale as the plurality of first prediction scores (Col. 7 lines 57-58, Referring again to FIG. 4, multiple predictive models, which can be all or a subset of the available predictive models, the models can be identical which means the scores are on the same scale). Zhang teaches, which Breckenridge fails to teach, applying an operating point to the plurality of first prediction scores to determine a plurality of first positive predictions and a plurality of first negative predictions; applying the operating point to the plurality of aligned second prediction scores to determine a plurality of second positive predictions and a plurality of second negative predictions (Section 4.1.1, Each point in the coordinate system represents one input instance and the coordinate on the X (Y) axis indicates the prediction score generated by the model Mi (Mj) on the class Ci. Hence, the points that are close to the origin indicate lower prediction confidence than those far from the origin. Since the prediction score is non-negative, we use the positive half and the negative half of the coordinate system to encode whether the prediction result on the instance is Ci or not, Mi and Mj are the two models); and generating, based on the plurality of first positive predictions, the plurality of first negative predictions, the plurality of second positive predictions, the plurality of second negative predictions, the plurality of labels, (Section 4.1.1, For example, instances in the positive half of the X axis (Q1 and Q4 in Figure 2) indicates that they are predicted by Mi as Ci. In contrast, instances in the negative half of the X axis (Q2 and Q3 in Figure 2) indicates that they are predicted by Mi as another class rather than Ci. Moreover, instances in the fourth quadrant (Q4 in Figure 2) indicates that they are predicted by Mi as Ci, but predicted by Mj as another class instead (since Q4 is within the negative half of the Y axis)).
Breckenridge and Zhang are analogous to the claimed invention because they are in the field of predictive models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the manifold system in Zhang to diagnose the model’s performance on different classes and to allow for drilling down to a subset of data points based on user-defined conditions (Zhang Section 4.1.1, Section 3.1).
Regarding claim 7, Breckenridge in view of Zhang teaches the method of claim 1, further Zhang teaches wherein generating the plurality of groups of samples of the plurality of samples includes: determining, with the at least one processor, a first group of samples of the plurality of samples for which a first prediction of the plurality of first predictions matches a label of the plurality of labels and a second prediction of the plurality of second predictions matches the label of the plurality of labels (Fig 2, Q1 instances in blue, first prediction is M_i and second is M_j); determining, with the at least one processor, a second group of samples of the plurality of samples for which the second prediction of the plurality of second predictions matches the label of the plurality of labels and the first prediction of the plurality of first predictions does not match the label of the plurality of labels (Fig 2, Q2 instances in blue); determining, with the at least one processor, a third group of samples of the plurality of samples for which the first prediction of the plurality of first predictions matches the label of the plurality of labels and the second prediction of the plurality of second predictions does not match the label of the plurality of labels (Fig 2, Q4 instances in blue); determining, with the at least one processor, a fourth group of samples of the plurality of samples for which the second prediction of the plurality of second predictions does not match the label of the plurality of labels and the first prediction of the plurality of first predictions does not match the label of the plurality of labels (Fig 2, Q3 instances in blue); determining, with the at least one processor, a fifth group of samples of the plurality of samples for which the first prediction of the plurality of first predictions does not match the label of the plurality of labels and the second prediction of the plurality of second predictions matches the label of the plurality of labels (Fig 2, Q4 instances in red); and determining, with the at least one processor, a sixth group of samples of the plurality of samples for which the second prediction of the plurality of second predictions does not match the label of the plurality of labels and the first prediction of the plurality of first predictions matches the label of the plurality of labels (Fig 2, Q2 instances in red).
Regarding claims 9-11, 13, 15, 17-19, the inventive concept is essentially the same with the addition of a system, comprising: at least one processor or a computer program product comprising at least one non-transitory computer-readable medium including program instructions which are taught by Breckenridge (Col. 1 lines 41-43, can be embodied in a computer-implemented system that includes one or more computers and one or more data storage devices, Col. 22 lines 11-15, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and/or data to a programmable processor).
Claim(s) 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Breckenridge in view of Zhang as applied to claims 1 and 9 above, and further in view of Koch (US20180240041A1).
Regarding claim 4, Breckenridge in view of Zhang teaches the method of claim 1, further Breckenridge teaches wherein a first set of hyperparameters for a machine learning algorithm used to generate the first machine learning model is different than a second set of hyperparameters for a same machine learning algorithm used to generate the second machine learning model (Col. 9 lines 13-16, For a given training function, multiple different hyper parameter configurations can be applied to the training function, again generating multiple different trained predictive models).Breckenridge in view of Zhang fails to teach and wherein identifying the weak point in the second machine learning model further includes: determining a difference in hyperparameters between the first set of hyperparameters and the second set of hyperparameters; determining, based on the same first feature included in the first portion of samples and the difference in the hyperparameters, one or more hyperparameters; adjusting the second set of hyperparameters based on the selected one or more hyperparameters; and generating, using the adjusted second set of hyperparameters, an updated second machine learning model.
Koch teaches and wherein identifying the weak point in the second machine learning model further includes: determining a difference in hyperparameters between the first set of hyperparameters and the second set of hyperparameters; determining, based on the same first feature included in the first portion of samples and the difference in the hyperparameters, one or more hyperparameters (Abs, For each session of the plurality of sessions, training of a model of the model type is requested using a training dataset and the assigned hyperparameter configuration, scoring of the trained model using a validation dataset and the assigned hyperparameter configuration is requested to compute an objective function value, and the received objective function value and the assigned hyperparameter configuration are stored. A best hyperparameter configuration is identified based on an extreme value of the stored objective function values); adjusting the second set of hyperparameters based on the selected one or more hyperparameters; and generating, using the adjusted second set of hyperparameters, an updated second machine learning model (Par 0073, For example, the trained model output includes information to execute the model generated using the input dataset with the best hyperparameter configuration).
Koch is analogous to the claimed invention because it is in the field of predictive models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the hyperparameter adjusting system in Koc alongside the combination of Breckenridge and Zhang to find the ideal values for hyperparameters and tune predictive models for particular datasets (Par 0004).
Regarding claim 12, the inventive concept is essentially the same as claim 4 with the addition of a system, comprising: at least one processor which is taught by Breckenridge (Col. 1 lines 41-43, can be embodied in a computer-implemented system that includes one or more computers and one or more data storage devices).
Claim(s) 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Breckenridge in view of Zhang as applied to claims 5 and 13 above, and further in view of Ren (Squares: Supporting Interactive Performance Analysis for Multiclass Classifiers).
Regarding claim 6, Breckenridge in view of Zhang teaches the method of claim 5, but fails to teach wherein aligning the plurality of second prediction scores to the same scale as the plurality of first prediction scores includes: assigning each first prediction score of the plurality of prediction scores to a first bucket of a plurality of first buckets according to a value of that first prediction score; determining, for each first bucket, a rate of positive first predictions up to the value of the first prediction score assigned to that first bucket; assigning each second prediction score of the plurality of prediction scores to a second bucket of a plurality of second buckets according to a value of that second prediction score; determining, for each second bucket, a rate of positive second predictions up to the value of the second prediction score assigned to that second bucket; and determining, for each second prediction score, an aligned score aligned to the same scale as the plurality of first predictions scores, the value of the first prediction score assigned to the first bucket of the plurality of first buckets for which the rate of positive first predictions is a same rate as the rate of positive second predictions of the second bucket to which that prediction score is assigned.
Ren teaches wherein aligning the plurality of second prediction scores to the same scale as the plurality of first prediction scores includes: assigning each first prediction score of the plurality of prediction scores to a first bucket of a plurality of first buckets according to a value of that first prediction score (Fig 1, the top distribution, buckets are the bars at each prediction score); determining, for each first bucket, a rate of positive first predictions up to the value of the first prediction score assigned to that first bucket (Fig 1, the size of the bar at a given prediction score); assigning each second prediction score of the plurality of prediction scores to a second bucket of a plurality of second buckets according to a value of that second prediction score (Fig 1, the bottom distribution); determining, for each second bucket, a rate of positive second predictions up to the value of the second prediction score assigned to that second bucket (Fig 1, the size of the bar at a given prediction score); and determining, for each second prediction score, an aligned score aligned to the same scale as the plurality of first predictions scores (Fig 1, scores for both distributions are aligned), the value of the first prediction score assigned to the first bucket of the plurality of first buckets for which the rate of positive first predictions is a same rate as the rate of positive second predictions of the second bucket to which that prediction score is assigned (Fig 1, bars for some prediction scores like 0.0 are the same across both distributions).
Ren is analogous to the claimed invention because it is in the field of predictive modeling. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have used the squares method described in Ren alongside the combination of Breckenridge and Zhang to “support estimating common performance metrics while displaying instance-level distribution information necessary for helping practitioners prioritize efforts and access data” (Ren Abs).
Conclusion
Regarding claims 8, 16, and 20, the claims were searched for but are not rejected with prior art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NATNAEL A ASEGDEW whose telephone number is (571)270-0407. The examiner can normally be reached 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NATNAEL A ASEGDEW/ Examiner, Art Unit 2122
/KAKALI CHAKI/ Supervisory Patent Examiner, Art Unit 2122