Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This action is in response to the amendments filed 06/23/2026. Claims 1, 4, and 7 have been amended, claim 10 has been added. Claims 1, 4, 7, and 10 are currently pending.
Response to Arguments
Applicant’s arguments regarding the 101 rejection have been fully considered but they are not persuasive. Applicant argues on page 13 that “the claimed subject matter includes process steps executed using one or more hardware processors; thus the claim does not recite a mental process because the steps are not practically performed in the human mind”. Applicant argues similarly on page 15 that the labeling function generator and feature subset generator constitute specific technical components. Examiner respectfully notes that the hardware processors, labeling function generator, and feature subset generator have been interpreted as generic computer components, as no special purpose hardware or technical description of these components is claimed. Limitations implemented by these processors are interpreted under MPEP 2106.05(f) as judicial exceptions merely applied by generic computer components.
Applicant argues on page 13 that “the claimed steps as a whole are providing a technical solution of determining the adequate amount of labeled dataset required for training the one or more machine learning models”. Examiner respectfully notes that “determining the adequate amount of labeled dataset required for training” was interpreted as a mental step, as a person could determine a required amount of training data. Similarly, Applicant’s argument on pages 14-16, 18-19, and 23 that the claim limitations directed to generating labelling functions, the generative model, and the sparse matrix, performing real-time labelling of new dynamic data, and providing personalized recommendations achieve an improvement to the functioning of a computer. Examiner notes that the “generative model” is not claimed in such a way that requires it to be interpreted as a machine learning or artificial intelligence model, and that a person could mentally constructive a generative model and learn from observed noise from labelling functions. Applicant further argues on pages 19 and 21-22 that the claims enable a reduction in latency and a reduction in training time as improvements to the functioning of a computer. Examiner notes that these alleged improvements all rely on limitations that were interpreted as judicial exceptions, at most merely implemented by generic computer components. As per MPEP 2106.05(a), a judicial exception cannot provide the sole basis of an improvement to the functioning of a computer. Therefore, these arguments are considered unpersuasive.
Applicant argues on pages 17 and 22-23 that the limitation “wherein the iterative comparison automatically controls a training loop” achieves an improvement to the functioning of a computer by “controlling when training must stop or expand”. Examiner respectfully disagrees and notes that the limitation directed to the iterative comparison was interpreted as a mental step, and that a person could mentally decide to terminate training or continue invoking a potentially mental generative model as interpreted above, to continue mentally labelling unlabelled data.
Lastly, Applicant compares the claims to Example 39 of the Subject Matter Eligibility examples and the Desjardins memo to argue that the claims should be considered eligible. Examiner respectfully notes that the claim in Example 39 was found to not recite any judicial exceptions, but that following the PEG analysis as required in MPEP 2106, Applicant’s claims recite more than one limitation directed to a judicial exception. Applicant has not clearly shown an improvement similar to the Desjardins memo, since the limitations on which Applicant relies to show an improvement were interpreted as judicial exceptions as explained above and in the more detailed analysis below. The 101 rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended.
Applicant’s arguments regarding the prior art rejection have been fully considered but are moot because of the new ground(s) of rejection. Applicant argues that the Walters reference does not teach a principal component analysis method. Examiner agrees but notes that the Gopalan reference has been brought in to teach this limitation in combination with the previously cited Ratner and Walters references. Examiner notes that the Gopalan reference is also relied upon to teach the amended limitation directed to customer recommendations. The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 4, 7, and 10 are rejected under 35 U.S.C. 101. Claims 1 and 10 are directed to a method, claim 4 is directed to a system, and claim 7 is directed to a non-transitory machine-readable information storage medium; therefore, claims 1, 4, 7, and 10 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). However, claims 1, 4, 7, and 10 fall within the judicial exception of an abstract idea, specifically the abstract ideas of “Mental Processes” (including observation, evaluation, and opinion) and “Mathematical Concepts (including mathematical calculations and relationships)”.
Claim 1:
Claim 1 is directed to a method; therefore, the claim does fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Claim 1 recites the following abstract ideas:
extracting, via the one or more hardware processors, a plurality of feature subsets from the labelled dataset by processing the labelled dataset using principal component analysis (mental step directed to observation, evaluation – a person could identify, or extract, a plurality of feature subsets in their mind from an observed labelled dataset using principal component analysis, potentially assisted by pen and paper. Examiner notes that the broadest reasonable interpretation of using principal component analysis also includes a mathematical calculation. Using a generic hardware processor to perform this step is interpreted as mere instructions to apply an abstract idea using a generic computer component (see MPEP 2106.05(f));
automatically generating, via the one or more hardware processors, a plurality of labelling functions for the labelled dataset using the one or more trained machine learning models, wherein the labels are auto-generated for required amount of training data (mental step directed to observation, evaluation – a person could generate labelling functions for an observed dataset and generate labels for a required amount of a training dataset in their mind. As the claim does not recite any particular technical details related to the trained machine learning model or auto-generating labels, using a generic trained machine learning model to automatically generate labels is interpreted as mere instructions to apply an abstract idea using a generic computer component (see MPEP 2106.05(f));
executing, via the one or more hardware processors, the plurality of labelling functions for processing the unlabelled dataset to generate a sparse matrix, wherein the sparse matrix includes one or more rows and one or more columns, wherein each row accommodates unlabelled data and each column accommodates the labelling functions corresponding to the one or more machine learning models (mental step directed to observation, evaluation – a person could generate a sparse matrix with rows for unlabelled data and columns for labelling functions by executing labelling functions in their mind to process an observed unlabelled dataset in their mind, potentially assisted by pen and paper (see MPEP 2106.04(a)(2)(III). Using a generic hardware processor to perform this step is interpreted as mere instructions to apply an abstract idea using a generic computer component (see MPEP 2106.05(f));
constructing, via the one or more hardware processors, a generative model for the sparse matrix to label the unlabelled dataset, wherein the generative model learns from noise of the plurality of labelling functions and outputs labels for the unlabelled dataset (Examiner notes that Applicant’s disclosure does not define a “generative model” as a machine learning or artificial intelligence model. Given this interpretation, using a generative model to construct a sparse matrix to label an unlabelled dataset is interpreted as a mental step directed to observation, evaluation – a person could learn from observed noise of labelling function in their mind and create a generative model in their mind to construct a sparse matrix, label an unlabelled dataset, and output labels in their mind, potentially assisted by pen and paper. Using a generic hardware processor to perform this step is interpreted as mere instructions to apply an abstract idea using a generic computer component (see MPEP 2106.05(f));
and [training, the one or more machine learning models, via the one or more hardware processors, with required amount of labelled dataset for labelling the unlabelled dataset based on] a labelled data prediction threshold which is determined using a training data recommender technique (Examiner notes that the broadest reasonable interpretation of a “training data recommender technique” includes a mental step directed to observation, evaluation – as a person could determine, or recommend, a labelled data prediction threshold in their mind),
wherein the training data recommender technique is constructed with a parameter including the user predefined labelled data prediction threshold to determine an amount of labelled training data required for training the one or more machine learning models (mental step directed to observation, evaluation – a person could construct a training data recommender technique in their mind based on an mentally predetermined labelled data prediction threshold in order to determine an amount of required labelled training data in their mind), and the user predefined labelled data prediction threshold leads to reduction in a training time while performing or executing the one or more machine learning models (this limitation is interpreted as the intended use or necessary outcome of constructing a training data recommender technique to determine a required amount of training data and does not provide further patentable weight to this limitation (see MPEP 2103));
wherein the required amount of labelled dataset for training the one or more machine learning models using the training data recommender technique is determined by: determining a plurality of prediction accuracy metrics of the test data associated with the labelled dataset based on using the one or more machine learning models (mental step directed to observation, evaluation – a person could determine prediction accuracy metrics in their mind for observed test data associated with an observed labelled dataset based on the observed output of one or more used machine learning models and determine a required amount of labelled training data for one or more machine learning models in their mind using a mentally determined or observed recommendation technique);
computing, a selected labelled data, for each machine learning model based on the initial labelled dataset, and the reduction factor (mental step directed to observation, evaluation – a person could compute selected labelled data for a given machine learning model in their mind based on an observed initial labelled dataset and an observed reduction factor, as the claim does not actively require that a machine learning model must be used for computing the selected label data);
and determining the required amount of the labelled dataset for training the one or more machine learning models based on (i) the selected labelled data, (ii) the prediction accuracy metrics of the test data, and (iii) the labelled data prediction threshold (mental step directed to observation, evaluation – a person could determine a required amount of labelled data for training a machine model in their mind based on observed selected labelled data, observed prediction accuracy metrics, and an observed labelled data prediction threshold);
comparing iteratively, for each machine learning model, the difference between the prediction accuracy metrics with the user predefined labelled data prediction threshold to determine whether amount of labelled data is sufficient to train the one or more machine learning models and labelling unlabelled dataset using additional labelled dataset when the amount of labelled data is not sufficient to train the one or more machine learning models (mental step directed to observation, evaluation – a person could compare the difference between observed or mentally determined prediction accuracy metrics with a user predefined labelled data prediction threshold in their mind, determine whether an amount of labelled training data is sufficient to train a model in their mind, and label previously unlabelled data using a labelled dataset in their mind, potentially assisted by pen and paper, having mentally determined that the amount of labelled training data is insufficient);
wherein the iterative comparison automatically controls a training loop of the one or more machine learning models by: (i) terminating the training when the user defined labelled data prediction threshold is satisfied; and (ii) automatically invoking the generative model to label additional unlabelled data to generate the additional labelled dataset for the training when the user defined labelled data prediction threshold is not satisfied (mental step directed to evaluation, judgement – a person could decide in their mind to terminate a training process or generate additional labelled data based on a mental iterative comparison of a prediction accuracy metric to a predefined threshold. Wherein this limitation may be automatically performed is interpreted as merely implementing this step using generic computer components (see MPEP 2106.05(f)), thereby reducing training time (this limitation is interpreted as the intended use or necessary outcome of the iterative comparison and does not provide further patentable weight to this limitation (see MPEP 2103));
and using the trained one or more machine learning models to perform real-time labelling of new dynamic data based on the user predefined labelled data prediction threshold and provide personalized recommendations to customers (mental step directed to observation, evaluation – a person could perform real-time labelling of newly observed dynamic data in their mind based on a mentally predefined labelled data prediction threshold and provide personalized recommendations to customers in their mind, potentially assisted by pen and paper. Using a generic trained machine learning model to perform this step is interpreted as mere instructions to apply an abstract idea using a generic computer component (see MPEP 2106.05(f)),
wherein the processor implemented method is implemented with reduction in a latency for training the one or more machine learning models using the user predefined labelling data prediction threshold (this limitation is interpreted as the intended use or necessary outcome of implementing the real-time labelling and does not provide further patentable weight to this limitation (see MPEP 2103)).
Claim 1 recites the following additional elements:
receiving, by a labelling function generator, via one or more hardware processors, (i) an unlabelled dataset, and (ii) a labelled dataset comprising a training data and a test data, wherein labelling function generator comprises a feature subset generator and one or more machine learning models;
feeding, via the one or more hardware processors, the plurality of feature subsets extracted from the labelled dataset to the one or more machine learning models;
training, the one or more machine learning models, via the one or more hardware processors, with required amount of labelled dataset for labelling the unlabelled dataset based on a user predefined labelled data prediction threshold [which is determined using a training data recommender technique];
obtaining a plurality of labelled dataset threshold parameters comprising (i) an initial labelled dataset, (ii) the reduction factor, R (iii) the test data, and (iv) a labelled data prediction threshold;
reducing the labelled training dataset by a reduction factor R that results in reduced amount of training labelled dataset of the one or more machine learning models, via the one or more hardware processors, wherein training time for the labelled datasets is reduced;
and retraining each of the machine learning models with reduced amount of training labelled dataset, wherein amount of labelled dataset is decreased by the reduction factor R to obtain performance of area under the curve (AUC), via the one or more hardware processors.
The one or more hardware processors and the labelling function generator comprising a feature subset generator are interpreted as generic computer components, as specific technical components are not required in the claim for the “labelling function generator”.
Receiving an unlabelled dataset, and (ii) a labelled dataset comprising a training data and a test data, and feeding extracted feature subsets to a machine learning model are interpreted as transmitting and receiving data over a network.
As no particular machine learning model nor particular technical steps related to training a model are claimed, the machine learning models are interpreted as generic computer components and training and retraining a generic machine learning model with a required amount of labelled training data are interpreted as generic computer activity.
Obtaining a plurality of labelled dataset threshold parameters comprising (i) an initial labelled dataset, (ii) a reduction factor, (iii) the test data, and (iv) a labelled data prediction threshold is interpreted as an additional element directed to transmitting and receiving data over a network.
Reducing the labelled training dataset by a reduction factor R that results in reduced amount of training labelled dataset and reduced training time of the one or more machine learning models is interpreted as extra-solution activity directed to selecting a particular type of data to be manipulated.
Wherein the output of a given model may be reflected using an Area Under the Curve (AUC) metric is interpreted as displaying, or transmitting data over a network. These additional elements, when considered as a whole with the aforementioned abstract ideas, do not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea (see MPEP 2106.05(d) and MPEP 2106.05(g)).
Claim 4 is a system claim and its limitation is included in claim 1. The only difference is that claim 4 requires a system. Therefore, claim 4 is rejected for the same reasons as claim 1.
Claim 7 is a non-transitory machine-readable information storage medium claim and its limitation is included in claim 1. The only difference is that claim 7 requires a non-transitory machine-readable information storage medium. Therefore, claim 7 is rejected for the same reasons as claim 1.
Claim 10 recites wherein the labelled dataset is fully labelled and includes a separate portion marked as the test data and X % of available labelled data referred as a gold data is considered to automatically generate the plurality of labelling functions which are then executed over increasing portions of remaining (100-X) % data of the unlabelled dataset to generate increasing portions of labelled data as (X+d1) %, (X+d2) %, where d1 and d2 are arbitrarily chosen, and the one or more machine learning models are trained over the increasing portions of the labelled data (X+d1) %, (X+d2) % to measure accuracy metrics over the test data.
Wherein the labelled dataset is fully labelled and includes a test portion is interpreted as further description of the kind of data used in the mental steps as described in the abstract ideas in claim 10.
Wherein X% of available labelled data, or gold data, is considered to generate the labelling functions executed over increasing arbitrarily chosen portions of unlabelled data to label those portions is interpreted as a mental step directed to observation, evaluation – a person could consider some percentage of observed gold data in their mind, arbitrarily choose portions of observed unlabelled data in their mind, and mentally consider the gold when mentally executing labelling functions over the observed unlabelled data. Examiner notes that the broadest reasonable interpretation of measuring accuracy metrics over test data includes a mental step directed to observation, evaluation – a person could evaluate, or measure, in their mind accuracy metrics of an observed machine learning model using test data.
As no particular machine learning model nor particular technical steps related to training a model are claimed, the machine learning models are interpreted as generic computer components and wherein one or more generic machine learning models are trained over increasing portions of labelled data training is interpreted as generic computer activity which, when considered as a whole with the claimed abstract ideas, does not integrate those claimed abstract ideas into a practical application or amount to significantly more than those claimed abstract ideas (see MPEP 2106.05(h) and MPEP 2106.05(d)).
Viewed as a whole, these additional claim elements do not provide meaningful limitations to transform the abstract idea into a patent eligible application of the abstract idea such that the claims amount to significantly more than the abstract idea itself. Therefore, the claims are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4, 7, and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Ratner et al (“Snorkel: Rapid Training Data Creation with Weak Supervision”, herein Ratner) in view of Walters et al (US 20200302234 A1, herein Walters), in further view of Gopalan (US 20180285772 A1, herein Gopalan).
Regarding claim 1, Ratner teaches a processor implemented method for generating labelled dataset using a training data recommender technique, comprising: receiving, by a labelling function generator, via one or more hardware processors, (i) an unlabelled dataset, and (ii) a labelled dataset comprising a training data and a test data (section 2 para. 5 recites “Our goal is to learn a parameterized classification model hθ that, given a data point x ∈ X, predicts its label y ∈ Y, where the set of possible labels Y is discrete. For example, x might be a medical image, and y a label indicating normal versus abnormal. In the relation extraction examples we look at, we often refer to x as a candidate. In a traditional supervised learning setup, we would learn hθ by fitting it to a training set of labelled data points. However, in our setting, we assume that we only have access to unlabelled data for training. We do assume access to a small set of labelled data used during development, called the development set, and a blind, held-out labelled test set for evaluation” (i.e., receiving unlabelled data, labelled training, or development, data, and test data). Section 2 para. 6 recites “The user of Snorkel aims to generate training labels by providing a set of labelling functions, which are black-box functions, λ: → ∪ {∅}, that take in a data point and output a label where we use ∅ to denote that the labelling functions abstains. Given m unlabelled data points and n labelling functions, Snorkel applies the labelling functions over the unlabelled data to produce a matrix of labelling function outputs” (i.e., a labelling function generator which operates on labelled and unlabelled data)),
wherein labelling function generator comprises a feature subset generator and one or more machine learning models (section 2.1 para. 7-9 recite “Weak classifiers: Classifiers that are insufficient for our task—e.g., limited coverage, noisy, biased, and/or trained on a different dataset—can be used as labeling functions. Labeling function generators: One higher-level abstraction that we can build on top of labeling functions in Snorkel is labeling function generators, which generate multiple labeling functions from a single resource, such as crowdsourced labels and distant supervision from structured knowledge bases. Example 2.4: A challenge in traditional distant supervision is that different subsets of knowledge bases have different levels of accuracy and coverage. In our running example, we can use the Comparative Toxicogenomics Database (CTD)4 as distant supervision, separately modeling different subsets of it with separate labeling functions. For example, we might write one labeling function to label a candidate True if it occurs in the “Causes” subset, and another to label it False if it occurs in the “Treats” subset. We can write this using a labeling function generator” (i.e., the labelling function generator can generate subsets of features and weak classifier machine learning models));
automatically generating, via the one or more hardware processors, a plurality of labelling functions for the labelled dataset using [the] one or more trained machine learning models, wherein the labels are auto-generated for required amount of training dataset (section 2 para. 6 recites “The user of Snorkel aims to generate training labels by providing a set of labelling functions, which are black-box functions, λ: → ∪ {∅}, that take in a data point and output a label where we use ∅ to denote that the labelling functions abstains”. Section 2.1 para. 4 recites “Snorkel includes a library of declarative operators that encode the most common weak supervision function types. These functions capture a range of common forms of weak supervision, for example: Weak classifiers: Classifiers that are insufficient for our task – e.g., limited coverage, noisy, biased, and/or trained on a different dataset – can be used as labelling functions” (i.e., using trained machine learning models to automatically generate labelling functions and automatically outputting labels for the required amount of the training dataset, which can be used for a labelled data set like the one described in at least section 2 para. 5 of Ratner));
executing, via the one or more hardware processors, the plurality of labelling functions for processing the unlabelled dataset to generate a sparse matrix, wherein the sparse matrix includes one or more rows and one or more columns, wherein each row accommodates unlabelled data and each column accommodates the labelling functions corresponding to the one or more machine learning models (section 2 para. 6 recites “The user of Snorkel aims to generate training labels by providing a set of labelling functions, which are black-box functions, λ: → ∪ {∅}, that take in a data point and output a label where we use ∅ to denote that the labelling functions abstains. Given m unlabelled data points and n labelling functions, Snorkel applies the labelling functions over the unlabelled data to produce a matrix of labelling function outputs Λ ∈ ( ∪ {∅})m×n. The goal of the remaining Snorkel pipeline is to synthesize this label matrix Λ—which may contain overlapping and conflicting labels for each data point—into a single vector of probabilistic training labels Ỹ = (ỹ1, …, ỹm), where ỹi ∈ [0, 1]. These training labels can then be used to train a discriminative model” (i.e., executing labelling functions to create a matrix, and using a generative model to label the dataset contained in the matrix). Section 3.1.1 para. 1 recites “We start by considering the label density dΛ of the label matrix Λ, defined as the mean number of non-abstention labels per data point. In the low-density setting, sparsity of labels will mean that there is limited room for even an optimal weighting of the labelling functions to diverge much from the majority vote” (i.e., training data with few labels, or low label density would produce a sparse matrix of training data and labelling functions in the method described in section 2 para. 6));
constructing, via the one or more hardware processors, a generative model for the sparse matrix to label the unlabelled dataset, wherein the generative model learns from noise of the plurality of labelling functions and outputs labels for the unlabelled dataset (section 2 para. 3 recites “Snorkel automatically learns a generative model over the labeling functions, which allows it to estimate their accuracies and correlations. This step uses no ground-truth data, learning instead from the agreements and disagreements of the labeling functions”. Section 2.3 para. 1 recites “We train a discriminative model hϴ on our probabilistic labels Y by minimizing a noise-aware variant of the loss, i.e., the expected loss with respect to Y” (i.e., constructing a generative model from the labelling functions that can learn from noise)),
and training, the one or more machine learning models, via the one or more hardware processors, with required amount of labelled dataset for labelling the unlabelled dataset based on a user predefined labelled data prediction threshold which is determined using a training data recommender technique, wherein the training data recommender technique is constructed with a parameter including the user predefined labelled data prediction threshold to determine an amount of labelled training data required for training the one or more machine learning models, (section 2 para. 6 recites “The goal of the remaining Snorkel pipeline is to synthesize this label matrix Λ—which may contain overlapping and conflicting labels for each data point—into a single vector of probabilistic training labels Ỹ = (ỹ1, …, ỹm), where ỹi ∈ [0, 1]. These training labels can then be used to train a discriminative model” (i.e., training a model on labelled data). Section 3.2.1 para. 1 recites “structure learning methods, whether pseudolikelihood or likelihood-based, crucially depend on a selection threshold ε for deciding which dependencies to add to the generative model” (i.e., a labelled data prediction threshold parameter). Section 3.2.1 para. 2 recites “Figure 5, middle, shows the effect of varying ε for the CDR task. Predictive performance improves as ε decreases until the model overfits”. Section 3.2.2. para. 1 recites “We therefore want to choose ε before other hyperparameters, without performing any parameter estimation. We propose using the number of correlations selected at each value of ε as an inexpensive indicator. The dashed lines in Figure 5 show that as ε decreases, the number of selected correlations follows a pattern. Generally, the number of correlations grows slowly at first, then hits an “elbow point” beyond which the number explodes, which fits the assumption that the correlation structure is sparse” (i.e., using a required amount of training data based on a threshold determined by a recommendation technique)),
and the user predefined labelled data prediction threshold leads to reduction in a training time while performing or executing the one or more machine learning models (this limitation is interpreted as the intended use or necessary outcome of constructing a training data recommender technique to determine a required amount of training data and does not provide further patentable weight to this limitation (see MPEP 2103));
wherein the required amount of labelled dataset for training the one or more machine learning models using the training data recommender technique is determined by: obtaining a plurality of labelled dataset threshold parameters comprising (i) an initial labelled dataset, (ii) [the reduction factor R], (iii) the test data, and (iv) the user predefined labelled data prediction threshold (Ratner section 2 para. 5 recites “We do assume access to a small set of labelled data used during development, called the development set, and a blind, held-out labelled test set for evaluation”. Ratner section 3.2.1 para. 1 recites “structure learning methods, whether pseudolikelihood or likelihood-based, crucially depend on a selection threshold ε for deciding which dependencies to add to the generative model” (i.e., obtaining an initial labelled dataset, test data, and a labelled data prediction threshold));
and labelling unlabelled datasets using additional labelled dataset when the amount of labelled data is not sufficient to train the one or more machine learning models (section 2 para. 6 recites “The user of Snorkel aims to generate training labels by providing a set of labelling functions, which are black-box functions, λ: → ∪ {∅}, that take in a data point and output a label where we use ∅ to denote that the labelling functions abstains. Given m unlabelled data points and n labelling functions, Snorkel applies the labelling functions over the unlabelled data to produce a matrix of labelling function outputs”. Section 3.2.2 para. 1 recites “Based on our observations, we seek to automatically choose a value of ε that trades off between predictive performance and computational cost using the labelling functions’ outputs alone. We therefore want to choose ε before other hyperparameters, without performing any parameter estimation. We propose using the number of correlations selected at each value of ε as an inexpensive indicator. The dashed lines in Figure 5 show that as ε decreases, the number of selected correlations follows a pattern. Generally, the number of correlations grows slowly at first, then hits an “elbow point” beyond which the number explodes, which fits the assumption that the correlation structure is sparse. In all three cases, setting ε to this elbow point is a safe tradeoff between predictive performance and computational cost” (i.e., labelling more unlabelled data when a training outcome is determined to be insufficient));
and obtaining the performance of the model using area under the curve (AUC) (section 4 and tables 4-5 recite “We evaluate Snorkel by drawing on deployments developed in collaboration with users. We report on two real-world deployments and four tasks on open-source data sets representative of other deployments. We see that by writing tens of labelling functions, we were able to approach or match results using hand-labelled training data which took weeks or months to assemble, coming within 2.11% of the F1 score of hand supervision on relation extraction tasks and an average 5.08% accuracy or AUC on cross-modal tasks, for an average 3.60% across all tasks” (i.e., using an area under the curve metric to reflect the model performance. Examiner notes that this limitation is interpreted in light of the 112(a) rejection of claim 1)).
However, while Ratner describes a context hierarchy where candidate data points x represent context types, or feature subsets, of labelled training data (see at least section 2 para. 8), Ratner does not explicitly teach extracting, via the one or more hardware processors, a plurality of feature subsets from the [labelled] dataset; feeding, via the one or more hardware processors, the plurality of feature subsets extracted from the [labelled] dataset to a one or more machine learning models; determining a plurality of prediction accuracy metrics of the test data associated with the [labelled] dataset based on using the one or more machine learning models; computing a selected [labelled] data for each machine learning model based on the initial [labelled] dataset and the reduction factor; and determining the required amount of the [labelled] dataset for training the one or more machine learning models based on (i) the selected [labelled] data, (ii) the prediction accuracy metrics of the test data, and (iii) the [labelled] data prediction threshold; comparing iteratively, for each machine learning model, the difference between the prediction accuracy metrics with the user predefined [labelled] data prediction threshold to determine whether the amount of [labelled] data is sufficient to train the one or more machine learning models; reducing the [labelled] training dataset by a reduction factor R that results in reduced amount of training [labelled] dataset of the one or more machine learning models, via the one or more hardware processors, wherein training time for the [labelled] datasets is reduced; and retraining each of the machine learning models with reduced amount of training [labelled] dataset, wherein the amount of [labelled] dataset is decreased by the reduction factor R [to obtain performance of area under the curve (AUC)], via the one or more hardware processors.
Walters teaches extracting, via the one or more hardware processors, a plurality of feature subsets from the [labelled] dataset by processing the [labelled] dataset using a principal component analysis (para. [0038] recites “Data profiler 110 may include one or more computing systems configured to perform operations consistent with extracting features from a dataset and/or comparing data profiles to assess similarity between datasets” (i.e., extracting features from a dataset));
feeding, via the one or more hardware processors, the plurality of feature subsets extracted from the [labelled] dataset to the one or more machine learning models (para. [0043] recites “Model generator 120 may generate models selected from multiple types of models. The models may include convolutional neural networks (CNN), recurrent neural networks (RNN), multilayer perceptrons (MLP), Random Forest (RP), and/or linear regressions. For example, identification models may be neural networks that determine attributes in a dataset based on extracted parameters” (i.e., extracted features are input to machine learning models));
determining a plurality of prediction accuracy metrics of the test data associated with the [labelled] dataset based on using the one or more machine learning models (Walters para. [0106] recites “In step 714 optimization system 105 may assess the accuracy of each one of the secondary models”. Walters para. [0108] recites “In step 716, based on the secondary models and their estimated accuracy, optimization system 105 may identify a minimum viable model from each sequence of secondary models. The minimum viable model may include the model that uses the smallest training dataset, identified from the progressively smaller datasets, and still achieves the desired accuracy threshold as evaluated in step 718” (i.e., determining prediction accuracy metrics for the machine learning models));
computing a selected [labelled] data for each machine learning model based on the initial [labelled] dataset and the reduction factor (Walters para. [0104] recites “In step 710, to search for minimum data requirements, optimization system 105 may generate a sequence of secondary models for each one of the primary models. The models in the sequence of secondary models may be trained with progressively reduced training datasets to figure out a minimum number of samples that is necessary to recreate (or approximate) the predictability and/or classification accuracy of the primary model. In some embodiments, primary models may be retrained using progressively less data but having the same hyper-parameters as the first user model. The reduction in the sample size may be linear (i.e., reduce by one million samples for every iteration) and/or exponential (i.e., reduce by a factor of 2 in every iteration)” (i.e., computing the size of a training dataset based on the initial size and a reduction factor));
and determining the required amount of the [labelled] dataset for training the one or more machine learning models based on (i) the selected [labelled] data, (ii) the prediction accuracy metrics of the test data, and (iii) the [labelled] data prediction threshold (Walters para. [0108] recites “In step 716, based on the secondary models and their estimated accuracy, optimization system 105 may identify a minimum viable model from each sequence of secondary models. The minimum viable model may include the model that uses the smallest training dataset, identified from the progressively smaller datasets, and still achieves the desired accuracy threshold as evaluated in step 718” (i.e., determining the required amount of training data based on the accuracy metrics and a predefined threshold));
comparing iteratively, for each machine learning model, the difference between the prediction accuracy metrics with the user predefined [labelled] data prediction threshold to determine whether the amount of [labelled] data is sufficient to train the one or more machine learning models (para. [0045] recites “Hyper-parameter optimizer 130 may include one or more computing systems that performs iterations in models to tune hyper-parameters”. Para. [0104] recites “In step 710, to search for minimum data requirements, optimization system 105 may generate a sequence of secondary models for each one of the primary models. The models in the sequence of secondary models may be trained with progressively reduced training datasets to figure out a minimum number of samples that is necessary to recreate (or approximate) the predictability and/or classification accuracy of the primary model. The reduction in the sample size may be linear (i.e., reduce by one million samples for every iteration) and/or exponential (i.e., reduce by a factor of 2 in every iteration)”. Para. [0106] recites “In step 714 optimization system 105 may assess the accuracy of each one of the secondary models”. Para. [0108] recites “In step 716, based on the secondary models and their estimated accuracy, optimization system 105 may identify a minimum viable model from each sequence of secondary models. The minimum viable model may include the model that uses the smallest training dataset, identified from the progressively smaller datasets, and still achieves the desired accuracy threshold as evaluated in step 718”. Para. [0122] recites “The accuracy threshold may be defined by a user” (i.e., iteratively comparing prediction accuracy to a threshold to determine a minimum required, or sufficient amount of training data)),
wherein the iterative comparison automatically controls a training loop of the one or more machine learning models by: (i) terminating the training when the user defined labelled data prediction threshold is satisfied; and (ii) automatically invoking the generative model to label additional unlabelled data to generate the additional labelled dataset for the training when the user defined labelled prediction threshold is not satisfied (section 3.2.2 para. 2 recites “On the large number of labeling functions in the Spouses task, structure learning for 25 values ofεtakes 14 minutes. On CDR, with a smaller number of labeling functions, it takes 30 seconds. Further, if the search is started at a low value of ε and increased, it can often be terminated early, when the number of selected correlations reaches a low value. Selecting the elbow point itself is straightforward. We use the point with greatest absolute difference from its neighbors, but more sophisticated schemes can also be applied” (i.e., terminating training when a threshold is met or continuing an iterative training loop to generate labelled data otherwise)),
thereby reducing training time (this limitation is interpreted as the intended use or necessary outcome of the iterative comparison and does not provide further patentable weight to this limitation (see MPEP 2103));
reducing the [labelled] training dataset by a reduction factor R that results in reduced amount of training [labelled] dataset of the one or more machine learning models, via the one or more hardware processors, wherein the training time for [labelled] datasets is reduced; retraining each of the machine learning models with reduced amount of training [labelled] dataset, wherein amount of [labelled] dataset is decreased by the reduction factor R [to obtain performance of area under the curve (AUC)], via the one or more hardware processors (para. [0104] recites “The models in the sequence of secondary models may be trained with progressively reduced training datasets to figure out a minimum number of samples that is necessary to recreate (or approximate) the predictability and/or classification accuracy of the primary model. In some embodiments, primary models may be retrained using progressively less data but having the same hyper-parameters as the first user model. If this CNN model was initially trained with fifty million records, the secondary models may reduce the number of records in the training data for every element of the sequence. The reduction in the sample size may be linear (i.e., reduce by one million samples for every iteration) and/or exponential (i.e., reduce by a factor of 2 in every iteration). Alternatively, the sequence may be generated using search algorithms such as a Fibonacci search, a jump search, an interpolation search, a recursive function, or a binary search to determine a testing sample size for training” (i.e., retraining a model with a dataset that was reduced by a reduction factor. Examiner notes that wherein the training time is reduced is interpreted as the necessary outcome of reducing the amount of labelled training data used and does not provide further patentable weight to this limitation));
wherein the processor implemented method is implemented with reduction in a latency for training the one or more machine learning models using the user predefined labelling data prediction threshold (this limitation is interpreted as the intended use or necessary outcome of implementing the real-time labelling and does not provide further patentable weight to this limitation (see MPEP 2103)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by using the feature extraction methods from Walters to supplement the preprocessing techniques supported in Ratner. Ratner and Walters are both directed to methods of determining the appropriate size of training data needed to train a model, wherein Ratner uses the selection threshold ε and Walters uses the method described in at least figure 7, which includes hyperparameter tuning and prediction accuracy assessment. As Ratner notes in Appendix C paragraph 7, the Snorkel framework “supports automatically defining candidates using their named-entity recognition features”, one of ordinary skill in the art would be motivated to supplement these recognition features with the feature extraction method from Walters.
However, the combination of Ratner and Walters does not explicitly teach extracting, [via the one or more hardware processors], a plurality of feature subsets from the [labelled] dataset by processing the [labelled] dataset using a principal component analysis, or using the trained one or more machine learning models to perform real-time labelling of new dynamic data [based on the user predefined labelled data prediction threshold] and provide personalized recommendations to customers.
Gopalan teaches extracting, [via the one or more hardware processors], a plurality of feature subsets from the [labelled] dataset by processing the [labelled] dataset using a principal component analysis (para. [0043] recites “the machine learning model may utilize a feature space comprising a reduced feature space that may be determined by first performing feature selection at step 210. For example, a feature selection process may include reducing the number relevant features to those which are most useful in a classification task. Alternatively, or in addition, a principal component analysis (PCA) may be applied to the training data set” (i.e., extracting features from a dataset using principal component analysis)),
and using the trained one or more machine learning models to perform real-time labelling of new dynamic data [based on the user predefined labelled data prediction threshold] and provide personalized recommendations to customers (para. [0024] recites “centralized system components may also include devices and/or servers for implementing machine learning models in accordance with the present disclosure for various services such as: traffic analysis, traffic shaping, firewall functions, malware detection, intrusion detection, customer churn prediction, content recommendation generation, and so forth”. Para. [0033] recites “At step 230, the processor processes a stream of new data to determine a likelihood of the new data from the data distribution that is computed at step 220” (i.e., determining a likelihood, or classification label, for new data, such as a content recommendation to a customer)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by adapting the feature extraction method from Walters (which modifies Ratner) to utilize the principal component analysis method from Gopalan. Walters teaches in at least paragraph [0039] that statistical analysis can be used to process objects in the dataset. One of ordinary skill in the art would recognize that Gopalan teaches a related statistical analysis in its principal component analysis that could be utilized by Walters to achieve a predictable result.
Claim 4 is a system claim and its limitation is included in claim 1. The only difference is that claim 4 requires a system. Therefore, claim 4 is rejected for the same reasons as claim 1.
Claim 7 is a non-transitory machine-readable information storage medium claim and its limitation is included in claim 1. The only difference is that claim 7 requires a non-transitory machine-readable information storage medium. Therefore, claim 7 is rejected for the same reasons as claim 1.
Regarding claim 10, the combination of Ratner, Walters, and Gopalan teaches the processor implemented method as claimed in claim 1, wherein the labelled dataset is fully labelled and includes a separate portion marked as the test data and X % of available labelled data referred as a gold data is considered to automatically generate the plurality of labelling functions (Ratner section 2 para. 5 recites “We do assume access to a small set of labeled data used during development, called the development set, and a blind, held-out labeled test set for evaluation. These sets can be orders of magnitudes smaller than a training set, making them economical to obtain” (i.e., fully labelled, or gold, data is used as test data is used as a development set when training the labelling functions)) which are then executed over increasing portions of remaining (100-X) % data of the unlabelled dataset to generate increasing portions of labelled data as (X+d1) %, (X+d2) %, where d1 and d2 are arbitrarily chosen, (Ratner section 2.3 para. 1 recites “A formal analysis shows that as we increase the amount of unlabeled data, the generalization error of discriminative models trained with Snorkel will decrease at the same asymptotic rate as traditional supervised learning models do with additional hand-labeled data, allowing us to increase predictive performance by adding more unlabeled data” (i.e., executing the labelling functions on increasing arbitrary portions of unlabelled data to generate labels for that unlabelled data)),
and the one or more machine learning models are trained over the increasing portions of the labelled data (X+d1) %, (X+d2) % to measure accuracy metrics over the test data (Ratner section 2.2 para. 3 recites “We then encode the generative model pw(Λ, Y) using three factor types, representing the labeling propensity, accuracy, and pairwise correlations of labeling functions”. Ratner section 5 para. 2 recites “Snorkel is distinguished from these approaches because its generative model supports a wide range of weak supervision sources, and it learns the accuracies and correlation structure among weak supervision sources without ground truth data” (i.e., measuring an accuracy metric measured from training the weak classifier machine learning models on the labelled data from the labelling functions)).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20190147297 A1 (Rogers et al) teaches a visualization system for a weakly supervised machine learning model that generates labels for data and utilizes principal component analysis to process the feature matrix of the training data.
US 11620574 B2 (Parameswaran et al) teaches a method for re-computing a machine learning or data processing workflow in response to adjustments to the workflow parameters, in order to assess the benefit of such adjustments to satisfy accuracy or other constraints.
“Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise” (Hendrycks et al) teaches a method utilizing trusted examples in a data-efficient manner to mitigate the effects of label noise on deep neural network classifiers.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEAH M FEITL whose telephone number is (571) 272-8350. The examiner can normally be reached on M-F 0900-1700 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/L.M.F./ Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147