DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is made non-final.
Claims 1-20 are pending. Claims 1, 11 and 17 are independent claims.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any statutory category. See MPEP 2106.03. Claim 1 recites: A method, comprising… Claim 1 is directed to a method (Step 1: YES).
Step 2A prong 1: Does the claim recite a judicial exception? Claim 1 recites: determining that there are deterministic relations in the training dataset between the plurality of input variables and the plurality of output variables (determining if there are relations in a dataset is a mental process or mathematical calculation)… determining whether the trained machine learning model has captured all of the deterministic relations within the training dataset (determining if a model captures all of the relations in a dataset can be performed mentally given model outputs and expected outputs, and/or can be performed by using mathematical calculations, i.e. statistical analysis). These steps can be performed mentally or are mathematical calculations (Step 2A prong 1: YES).
Step 2A prong 2: Does the claim recite additional elements? Do those additional elements, considered individually and in combination, integrate the judicial exception into a practical application? Claim 1 recites: retrieving, from a database, a training dataset containing a plurality of entries, each entry including a plurality of input variables and a plurality of output variables… in response to determining that there are deterministic relations, training a machine learning model to capture the deterministic relations within the training dataset… and retraining the trained machine learning model when it is determined that the trained machine learning model has not captured all of the deterministic relations; and returning the trained machine learning model. Retrieving training data entries from a database is insignificant extra-solution activity of data gathering. Both training and possibly retraining a machine learning model are recited at a high level of generality and lack details on how the model operates and how it solves a technical problem. Thus, the limitation represents no more than mere instructions to implement the abstract idea which is equivalent to adding the words “apply it” to the recited judicial exception. See MPEP 2106.05(f) (Step 2A prong 2: NO).
Step 2B: These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)), or provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)). These limitations, taken either alone or in combination, fail to provide an inventive concept (Step 2B: NO). Thus, the claim is not patent eligible.
Regarding claims 2-10, they recite limitations which further narrow the abstract idea by specifying more details of the mental and mathematical process that occurs (Claim 2, using the machine learning model by inputting data and receiving predicted outputs is still recited at a high level of generality, and the determining of residuals for output data or correlation between inputs/residuals are mathematical calculation; Claim 3, calculating that residuals and input variables are correlated is a mathematical calculation, and determining that the relations remain to be captured based on correlations is a mental process; Claim 4, calculating a difference between output variables and predicted outputs for ordinal data is a mathematical calculation; Claim 5, setting a residual to zero when data is nominal and predicted outputs match output variables is a mental process; Claim 6, setting a residual to a value representative of an output variable when there is a mismatch is a mental process; Claim 7, calculating a Pearson correlation coefficient is a mathematical calculation; Claim 8, calculating a mutual information value is a mathematical calculation, Claim 9, calculating a stochastic independence value is a mathematical calculation; Claim 10, the recited retraining steps are generically recited steps that still are interpreted to be mere instructions to apply the judicial exception – hyperparameter tuning, modifying a loss function, and modifying a model architecture can be performed on many if not all machine learning models).
Regarding claim 11, it is a system (i.e., an apparatus) that implements a method similar to the method of claim 1 and is rejected on the same grounds – see above.
Regarding claims 12-15 and 16, they recite similar limitations to claims 2-5 and 10, respectively, and are rejected on the same grounds.
Regarding claim 17, it is an apparatus that implements a method similar to the method of claim 1 and is rejected on the same grounds – see above.
Regarding claims 18-20, they recite similar limitations to claims 2-4, respectively, and are rejected on the same grounds.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 7, 8, 9, 11 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beregovaya et al. (US 20260057195 A1), herein Beregovaya in view of Volkovs et al. (US 20230244962 A1), herein Volkovs.
Regarding claim 1, Beregovaya teaches: A method, comprising: retrieving, from a database, a training dataset containing a plurality of entries, each entry including a plurality of input variables and a plurality of output variables (¶98, In step 802, the application program 202 (or another suitable logic running on the computing instance 104) trains (e.g., via Python or other libraries that employ machine learning algorithms) a set of machine learning models on a set of unstructured texts recited in a source language and a set of user engagement analytic parameters. As further explained below, this training occurs by a set of supervised machine learning algorithms (e.g., a classification algorithm, a linear regression algorithm). The set of unstructured texts is recited in the source language (e.g., English, Russian, Spanish, Arabic, Cantonese, Hebrew) and contains the set of linguistic features); determining that there are deterministic relations in the training dataset between the plurality of input variables and the plurality of output variables (¶130, performs feature selection by identifying importance of each feature in machine learning algorithms and removing (or ignoring) unnecessary features… Some techniques may include Wrapper methods (e.g., forward, backward, and stepwise selection), Filter methods (e.g., ANOVA, Pearson correlation, variance thresholding, Minimum-Redundancy-Maximum-Relevance (MRMR)), Embedded methods (e.g., Lasso, Ridge, Decision Tree), or other suitable techniques… For example, the engineer user profile may run a script…to find (a) all the pairs of highly correlated variables exceeding a correlation threshold such as 0.75 or (b) a mutual information score (MIS) of each feature to a target variable); in response to determining that there are deterministic relations, training a machine learning model to capture the deterministic relations within the training dataset (¶130, The remaining, loosely correlated features are more salient and relevant are therefore used in step 7, model training – here, the “loosely correlated” refers to input variables that are loosely correlated with each other, and not output variables)…
Beregovaya fails to teach: determining whether the trained machine learning model has captured all of the deterministic relations within the training dataset; and retraining the trained machine learning model when it is determined that the trained machine learning model has not captured all of the deterministic relations; and returning the trained machine learning model.
However, in the same field of endeavor, Volkovs teaches: determining whether the trained machine learning model has captured all of the deterministic relations within the training dataset; and retraining the trained machine learning model when it is determined that the trained machine learning model has not captured all of the deterministic relations; and returning the trained machine learning model (¶26, The importance scores for features at particular times for a data record (or for multiple data records) may be compared with these “expected” feature importance values. When the importance scores are consistent with the expected importance, a trained model may be considered validated and used in additional circumstances. When the importance scores are inconsistent with the expected importance, the trained model may be retrained).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to retrain models that do not fully capture deterministic relations within a training data set as disclosed by Volkovs in the method disclosed by Beregovaya to automate the retraining process and enforce constraints (¶26, module 140 may also be configured to automatically validate or retrain a model based on the importance scores determined by the model for particular predictions… the relative importance of a feature may be associated in some circumstances with human-designated or experimental (e.g., controlled scientific experiments) values).
Regarding claim 7, Beregovaya further teaches: The method as in claim 1, wherein determining that there are deterministic relations in the training dataset includes calculating a Pearson correlation coefficient between the plurality of input variables and the plurality of output variables of the training dataset (¶130, performs feature selection by identifying importance of each feature in machine learning algorithms and removing (or ignoring) unnecessary features… techniques may include Wrapper methods (e.g., forward, backward, and stepwise selection), Filter methods (e.g., ANOVA, Pearson correlation, variance thresholding… run a script… to find (a) all the pairs of highly correlated variables exceeding a correlation threshold such as 0.75).
Regarding claim 8, Beregovaya further teaches: The method as in claim 1, wherein determining that there are deterministic relations in the training dataset includes calculating a mutual information value between the plurality of input variables and the plurality of output variables in the training dataset (¶130, performs feature selection by identifying importance of each feature in machine learning algorithms and removing (or ignoring) unnecessary features… For example, the engineer user profile may run a script… to find (a) all the pairs of highly correlated variables exceeding a correlation threshold such as 0.75 or (b) a mutual information score (MIS) of each feature to a target variable).
Regarding claim 9, Beregovaya further teaches: The method as in claim 1, wherein determining that there are deterministic relations in the training dataset includes calculating a stochastic independence value between the plurality of input variables and the plurality of output variables in the training dataset (¶130, performs feature selection by identifying importance of each feature in machine learning algorithms and removing (or ignoring) unnecessary features… For example, the engineer user profile may run a script… to find (a) all the pairs of highly correlated variables exceeding a correlation threshold such as 0.75 or (b) a mutual information score (MIS) of each feature to a target variable – a low mutual information score is representative of an independent relationship, thus a mutual information score falls under the broadest reasonable interpretation (BRI) of a “stochastic independence value” between inputs and outputs).
Regarding claim 11, it is a system (i.e., an apparatus) that implements a method similar to the method of claim 1 and is rejected on the same grounds – see above.
Regarding claim 17, it is an apparatus that implements a method similar to the method of claim 1 and is rejected on the same grounds – see above.
Claim(s) 2, 3, 4, 12, 13, 14, 18, 19 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beregovaya in view of Volkovs as applied to claims 1, 11 and 17 above, and further in view of Rajarathinam et al. (US 20200202179 A1), herein Rajarathinam.
Regarding claim 2, Beregovaya further teaches: The method as in claim 1, wherein determining whether the trained machine learning model has captured all of the deterministic relations within the training dataset comprises: for each entry in the training dataset: providing the plurality of input variables as input to the trained machine learning model to generate a plurality of predicted outputs, each of the plurality of predicted outputs associated with one of the plurality of output variables from the training dataset (¶102, the set of machine learning models may be trained by the set of supervised machine learning algorithms based on mutual information between the set of linguistic features identified in the set of unstructured texts and the set of user engagement analytic parameters measured for the set of unstructured texts to correlate how the set of linguistic features identified in the set of unstructured texts is predicted to impact the set of user engagement analytic parameters)…
Beregovaya in view of Volkovs fails to explicitly teach: and generating a plurality of residuals, each residual generated by comparing one of the plurality of predicted outputs and its associated output variable; determining whether there is correlation between the plurality of input variables in the training dataset and the plurality of residuals; and determining that the trained machine learning model has not captured all of the deterministic relations when there is correlation between the plurality of input variables in the training dataset and the plurality of residuals.
However, in the same field of endeavor, Rajarathinam teaches: and generating a plurality of residuals, each residual generated by comparing one of the plurality of predicted outputs and its associated output variable (¶5, The logic circuitry may perform residual modeling on each model in the set during the monitoring period, to determine a list of input features that contribute to a residual for each model of the set to two or more models. The residual may comprise a difference between a result predicted by each model and an expected result); determining whether there is correlation between the plurality of input variables in the training dataset and the plurality of residuals (¶45, the monitor logic circuitry 1015 may compare a correlation value that results from a correlation of a residual with an input feature to a correlation threshold. If the correlation value exceeds the correlation threshold or otherwise indicates that the correlation is higher than the correlation indicated by the correlation threshold, the monitor logic circuitry 1015 may determine that the input feature contributes to the residual); and determining that the trained machine learning model has not captured all of the deterministic relations when there is correlation between the plurality of input variables in the training dataset and the plurality of residuals (¶44, correlation of the residuals with current and past tensors of data for input features of a good model should show that the input features from the monitoring period data do not correlate with the residuals. Thus, correlation between a residual and current or past tensors of input data during the monitoring period may indicate that the model does not properly use an input feature).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to evaluate models via residual analysis as disclosed by Rajarathinam in the method disclosed by Beregovaya in view of Volkovs to enable model evaluation and repair (¶19, Identification of features that contribute to the residuals can describe key features associated with the breakdown of models. The description of the key features associated with the breakdown of models may facilitate creation of a new model that accounts for these key features and/or identify features to target with corrective measures to improve or repair existing models).
Regarding claim 3, Beregovaya in view of Volkovs fails to teach: The method as in claim 2, wherein deterministic relations remain to be captured for an output variable when the residual associated with the output variable is correlated with at least one of the plurality of input variables.
However, in the same field of endeavor, Rajarathinam teaches: wherein deterministic relations remain to be captured for an output variable when the residual associated with the output variable is correlated with at least one of the plurality of input variables (¶44, correlation of the residuals with current and past tensors of data for input features of a good model should show that the input features from the monitoring period data do not correlate with the residuals. Thus, correlation between a residual and current or past tensors of input data during the monitoring period may indicate that the model does not properly use an input feature – Rajarathinam operates on a per-input feature basis – ¶45, the monitor logic circuitry 1015 may compare a correlation value that results from a correlation of a residual with an input feature to a correlation threshold).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to evaluate if deterministic relations still remain by using residual analysis on individual input features as disclosed by Rajarathinam in the method disclosed by Beregovaya in view of Volkovs to enable model evaluation and repair (¶19, Identification of features that contribute to the residuals can describe key features associated with the breakdown of models. The description of the key features associated with the breakdown of models may facilitate creation of a new model that accounts for these key features and/or identify features to target with corrective measures to improve or repair existing models).
Regarding claim 4, Beregovaya in view of Volkovs fails to teach: The method as in claim 2, wherein generating the plurality of residuals includes calculating the difference between an output variable and the predicted output associated with the output variable when the output variable is ordinal data.
However, in the same field of endeavor, Rajarathinam teaches: wherein generating the plurality of residuals includes calculating the difference between an output variable and the predicted output associated with the output variable when the output variable is ordinal data (Abstract, Logic may perform residual modeling on each model in the set during the monitoring period and may determine a list of input features that contribute to a residual of each model of the set. A residual comprises a difference between a predicted result and an expected result).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to calculate differences between predicted and actual output variables as disclosed by Rajarathinam in the method disclosed by Beregovaya in view of Volkovs to enable model evaluation and repair (¶19, Identification of features that contribute to the residuals can describe key features associated with the breakdown of models. The description of the key features associated with the breakdown of models may facilitate creation of a new model that accounts for these key features and/or identify features to target with corrective measures to improve or repair existing models).
Regarding claims 12, 13 and 14, they recite similar limitations to claims 2, 3 and 4 respectively and are rejected on the same grounds – see above.
Regarding claims 18, 19 and 20, they recite similar limitations to claims 2, 3 and 4 respectively and are rejected on the same grounds – see above.
Claim(s) 5, 6 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beregovaya in view of Volkovs and Rajarathinam as applied to claims 2 and 12 above, and further in view of Ciampiconi et al. (“A survey and taxonomy of loss functions in machine learning”, 2023), herein Ciampiconi.
Regarding claim 5, Beregovaya in view of Volkovs and Rajarathinam fails to teach: The method as in claim 2, wherein generating the plurality of residuals includes setting the residual to a value zero when output variable is nominal data and the predicted output associated with the output variable accurately predicts the output variable.
However, in the same field of endeavor, Ciampiconi teaches: wherein generating the plurality of residuals includes setting the residual to a value zero when output variable is nominal data and the predicted output associated with the output variable accurately predicts the output variable (pg. 12, 5.2, Zero-One loss, The basic and more intuitive margin based classification loss is the Zero-One loss. It assigns 1 to a misclassified observation and 0 to a correctly classified one).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to set residuals to a value zero when the predicted output matches the output variable as disclosed by Ciampiconi in the method disclosed by Beregovaya in view of Volkovs and Rajarathinam to create simple and intuitive loss values (pg. 12, 5.2, the most basic and intuitive one, the Zero-One loss).
Regarding claim 6, Beregovaya in view of Volkovs and Rajarathinam fails to teach: The method as in claim 2, wherein generating the plurality of residuals includes setting the residual to a value representative of the output variable when the output variable is nominal data and the predicted output associated with the output variable inaccurately predicting the output variable.
However, in the same field of endeavor, Ciampiconi teaches: wherein generating the plurality of residuals includes setting the residual to a value representative of the output variable when the output variable is nominal data and the predicted output associated with the output variable inaccurately predicting the output variable (pg. 12, 5.2, Zero-One loss, The basic and more intuitive margin based classification loss is the Zero-One loss. It assigns 1 to a misclassified observation and 0 to a correctly classified one – in this case, “1” represents a difficult output).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to set residuals to a value representative of an output variable when the predicted output does not match the output variable as disclosed by Ciampiconi in the method disclosed by Beregovaya in view of Volkovs and Rajarathinam to create simple and intuitive loss values (pg. 12, 5.2, the most basic and intuitive one, the Zero-One loss).
Regarding claim 15, it recites similar limitations to claim 5 and is rejected on the same grounds – see above
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beregovaya in view of Volkovs as applied to claim 1 above, and further in view of Durvasula et al. (US 20250139378 A1), herein Durvasula.
Regarding claim 10, Beregovaya in view of Volkovs fails to explicitly teach: The method as in claim 1, wherein retraining the trained machine learning model includes at least one of hyperparameter tuning, modifying the loss function, and modifying the model architecture.
However, in the same field of endeavor, Durvasula teaches: wherein retraining the trained machine learning model includes at least one of hyperparameter tuning, modifying the loss function, and modifying the model architecture (¶126, In some implementations, tuning the trained ML model can include retraining the ML model using additional training data, similar to block 530. In some implementations, tuning the trained ML model can include tuning at least one hyperparameter of the trained ML model. The tuned ML model can then be evaluated at operation 540 to determine whether it is ready for deployment).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to retrain a model with hyperparameter tuning as disclosed by Durvasula in the method disclosed by Beregovaya in view of Volkovs to improve model performance metrics (¶126, If the trained ML model is determined not to be ready for deployment (e.g., the at least one performance metric does not satisfy the threshold performance condition), then processing logic can tune the trained ML model to obtain a tuned ML model).
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Beregovaya in view of Volkovs and Rajarathinam as applied to claim 12 above, and further in view of Durvasula.
Regarding claim 16, Beregovaya in view of Volkovs and Rajarathinam fails to explicitly teach: The system of claim 12, wherein retraining the trained machine learning model includes at least one of hyperparameter tuning, modifying the loss function, and modifying the model architecture.
However, in the same field of endeavor, Durvasula teaches: wherein retraining the trained machine learning model includes at least one of hyperparameter tuning, modifying the loss function, and modifying the model architecture (¶126, In some implementations, tuning the trained ML model can include retraining the ML model using additional training data, similar to block 530. In some implementations, tuning the trained ML model can include tuning at least one hyperparameter of the trained ML model. The tuned ML model can then be evaluated at operation 540 to determine whether it is ready for deployment).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to retrain a model with hyperparameter tuning as disclosed by Durvasula in the method disclosed by Beregovaya in view of Volkovs and Rajarathinam to improve model performance metrics (¶126, If the trained ML model is determined not to be ready for deployment (e.g., the at least one performance metric does not satisfy the threshold performance condition), then processing logic can tune the trained ML model to obtain a tuned ML model).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Patel et al. (US 20220309391 A1), discusses determining trends in residuals that can indicate performance deficiencies.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HARRISON CHAN YOUNG KIM whose telephone number is (571)272-0713. The examiner can normally be reached Monday - Friday 10:00 am - 7:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CESAR PAULA can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HARRISON C KIM/ Examiner, Art Unit 2145
/CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145