DETAILED ACTION
Notice of AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Regarding French Patent App. No. FR2303291 (filed 4/3/2023), receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement submitted on 4/1/2024 has been considered.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Step 1 of the Alice/Mayo framework, Claims 1-10 and 18-20 are directed to a method (a process), Claim 11 is directed to a non-transitory memory (an article of manufacture), and Claims 12-17 are directed to a computing system (a machine), which each fall within one of the four statutory categories of inventions.
Regarding Claim 1
Step 2A, prong 1 (Is the claim directed to a law of nature, a natural phenomenon or an abstract idea).
Claim 1 recites the following mental processes, that in each case under the broadest reasonable interpretation, covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components (e.g., “computer-implemented”).
A ... method for searching for an optimal hyperparameter combination for defining a machine learning model, the method comprising (under the broadest reasonable interpretation, a human can mentally search for an optimal set of hyperparameters for a machine learning model, such as for a simple neural network having 1 input node, 1 output node, and 1 hidden layer/node, determining a simple set of optimized learning rate, batch size, etc., based on experimentation)
calculate a performance score associated with the hyperparameter combination tested from test data, the optimal hyperparameter combination corresponding to the hyperparameter combination having obtained the best performance score among the hyperparameter combinations tested (under the broadest reasonable interpretation, a human can mentally calculate performance scores associated with each test run, where the highest performance score corresponds to the optimal hyperparameter combination)
defining a weighting coefficient for adjusting an amount of training data used for the training phase, the weighting coefficient being dynamically adapted during different tests of the hyperparameter combinations (under the broadest reasonable interpretation, a human can mentally define such a weighting coefficient, such as by using the mathematical equation set forth at para. 0084 of the instant specification)
Step 2A, prong 2 (Does the claim recite additional elements that integrate the judicial exception into a practical application?).
The judicial exception is not integrated into a practical application.
Regarding the “computer-implemented” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a computer. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a computer). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)).
Regarding the “performing a plurality of tests of hyperparameter combination, each test of hyperparameter combination including a training phase and a test phase, wherein the training phase is adapted to train the machine learning model from training data and the test phase is adapted to ...” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of training and testing a generic machine learning model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (training and testing a generic machine learning model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)).
Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
Step 2B (Does the claim recite additional elements that amount to significantly more than the judicial exception?)
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding the “computer-implemented” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)).
Regarding the “performing a plurality of tests of hyperparameter combination, each test of hyperparameter combination including a training phase and a test phase, wherein the training phase is adapted to train the machine learning model from training data and the test phase is adapted to ...” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)).
Accordingly, at Step 2B after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application.
Regarding Claim 2
Step 2A, Prong 1
wherein the weighting coefficient is initialized to an initial weighting coefficient. (under the broadest reasonable interpretation, a human can mentally initialize a weighting coefficient to an initial value, such as 1)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 3
Step 2A, Prong 1
wherein the initial weighting coefficient is less than or equal to 1%. (under the broadest reasonable interpretation, a human can mentally initiate the weighting coefficient at 1%)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 4
Step 2A, Prong 1
wherein the weighting coefficient is updated for each test of hyperparameter combination. (under the broadest reasonable interpretation, a human can mentally update the weighting coefficient for each test of the hyperparameter combination)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 5
Step 2A, Prong 1
wherein updating the weighting coefficient to be used for a given test of hyperparameter combination comprises calculating a new weighting coefficient from an old weighting coefficient used during the test of hyperparameter combination directly preceding the given test of hyperparameter combination. (under the broadest reasonable interpretation, a human can mentally calculate a new weighting coefficient directly using the previous coefficient, such as by using the mathematical formula in para. 0084 of the instant specification)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 6
Step 2A, Prong 1
wherein the new weighting coefficient is calculated by a formula k*α, where α is the old weighting coefficient and k is a coefficient greater than 1. (under the broadest reasonable interpretation, a human can mentally calculate a new weighting coefficient directly using the previous coefficient, such as by using the mathematical formula claimed herein)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 7
Step 2A, Prong 1
further comprising defining a dynamically defined best weighting coefficient, this best weighting coefficient corresponding to the weighting coefficient used for the training phase of the test of the hyperparameter combination having obtained the best performance score among the hyperparameter combinations already tested. (under the broadest reasonable interpretation, a human can mentally identify the weighting coefficient corresponding to the best hyperparameter combination as the best weighting coefficient)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 8
Step 2A, Prong 1
wherein updating the weighting coefficient comprises comparing a new weighting coefficient with the value 100% and with the value w*A, where A is the best weighting coefficient defined and w is a coefficient greater than 1, the weighting coefficient being updated to the value of the new weighting coefficient if the new weighting coefficient calculated is less than or equal to the value 100% or to the value w*A, or updated to the value of an initial weighting coefficient otherwise. (under the broadest reasonable interpretation, a human can mentally update the weighting coefficient using the mathematical relationships set forth in this limitation)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 9
Step 2A, Prong 2
Regarding the “training a machine learning model defined by the optimal combination of hyperparameter with all the training data” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of generic machine learning model training. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (generic machine learning model training). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)).
Step 2B
Regarding the “training a machine learning model defined by the optimal combination of hyperparameter with all the training data” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)).
Regarding Claim 10
Step 2A, Prong 2
Regarding the “for each machine learning model defined by a combination of hyperparameters having made it possible to obtain a better performance score among the combinations of hyperparameters already tested, training of this model with all the data each time a better performance score is obtained” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of generic machine learning model training. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (generic machine learning model training). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)).
Step 2B
Regarding the “for each machine learning model defined by a combination of hyperparameters having made it possible to obtain a better performance score among the combinations of hyperparameters already tested, training of this model with all the data each time a better performance score is obtained” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)).
Regarding Claim 11
Step 2A, Prong 1
Claim 11 recites a non-transitory memory that corresponds to the method of claim 1, and therefore the analysis under Step 2A, Prong 1 with respect to claim 1 also applies to this claim 11. While claim 11 recites additional generic computing components (“non-transitory memory”, “computer”, and “computer program”), such additional generic computing components do not change the analysis under Step 2A, Prong 1.
Step 2A, Prong 2
Claim 11 recites a non-transitory memory that corresponds to the method of claim 1, and therefore the analysis under Step 2A, Prong 2 with respect to claim 1 also applies to this claim 11. While claim 11 recites additional generic computing components (“non-transitory memory”, “computer”, and “computer program”), such additional generic computing components do not change the analysis under Step 2A, Prong 2. (See MPEP 2106.05(f)).
Step 2B
Claim 11 recites a non-transitory memory that corresponds to the method of claim 1, and therefore the analysis under Step 2B with respect to claim 1 also applies to this claim 11. While claim 11 recites additional generic computing components (“non-transitory memory”, “computer”, and “computer program”), such additional generic computing components do not change the analysis under Step 2B. (See MPEP 2106.05(f)).
Regarding Claim 12
Step 2A, Prong 1
Claim 12 recites a computing system that corresponds to the method of claim 1, and therefore the analysis under Step 2A, Prong 1 with respect to claim 1 also applies to this claim 12. While claim 12 recites additional generic computing components (“non-transitory memory”, “processing unit”, “computer”, and “computer program”), such additional generic computing components do not change the analysis under Step 2A, Prong 1.
Step 2A, Prong 2
Claim 12 recites a computing system that corresponds to the method of claim 1, and therefore the analysis under Step 2A, Prong 2 with respect to claim 1 also applies to this claim 12. While claim 12 recites additional generic computing components (“non-transitory memory”, “processing unit”, “computer”, and “computer program”), such additional generic computing components do not change the analysis under Step 2A, Prong 2. (See MPEP 2106.05(f)).
Step 2B
Claim 12 recites a computing system that corresponds to the method of claim 1, and therefore the analysis under Step 2B with respect to claim 1 also applies to this claim 12. While claim 12 recites additional generic computing components (“non-transitory memory”, “processing unit”, “computer”, and “computer program”), such additional generic computing components do not change the analysis under Step 2B. (See MPEP 2106.05(f)).
Claims 13-17 depend from claim 12 and correspond to the methods of claims 2-6, and are therefore rejected for the same reasons explained above with respect to claim 12 and claims 2-6, respectively.
Regarding Claim 18
Step 2A, prong 1 (Is the claim directed to a law of nature, a natural phenomenon or an abstract idea).
Claim 18 recites the following mental processes, that in each case under the broadest reasonable interpretation, covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components (e.g., “hardware processor”, “machine-learning model”).
A ... method for searching for an optimal hyperparameter combination for defining an automatic learning model, the method comprising: (under the broadest reasonable interpretation, a human can mentally search for an optimal set of hyperparameters for a machine learning model, such as for a simple neural network having 1 input node, 1 output node, and 1 hidden layer/node, determining a simple set of optimized learning rate, batch size, etc., based on experimentation)
initializing a weighting coefficient; (under the broadest reasonable interpretation, a human can mentally initialize a weighting coefficient to an initial value, such as 1)
evaluating a performance of a hyperparameter combination, the evaluating being performed in a training phase using a portion of the training data based on the weighting coefficient and a test phase using the test data; (under the broadest reasonable interpretation, a human can mentally evaluate how well a hyperparameter combination performed, such as by mentally assigning a score from 1-100 based on the human’s mental evaluation of the results)
calculating a performance score associated with the hyperparameter combination; (under the broadest reasonable interpretation, a human can mentally evaluate how well a hyperparameter combination performed, such as by mentally calculating a score from 1-100 based on the human’s mental evaluation of the results)
calculating a new weighting coefficient that is greater than the initial weighting coefficient; (under the broadest reasonable interpretation, a human can mentally calculate a new weighting coefficient, making sure that the new coefficient is greater than the initial weighting coefficient)
repeating the evaluating for a new hyperparameter combination, the repeated evaluating performed with a portion of the training data based on the new weighting coefficient and the test data; (under the broadest reasonable interpretation, a human can mentally repeat the evaluation of a new hyperparameter combination)
calculating a new performance score associated with the hyperparameter combination; and (under the broadest reasonable interpretation, a human can mentally evaluate how well a hyperparameter combination performed, such as by mentally calculating a score from 1-100 based on the human’s mental evaluation of the results)
comparing the performance score with the new performance score. (under the broadest reasonable interpretation, a human can compare performance scores to see which one is higher)
Step 2A, prong 2 (Does the claim recite additional elements that integrate the judicial exception into a practical application?).
The judicial exception is not integrated into a practical application.
Regarding the “computer-implemented” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a computer. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a computer). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)).
Regarding the “receiving training data; receiving test data” limitation, such additional element of a data gathering step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. pre-solution activity of gathering data for use in the claimed process (see MPEP 2106.05(g)).
Step 2B (Does the claim recite additional elements that amount to significantly more than the judicial exception?)
In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception.
Regarding the “computer-implemented” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)).
Regarding the “receiving training data; receiving test data” limitation, as discussed above, the additional element of a data gathering step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. pre-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
Regarding Claim 19
Step 2A, Prong 1
wherein the steps of calculating a new weighting coefficient, evaluating a new hyperparameter combination and, calculating a new performance score are repeated until an optimal hyperparameter combination is obtained, the optimal hyperparameter combination corresponding to the hyperparameter combination having obtained the best performance score among the hyperparameter combinations evaluated. (under the broadest reasonable interpretation, a human can mentally repeat these steps of calculating a new weight coefficient, evaluating a new hyperparameter combination and, calculating a new performance score)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Regarding Claim 20
Step 2A, Prong 1
wherein the initial weighting coefficient is less than or equal to 1%. (under the broadest reasonable interpretation, a human can mentally initiate the weighting coefficient at 1%)
Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 4-5, 11-13, 15-16, and 18-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 20200065712 A1, hereinafter referenced as WANG.
Regarding Claim 1
WANG teaches:
A computer-implemented method for searching for an optimal hyperparameter combination for defining a machine learning model, the method comprising: (WANG, para. 0017: “Accordingly, different configurations within the set 102 generally differ in one or more of the types of predictive mode, the learning algorithm, the hyperparameters associated with the model or learning algorithm, and the computation and selection of input features.”;
WANG, para. 0018: “Machine-learning configurations for neural-network models, in turn, may specify various associated learning algorithms (e.g., backpropagation of errors, or reinforcement learning with various rewards), and differ in hyperparameters such as the number of layers within a network, or the step size used when adjusting network weights in the learning process.”;
WANG, para. 0021: “With renewed reference to FIG. 1, the computing system 100 may include multiple processing components to select an approximate best configuration 106 from the candidate set 102”;
Examiner’s Note: WANG teaches computer-implemented techniques for selecting an optimal configuration, including configuration of selected hyperparameters, for a neural network model)
performing a plurality of tests of hyperparameter combination, each test of hyperparameter combination including a training phase and a test phase, wherein the training phase is adapted to train the machine learning model from training data and the test phase is adapted to calculate a performance score associated with the hyperparameter combination tested from test data, (WANG, para. 0015: “The dataset 104 may be divided (e.g., randomly) into a training dataset 107 used to train the candidate configurations, and a test dataset 108 used to validate, that is, test the performance of, the trained configurations.”;
WANG, para. 0029: “In each loop of the iterative process, the training and test datasets are sampled (e.g., by data sampler 114), at 410, based on the determined sample sizes. At 412, the probe configuration Cprob is trained on the sampled training dataset, and then evaluated on the sampled test dataset (or, in some embodiments, on the full test dataset) (e.g., by training and test component 112). In the course of training and testing, a quality metric characterizing the performance of the trained probe configuration Cprob is evaluated on the sampled training and test datasets. For example, if predictive accuracy is used as the quality metric, training and test accuracies are computed. At 414, the estimated confidence interval associated with the probe configuration Cprob is updated (e.g., by scheduling and pruning component 116) based on the training and test accuracies (or training and test values of some other quality metric), optionally in conjunction with other parameters. The confidence interval provides estimated bounds for the real performance of the probe configuration Cprob, that is, the accuracy (or other quality metric) the configuration would achieve if trained on the full training dataset 107 and tested on the full test dataset 108. ”;
Examiner’s Note: WANG discloses that at step 412, the configuration is tested and validated using test and training datasets, wherein a quality metric is evaluated on the test dataset to determine a confidence interval (corresponding to recited “performance score”))
the optimal hyperparameter combination corresponding to the hyperparameter combination having obtained the best performance score among the hyperparameter combinations tested; and (WANG, para. 0031: “The iterative process of training and testing the selected configuration on sampled datasets, updating the confidence interval, pruning the candidate set Ω of remaining configurations, and selecting a new probe configuration (operations 410-420) may continue as long as more than one candidate configuration remain in the set Ω. Once only one configuration is left within the set Ω, that configuration is returned as the approximate best configuration (at 422). Alternatively, in some embodiments (not illustrated in FIG. 4), the iterative process may terminate when a specified time limit has been reached; and among the configurations then still remaining within the set Ω, one may be selected (e.g., based on a highest associated upper or lower bound) and returned as the approximate best configuration.”;
Examiner’s Note: WANG teaches returning the best configuration (corresponding to recited “optimal hyperparameter combination”) remaining based on the confidence interval)
defining a weighting coefficient for adjusting an amount of training data used for the training phase, the weighting coefficient being dynamically adapted during different tests of the hyperparameter combinations. (WANG, para. 0031: “Following pruning (at 418), a configuration for the next probe is selected from the remaining configurations set Ω, and the associated sample sizes for sampling the training and test datasets are determined (e.g., by scheduling and pruning component 116) at 420.”
WANG, para. 0041: “Once the configuration for the next probe has been selected, the associated sample size for the next probe is determined (at 514), e.g., based on a sampling schedule associated with the selected configuration. The selected configuration and associated sample size are then output (at 516) back into the iterative method 400.”
WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: WANG discloses a factor (c) for determining how much to increase a testing or training dataset size (corresponding to recited “weighting coefficient”), where as shown in Fig. 4 (at step 420), and Fig. 5 (at step 514), the size of the testing dataset is increased for each testing iteration (corresponding to recited “the weighting coefficient being dynamically adapted during different tests of the hyperparameter combinations” limitation)
Regarding Claim 2
WANG teaches the method of claim 1 as explained above. WANG further teaches:
wherein the weighting coefficient is initialized to an initial weighting coefficient. (WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: when alpha = 1, the initial step size is to double the sample size for each successive probe)
Regarding Claim 4
WANG teaches the method of claim 1 as explained above. WANG further teaches:
wherein the weighting coefficient is updated for each test of hyperparameter combination. (WANG, para. 0031: “Following pruning (at 418), a configuration for the next probe is selected from the remaining configurations set Ω, and the associated sample sizes for sampling the training and test datasets are determined (e.g., by scheduling and pruning component 116) at 420.”
WANG, para. 0041: “Once the configuration for the next probe has been selected, the associated sample size for the next probe is determined (at 514), e.g., based on a sampling schedule associated with the selected configuration. The selected configuration and associated sample size are then output (at 516) back into the iterative method 400.”
WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: as shown in Fig. 4 (at step 420), and Fig. 5 (at step 514), the size of the testing dataset is increased for each testing iteration)
Regarding Claim 5
WANG teaches the method of claim 4 as explained above. WANG further teaches:
wherein updating the weighting coefficient to be used for a given test of hyperparameter combination comprises calculating a new weighting coefficient from an old weighting coefficient used during the test of hyperparameter combination directly preceding the given test of hyperparameter combination. (WANG, para. 0031: “Following pruning (at 418), a configuration for the next probe is selected from the remaining configurations set Ω, and the associated sample sizes for sampling the training and test datasets are determined (e.g., by scheduling and pruning component 116) at 420.”
WANG, para. 0041: “Once the configuration for the next probe has been selected, the associated sample size for the next probe is determined (at 514), e.g., based on a sampling schedule associated with the selected configuration. The selected configuration and associated sample size are then output (at 516) back into the iterative method 400.”
WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: as shown in Fig. 4 (at step 420), and Fig. 5 (at step 514), the size of the testing dataset is increased for each testing iteration based on the size of the factor (c) in the previous iteration)
Regarding Claim 11
WANG teaches:
A non-transitory memory storing a computer program comprising instructions which, when the program is executed by a computer, cause the computer to implement a method comprising: (WANG, para. 0045: “In general, the operations, algorithms, and methods described herein may be implemented in any suitable combination of software, hardware, and/or firmware, and the provided functionality may be grouped into a number of components, modules, or mechanisms. Modules and components can constitute either software components (e.g., code embodied on a non-transitory machine-readable medium) or hardware-implemented components.”)
performing a plurality of tests of hyperparameter combination, each test of hyperparameter combination including a training phase and a test phase, wherein the training phase is adapted to train a machine learning model from training data and the test phase is adapted to calculate a performance score associated with the hyperparameter combination tested from test data; and (WANG, para. 0015: “The dataset 104 may be divided (e.g., randomly) into a training dataset 107 used to train the candidate configurations, and a test dataset 108 used to validate, that is, test the performance of, the trained configurations.”;
WANG, para. 0029: “In each loop of the iterative process, the training and test datasets are sampled (e.g., by data sampler 114), at 410, based on the determined sample sizes. At 412, the probe configuration Cprob is trained on the sampled training dataset, and then evaluated on the sampled test dataset (or, in some embodiments, on the full test dataset) (e.g., by training and test component 112). In the course of training and testing, a quality metric characterizing the performance of the trained probe configuration Cprob is evaluated on the sampled training and test datasets. For example, if predictive accuracy is used as the quality metric, training and test accuracies are computed. At 414, the estimated confidence interval associated with the probe configuration Cprob is updated (e.g., by scheduling and pruning component 116) based on the training and test accuracies (or training and test values of some other quality metric), optionally in conjunction with other parameters. The confidence interval provides estimated bounds for the real performance of the probe configuration Cprob, that is, the accuracy (or other quality metric) the configuration would achieve if trained on the full training dataset 107 and tested on the full test dataset 108. ”;
Examiner’s Note: WANG discloses that at step 412, the configuration is tested and validated using test and training datasets, wherein a quality metric is evaluated on the test dataset to determine a confidence interval (corresponding to recited “performance score”))
defining a weighting coefficient for adjusting the amount of training data used for the training phase, the weighting coefficient being dynamically adapted during different tests of the hyperparameter combinations (WANG, para. 0031: “Following pruning (at 418), a configuration for the next probe is selected from the remaining configurations set Ω, and the associated sample sizes for sampling the training and test datasets are determined (e.g., by scheduling and pruning component 116) at 420.”
WANG, para. 0041: “Once the configuration for the next probe has been selected, the associated sample size for the next probe is determined (at 514), e.g., based on a sampling schedule associated with the selected configuration. The selected configuration and associated sample size are then output (at 516) back into the iterative method 400.”
WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: WANG discloses a factor (c) for determining how much to increase a testing or training dataset size (correspond go recited “weighting coefficient”), where as shown in Fig. 4 (at step 420), and Fig. 5 (at step 514), the size of the testing dataset is increased for each testing iteration (corresponding to recited “the weighting coefficient being dynamically adapted during different tests of the hyperparameter combinations” limitation)
to determine an optimal hyperparameter combination corresponding to the hyperparameter combination having obtained the best performance score among the hyperparameter combinations tested. (WANG, para. 0031: “The iterative process of training and testing the selected configuration on sampled datasets, updating the confidence interval, pruning the candidate set Ω of remaining configurations, and selecting a new probe configuration (operations 410-420) may continue as long as more than one candidate configuration remain in the set Ω. Once only one configuration is left within the set Ω, that configuration is returned as the approximate best configuration (at 422). Alternatively, in some embodiments (not illustrated in FIG. 4), the iterative process may terminate when a specified time limit has been reached; and among the configurations then still remaining within the set Ω, one may be selected (e.g., based on a highest associated upper or lower bound) and returned as the approximate best configuration.”;
Examiner’s Note: WANG teaches returning the best configuration (corresponding to recited “optimal hyperparameter combination”) remaining based on the confidence interval)
Regarding Claim 12
WANG teaches:
A computing system comprising: the memory according to claim 11; and a processing unit coupled to the memory and configured to execute the computer program. (WANG, para. 0045: “In general, the operations, algorithms, and methods described herein may be implemented in any suitable combination of software, hardware, and/or firmware, and the provided functionality may be grouped into a number of components, modules, or mechanisms. Modules and components can constitute either software components (e.g., code embodied on a non-transitory machine-readable medium) or hardware-implemented components.”;
WANG, para. 0058: “The instructions 624 can also reside, completely or at least partially, within the main memory 604 and/or within the processor 602 during execution thereof by the computer system 600, with the main memory 604 and the processor 602 also constituting machine-readable media.”;
Examiner’s Note: see analysis of claim 11 above)
Claim 13 depends from claim 12 and corresponds to the method of claim 2, and is therefore rejected for the same reasons explained above with respect to claims 2 and 12.
Claim 15 depends from claim 12 and corresponds to the method of claim 4, and is therefore rejected for the same reasons explained above with respect to claims 4 and 12.
Claim 16 depends from claim 15 and corresponds to the method of claim 5, and is therefore rejected for the same reasons explained above with respect to claims 5 and 15.
Regarding Claim 18
WANG teaches:
A computer-implemented method for searching for an optimal hyperparameter combination for defining an automatic learning model, the method comprising: (WANG, para. 0017: “Accordingly, different configurations within the set 102 generally differ in one or more of the types of predictive mode, the learning algorithm, the hyperparameters associated with the model or learning algorithm, and the computation and selection of input features.”;
WANG, para. 0018: “Machine-learning configurations for neural-network models, in turn, may specify various associated learning algorithms (e.g., backpropagation of errors, or reinforcement learning with various rewards), and differ in hyperparameters such as the number of layers within a network, or the step size used when adjusting network weights in the learning process.”;
WANG, para. 0021: “With renewed reference to FIG. 1, the computing system 100 may include multiple processing components to select an approximate best configuration 106 from the candidate set 102”;
Examiner’s Note: WANG teaches computer-implemented techniques for selecting an optimal configuration, including configuration of selected hyperparameters, for a neural network model)
initializing a weighting coefficient; (WANG, para. 0031: “Following pruning (at 418), a configuration for the next probe is selected from the remaining configurations set Ω, and the associated sample sizes for sampling the training and test datasets are determined (e.g., by scheduling and pruning component 116) at 420.”
WANG, para. 0041: “Once the configuration for the next probe has been selected, the associated sample size for the next probe is determined (at 514), e.g., based on a sampling schedule associated with the selected configuration. The selected configuration and associated sample size are then output (at 516) back into the iterative method 400.”
WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: WANG discloses a factor (c) for determining how much to increase a testing or training dataset size (corresponding to recited “weighting coefficient”), where the first use of such factor (c) is the initialized coefficient)
receiving training data; (WANG, para. 0015: “The dataset 104 may be divided (e.g., randomly) into a training dataset 107 used to train the candidate configurations, and a test dataset 108 used to validate, that is, test the performance of, the trained configurations.”;
WANG, para. 0022: “The data sampler 114 is configured to generate the sampled training and test datasets 118, 119 by sampling (e.g., randomly) from the full training and test datasets 107, 108, respectively, using sample sizes 122 determined by the scheduling and pruning component 116 and communicated to the data sampler 114”)
receiving test data; (WANG, para. 0015: “The dataset 104 may be divided (e.g., randomly) into a training dataset 107 used to train the candidate configurations, and a test dataset 108 used to validate, that is, test the performance of, the trained configurations.”;
WANG, para. 0022: “The data sampler 114 is configured to generate the sampled training and test datasets 118, 119 by sampling (e.g., randomly) from the full training and test datasets 107, 108, respectively, using sample sizes 122 determined by the scheduling and pruning component 116 and communicated to the data sampler 114”)
evaluating a performance of a hyperparameter combination, the evaluating being performed in a training phase using a portion of the training data based on the weighting coefficient and a test phase using the test data; (WANG, para. 0029: “In each loop of the iterative process, the training and test datasets are sampled (e.g., by data sampler 114), at 410, based on the determined sample sizes. At 412, the probe configuration Cprob is trained on the sampled training dataset, and then evaluated on the sampled test dataset (or, in some embodiments, on the full test dataset) (e.g., by training and test component 112). In the course of training and testing, a quality metric characterizing the performance of the trained probe configuration Cprob is evaluated on the sampled training and test datasets. For example, if predictive accuracy is used as the quality metric, training and test accuracies are computed. At 414, the estimated confidence interval associated with the probe configuration Cprob is updated (e.g., by scheduling and pruning component 116) based on the training and test accuracies (or training and test values of some other quality metric), optionally in conjunction with other parameters. The confidence interval provides estimated bounds for the real performance of the probe configuration Cprob, that is, the accuracy (or other quality metric) the configuration would achieve if trained on the full training dataset 107 and tested on the full test dataset 108. ”;
Examiner’s Note: WANG discloses that at step 412, the configuration is tested and validated using test and training datasets)
calculating a performance score associated with the hyperparameter combination; WANG, para. 0029: “In each loop of the iterative process, the training and test datasets are sampled (e.g., by data sampler 114), at 410, based on the determined sample sizes. At 412, the probe configuration Cprob is trained on the sampled training dataset, and then evaluated on the sampled test dataset (or, in some embodiments, on the full test dataset) (e.g., by training and test component 112). In the course of training and testing, a quality metric characterizing the performance of the trained probe configuration Cprob is evaluated on the sampled training and test datasets. For example, if predictive accuracy is used as the quality metric, training and test accuracies are computed. At 414, the estimated confidence interval associated with the probe configuration Cprob is updated (e.g., by scheduling and pruning component 116) based on the training and test accuracies (or training and test values of some other quality metric), optionally in conjunction with other parameters. The confidence interval provides estimated bounds for the real performance of the probe configuration Cprob, that is, the accuracy (or other quality metric) the configuration would achieve if trained on the full training dataset 107 and tested on the full test dataset 108. ”;
Examiner’s Note: a quality metric is evaluated on the test dataset to determine a confidence interval (corresponding to recited “performance score”))
calculating a new weighting coefficient that is greater than the initial weighting coefficient; (WANG, para. 0031: “Following pruning (at 418), a configuration for the next probe is selected from the remaining configurations set Ω, and the associated sample sizes for sampling the training and test datasets are determined (e.g., by scheduling and pruning component 116) at 420.”
WANG, para. 0041: “Once the configuration for the next probe has been selected, the associated sample size for the next probe is determined (at 514), e.g., based on a sampling schedule associated with the selected configuration. The selected configuration and associated sample size are then output (at 516) back into the iterative method 400.”
WANG, para. 0042: “In accordance with various embodiments, sample schedules associated with the configurations are predetermined, such that the sample size(s) for the next probe (i.e., the size of the sampled training dataset and, if the test dataset is likewise sampled, the sample size of the sampled test dataset) can simply be looked up during the iterative process. The sample schedules may be geometric, meaning that, between any two consecutive probes of the same configuration, the sample size (for training or testing) increases by a constant factor c. The optimal value of c is generally dependent on certain aspects of the configuration e.g., the model and/or learning algorithm). It can be shown that, when the probing time Ti(s) for configuration Ci is a power function of the training sample size s, i.e., Ti(s)=sα (where α is a real number), the optimal step size follows.
PNG
media_image1.png
38
62
media_image1.png
Greyscale
For example, if the time to probe a configuration is proportional to the sample size (i.e., α=1), the sample size may be doubled for each successive probe of that configuration. A progressive test sample schedule can be similarly determined based on the functional dependence of the test time on the size of the sampled test dataset.”;
Examiner’s Note: WANG discloses a factor (c) for determining how much to increase a testing or training dataset size such that each successive iteration uses additional samples)
repeating the evaluating for a new hyperparameter combination, the repeated evaluating performed with a portion of the training data based on the new weighting coefficient and the test data; (WANG, para. 0028: “The method 400 involves an iterative process of training and testing (herein also “probing”) a selected candidate configuration (herein also the “probe configuration”) and pruning the candidate set of remaining configurations, Ω, based on these probes.”;
WANG, para. 0031: “The iterative process of training and testing the selected configuration on sampled datasets, updating the confidence interval, pruning the candidate set Ω of remaining configurations, and selecting a new probe configuration (operations 410-420) may continue as long as more than one candidate configuration remain in the set Ω. Once only one configuration is left within the set Ω, that configuration is returned as the approximate best configuration (at 422). Alternatively, in some embodiments (not illustrated in FIG. 4), the iterative process may terminate when a specified time limit has been reached; and among the configurations then still remaining within the set Ω, one may be selected (e.g., based on a highest associated upper or lower bound) and returned as the approximate best configuration.”;
Examiner’s Note: As shown by Figs. 4 and 5, such steps of evaluating configurations based on successive sample sizes are iterative)
calculating a new performance score associated with the hyperparameter combination; and (WANG, para. 0029: “At 414, the estimated confidence interval associated with the probe configuration Cprob is updated (e.g., by scheduling and pruning component 116) based on the training and test accuracies (or training and test values of some other quality metric), optionally in conjunction with other parameters. The confidence interval provides estimated bounds for the real performance of the probe configuration Cprob, that is, the accuracy (or other quality metric) the configuration would achieve if trained on the full training dataset 107 and tested on the full test dataset 108.”)
comparing the performance score with the new performance score. (WANG, para. 0030: “At 416, the updated lower bound Cprob.l is compared against the lower bound C.i′.l of the current presumed best configuration Ci′, and if Cprob.l > C.i′.l the presumed best configuration Ci′ is updated to the probe configuration Cprob (and the lower bound C.i′.l is, accordingly, updated to Cprob.l).”)
Regarding Claim 19
WANG teaches the method of claim 8 as explained above. WANG further teaches:
wherein the steps of calculating a new weighting coefficient, evaluating a new hyperparameter combination and, calculating a new performance score are repeated until an optimal hyperparameter combination is obtained, the optimal hyperparameter combination corresponding to the hyperparameter combination having obtained the best performance score among the hyperparameter combinations evaluated. (WANG, para. 0028: “The method 400 involves an iterative process of training and testing (herein also “probing”) a selected candidate configuration (herein also the “probe configuration”) and pruning the candidate set of remaining configurations, Ω, based on these probes.”;
WANG, para. 0031: “The iterative process of training and testing the selected configuration on sampled datasets, updating the confidence interval, pruning the candidate set Ω of remaining configurations, and selecting a new probe configuration (operations 410-420) may continue as long as more than one candidate configuration remain in the set Ω. Once only one configuration is left within the set Ω, that configuration is returned as the approximate best configuration (at 422). Alternatively, in some embodiments (not illustrated in FIG. 4), the iterative process may terminate when a specified time limit has been reached; and among the configurations then still remaining within the set Ω, one may be selected (e.g., based on a highest associated upper or lower bound) and returned as the approximate best configuration.”;
Examiner’s Note: As shown by Figs. 4 and 5, such steps of evaluating configurations based on successive sample sizes are iterative until a single best configuration remains)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 3, 6, 14, 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over WANG in view of US 20220147672 A1, hereinafter referenced as NISTALA.
Regarding Claim 3
WANG teaches the method of claim 2 as explained above. However, WANG fails to explicitly teach:
wherein the initial weighting coefficient is less than or equal to 1%.
However, in a related field of endeavor (machine learning models, see para. 0039), NISTALA teaches and makes obvious:
wherein the initial weighting coefficient is less than or equal to 1%. (NISTALA, para. 0074: “In the FIG. 9, the grid search technique for identifying optimum combination of the first set of data and the second set of data during model re-tuning is depicted. According to this technique, ‘n’ training datasets are obtained by combining α % of the second set of data and β% of the first set of data (e.g. 90% of the second set of data and 20% of the first set of data). Here, α and β can take any value between 0 and 100.”;
Examiner’s Note: NISTALA discloses using between 0-100% of a dataset for model re-tuning; the WANG-NISTALA combination now starts off with an initial weighting of 1% of a training set as in NISTALA; the examiner notes that pursuant to MPEP 2131.03, a specific example in the prior art which is within a claimed range anticipates the range, and 1% is within the range cited in NISTALA)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of WANG and NISTALA as explained above. As explained by NISTALA, one of ordinary skill would have been motivated to do so in order to separate training data based on recency, and then identifying an optimum combination of such training data to result in an optimized model. (para. 0073).
Regarding Claim 6
WANG teaches the method of claim 5 as explained above. However, WANG fails to explicitly teach:
wherein the new weighting coefficient is calculated by a formula k*α, where α is the old weighting coefficient and k is a coefficient greater than 1.
However, in a related field of endeavor (machine learning models, see para. 0039), NISTALA teaches and makes obvious:
wherein the new weighting coefficient is calculated by a formula k*α, where α is the old weighting coefficient and k is a coefficient greater than 1. (NISTALA, para. 0074: “In the FIG. 9, the grid search technique for identifying optimum combination of the first set of data and the second set of data during model re-tuning is depicted. According to this technique, ‘n’ training datasets are obtained by combining α % of the second set of data and β% of the first set of data (e.g. 90% of the second set of data and 20% of the first set of data). Here, α and β can take any value between 0 and 100.”;
Examiner’s Note: NISTALA discloses using between 0-100% of a dataset for model re-tuning; the WANG-NISTALA combination now takes the previous factor (c) (corresponding to recited α) for the previous iteration, and multiples by 1 + 0-100% as disclosed by NISTALA (corresponding to recited “k”))
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of WANG and NISTALA as explained above. As explained by NISTALA, one of ordinary skill would have been motivated to do so in order to separate training data based on recency, and then identifying an optimum combination of such training data to result in an optimized model. (para. 0073).
Claim 14 depends from claim 13 and corresponds to the method of claim 3, and is therefore rejected for the same reasons explained above with respect to claims 3 and 13.
Claim 17 depends from claim 16 and corresponds to the method of claim 6, and is therefore rejected for the same reasons explained above with respect to claims 6 and 16.
Claim 20 depends from claim 18 and corresponds to the method of claim 3, and is therefore rejected for the same reasons explained above with respect to claims 3 and 18.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over WANG in view of US 20180018586 A1, hereinafter referenced as KOBAYASHI.
Regarding Claim 7
WANG teaches the method of claim 1 as explained above. However, WANG fails to explicitly teach:
defining a dynamically defined best weighting coefficient, this best weighting coefficient corresponding to the weighting coefficient used for the training phase of the test of the hyperparameter combination having obtained the best performance score among the hyperparameter combinations already tested.
However, in a related field of endeavor (machine learning, see para. 0002), KOBAYASHI teaches and makes obvious:
defining a dynamically defined best weighting coefficient, this best weighting coefficient corresponding to the weighting coefficient used for the training phase of the test of the hyperparameter combination having obtained the best performance score among the hyperparameter combinations already tested. (KOBAYASHI, para. 0052: “As for the machine learning algorithm 13a having generated a model with the maximum prediction performance score 14, the training dataset size 17a to be used next is determined based on the maximum prediction performance score 14, the estimated prediction performance scores 15a and 15b, and the estimated runtimes 16a and 16b. In addition, as for the machine learning algorithm 13b, the training dataset size 17b to be used next is determined based on the maximum prediction performance score 14 achieved by the machine learning algorithm 13a, the estimated prediction performance scores 15c and 15d, and the estimated runtimes 16c and 16d..”;
Examiner’s Note: the WANG-KOBAYASHI combination now dynamically tunes the factor (c) of WANG to determine a factor (c) that is used with the best configuration as in KOBAYASHI)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of WANG and KOBAYASHI as explained above. As explained by KOBAYASHI, one of ordinary skill would have been motivated to do so in order to “skip fruitless intermediate learning steps taking place.” (para. 0053).
Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over US 20220383038 A1, hereinafter referenced as HINES.
Regarding Claim 9
WANG teaches the method of claim 1 as explained above. However, WANG fails to explicitly teach:
further comprising training a machine learning model defined by the optimal combination of hyperparameter with all the training data.
However, in a related field of endeavor (machine learning model training, see para. 0002), HINES teaches and makes obvious:
further comprising training a machine learning model defined by the optimal combination of hyperparameter with all the training data. (HINES, para. 0027: “The data preprocessor 114 can receive training data (e.g., unstructured data, semi-structured data, unstructured data, multidimensional data, etc.) and select reference data (also known as the “first set of data”) from the training data, to train the machine learning model 115 during a training phase. Each item in the reference data can be considered as a singular input to the machine learning model 115 during the training phase. In some instances, the reference data can include the entire training data or a sufficient subset of data (sufficient to train the machine learning model 115 with a desired accuracy) from the training data.”;
Examiner’s Note: the WANG-HINES combination now trains the machine learning models of WANG, using the best configuration determined by WANG, using the entire training dataset as in HINES)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of WANG and HINES as explained above. As explained by HINES, one of ordinary skill would have been motivated to do so in order to sufficiently train a machine learning model with a desired accuracy. (para. 0027). One of ordinary skill in the art would understand that in general, using more training data results in a higher accuracy trained model.
Regarding Claim 10
WANG teaches the method of claim 1 as explained above. However, WANG fails to explicitly teach:
for each machine learning model defined by a combination of hyperparameters having made it possible to obtain a better performance score among the combinations of hyperparameters already tested, training of this model with all the data each time a better performance score is obtained.
However, in a related field of endeavor (machine learning model training, see para. 0002), HINES teaches and makes obvious:
for each machine learning model defined by a combination of hyperparameters having made it possible to obtain a better performance score among the combinations of hyperparameters already tested, training of this model with all the data each time a better performance score is obtained. (HINES, para. 0027: “The data preprocessor 114 can receive training data (e.g., unstructured data, semi-structured data, unstructured data, multidimensional data, etc.) and select reference data (also known as the “first set of data”) from the training data, to train the machine learning model 115 during a training phase. Each item in the reference data can be considered as a singular input to the machine learning model 115 during the training phase. In some instances, the reference data can include the entire training data or a sufficient subset of data (sufficient to train the machine learning model 115 with a desired accuracy) from the training data.”;
Examiner’s Note: the WANG-HINES combination now trains the machine learning models of WANG, at each iteration using the then-best configuration determined by WANG, using the entire training dataset as in HINES)
Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of WANG and HINES as explained above. As explained by HINES, one of ordinary skill would have been motivated to do so in order to sufficiently train a machine learning model with a desired accuracy. (para. 0027). One of ordinary skill in the art would understand that in general, using more training data results in a higher accuracy trained model.
Allowable Subject Matter
Claim 8 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, provided that the rejection under 35 U.S.C. 101 is overcome.
The following is a statement of reasons for the indication of allowable subject matter:
Claim 8 would be considered allowable because none of the references of record either alone or in combination fairly disclose or suggest the combination of limitations specified in claim 8, including at least:
wherein updating the weighting coefficient comprises comparing a new weighting coefficient with the value 100% and with the value w*A, where A is the best weighting coefficient defined and w is a coefficient greater than 1, the weighting coefficient being updated to the value of the new weighting coefficient if the new weighting coefficient calculated is less than or equal to the value 100% or to the value w*A, or updated to the value of an initial weighting coefficient otherwise.
The closest prior art of record discloses:
US 20200065712 A1, hereinafter referenced as WANG discloses a system for determining a best configuration for a machine learning model, including best hyperparameters. (paras. 0017-0018, 0021).
US 20180018586 A1, hereinafter referenced as KOBAYASHI, discloses techniques for determining an optimum number of training dataset samples to use during training. (para. 0052).
However, the examiner has found that the distinct feature of the Applicant's claimed invention over the prior art is the explicit claiming of the aforementioned limitations in combination with all the other limitations as specified in claim 8. Therefore, because the prior art of record does not anticipate nor make obvious the limitations of claim 8, claim 8 would be allowed if rewritten in independent form including all of the limitations of the base claim and any intervening claims, provided that the rejection under 35 U.S.C. 101 is overcome.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20210089937 A1 (Zhang). “Turning to FIG. 2, during a configuration process of modifying performance evaluation scheme, a validation component 202 adapts a validation configuration decision to machine learning algorithms and adjusts ratio of size of training data set relative to size of a validation data set. Moreover, the validation component 202 can use a procedure for selecting Holdout size and a scheme to reduce variance in a generalization error estimate in Holdout validation via Bootstrapping. Along with adjusting ratio of size of training data by validation component 202, configuration component 110 as discussed previously can generate a set of samples of the ratio and an associated metric of performance evaluation accuracy.” (para. 0030).
US 20200226496 A1 (Basu). “The evaluation application 216 is configured to instruct the master server 112 to optimize one or more of the hyperparameters 226 for a selected machine-learning model 232. In addition, the evaluation application 216 may communicate one or more tuning parameter(s) 228 to the master server 112 that define the optimization parameters for optimizing the hyperparameters 226. For example, the tuning parameters 228 may include the predetermined sampling percentage value (e.g., the value of η), an upper limit value (e.g., the value of N), the performance metric and/or quality metric function to determine (e.g., f(x)), one or more values that define X, the kernel function k.sub.θ, or any combination thereof.” (para. 0066).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C LEE whose telephone number is (571)272-4933. The examiner can normally be reached M-F 12:00 pm - 8:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571-272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL C. LEE/Examiner, Art Unit 2128