DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over US Pub No. US20200372342A1 Nair et al. (“Nair”) in view of Ying, Xue. "An overview of overfitting and its solutions." Journal of physics: Conference series. Vol. 1168. No. 2. IOP Publishing, 2019. (“Ying”)
In regards to claim 1 and analogous claims 8 and 15,
Nair teaches A method for mitigating overfitting during training of an artificial intelligence (AI) model, the method comprising: initiating, by AI model training circuitry, a model training session for the AI model;
(Nair, “[0068] In operation 430, training may begin or continue (e.g. for a next training interval or epoch) on a NN [initiating, by AI model training circuitry, a model training session for the AI model], e.g. by a system such as shown in FIG. 3, where training data is presented to a NN over epochs.”)
(Nair, “[0056] FIG. 2 shows a high-level block diagram of an exemplary computing device which may be used with embodiments of the present invention. Computing device 100 may include a controller or processor 105 that may be or include, for example, one or more central processing unit processor(s) (CPU), one or more Graphics Processing Unit(s) (GPU or GPGPU), a chip or any suitable computing or computational device, an operating system 115, a memory 120, a storage 130, input devices 135 and output devices 140. Each of modules and equipment discussed elsewhere (e.g. FIG. 3) such as server 20, model(s) 24, computers 10, NNs 12, training software 14, stopping module 14′, and other equipment and modules mentioned herein may be, include, or be executed by a computing device such as included in FIG. 2, although various units among these entities may be combined into one computing device.”)
Nair teaches and for each model training epoch of a plurality of model training epochs associated with the model training session: determining, by the AI model training circuitry, an overfitting metric value associated with an overfitting metric, [wherein the overfitting metric indicates an overfitting condition associated with the AI model],
(Nair, “[0017] In one embodiment, over a series of NN training epochs, where in each epoch the NN undergoes training [and for each model training epoch of a plurality of model training epochs associated with the model training session] and a loss is computed [determining, by the AI model training circuitry, an overfitting metric value ex. training loss associated with an overfitting metric ie validation loss and training loss; wherein training loss is one value of the overfitting metric; the other being validation loss], an expected training loss may be determined, using a set of model parameters, data describing training loss of the NN, and a model that has been trained using training losses of a plurality of NNs other than the NN.”)
Nair further discloses early stopping in consideration to validation loss
(Nair, [0103], “An alternate embodiment may cause stopping of training as soon as a minimum predicted validation loss is achieved (possibly within some tolerance or accuracy) or when a wait or patience value exceeds a threshold.”)
Nair teaches determining, by the AI model training circuitry, whether the overfitting metric value satisfies a first overfitting metric threshold,
(Nair, “[0032] Embodiments of the invention include systems and methods that may stop training of a NN if it is determined that likelihood of improvement in loss over a best loss seen is below a threshold, or the training has proceeded to the mean expected or final value of training (typically after a certain number of initial training cycles or epochs), or if training has “stalled” and does not improve (e.g. over a historic best value) [determining, by the AI model training circuitry, whether the overfitting metric value satisfies a first overfitting metric threshold; ie the threshold of if training has “stalled” or not improved indicated by comparing the current loss value over a historic best value (threshold)] over a certain number of training cycles or epochs.”)
Nair teaches and in response to determining that the overfitting metric value satisfies the first overfitting metric threshold: determining, by the AI model training circuitry and based on the overfitting metric value, whether the overfitting condition associated with the AI model has deteriorated; and in response to determining that the overfitting condition associated with the AI model has deteriorated: incrementing, by the AI model training circuitry, a patience counter value associated with a patience counter,
(Nair, “[0076] In operation 470, if the actual training loss is not less than a historic minimum training loss [and in response to determining that the overfitting metric value satisfies the first overfitting metric threshold: determining, by the AI model training circuitry and based on the overfitting metric value, whether the overfitting condition associated with the AI model has deteriorated wherein Examiner interprets deteriorated as the loss has not improved (ie the actual training loss is not less than a historic minimum training loss); and in response to determining that the overfitting condition associated with the AI model has deteriorated], a wait value may be incremented or increased, e.g. by a convenient integer count value such as 1 [incrementing, by the AI model training circuitry, a patience counter value ie wait value associated with a patience counter].”)
Nair teaches determining, by the AI model training circuitry, whether the patience counter value satisfies a patience threshold, and in response to determining that the patience counter value satisfies the patience threshold, terminating, by the AI model training circuitry, the model training session for the AI model.
(Nair, [0033], “In one embodiment, if an expected or predicted training loss (e.g. that returned by a model) is greater than or equal to the actual training loss of the target NN, or if patience has been exhausted (e.g., if a wait value, which is incremented if no improvement is seen, is greater than a wait threshold) training may be stopped [determining, by the AI model training circuitry, whether the patience counter value ie wait value satisfies a patience threshold ie wait threshold (patience exhausted), and in response to determining that the patience counter value satisfies the patience threshold, terminating, by the AI model training circuitry, the model training session for the AI model].”)
However, Nair does not explicitly teach wherein the overfitting metric indicates an overfitting condition associated with the AI model
Ying teaches wherein the overfitting metric indicates an overfitting condition associated with the AI model
(Ying, Section 2., “As shown in Figure 1, where the horizontal axis is epoch, and the vertical axis is error, the blue line shows the training error and the red line shows the validation error. If the model continues learning after the point, the validation error will increase while the training error will continue decreasing.
PNG
media_image1.png
170
250
media_image1.png
Greyscale
If we stop learning before the point, it’s under-fitting. If we stop after the point, we get over-fitting [wherein the overfitting metric ie proportion of validation loss to training loss indicates an overfitting condition associated with the AI model]. So the aim is to find the exact point to stop training.”)
Nair and Ying are both considered to be analogous to the claimed invention because they are in the same field of improving machine learning models through early stopping. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Nair to incorporate the teachings of Ying in order to consider early stopping in view of overfitting as doing so would prevent overfitting and improve generalization of the model (Nair, “[0004] Machine learning algorithms, using e.g. neural network (NN) or connectionist systems, are trained or learn by iteratively adjusting model parameters to optimize an objective function, also known as a loss function. The goal of training a model may be to determine a set of hyperparameters, and/or weights of links or functions connecting neurons, and an optimization procedure that minimizes the training loss and allows the model to achieve generalization on unseen data.”) (Ying, Abstract, “Overfitting is a fundamental issue in supervised machine learning which prevents us from perfectly generalizing the models to well fit observed data on training data, as well as unseen data on testing set. Because of the presence of noise, the limited size of training set, and the complexity of classifiers, overfitting happens. This paper is going to talk about overfitting from the perspectives of causes and solutions. To reduce the effects of overfitting, various strategies are proposed to address to these causes: 1) “early-stopping” strategy is introduced to prevent overfitting by stopping training before the performance stops optimize; 2) “network-reduction” strategy is used to exclude the noises in training set; 3) “data-expansion” strategy is proposed for complicated models to fine-tune the hyper-parameters sets with a great amount of data; and 4) “regularization” strategy is proposed to guarantee models performance to a great extent while dealing with real world issues by feature-selection, and by distinguishing more useful and less useful features.”)
In regards to claim 2 and analogous claims 9 and 16,
Nair and Ying teaches The method of claim 1, Nair teaches wherein determining whether the overfitting condition associated with the AI model has deteriorated comprises: determining, by the AI model training circuitry, whether a gradient value associated with a gradient corresponding to the overfitting metric is a positive gradient value or a negative gradient value for a given model training epoch of the plurality of model training epochs associated with the model training session, wherein a positive gradient value indicates a deterioration of the overfitting condition, and wherein a negative gradient value indicates an improvement of the overfitting condition.
PNG
media_image2.png
347
449
media_image2.png
Greyscale
(Nair, fig. 5B., “[0093] FIGS. 5A and 5B depict example prediction loss curves [determining, by the AI model training circuitry, whether a gradient value ie the change in the loss over epochs associated with a gradient corresponding to the overfitting metric is a positive gradient value ie positive slope of the loss curve or a negative gradient value ie negative slope of loss curve for a given model training epoch ie the slope at the given epoch of the plurality of model training epochs associated with the model training session], according to one embodiment of the invention.”)
(Nair, “[0032] Embodiments of the invention include systems and methods that may stop training of a NN if it is determined that likelihood of improvement in loss over a best loss seen is below a threshold, or the training has proceeded to the mean expected or final value of training (typically after a certain number of initial training cycles or epochs), or if training has “stalled” and does not improve (e.g. over a historic best value) [wherein a positive gradient value indicates a deterioration of the overfitting condition] over a certain number of training cycles or epochs.”)
(Nair, [0076], fig. 5B. shows, “Typically, if the loss function of a NN is decreasing it indicates an improvement of the functionality of the NN during training [wherein a negative gradient value indicates an improvement of the overfitting condition]. A historic minimum training loss may be for example the lowest training loss value in the loss history for the NN.”)
In regards to claim 3 and analogous claims 10 and 17,
Nair and Ying teaches The method of claim 2, Nair teaches wherein the patience counter value associated with the patience counter is incremented based on determining that the gradient value associated with the gradient corresponding to the overfitting metric is a positive gradient value.
(Nair, “[0076] In operation 470, if the actual training loss is not less than a historic minimum training loss [based on determining that the gradient value associated with the gradient corresponding to the overfitting metric is a positive gradient value], a wait value may be incremented or increased, e.g. by a convenient integer count value such as 1 [wherein the patience counter value associated with the patience counter is incremented].”)
In regards to claim 4 and analogous claims 11 and 18,
Nair and Ying teaches The method of claim 1,
Nair teaches the method further comprising: in response to determining that the overfitting metric value does not satisfy the first overfitting metric threshold: resetting, by the AI model training circuitry, the patience counter value associated with the patience counter.
(Nair, “[0075] In operation 460, if the current or actual training loss is less than a historic minimum training loss the wait value is not increased [in response to determining that the overfitting metric value does not satisfy the first overfitting metric threshold], and may be reset or set to zero [resetting, by the AI model training circuitry, the patience counter value associated with the patience counter], or to a value indicating waiting should begin at the beginning or for the maximum of the wait period.”)
In regards to claim 5 and analogous claims 12 and 19,
Nair and Ying teaches The method of claim 1,
Nair teaches the method further comprising: for each model training epoch of the plurality of model training epochs associated with the model training session: determining, by the AI model training circuitry, whether the overfitting metric value satisfies a second overfitting metric threshold; and in response to determining that the overfitting metric value satisfies the second overfitting metric threshold: terminating, by the AI model training circuitry, the model training session for the AI model.
(Nair, “[0018] One embodiment of a predictive early stopping tool may estimate the value of a local minima [a second overfitting metric threshold] beforehand, and may terminate the model training once it is achieved [in response to determining that the overfitting metric value satisfies the second overfitting metric threshold: terminating, by the AI model training circuitry, the model training session for the AI model].” Wherein the model training iterates over the epochs and “once it is achieved” is the epoch in which the estimated local minima is reached)
In regards to claim 6 and analogous claims 13 and 20,
Nair and Ying teaches The method of claim 1, Ying teaches wherein the overfitting metric is a ratio of validation loss to training loss,
(Ying, Section 2., “As shown in Figure 1, where the horizontal axis is epoch, and the vertical axis is error, the blue line shows the training error and the red line shows the validation error. If the model continues learning after the point, the validation error will increase while the training error will continue decreasing.
PNG
media_image1.png
170
250
media_image1.png
Greyscale
If we stop learning before the point, it’s under-fitting. If we stop after the point, we get over-fitting. So the aim is to find the exact point to stop training. In this way, we got a perfect fit between under-fitting and over-fitting [wherein the overfitting metric is a ratio of validation loss to training loss; ie the proportion of validation loss to training loss indicates overfitting/underfitting/a perfect fit wherein overfitting is identified with as validation error getting disproportionally larger than the training error ex. from 1:1 [validation error]:[training error] to 2:1 [validation error]:[training error]].”)Ying teaches wherein the validation loss is an error value quantifying one or more errors made by the AI model when using model validation data as input data,
and wherein the training loss is an error value quantifying one or more errors made by the AI model when using training data as input data.
(Ying, Abstract, “Overfitting is a fundamental issue in supervised machine learning which prevents us from perfectly generalizing the models to well fit observed data on training data [wherein the training loss is an error value quantifying one or more errors made by the AI model when using training data as input data], as well as unseen data on testing set.”)
(Ying, Section 2., “More generally, we can track the accuracy on validation set [wherein the validation loss is an error value quantifying one or more errors made by the AI model when using model validation data as input data] instead of test set in order to determine when to stop training. In another word, we use validation set to figure out a perfect set of values for the hyper-parameters, and later use the test set to do the final evaluation of accuracy. In this way, a higher generalization level can be guaranteed, compared to directly using test data to find out hyper-parameters’ values.”)
In regards to claim 7 and analogous claim 14,
Nair and Ying teaches The method of claim 1,
Nair teaches the method further comprising: receiving, by communications hardware and prior to initiating the model training session for the AI model, a model training session configuration request associated with the model training session for the AI model, wherein the model training session configuration request comprises one or more of a preferred number of model training epochs for the plurality of model training epochs, a first defined value for the first overfitting metric threshold, a second defined value for a second overfitting metric threshold, or a third defined value for the patience threshold,
Examiner’s note: Examiner interprets the BRI of “a model training session configuration request associated with the model training session for the AI model” as a call procedure to configure parameters of the model training session wherein Examiner notes the call procedure must be defined in order for the training call procedure to call it.
(Nair, [0112-0113] Table 1, Algorithm 1 Early stopping, “
Input: …. patience = Z > 0 // threshold for wait [a third defined value for the patience threshold]
Output :
True / False // True tells calling procedure to stop training ; False = keep training
1 : procedure EARLY STOPPING ( loss_history , current , 0 , threshold , patience ) // call
procedure for local stopping with input [receiving, by communications hardware and prior to initiating the model training session for the AI model, a model training session configuration request associated with the model training session for the AI model]
Nair teaches and wherein the model training session is initiated based on the model training session configuration request.
(Nair, [0113], “The example procedures shown in Tables 1 and 2 call or execute an example procedure such as shown in Table 3, to return a probability for improvement. The example procedures shown in Tables 1 and 2 may in turn be called or executed by [based on the model training session configuration request] a procedure choosing hyperparameters for a NN, or a procedure training a NN for which hyperparameters have been chosen [wherein the model training session is initiated; wherein the procedure training calls the procedure shown in Table 1 ie Algorithm 1].”)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US Pub No. US20230061222A1 Cmielowski et al. teaches Early stopping of artificial intelligence model training using control limits
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JASMINE THAI whose telephone number is (703)756-5904. The examiner can normally be reached M-F 8-4.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.T.T./Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129