Prosecution Insights
Last updated: October 02, 2026
Application No. 17/971,204

DOMAIN GENERALIZABLE CONTINUAL LEARNING USING COVARIANCES

Final Rejection §101§102§103
Filed
Oct 21, 2022
Priority
Nov 12, 2021 — provisional 63/278,512
Examiner
PARK, GRACE A
Art Unit
2144
Tech Center
2100 — Computer Architecture & Software
Assignee
NEC Laboratories America Inc.
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
437 granted / 573 resolved
+21.3% vs TC avg
Strong +18% interview lift
Without
With
+17.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
18 currently pending
Career history
596
Total Applications
across all art units

Statute-Specific Performance

§101
11.2%
-28.8% vs TC avg
§103
56.6%
+16.6% vs TC avg
§102
15.5%
-24.5% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 573 resolved cases

Office Action

§101 §102 §103
Detailed Action Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 2, 11, & 20 are objected to because of the following informalities: “Riemmanian” should read as “Riemannian”. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mathematical process) without significantly more. Regarding claim 1, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A computer-implemented method for model training”. A method is one of the four statutory categories of invention. In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, is directed to a mathematical process: “training… a task-based neural network classifier having a center and a covariance matrix for each of a plurality of classes in a last layer of the task- based neural network classifier and a plurality of convolutional layers preceding the last layer, by using a similarity between an image feature of a last convolutional layer from among the plurality of convolutional layers and the center and the covariance matrix for a given one of the plurality of classes, the similarity minimizing an impact of a data model forgetting problem” (In view of the specification at [0061-0071], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea. In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “A computer-implemented method for model training” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) “receiving, by a hardware processor, sets of images, each set corresponding to a respective task” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) “…by the hardware processor…” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements (ii) & (iv) recite use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional element (iii) recites an insignificant extra-solution activity. Further, element (iii) recites steps that store and retrieve information in memory which has been determined by the courts to recite a well-understood, routine, and conventional activity which is not indicative of significantly more (Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)). Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 2 recites the following additional abstract idea: “wherein the similarity is measured based on a distance calculation made using a Mahalanobis distance that induces a Riemmanian geometry” (In view of the specification at [0061-0071], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 3, it is dependent upon claim 2, and thereby incorporates the limitations of, and corresponding analysis applied to claim 2. Further, claim 3 recites the following additional abstract idea: “wherein the distance calculation for a given one of the plurality of classes is a squared distance of a difference between a sample representation of the image feature and a class-center of the given one of the plurality of classes multiplied by a decomposed covariance” (In view of the specification at [0061-0071], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 4, it is dependent upon claim 3, and thereby incorporates the limitations of, and corresponding analysis applied to claim 3. Further, claim 4 recites the following additional abstract idea: “wherein the distance calculation is a prediction” (In view of the specification at [0061-0071], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 5, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 5 recites the following additional abstract idea: “training the neural network classifier to minimize a cross-entropy loss between a prediction and a class label” (In view of the specification at [0075-0080], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 6, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 6 recites “adding a new center and covariance for the at least one of the plurality of new classes” (In step2A, prong 2, this recites mere instructions to apply the judicial exception (MPEP 2106.05(f).) In step 2B, mere instructions to apply the judicial exception is not indicative of significantly more.) Further, claim 6 recites “training the model using the new task to recognize the new task in the future” (In step2A, prong 2, this recites mere instructions to apply the judicial exception (MPEP 2106.05(f).) In step 2B, mere instructions to apply the judicial exception is not indicative of significantly more.) Further, claim 6 recites “receiving a new task to classify into at least one of a plurality of new classes” (In step 2A, prong 2, this recites insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g).) In step 2B, this recites transmitting/receiving data over a network, which the courts have found to be a well-understood, routine, and conventional activity (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362). Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 7, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 7 recites “wherein the neural network is trained using training data comprising respective pluralities of images pertaining to respective given tasks with distinct classes and domains” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 8, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 8 recites the following additional abstract ideas: “calculating a knowledge distillation loss by calculating a smooth transition coefficient between a current task-based neural network classifier and a prior task-based neural network classifier for a given task” (In view of the specification at [0081-0091], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) “further calculating an exponential moving average” (In view of the specification at [0081-0091], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 9, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 9 recites the following additional abstract idea: “wherein the task-based neural network classifier uses covariance to estimate a curvature between a mean in the task-based neural network classifier and an image feature from a new task” (In view of the specification at [0092-0098], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Regarding claim 10, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A computer program product for model training, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith”. A non-transitory computer readable medium is within one of the four statutory categories of invention. In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, is directed to a mathematical process: “training… a task-based neural network classifier having a center and a covariance matrix for each of a plurality of classes in a last layer of the task- based neural network classifier and a plurality of convolutional layers preceding the last layer, by using a similarity between an image feature of a last convolutional layer from among the plurality of convolutional layers and the center and the covariance matrix for a given one of the plurality of classes, the similarity minimizing an impact of a data model forgetting problem” (In view of the specification at [0061-0071], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea. In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “A computer program product for model training, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) “receiving, by a hardware processor, sets of images, each set corresponding to a respective task” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) “…by the hardware processor…” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements (ii) & (iv) recite use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional element (iii) recites an insignificant extra-solution activity. Further, element (iii) recites steps that store and retrieve information in memory which has been determined by the courts to recite a well-understood, routine, and conventional activity which is not indicative of significantly more (Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)). Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Regarding claims 11-18, they are dependent upon claim 10 and thereby incorporate the limitations of, and corresponding analysis applied to claim 10. Further, claims 11-18 comprise similar additional limitations as claims 2-9, respectively, and are rejected under the same rationale. Regarding claim 19, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A computer processing system for model training, comprising: a memory device for storing program code; and a hardware processor operatively coupled to the memory device for running the program code”. A system of the described configuration is within one of the four statutory categories of invention. In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, is directed to a mathematical process: “train a task-based neural network classifier having a center and a covariance matrix for each of a plurality of classes in a last layer of the task-based neural network classifier and a plurality of convolutional layers preceding the last layer, by using a similarity between an image feature of a last convolutional layer from among the plurality of convolutional layers and the center and the covariance matrix for a given one of the plurality of classes, the similarity minimizing an impact of a data model forgetting problem” (In view of the specification at [0061-0071], this seems to be directed to a mathematical process made up of mathematical equations and calculations, which recites an abstract idea (MPEP 2106).) If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea. In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: “A computer processing system for model training, comprising: a memory device for storing program code; and a hardware processor operatively coupled to the memory device for running the program code” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).) “receive sets of images, each set corresponding to a respective task” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional element (ii) recites use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional element (iii) recites an insignificant extra-solution activity. Further, element (iii) recites steps that store and retrieve information in memory which has been determined by the courts to recite a well-understood, routine, and conventional activity which is not indicative of significantly more (Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)). Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Regarding claim 20, it is dependent upon claim 19 and thereby incorporates the limitations of, and corresponding analysis applied to claim 19. Further, claim 20 comprises similar additional limitations as claim 2, and is rejected under the same rationale. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-7, & 9 are rejected under 35 U.S.C. 102(a)(1) as being clearly anticipated by Lee, K. et al. “A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks.” Available at https://arxiv.org/abs/1807.03888 on 27 October 2018 (hereafter, LEE) Regarding claim 1, LEE teaches “A computer-implemented method for model training, comprising receiving, by a hardware processor, sets of images, each set corresponding to a respective task”: ([Page 2, Contribution, Paragraphs 1-2] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs (Deep Neural networks being the models to be trained) utilizing the concept of a “generative” (distance-based) classifier… We demonstrate the effectiveness of the proposed method using deep convolutional neural networks, such as DenseNet [14] and ResNet [12] trained for image classification tasks (a neural network classifier trained by receiving sets of images that are each based on various tasks such as CIFAR which contains object recognition tasks such as birds (https://www.cs.toronto.edu/~kriz/cifar.html) or SVHN, which contains number recognition tasks on street-view homes (https://www.kaggle.com/datasets/stanfordu/street-view-house-numbers)”) Additionally, neural networks are by definition, “computer program products” which by nature, computer-implemented. Further, LEE teaches “training, by the hardware processor, a task-based neural network classifier having a center and a covariance matrix for each of a plurality of classes in a last layer of the task-based neural network classifier and a plurality of convolutional layers preceding the last layer, by using a similarity between an image feature of a last convolutional layer from among the plurality of convolutional layers and the center and the covariance matrix for a given one of the plurality of classes, the similarity minimizing an impact of a data model forgetting problem”: ([Page 2, Contribution, Paragraphs 1-2] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs utilizing the concept of a “generative” (distance-based) classifier… Under this assumption, we define the confidence score using the Mahalanobis distance with respect to the closest class-conditional distribution (a well-known mathematical calculation that calculates similarity using the mean (center) and the covariance matrix), where its parameters are chosen as empirical class means and tied empirical covariance of training samples… We demonstrate the effectiveness of the proposed method using deep convolutional neural networks (neural networks with a plurality of convolutional layers), such as DenseNet [14] and ResNet [12] trained for image classification tasks (a neural network classifier trained using image features) on various datasets including CIFAR [15], SVHN [28], ImageNet [5] and LSUN [32]…”) And further: ([Pages 2-3, Why Mahalanobis distance-based score?] “Let PNG media_image1.png 15 48 media_image1.png Greyscale be an input and PNG media_image2.png 13 26 media_image2.png Greyscale PNG media_image3.png 18 102 media_image3.png Greyscale be its label. Suppose that a pre-trained softmax neural classifier is given (a pre-trained classifier based on Mahalanobis distance would then have a center and a covariance matrix for its known classes): PNG media_image4.png 30 205 media_image4.png Greyscale where wc and bc are the weight and the bias of the softmax classifier for class c, and PNG media_image5.png 16 26 media_image5.png Greyscale denotes the output of the penultimate layer of DNNs (the last layer preceded by a plurality of convolutional layers) … To estimate the parameters of the generative classifier from the pre-trained softmax neural classifier, we compute the empirical class mean and covariance of training samples PNG media_image6.png 17 141 media_image6.png Greyscale PNG media_image7.png 37 463 media_image7.png Greyscale where Nc is the number of training samples with label c. This is equivalent to fitting the class-conditional Gaussian distributions with a tied covariance to training samples under the maximum likelihood estimator. Mahalanobis distance-based confidence score. Using the above induced class-conditional Gaussian distributions, we define the confidence score M(x) using the Mahalanobis distance between test sample x and the closest class-conditional Gaussian distribution, i.e., PNG media_image8.png 22 278 media_image8.png Greyscale (Here, we see the Mahalanobis distance being used, which compares the similarities of a test sample (image feature) using the center and covariance to the known plurality of classes’ center and covariance matrices. Here, f(x) is a sample representation of the image feature. PNG media_image9.png 15 16 media_image9.png Greyscale is the class mean/center. Σ is the tied covariance estimated from the entire training set, which is explicitly multiplied with the squared difference vector via its inverse Σ-1 which is algebraically equivalent to computing a squared norm in a whitened space (i.e., multiplication by a decomposed covariance factor such as Σ-1/2) and thus the result, is an estimated curvature between the mean and the image feature.)”) And further: ([Page 4, Algorithm 1] “ PNG media_image10.png 199 531 media_image10.png Greyscale ”) This algorithm illustrates the process of computing the confidence score using Mahalanobis distance, covariance and mean (center) And further: ([Page 5, Algorithm 2] PNG media_image11.png 153 534 media_image11.png Greyscale ”) This algorithm illustrates the process when used with class-incremental learning, to minimize data forgetting, by comparing the similarity of each training sample of the preceding layer with the known classes. Regarding claim 2, LEE teaches the limitations of claim 1. Further, LEE teaches “wherein the similarity is measured based on a distance calculation made using a Mahalanobis distance that induces a Riemmanian geometry”: ([Page 2, Contribution, Paragraph 1] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs utilizing the concept of a “generative” (distance-based) classifier… Under this assumption, we define the confidence score using the Mahalanobis distance with respect to the closest class-conditional distribution (a well-known mathematical calculation used in the last layer that calculates similarity using the mean (center) and the covariance matrix), where its parameters are chosen as empirical class means and tied empirical covariance of training samples…”) This shows that the similarity measure/confidence score is based on a distance calculation made using a Mahalanobis distance. And further: ([Page 5, Algorithm 2] PNG media_image11.png 153 534 media_image11.png Greyscale ”) The above equations show a covariance matrix Σ that varies smoothly, thus inducing a non-flat Riemannian metric / a Riemannian geometry. Regarding claim 3, LEE teaches the limitations of claim 2. Further, LEE teaches “wherein the distance calculation for a given one of the plurality of classes is a squared distance of a difference between a sample representation of the image feature and a class-center of the given one of the plurality of classes multiplied by a decomposed covariance”: ([Page 3, Equation 2] “ PNG media_image8.png 22 278 media_image8.png Greyscale ”) Here, f(x) is a sample representation of the image feature. PNG media_image9.png 15 16 media_image9.png Greyscale is the class center. The subtraction and squaring via quadratic form matches the claimed “squared distance of a difference between a sample representation of the image feature and a class-center of the given one of the plurality of classes”, where Σ is the tied covariance estimated from the entire training set, which is explicitly multiplied with the squared difference vector via its inverse Σ-1 which is algebraically equivalent to computing a squared norm in a whitened space (i.e., multiplication by a decomposed covariance factor such as Σ-1/2) Regarding claim 4, LEE teaches the limitations of claim 3. Further LEE teaches “wherein the distance calculation is a prediction”: ([Page 4, Paragraph 1] “We remark that this corresponds to predicting a class label using the posterior distribution from generative classifier with the uniform class prior. Interestingly, we found that the softmax accuracy (red bar) is also achieved by the Mahalanobis distance-based classifier (blue bar), while conventional knowledge is that a generative classifier trained from scratch typically performs much worse than a discriminative classifier such as softmax.”) This citation simply shows that a classifying data with a neural network classifier inherently corresponds with making a prediction, and thus, using a Mahalanobis distance to train the classifier is a prediction. Regarding claim 5, LEE teaches the limitations of claim 1. Further, LEE teaches “training the neural network classifier to minimize a cross-entropy loss between a prediction and a class label”: ([Page 7, Comparison of robustness, sentence 8] “…we remark that our method using softmax neural classifier trained by standard cross entropy loss (cross-entropy is the method used, which is the standard method for neural networks, aiming to minimize the loss between predictions and class labels) typically outperforms the ODIN using softmax neural classifier trained by confidence loss…”) Regarding claim 6, LEE teaches the limitations of claim 1. Further LEE teaches “receiving a new task to classify into at least one of a plurality of new classes; adding a new center and covariance for the at least one of the plurality of new classes; and training the model using the new task to recognize the new task in the future”: ([Page 5, Algorithm 2] “ PNG media_image11.png 153 534 media_image11.png Greyscale ”) This shows the use of the methods within class-incremental learning where new tasks are classified into new classes. You can also see the new class mean (center) and the covariance for the new class explicitly being calculated here. Finally, updates the shared covariance and returns the mean and covariance of all classes, thereby training the model using the new task to recognize the new task in the future. Regarding claim 7, LEE teaches the limitations of claim 1. Further LEE teaches “wherein the neural network is trained using training data comprising respective pluralities of images pertaining to respective given tasks with distinct classes and domains”: ([Page 2, Contribution, paragraph 2] “We demonstrate the effectiveness of the proposed method using deep convolutional neural networks, such as DenseNet [14] and ResNet [12] trained for image classification tasks on various datasets including CIFAR [15], SVHN [28], ImageNet [5] and LSUN [32]…”) Here, we see the datasets used as training data, including CIFAR, which comprises 10 classes of a specific domain (natural images) (https://www.cs.toronto.edu/~kriz/cifar.html), the SVHN dataset, which comprises 10 classes of a different domain (street numbers of houses) (https://www.kaggle.com/datasets/stanfordu/street-view-house-numbers) in addition to ImageNet and LSUN, thus providing evidence of training data comprising pluralities of images pertaining to respective given tasks (number recognition, object recognition such as birds or boats) with distinct classes and domains. Regarding claim 9, LEE teaches the limitations of claim 1. Further LEE teaches “wherein the task-based neural network classifier uses covariance to estimate a curvature between a mean in the task-based neural network classifier and an image feature from a new task”: ([Page 2, Contribution, Sentence 3] “Given deep neural networks (DNNs) with the softmax classifier, we propose a simple yet effective method for detecting abnormal samples such as out of-distribution (OOD) and adversarial ones. We first present the proposed confidence score based on an induced generative classifier under Gaussian discriminant analysis (GDA) (it can be understood by one skilled in the art that Gaussian discriminant analysis is a known technique that uses a covariance matrix to explicitly model the curvature of the feature distribution around the mean (center)), and then introduce additional techniques to improve its performance. We also discuss how the confidence score is applicable to incremental learning.”) And further: ([Page 3, Equation 2] “ PNG media_image8.png 22 278 media_image8.png Greyscale ”) Here, f(x) is a sample representation of the image feature. PNG media_image9.png 15 16 media_image9.png Greyscale is the class mean/center. Σ is the tied covariance estimated from the entire training set, which is explicitly multiplied with the squared difference vector via its inverse Σ-1 which is algebraically equivalent to computing a squared norm in a whitened space (i.e., multiplication by a decomposed covariance factor such as Σ-1/2) and thus the result, is an estimated curvature between the mean and the image feature. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over LEE, as applied to claims above, and further in view of Tarvainen, A. et al. “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results.” Available at https://arxiv.org/abs/1703.01780 on 16 April 2018 (hereafter, TARVAINEN) Regarding claim 8, LEE teaches the limitations of claim 1. LEE fails to explicitly teach “calculating a knowledge distillation loss by calculating a smooth transition coefficient between a current task-based neural network classifier and a prior task-based neural network classifier for a given task and further calculating an exponential moving average.” However, analogous art, TARVAINEN, does teach this: ([Pages 2-3, Section 2. Mean Teacher] “To overcome the limitations of Temporal Ensembling, we propose averaging model weights instead of predictions. Since the teacher model (a prior task-based neural network classifier) is an average of consecutive student models, we call this the Mean Teacher method (Figure 2). Averaging model weights over training steps tends to produce a more accurate model than using the final weights directly [19]. We can take advantage of this during training to construct better targets. Instead of sharing the weights with the student model (a current task-based neural network classifier), the teacher model uses the EMA weights of the student model (explicit calculation of an exponential moving average). Now it can aggregate information after every step instead of every epoch. In addition, since the weight averages improve all layer outputs, not just the top output, the target model has better intermediate representations. These aspects lead to two practical advantages over Temporal Ensembling: First, the more accurate target labels lead to a faster feedback loop between the student and the teacher models, resulting in better test accuracy. Second, the approach scales to large datasets and on-line learning. More formally, we define the consistency cost J (the knowledge distillation loss) as the expected distance between the prediction of the student model (with weights PNG media_image12.png 14 9 media_image12.png Greyscale and noise PNG media_image13.png 13 7 media_image13.png Greyscale ) and the prediction of the teacher model (with weights PNG media_image14.png 16 11 media_image14.png Greyscale and noise PNG media_image15.png 17 12 media_image15.png Greyscale ). PNG media_image16.png 29 260 media_image16.png Greyscale The difference between the II model, Temporal Ensembling, and Mean teacher is how the teacher predictions are generated. Whereas the II model uses PNG media_image14.png 16 11 media_image14.png Greyscale = PNG media_image12.png 14 9 media_image12.png Greyscale , and Temporal Ensembling approximates PNG media_image17.png 16 66 media_image17.png Greyscale with a weighted average of successive predictions, we define PNG media_image18.png 17 15 media_image18.png Greyscale at training step t as the EMA (explicit calculation of expected moving average) of successive PNG media_image12.png 14 9 media_image12.png Greyscale weights: PNG media_image19.png 19 143 media_image19.png Greyscale where PNG media_image20.png 12 11 media_image20.png Greyscale is a smoothing coefficient hyperparameter ( a smooth transition coefficient). An additional difference between the three algorithms is that the II model applies training to PNG media_image14.png 16 11 media_image14.png Greyscale whereas Temporal Ensembling and Mean Teacher treat it as a constant with regards to optimization.”) Here, the “consistency cost”, which is calculated as the mean squared error between teacher and student predictions, is equivalent to the claimed “knowledge distillation loss” calculated from “a smooth transition coefficient” ( PNG media_image20.png 12 11 media_image20.png Greyscale ) between a current task-based neural network classifier (student model PNG media_image21.png 17 11 media_image21.png Greyscale ) and a prior task-based neural network classifier (teacher model PNG media_image18.png 17 15 media_image18.png Greyscale ) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of LEE with the teachings of TARVAINEN because both references seek to improve the efficiency and accuracy of neural network model training. One of ordinary skill in the art would be motivated to do so because, as TARVAINEN points out, ([Section 2. Mean Teacher] “These aspects lead to two practical advantages over Temporal Ensembling: First, the more accurate target labels lead to a faster feedback loop between the student and the teacher models, resulting in better test accuracy. Second, the approach scales to large datasets and on-line learning”.) Claims 10-16 & 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over LEE, as applied to claims above, and further in view of Singh, H. “Everything you Need to Know About Hardware Requirements for Machine Learning.” Available at https://www.einfochips.com/blog/everything-you-need-to-know-about-hardware-requirements-for-machine-learning/ on 24 February 2019 (hereafter, SINGH) Regarding claim 10, LEE teaches “A computer program product for model training, the computer program product comprising… receiving, by a hardware processor, sets of images, each set corresponding to a respective task”: ([Page 2, Contribution, Paragraphs 1-2] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs (Deep Neural networks being the models to be trained) utilizing the concept of a “generative” (distance-based) classifier… We demonstrate the effectiveness of the proposed method using deep convolutional neural networks, such as DenseNet [14] and ResNet [12] trained for image classification tasks (a neural network classifier trained by receiving sets of images that are each based on various tasks such as CIFAR which contains object recognition tasks such as birds (https://www.cs.toronto.edu/~kriz/cifar.html) or SVHN, which contains number recognition tasks on street-view homes (https://www.kaggle.com/datasets/stanfordu/street-view-house-numbers)”) Additionally, Deep Neural Networks are, by definition, computer program products for model training. Further, LEE teaches “training, by the hardware processor, a task-based neural network classifier having a center and a covariance matrix for each of a plurality of classes in a last layer of the task-based neural network classifier and a plurality of convolutional layers preceding the last layer, by using a similarity between an image feature of a last convolutional layer from among the plurality of convolutional layers and the center and the covariance matrix for a given one of the plurality of classes, the similarity minimizing an impact of a data model forgetting problem”: ([Page 2, Contribution, Paragraphs 1-2] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs utilizing the concept of a “generative” (distance-based) classifier… Under this assumption, we define the confidence score using the Mahalanobis distance with respect to the closest class-conditional distribution (a well-known mathematical calculation that calculates similarity using the mean (center) and the covariance matrix), where its parameters are chosen as empirical class means and tied empirical covariance of training samples… We demonstrate the effectiveness of the proposed method using deep convolutional neural networks (neural networks with a plurality of convolutional layers), such as DenseNet [14] and ResNet [12] trained for image classification tasks (a neural network classifier trained using image features) on various datasets including CIFAR [15], SVHN [28], ImageNet [5] and LSUN [32]…”) And further: ([Pages 2-3, Why Mahalanobis distance-based score?] “Let PNG media_image1.png 15 48 media_image1.png Greyscale be an input and PNG media_image2.png 13 26 media_image2.png Greyscale PNG media_image3.png 18 102 media_image3.png Greyscale be its label. Suppose that a pre-trained softmax neural classifier is given (a pre-trained classifier based on Mahalanobis distance would then have a center and a covariance matrix for its known classes): PNG media_image4.png 30 205 media_image4.png Greyscale where wc and bc are the weight and the bias of the softmax classifier for class c, and PNG media_image5.png 16 26 media_image5.png Greyscale denotes the output of the penultimate layer of DNNs (the last layer preceded by a plurality of convolutional layers) … To estimate the parameters of the generative classifier from the pre-trained softmax neural classifier, we compute the empirical class mean and covariance of training samples PNG media_image6.png 17 141 media_image6.png Greyscale PNG media_image7.png 37 463 media_image7.png Greyscale where Nc is the number of training samples with label c. This is equivalent to fitting the class-conditional Gaussian distributions with a tied covariance to training samples under the maximum likelihood estimator. Mahalanobis distance-based confidence score. Using the above induced class-conditional Gaussian distributions, we define the confidence score M(x) using the Mahalanobis distance between test sample x and the closest class-conditional Gaussian distribution, i.e., PNG media_image8.png 22 278 media_image8.png Greyscale (Here, we see the Mahalanobis distance being used, which compares the similarities of a test sample (image feature) using the center and covariance to the known plurality of classes’ center and covariance matrices. Here, f(x) is a sample representation of the image feature. PNG media_image9.png 15 16 media_image9.png Greyscale is the class mean/center. Σ is the tied covariance estimated from the entire training set, which is explicitly multiplied with the squared difference vector via its inverse Σ-1 which is algebraically equivalent to computing a squared norm in a whitened space (i.e., multiplication by a decomposed covariance factor such as Σ-1/2) and thus the result, is an estimated curvature between the mean and the image feature.)”) And further: ([Page 4, Algorithm 1] “ PNG media_image10.png 199 531 media_image10.png Greyscale ”) This algorithm illustrates the process of computing the confidence score using Mahalanobis distance, covariance and mean (center) And further: ([Page 5, Algorithm 2] PNG media_image11.png 153 534 media_image11.png Greyscale ”) This algorithm illustrates the process when used with class-incremental learning, to minimize data forgetting, by comparing the similarity of each training sample of the preceding layer with the known classes. LEE fails to explicitly teach “a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method”. However, analogous art, SINGH, does teach this: ([Memory and Storage] “Machine learning models are getting larger, requiring more memory and storage capacity. High-bandwidth memory (HBM) and solid-state drives (SSDs) (non-transitory computer readable storage) are becoming critical for training and running (storing program instructions executable by a computer to perform the method) these models efficiently.”) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of LEE with the teachings of SINGH because LEE teaches optimal training and use methods for machine learning models while SINGH teaches the hardware requirements of machine learning models One of ordinary skill in the art would be motivated to do so because without meeting the hardware requirements for machine learning models, the operation of machine learning models will not be possible, and thus, the methods would not work. Regarding claims 11-16, & 18, LEE in view of SINGH teaches the limitations of claim 10. Further, claims 11-16, & 18 recite similar additional limitations as claims 2-7, & 9, respectively, and are rejected under the same rationale. Regarding claim 19, LEE teaches “receive sets of images, each set corresponding to a respective task”: ([Page 2, Contribution, Paragraphs 1-2] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs (Deep Neural networks being the models to be trained) utilizing the concept of a “generative” (distance-based) classifier… We demonstrate the effectiveness of the proposed method using deep convolutional neural networks, such as DenseNet [14] and ResNet [12] trained for image classification tasks (a neural network classifier trained by receiving sets of images that are each based on various tasks such as CIFAR which contains object recognition tasks such as birds (https://www.cs.toronto.edu/~kriz/cifar.html) or SVHN, which contains number recognition tasks on street-view homes (https://www.kaggle.com/datasets/stanfordu/street-view-house-numbers)”) Further, LEE teaches “train a task-based neural network classifier having a center and a covariance matrix for each of a plurality of classes in a last layer of the task-based neural network classifier and a plurality of convolutional layers preceding the last layer, by using a similarity between an image feature of a last convolutional layer from among the plurality of convolutional layers and the center and the covariance matrix for a given one of the plurality of classes, the similarity minimizing an impact of a data model forgetting problem”: ([Page 2, Contribution, Paragraphs 1-2] “…Our high-level idea is to measure the probability density of test sample on feature spaces of DNNs utilizing the concept of a “generative” (distance-based) classifier… Under this assumption, we define the confidence score using the Mahalanobis distance with respect to the closest class-conditional distribution (a well-known mathematical calculation that calculates similarity using the mean (center) and the covariance matrix), where its parameters are chosen as empirical class means and tied empirical covariance of training samples… We demonstrate the effectiveness of the proposed method using deep convolutional neural networks (neural networks with a plurality of convolutional layers), such as DenseNet [14] and ResNet [12] trained for image classification tasks (a neural network classifier trained using image features) on various datasets including CIFAR [15], SVHN [28], ImageNet [5] and LSUN [32]…”) And further: ([Pages 2-3, Why Mahalanobis distance-based score?] “Let PNG media_image1.png 15 48 media_image1.png Greyscale be an input and PNG media_image2.png 13 26 media_image2.png Greyscale PNG media_image3.png 18 102 media_image3.png Greyscale be its label. Suppose that a pre-trained softmax neural classifier is given (a pre-trained classifier based on Mahalanobis distance would then have a center and a covariance matrix for its known classes): PNG media_image4.png 30 205 media_image4.png Greyscale where wc and bc are the weight and the bias of the softmax classifier for class c, and PNG media_image5.png 16 26 media_image5.png Greyscale denotes the output of the penultimate layer of DNNs (the last layer preceded by a plurality of convolutional layers) … To estimate the parameters of the generative classifier from the pre-trained softmax neural classifier, we compute the empirical class mean and covariance of training samples PNG media_image6.png 17 141 media_image6.png Greyscale PNG media_image7.png 37 463 media_image7.png Greyscale where Nc is the number of training samples with label c. This is equivalent to fitting the class-conditional Gaussian distributions with a tied covariance to training samples under the maximum likelihood estimator. Mahalanobis distance-based confidence score. Using the above induced class-conditional Gaussian distributions, we define the confidence score M(x) using the Mahalanobis distance between test sample x and the closest class-conditional Gaussian distribution, i.e., PNG media_image8.png 22 278 media_image8.png Greyscale (Here, we see the Mahalanobis distance being used, which compares the similarities of a test sample (image feature) using the center and covariance to the known plurality of classes’ center and covariance matrices. Here, f(x) is a sample representation of the image feature. PNG media_image9.png 15 16 media_image9.png Greyscale is the class mean/center. Σ is the tied covariance estimated from the entire training set, which is explicitly multiplied with the squared difference vector via its inverse Σ-1 which is algebraically equivalent to computing a squared norm in a whitened space (i.e., multiplication by a decomposed covariance factor such as Σ-1/2) and thus the result, is an estimated curvature between the mean and the image feature.)”) And further: ([Page 4, Algorithm 1] “ PNG media_image10.png 199 531 media_image10.png Greyscale ”) This algorithm illustrates the process of computing the confidence score using Mahalanobis distance, covariance and mean (center) And further: ([Page 5, Algorithm 2] PNG media_image11.png 153 534 media_image11.png Greyscale ”) This algorithm illustrates the process when used with class-incremental learning, to minimize data forgetting, by comparing the similarity of each training sample of the preceding layer with the known classes. LEE fails to explicitly teach “A computer processing system for model training, comprising: a memory device for storing program code; and a hardware processor operatively coupled to the memory device for running the program code”. However, analogous art, SINGH, does teach this: ([Memory and Storage] “Machine learning models are getting larger, requiring more memory and storage capacity. High-bandwidth memory (HBM) and solid-state drives (SSDs) (a memory device) are becoming critical for training and running (storing and running program code) these models efficiently.”) And further: ([So how can we make the training model faster?, Paragraph 1] “This can be accomplished simply by performing all the operations at the same time, instead of taking them one after the other (all processes operatively coupled for running the program code). This is where the GPU (a hardware processor) comes into the picture, with several thousand cores designed to compute with almost 100% efficiency. Turns out these processors are suited to perform the computation of neural networks as well.”) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of LEE with the teachings of SINGH because LEE teaches optimal training and use methods for machine learning models while SINGH teaches the hardware requirements of machine learning models One of ordinary skill in the art would be motivated to do so because without meeting the hardware requirements for machine learning models, the operation of machine learning models will not be possible, and thus, the methods would not work. Regarding claim 20, LEE in view of SINGH teaches the limitations of claim 19. Further, claim 20 comprises similar additional limitations as claim 2, and is rejected under the same rationale. Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over LEE in view of SINGH, as applied to claims above, and further in view of TARVAINEN. Regarding claim 17, LEE in view of SINGH teaches the limitations of claim 10. LEE in view of SINGH fails to explicitly teach “calculating a knowledge distillation loss by calculating a smooth transition coefficient between a current task-based neural network classifier and a prior task-based neural network classifier for a given task and further calculating an exponential moving average.” However, analogous art, TARVAINEN, does teach this: ([Pages 2-3, Section 2. Mean Teacher] “To overcome the limitations of Temporal Ensembling, we propose averaging model weights instead of predictions. Since the teacher model (a prior task-based neural network classifier) is an average of consecutive student models, we call this the Mean Teacher method (Figure 2). Averaging model weights over training steps tends to produce a more accurate model than using the final weights directly [19]. We can take advantage of this during training to construct better targets. Instead of sharing the weights with the student model (a current task-based neural network classifier), the teacher model uses the EMA weights of the student model (explicit calculation of an exponential moving average). Now it can aggregate information after every step instead of every epoch. In addition, since the weight averages improve all layer outputs, not just the top output, the target model has better intermediate representations. These aspects lead to two practical advantages over Temporal Ensembling: First, the more accurate target labels lead to a faster feedback loop between the student and the teacher models, resulting in better test accuracy. Second, the approach scales to large datasets and on-line learning. More formally, we define the consistency cost J (the knowledge distillation loss) as the expected distance between the prediction of the student model (with weights PNG media_image12.png 14 9 media_image12.png Greyscale and noise PNG media_image13.png 13 7 media_image13.png Greyscale ) and the prediction of the teacher model (with weights PNG media_image14.png 16 11 media_image14.png Greyscale and noise PNG media_image15.png 17 12 media_image15.png Greyscale ). PNG media_image16.png 29 260 media_image16.png Greyscale The difference between the II model, Temporal Ensembling, and Mean teacher is how the teacher predictions are generated. Whereas the II model uses PNG media_image14.png 16 11 media_image14.png Greyscale = PNG media_image12.png 14 9 media_image12.png Greyscale , and Temporal Ensembling approximates PNG media_image17.png 16 66 media_image17.png Greyscale with a weighted average of successive predictions, we define PNG media_image18.png 17 15 media_image18.png Greyscale at training step t as the EMA (explicit calculation of expected moving average) of successive PNG media_image12.png 14 9 media_image12.png Greyscale weights: PNG media_image19.png 19 143 media_image19.png Greyscale where PNG media_image20.png 12 11 media_image20.png Greyscale is a smoothing coefficient hyperparameter ( a smooth transition coefficient). An additional difference between the three algorithms is that the II model applies training to PNG media_image14.png 16 11 media_image14.png Greyscale whereas Temporal Ensembling and Mean Teacher treat it as a constant with regards to optimization.”) Here, the “consistency cost”, which is calculated as the mean squared error between teacher and student predictions, is equivalent to the claimed “knowledge distillation loss” calculated from “a smooth transition coefficient” ( PNG media_image20.png 12 11 media_image20.png Greyscale ) between a current task-based neural network classifier (student model PNG media_image21.png 17 11 media_image21.png Greyscale ) and a prior task-based neural network classifier (teacher model PNG media_image18.png 17 15 media_image18.png Greyscale ) It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of LEE in view of SINGH with the teachings of TARVAINEN because both references seek to improve the efficiency and accuracy of neural network model training. One of ordinary skill in the art would be motivated to do so because, as TARVAINEN points out, ([Section 2. Mean Teacher] “These aspects lead to two practical advantages over Temporal Ensembling: First, the more accurate target labels lead to a faster feedback loop between the student and the teacher models, resulting in better test accuracy. Second, the approach scales to large datasets and on-line learning”.) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW LEE LEWIS whose telephone number is (571)272-1906. The examiner can normally be reached Monday: 9:30AM - 3:30PM and Tuesday - Friday: 9:30AM - 6PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Matthew Lee Lewis/Examiner, Art Unit 2144 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Oct 21, 2022
Application Filed
Oct 01, 2025
Non-Final Rejection mailed — §101, §102, §103
Dec 18, 2025
Interview Requested
Jan 02, 2026
Response Filed
Sep 30, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748944
SYSTEM AND METHOD FOR MACHINE LEARNING ARCHITECTURE FOR OUT-OF-DISTRIBUTION DATA DETECTION
4y 8m to grant Granted Sep 29, 2026
Patent 12748819
REDUCING UTILIZATION OF COMPUTATIONAL RESOURCES ASSOCIATED WITH SEGMENTING DATASETS VIA A CLUSTER- ENSEMBLE MODEL SYSTEMS AND METHODS
2y 11m to grant Granted Sep 29, 2026
Patent 12737646
TIERED ANOMALY DETECTION
3y 3m to grant Granted Sep 15, 2026
Patent 12711430
PROCESSORS AND METHODS FOR SELECTING A TARGET MODEL FOR AN UNLABELED DATASET
3y 7m to grant Granted Aug 18, 2026
Patent 12694259
LOW POWER MULTI-STAGE SELECTABLE NEURAL NETWORK SUPPRESSION
4y 10m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
94%
With Interview (+17.6%)
3y 4m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 573 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month