Prosecution Insights
Last updated: October 02, 2026
Application No. 18/189,039

MACHINE LEARNING MODEL TRAINING FOR IMPROVING ANOMALY DETECTION

Final Rejection §103
Filed
Mar 23, 2023
Examiner
KAPOOR, DEVAN
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
UnitedHealth Group Incorporated
OA Round
2 (Final)
7%
Grant Probability
At Risk
3-4
OA Rounds
9m
Est. Remaining
18%
With Interview

Examiner Intelligence

Grants only 7% of cases
7%
Career Allowance Rate
1 granted / 14 resolved
-47.9% vs TC avg
Moderate +11% lift
Without
With
+11.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
29 currently pending
Career history
47
Total Applications
across all art units

Statute-Specific Performance

§101
34.0%
-6.0% vs TC avg
§103
57.4%
+17.4% vs TC avg
§102
5.8%
-34.2% vs TC avg
§112
2.2%
-37.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the application filed on 05/21/026. Claims 1-20 are pending and have been examined. This action is Final Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments: Argument 1: The applicant argues on pages ~11-17 of the remarks that amended claims 1-20 are patent eligible because they recite a specific improvement to machine-learning training rather than merely mathematical concepts. Using claim 1 as exemplary, the applicant emphasizes generating normal and anomaly autoencoder losses, generating a global classification loss, combining those values through a weighted composite loss, and then adjusting parameters of the anomaly-detection model using that composite loss to improve performance on class-imbalanced datasets. Under Step 2A Prong One, the applicant contends these are concrete model-training operations that may involve mathematics but do not themselves recite a mathematical concept. Under Prong Two, the applicant relies heavily on Ex parte Desjardins and the 2025 Desjardins Memorandum, arguing that improving model training and system performance through ML-parameter adjustments constitutes a technological improvement and practical application; specifically, the claimed technique allegedly balances reconstruction accuracy and classification discrimination to provide more stable updates and better accuracy/reliability for underrepresented classes. Under Step 2B, the applicant alternatively argues that this particular combination of loss generation, weighting, and parameter optimization is unconventional and provides an inventive concept. Claims 8 and 15, and their dependents, are asserted eligible for substantially the same reasons because they recite corresponding system and computer-readable-medium implementations. Response to Argument 1: The argument has considered the argument set forth above. With the amendments and remarks considered, the examiner has considered the argument persuasive and thus the rejection is withdrawn. Argument 2: The applicant argues on pages 18-19 of the remarks traversal of the rejections of claims 1-3, 6-10, 13-17, and 20 over Ruff in view of Huang, and claims 4-5, 11-12, and 18-19 over Ruff/Huang/Liu, focusing primarily on the newly amended independent-claim limitation. The applicant argues that the references do not teach a global classification loss representing the difference between the labeled classification parameter and a predicted classification parameter where that predicted classification itself is based on the labeled training object, the normal prediction loss, and the anomaly prediction loss. In particular, the applicant attacks the Office's reliance on Huang, asserting that Huang's total loss merely combines reconstruction, consistency/normalization, and assistant losses and does not disclose the claimed classification loss architecture in which the prediction incorporates the normal and anomaly autoencoder losses. The applicant further argues that Ruff and Liu do not cure this deficiency. Because claims 8 and 15 contain substantially similar limitations to claim 1, the applicant extends the same argument to those independent claims and, through dependency, to all remaining claims, requesting withdrawal of all 103 rejections. Response to Argument 2: The applicant’s arguments are not persuasive because the present rejection no longer relies on Ruff and Huang to establish the newly amended relationship, in light of the amendments, between the normal and anomaly prediction losses and the predicted classification parameter. As set forth in the updated rejection, Lübbering teaches the underlying anomaly-detection architecture, including separate inlier and rest-sample reconstruction losses, mapping reconstruction error to a predicted class probability, calculating a binary cross-entropy classification loss between the target label and predicted probability, combining the reconstruction and classification losses into a weighted overall loss, and updating the autoencoder parameters using that loss in a class-imbalanced setting. Although Lübbering does not expressly teach that the predicted classification is based on both the normal and anomaly prediction loss parameters, Zhang cures that deficiency by teaching class-specific autoencoders and determining the predicted label by calculating the reconstruction error for each respective autoencoder and assigning the input to the class having the minimum reconstruction error. Thus, in the claimed normal/anomaly implementation, the predicted classification is determined using the respective normal and anomaly reconstruction-loss information. Accordingly, when Zhang’s class-specific reconstruction-error classification technique is incorporated into Lübbering’s system, Lübbering’s binary cross-entropy loss is calculated from a predicted classification that is itself based on the respective normal and anomaly reconstruction losses, thereby teaching the amended global-classification-loss relationship. The remaining dependent-claim limitations are separately taught by Ruff, Huang, and Liu as set forth in the updated mappings. Therefore, the applicant’s criticism of the prior Ruff/Huang/Liu mapping does not overcome the presently applied combination, and the 103 rejections of claims 1-20 are maintained. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 8 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over NPL reference “Bounding Open Space Risk with Decoupling Autoencoders in Open Set Recognition.”, by Lübbering et. al. (referred herein as Lübbering) in view of NPL reference “Robust Class-Specific Autoencoder for Data Cleaning and Classification in the Presence of Label Noise”, by Zhang et. al. (referred herein as Zhang). Regarding claim 1, Lübbering teaches: A computer-implemented method comprising: receiving, by one or more processors, a plurality of labeled training data objects, wherein (a) a labeled training data object of the plurality of labeled training data objects is associated with one of a plurality of labeled classification parameters, and (b) a labeled classification parameter of the plurality of labeled classification parameters indicates at least one of a normal classification label or an anomaly classification label; ([Lübbering, page 354, sec. 2] “However, there are common situations where OVR becomes relevant, e.g., if faced with only positively labeled samples and all remaining samples with potentially unknown sources are assigned to a single negative class [16,31,66] or if the goal, as in this paper, is to filter normal samples from abnormal ones [6].” AND [Lübbering, page 356-357, Sec. 4] “During training of DNNs, we minimize a surrogate loss function L, such as negative log-likelihood (NLL) instead of the non-differentiable 0-1 loss, over a given empirical data distribution…where N is the training set size and f the model with parameters Θ.”, wherein the examiner interprets the positively labeled samples and samples assigned to a single negative class to be the same as labeled training data objects associated with respective labeled classification parameters because they are both directed to training samples associated with respective classification labels. The examiner further interprets normal samples and abnormal ones to be the same as samples associated with a normal classification label and an anomaly classification label because they are both directed to distinguishing data belonging to a normal class from data belonging to an abnormal class. Lastly, the examiner interprets training the disclosed DNN model over the training set to be the same as computer-implemented processing by one or more processors because they are both directed to computational execution of machine-learning operations on training data). generating, by the one or more processors and using an anomaly detection machine learning model and the plurality of labeled training data objects, (a) a normal prediction loss parameter associated with the normal classification label and (b) an anomaly prediction loss parameter associated with the anomaly classification label, wherein the anomaly detection machine learning model comprises an autoencoder and the normal prediction loss parameter and the anomaly prediction loss parameter comprise autoencoder losses; ([Lübbering, page 355, sec. 3] “Similar to existing autoencoder-based approaches for outlier detection [8,26,45], the decoupling autoencoder (DAE) method learns the outlierness of a sample via its reconstruction error…From an architectural point of view, as displayed in Fig. 3, the network reconstructs a sample x ∈ Rn using the autoencoder ϕ(x) = d(e(x)) consisting of an encoder e(·) and a decoder d(·)…To minimize inlier reconstruction errors and maximize rest sample reconstruction errors, such that the inlier samples are easily distinguishable from rest samples within the one-dimensional reconstruction error space.”, wherein the examiner interprets the decoupling autoencoder that learns the outlierness of a sample via reconstruction error to be the same as the anomaly detection machine learning model comprising an autoencoder because they are both directed to an autoencoder model used to distinguish outlier or anomalous data; and the examiner interprets inlier reconstruction errors and rest sample reconstruction errors to be the same as the normal prediction loss parameter and anomaly prediction loss parameter, respectively, because they are both directed to autoencoder reconstruction losses associated with the respective normal/inlier and rest/anomalous classifications). generating, by the one or more processors and using a classification prediction machine learning model, a global classification loss parameter based on the plurality of labeled training data objects, wherein the global classification loss parameter indicates a difference level between the labeled classification parameter and a predicted classification parameter produced by the classification prediction machine learning model based on the labeled training data object; ([Lübbering, page 353] “To this end, we propose the decoupling autoencoder (DAE) method, a novel autoencoder-based architecture that learns a radial basis function (RBF) kernel mapping the reconstruction error to class probabilities.” AND [Lübbering, page 355, Sec. 3] “The reconstruction error is mapped to the inlier probability via Gaussian g…The overall loss function Lˆ incorporates these three training objectives by combining the adversarial loss function LR, binary cross-entropy (BCE) classification loss LBCE(yˆ, y) = -[y ln yˆ + (1 - y)ln(1 - ˆy)] and a regularizer term |t|, as follows:…where y and yˆ denote the target label and the model’s predicted COI probability of sample x, respectively.”, wherein the examiner interprets the radial basis function kernel mapping the reconstruction error to class probabilities to be the same as the classification prediction machine learning model because they are both directed to generating a predicted classification value for an input data object; the examiner interprets the target label y to be the same as the labeled classification parameter because they are both directed to the known classification associated with the data object. Further, the examiner interprets the model’s predicted COI probability yˆ to be the same as the predicted classification parameter because they are both directed to a classification value predicted for the data object; and that the binary cross-entropy classification loss calculated from y and yˆ to be the same as the global classification loss parameter indicating a difference level between the labeled classification parameter and the predicted classification parameter because they are both directed to measuring error between a target classification and a predicted classification). generating, by the one or more processors, a composite loss parameter based on the normal prediction loss parameter, the global classification loss parameter, a normal prediction weight parameter, and a global classification weight parameter; ([Lübbering, page 355, Sec. 3] “The overall loss function Lˆ incorporates these three training objectives by combining the adversarial loss function LR, binary cross-entropy (BCE) classification loss LBCE(yˆ, y) = -[y ln yˆ + (1 - y)ln(1 - ˆy)] and a regularizer term |t|, as follows:” AND “Factors λ2 and λ3 scale the classification loss term and |t| regularization.” AND [Lübbering, page 356, Sec. 3] “where scaling factors λ0 ∈ R+ and λ1 ∈ R+ determine the minimization/maximization magnitude, respectively.”, wherein the examiner interprets the overall loss function Lˆ to be the same as the composite loss parameter because they are both directed to a training loss formed from multiple loss components; the examiner interprets LR1 and λ0 for inlier samples to be the same as the normal prediction loss parameter and normal prediction weight parameter because they are both directed to the normal/inlier reconstruction loss and a scaling parameter applied to that reconstruction loss; and the examiner interprets LBCE and λ2 to be the same as the global classification loss parameter and global classification weight parameter because they are both directed to the classification loss and a scaling parameter applied to the classification loss within the overall training loss). initiating, by the one or more processors, the performance of one or more prediction-based operations based on the anomaly detection machine learning model and the composite loss parameter, wherein the one or more prediction-based operations comprise adjusting a parameter of the anomaly detection machine learning model based on the composite loss parameter in order to improve the performance of the anomaly detection machine learning model with respect to class-imbalanced training datasets. ([Lübbering, page 353, sec. 1] “The inlier and outlier distributions are separated by a decision boundary that is optimized end-to-end to be as close as possible to the inlier distribution. Thus, observed and unobserved rest samples can be effectively rejected, resulting in an increased robustness.” AND [Lübbering, page 356, Sec. 3] “where Θ are the network weights of autoencoder ϕ…. Thus, the expression in Eq. (8) zeroes out, which is why the gradient updates become ineffective for inliers with large eMSE. As LR is independent of Gaussian g, it is not affected by this problem and enforces convergence by minimizing these inliers.” AND [Lübbering, page 371, sec. 7.4] “the combination of all loss terms in Lˆ yields the best separation of inliers and outliers due to effective minimization of inliers and maximization of outliers and solves the vanishing gradient problems.”, wherein the examiner interprets the network weights Θ to be the same as a parameter of the anomaly detection machine learning model because they are both directed to trainable parameters of the autoencoder; the examiner interprets the gradient updates and end-to-end optimization using the disclosed overall loss function to be the same as adjusting a parameter of the anomaly detection machine learning model based on the composite loss parameter because they are both directed to updating trainable autoencoder parameters according to the combined training loss; and the examiner interprets the resulting increased robustness and improved separation of inliers and outliers to be the same as improving the performance of the anomaly detection machine learning model because they are both directed to improving the model's ability to distinguish normal/inlier data from anomalous/rest data). with respect to class-imbalanced training datasets. ([Lübbering, page 352, sec 1.] “In OSR, we usually deal with significant class imbalance. When the focus shifts toward outlier detection, for instance, in the case of computer virus detection, inlier samples outnumber instances of the rest class. Further, only some classes of the problem domain are usually known for the set of rest classes (RCs), and outliers or dataset shifts are only witnessed at test time, making OSR a highly imbalanced semi-supervised setting.”, wherein the examiner interprets the disclosed setting in which inlier samples outnumber instances of the rest class and which is characterized by significant class imbalance to be the same as class-imbalanced training datasets because they are both directed to datasets having unequal representation of the respective classes). Lübbering does not teach a predicted classification parameter produced by the classification prediction machine learning model based on the labeled training data object, the normal prediction loss parameter, and the anomaly prediction loss parameter;. Zhang teaches a predicted classification parameter produced by the classification prediction machine learning model based on the labeled training data object, the normal prediction loss parameter, and the anomaly prediction loss parameter; ([Zhang, page 5, sec. 3.1] “Assume that we have N labeled training data, {xi, yi},i = 1,..., N in K classes, and some portions of the training data contain label noise, i.e., their labels are incorrectly annotated. Our goal is to learn a classifier f that assigns a new data point x to one of the K categories, in spite of the data contamination by label noise.” AND [Zhang, page 6, sec. 3.2] “For N labeled training data, {xi, yi},i = 1,..., N in K classes, we first separate them based on their corresponding label (note that some of the samples actually do not belong to it due to label noise), and then learn separately autoencoder by the loss function shown in Eq. (3):” AND [Zhang, page 9, Sec. 3.3] “As the proposed class-specific autoencoder has effectively reduced the influence of the label noise, meanwhile, the reconstruction error can distinguish outliers from normal data, the data cleaning and classification tasks can be solved in a unified way based on the minimum reconstruction error criterion. Particularly, for the classification task, we predict the label yt of a test data xt by simply calculating the reconstruction error on each autoencoder and then assign it to the class with minimum reconstruction error, as follows:”, wherein the examiner interprets the classifier based on learned class-specific autoencoders to be the same as the classification prediction machine learning model because they are both directed to a learned machine-learning arrangement that generates a predicted class for an input data object. Furthermore, the examiner interprets the reconstruction error calculated on each class-specific autoencoder to be the same as the normal prediction loss parameter and anomaly prediction loss parameter when the respective classes comprise normal and anomaly classes because they are both directed to class-associated autoencoder reconstruction-loss values. Also, the examiner interprets predicting the label by calculating the reconstruction error on each autoencoder and assigning the data object to the class having the minimum reconstruction error to be the same as producing the predicted classification parameter based on the normal prediction loss parameter and the anomaly prediction loss parameter because they are both directed to determining a predicted classification using reconstruction-loss information associated with each respective class). Lübbering, Zhang, and the instant application are analogous art because they are both directed to machine-learning classification using autoencoders and reconstruction-error information. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the reconstruction-error-based classification technique disclosed by Lübbering to include the class-specific reconstruction-error classification technique disclosed by Zhang, such that reconstruction-loss information associated with the respective normal and anomaly classes is used in determining the predicted classification. One would have been motivated to make such a modification to improve class discrimination by utilizing reconstruction information learned for each respective class, as suggested by Zhang ([Zhang, page 9, sec. 3.3] “the reconstruction error can distinguish outliers from normal data, the data cleaning and classification tasks can be solved in a unified way based on the minimum reconstruction error criterion…we predict the label yt of a test data xt by simply calculating the reconstruction error on each autoencoder and then assign it to the class with minimum reconstruction error”). Claim 8 and 15 are analogous to claim 1, aside from claim type and minute differences, and thus the same rejection as above. Claim(s) 2-3, 9-10, and 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lübbering in view of Zhang further in view of NPL reference “Deep semi-supervised anomaly detection”, by Ruff et. al. (referred herein as Ruff). Regarding claim 2, Lübbering and Zhang teach the computer-implemented method of claim 1 (see rejection of claim 1). Lübbering and Zhang do not teach selecting a plurality of normal training data objects and a plurality of anomaly training data objects from the plurality of labeled training data objects, wherein (a) a normal training data object of the plurality of normal training data objects is associated with the normal classification label, and (b) an anomaly training data object of the plurality of anomaly training data objects is associated with the anomaly classification label. Ruff teaches selecting a plurality of normal training data objects and a plurality of anomaly training data objects from the plurality of labeled training data objects; ([Ruff, page 6] “In every setup, we set one of the ten classes to be the normal class and let the remaining nine classes represent anomalies. We use the original training data of the respective normal class as the unlabeled part of our training set. Thus we start with a clean AD setting that fulfills the assumption that most (in this case all) unlabeled samples are normal. The training data of the respective nine anomaly classes then forms the data pool from which we draw anomalies for training to create different scenarios.”, wherein the examiner interprets “use the original training data of the respective normal class” to be the same as selecting a plurality of normal training data objects because they are both directed to selecting normal training samples; and the examiner interprets “data pool from which we draw anomalies for training” to be the same as selecting a plurality of anomaly training data objects because they are both directed to selecting anomaly samples for training). Lübbering, Zhang, Ruff, and the instant application are analogous art because they are all directed to machine-learning classification and/or anomaly detection using training data associated with normal and anomalous classifications. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of claim 1 disclosed by Lübbering and Zhang to include the selection of normal training data and anomaly training data from the available training data disclosed by Ruff. One would have been motivated to make such a modification in order to provide respective normal and anomalous training examples for training an anomaly detection model to distinguish normal samples from anomalous samples, as suggested by Ruff ([Ruff, page 6] “In every setup, we set one of the ten classes to be the normal class and let the remaining nine classes represent anomalies. We use the original training data of the respective normal class as the unlabeled part of our training set. Thus we start with a clean AD setting that fulfills the assumption that most (in this case all) unlabeled samples are normal. The training data of the respective nine anomaly classes then forms the data pool from which we draw anomalies for training to create different scenarios.”) Lübbering, Zhang, and Ruff do not teach wherein (a) a normal training data object of the plurality of normal training data objects is associated with the normal classification label, and (b) an anomaly training data object of the plurality of anomaly training data objects is associated with the anomaly classification label. Huang teaches wherein (a) a normal training data object of the plurality of normal training data objects is associated with the normal classification label, and (b) an anomaly training data object of the plurality of anomaly training data objects is associated with the anomaly classification label ([Huang, page 4] “Given the input space X consisting of normal data XN and anomalous data XA, where X = XN ∪ XA. For semi-supervised anomaly detection (AD), we are given n unlabeled samples xu1,···,xun ∈ X and m labeled samples (xl1, y1),···,(xlm, ym) ∈ X ×Y with Y = {-1,1} where y = 1 denotes normal samples and y = -1 denotes anomalous samples.”, wherein the examiner interprets labeled samples having y = 1 to be the same as normal training data objects associated with the normal classification label because they are both directed to labeled training samples identified as normal; and labeled samples having y = -1 to be the same as anomaly training data objects associated with the anomaly classification label because they are both directed to labeled training samples identified as anomalous). Lübbering, Zhang, Ruff, Huang, and the instant application are analogous art because they are all directed to machine-learning classification or anomaly detection using training data associated with normal and anomalous classifications. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Lübbering and Zhang to include the training-data selection and label-association techniques disclosed by Ruff and Huang. One would have been motivated to do so to organize training samples according to normal and anomaly classes for training an anomaly-detection model, as suggested by Ruff ([Ruff, page 6] “we set one of the ten classes to be the normal class and let the remaining nine classes represent anomalies”), furthermore by Huang ([Huang, page 4] “y = 1 denotes normal samples and y = -1 denotes anomalous samples”). Claim 9 and 16 are analogous to claim 2, aside from claim type and minute differences, and thus the same rejection as above. Regarding claim 3, Lübbering, Zhang, Ruff, and Huang teach the computer-implemented method of claim 2 (see rejection of claim 2). Ruff further teaches wherein a normal training data object count associated with the plurality of normal training data objects is larger than an anomaly training data object count associated with the plurality of anomaly training data objects; ([Ruff, page 1] “Typically AD methods attempt to learn a ‘compact’ description of the data in an unsupervised manner assuming that most of the samples are normal (i.e., not anomalous).”, wherein the examiner interprets “most of the samples are normal (i.e., not anomalous)” to be the same as the normal training data object count being larger than the anomaly training data object count because they are both directed to a dataset containing more normal samples than anomalous samples). Lübbering, Zhang, Ruff, Huang, and the instant application are analogous art because they are all directed to machine-learning anomaly detection and/or classification using normal and anomalous training data. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of claim 2 disclosed by Lübbering, Zhang, Ruff, and Huang to include the number of normal training data objects is greater than the number of anomaly training data objects, as disclosed by Ruff. One would have been motivated to do so efficiently teach that anomaly detection commonly operates under the assumption that most samples are normal, thereby providing a training-data distribution suitable for learning a representation of normal data while detecting less frequently occurring anomalous data, as suggested by Ruff (([Ruff, page 1] “Typically AD methods attempt to learn a ‘compact’ description of the data in an unsupervised manner assuming that most of the samples are normal (i.e., not anomalous).”Claim 10 and 17 are analogous to claim 3, aside from claim type and minute differences, and thus the same rejection as above. Claim(s) 4-5, 11-12, and 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lübbering in view of Zhang in view of NPL reference “ESAD: End-to-End Semi-Supervised Anomaly Detection.”, by Huang et. al. (referred herein as Huang) further in view of NPL reference “Semi-Supervised Anomaly Detection with Dual Prototypes Autoencoder for Industrial Surface Inspection.”, by Liu et. al. (referred herein as Liu). Regarding claim 4, Lübbering and Zhang teach the computer-implemented method of claim 1 (see rejection of claim 1). Lübbering and Zhang do not teach from the plurality of labeled training data objects that is associated with the normal classification label; wherein the normal prediction loss parameter indicates a reconstruction loss measure associated with the anomaly detection machine learning model in reconstructing a plurality of normal training data objects from the plurality of labeled training data objects that is associated with the normal classification label. Huang teaches from the plurality of labeled training data objects that is associated with the normal classification label; ([Huang, page 4] “we are given n unlabeled samples xu1,···,xun ∈ X and m labeled samples (xl1, y1),···,(xlm, ym) ∈ X ×Y with Y = {-1,1} where y = 1 denotes normal samples and y = -1 denotes anomalous samples.”, wherein the examiner interprets labeled samples having y = 1 to be the same as labeled training data objects associated with the normal classification label because they are both directed to labeled samples identified as normal).Top of Form Lübbering, Zhang, Huang, and the instant application are analogous art because they are all directed to machine-learning anomaly detection using labeled normal and anomalous data and autoencoder-based loss information. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of claim 1 disclosed by Lübbering and Zhang to include the use of labeled normal training samples in generating the reconstruction-based loss associated with normal data as disclosed by Huang. One would have been motivated to make such a modification in order to expressly associate the reconstruction-loss calculation with training samples identified as normal, thereby enabling the anomaly detection model to learn the reconstruction characteristics of the normal class and distinguish normal data from anomalous data, as suggested by Huang ([Huang, page 4] “we are given n unlabeled samples xu1,···,xun ∈ X and m labeled samples (xl1, y1),···,(xlm, ym) ∈ X ×Y with Y = {-1,1} where y = 1 denotes normal samples and y = -1 denotes anomalous samples.”) Lübbering, Zhang, and Huang do not teach wherein the normal prediction loss parameter indicates a reconstruction loss measure associated with the anomaly detection machine learning model in reconstructing a plurality of normal training data objects. Liu teaches wherein the normal prediction loss parameter indicates a reconstruction loss measure associated with the anomaly detection machine learning model in reconstructing a plurality of normal training data objects; ([Liu, page 4, sec 3.3] “Reconstruction Loss. Given the training set D = {xi|i = 1, 2, ⋯ , M} containing M samples, we firstly consider the distance between the input image x and its reconstruction x̂. The reconstruction error on each sample is minimized as follows: Lrec(x, x̂) = ||x - x̂||²2”, wherein the examiner interprets the reconstruction error to be the same as the reconstruction loss measure because they are both directed to measuring error between an input and its reconstruction; [Liu, page 2] “In the anomaly detection task, AE is usually trained by minimizing the reconstruction error of defect-free samples, and then the reconstruction error is adopted as an indicator of anomalies.”; and [Liu, page 4] “Adopting autoencoder structure, anomaly detection models are expected to minimize the reconstruction error on normal images during training and induce large reconstruction error on anomalies at the test stage.”, wherein the examiner interprets defect-free/normal images to be the same as the plurality of normal training data objects because they are both directed to normal samples used to train an autoencoder-based anomaly detection model). Lübbering, Zhang, Huang, Liu, and the instant application are analogous art because they are all directed to machine-learning anomaly detection or classification using normal/anomalous data and autoencoder-based reconstruction information. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Lübbering and Zhang to include the labeled normal/anomalous training framework disclosed by Huang and Liu's normal-sample reconstruction-loss technique. One would have been motivated to do so to train an anomaly-detection autoencoder to characterize normal data through reconstruction error, as suggested by Liu ([Liu, page 2] “In the anomaly detection task, AE is usually trained by minimizing the reconstruction error of defect-free samples, and then the reconstruction error is adopted as an indicator of anomalies.”). Claim 11 and 18 are analogous to claim 4, aside from claim type and minute differences, and thus the same rejection as above. Regarding claim 5, Lübbering, Zhang, Huang, and Liu teach the computer-implemented method of claim 4 (see rejection of claim 4). Liu further teaches: generating, by the one or more processors, a plurality of encoded normal training data objects based on the plurality of normal training data objects; ([Liu, page 3, sec. 3-3.1] “Suppose the training set D = {xi|i = 1, 2, ⋯ , M} is given, where xi ∈ X is the ith normal sample (totally M samples) of the sample space X….Given an image sample x ∈ X, the encoder-1 network encodes it as a latent vector z ∈ Z”, wherein the examiner interprets encoding the normal samples as latent vectors to be the same as generating a plurality of encoded normal training data objects based on the plurality of normal training data objects because they are both directed to encoding normal training samples into latent representations). generating, by the one or more processors, a plurality of reconstructed normal training data objects based on the plurality of encoded normal training data objects; ([Liu, page 3, sec 3.1] “given an input, we first encode it as a latent vector using encoder-1, then the decoder is applied to get the reconstructed image” AND [Liu, page 4, sec 3.2] “The decoder network usually cooperates with the encoder network to reconstruct the input image from the latent vector z.”, wherein the examiner interprets reconstructing the input image from latent vector z to be the same as generating reconstructed normal training data objects based on encoded normal training data objects because they are both directed to decoding encoded representations of normal training inputs to produce reconstructions). generating, by the one or more processors, the normal prediction loss parameter based on the plurality of normal training data objects and the plurality of reconstructed normal training data objects; ([Liu, page 4, sec 3.3] “we firstly consider the distance between the input image x and its reconstruction x̂. The reconstruction error on each sample is minimized as follows: Lrec(x, x̂) = ||x - x̂||²2”, wherein the examiner interprets the reconstruction error calculated between input x and reconstruction x̂ to be the same as generating the normal prediction loss parameter based on the normal training data objects and reconstructed normal training data objects because they are both directed to calculating reconstruction loss from an original normal input and its reconstruction). Lübbering, Zhang, Huang, Liu, and the instant application are analogous art, because they are all directed to machine-learning anomaly detection using autoencoder architectures, reconstruction of training data, and reconstruction-loss information for distinguishing normal data from anomalous data. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of claim 4 disclosed by Lübbering, Zhang, Huang, and Liu to include the encoding of normal training samples into latent representations, reconstruction of the normal training samples from those latent representations, and calculation of reconstruction loss from the original and reconstructed normal samples disclosed by Liu. One would have been motivated to make such a modification in order to implement the normal-data reconstruction-loss determination through the conventional encoder-decoder operation of an autoencoder, thereby providing a reconstruction error that characterizes how accurately the anomaly detection model reconstructs normal training data, as suggested by Liu ([Liu, page 3, sec. 3-3.1] “Suppose the training set D = {xi|i = 1, 2, ⋯ , M} is given, where xi ∈ X is the ith normal sample (totally M samples) of the sample space X….Given an image sample x ∈ X, the encoder-1 network encodes it as a latent vector z ∈ Z AND ([Liu, page 3, sec 3.1] “given an input, we first encode it as a latent vector using encoder-1, then the decoder is applied to get the reconstructed image” AND [Liu, page 4, sec 3.2] “The decoder network usually cooperates with the encoder network to reconstruct the input image from the latent vector z.” AND ([Liu, page 4, sec 3.3] “we firstly consider the distance between the input image x and its reconstruction x̂. The reconstruction error on each sample is minimized as follows: Lrec(x, x̂) = ||x - x̂||²2”) Claim 12 and 19 are analogous to claim 5, aside from claim type and minute differences, and thus the same rejection as above.Bottom of Form Claim(s) 6-7, 13-14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lübbering in view of Zhang further in view of Huang. Regarding claim 6, Lübbering and Zhang teach the computer-implemented method of claim 1 (see rejection of claim 1). Lübbering and Zhang do not teach wherein the anomaly prediction loss parameter indicates a reconstruction loss measure associated with the anomaly detection machine learning model in reconstructing a plurality of anomaly training data objects from the plurality of labeled training data objects that is associated with the anomaly classification label. Huang teaches: wherein the anomaly prediction loss parameter indicates a reconstruction loss measure associated with the anomaly detection machine learning model in reconstructing a plurality of anomaly training data objects from the plurality of labeled training data objects that is associated with the anomaly classification label; ([Huang, page 5, sec. 3] “With unlabeled samples xu1,···,xun and labeled samples xl1,···,xlm, we want the autoencoder to well reconstruct the normal data but erroneously reconstruct the labeled anomalous data, thus the reconstruction likelihood is maximized for the normal data and minimized for the labeled anomalous data.” and “The reconstruction loss is defined as follows: Lrec-semi =” followed by the reconstruction-loss expression for the unlabeled and labeled samples; AND [Huang, page 4, sec. 3] “where y = 1 denotes normal samples and y = -1 denotes anomalous samples.”, wherein the examiner interprets the reconstruction loss applied to the labeled anomalous data to be the same as an anomaly prediction loss parameter indicating a reconstruction loss measure because they are both directed to measuring autoencoder reconstruction performance for anomalous training samples, and the examiner interprets labeled samples having y = -1 to be the same as anomaly training data objects associated with the anomaly classification label because they are both directed to labeled anomalous samples). Lübbering, Zhang, Huang, and the instant application are analogous art because they are all directed to machine-learning anomaly detection using normal and anomalous data and autoencoder reconstruction information. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Lübbering and Zhang to include Huang's labeled-anomaly reconstruction-loss technique. One would have been motivated to do so to improve discrimination between normal and anomalous samples by training the autoencoder differently with respect to labeled anomalous data, as suggested by Huang ([Huang, page 5, sec. 3] “we want the autoencoder to well reconstruct the normal data but erroneously reconstruct the labeled anomalous data”). Claim 13 and 20 are analogous to claim 6, aside from claim type and minute differences, and thus the same rejection as above. Regarding claim 7, Lübbering, Zhang, and Huang teach the computer-implemented method of claim 6 (see rejection of claim 6). Huang further teaches: further comprising: generating, by the one or more processors, a plurality of encoded anomaly training data objects based on the plurality of anomaly training data objects that is associated with the anomaly classification label; ([Huang, page 4, sec 3] “Given the input space X consisting of normal data XN and anomalous data XA, where X = XN ∪ XA. For semi-supervised anomaly detection (AD), we are given n unlabeled samples x u 1 ,··· ,x u n ∈ X and m labeled samples (x l 1 , y1),··· ,(x l m, ym) ∈ X ×Y with Y = {-1,1} where y = 1 denotes normal samples and y = -1 denotes anomalous samples.” AND [Huang, page 5, sec. 3] “Enc1(·) emphasizes mutual information optimization and the second encoder Enc2(·) focuses on entropy optimization, and in the meanwhile, the two encoders are enforced to share similar encoding via a consistent constraint on their latent representations. The encoder-decoder-encoder architecture can be expressed as: z = Enc1(x), xˆ = Dec(z), zˆ = Enc2(xˆ), (Equation 4) where xˆ is the output of the decoder, and z and zˆ are the latent representations from the first and second encoders, respectively.”, wherein the examiner interprets “m labeled samples … where y = -1 denotes anomalous samples” and “z = Enc1(x)” to be the same as “generating … a plurality of encoded anomaly training data objects … associated with the anomaly classification label” because they are both directed to processing multiple labeled anomalous (y = -1) training samples and encoding each such input x into an encoded representation z using an encoder (Enc1). generating, by the one or more processors, a plurality of reconstructed anomaly training data objects based on the plurality of encoded anomaly training data objects; [Huang, page 5, sec. 3] “Enc1(·) emphasizes mutual information optimization and the second encoder Enc2(·) focuses on entropy optimization, and in the meanwhile, the two encoders are enforced to share similar encoding via a consistent constraint on their latent representations. The encoder-decoder-encoder architecture can be expressed as: z = Enc1(x), xˆ = Dec(z), zˆ = Enc2(xˆ), (Equation 4) where xˆ is the output of the decoder, and z and zˆ are the latent representations from the first and second encoders, respectively.” wherein the examiner interprets “xˆ = Dec(z)” to be the same as “generating … a plurality of reconstructed anomaly training data objects based on the plurality of encoded anomaly training data objects” because they are both directed to producing reconstructed outputs (xˆ) by decoding encoded/latent representations (z) using a decoder (Dec).) and generating, by the one or more processors, the anomaly prediction loss parameter based on the plurality of anomaly training data objects and the plurality of reconstructed anomaly training data objects. ([Huang, page 5, sec. 3] “Lrec-semi = 1/n ∑ ||x̂ u i - x u i||² + 1/m ∑ ||x̂ l j - Φ(x l j)||² where, Φ(x l j) = φ(x l j), if yj = -1 (Equation 5)”, AND [Huang, page 5, sec.3] “we want the autoencoder to well reconstruct the normal data but erroneously reconstruct the labeled anomalous data, thus the reconstruction likelihood is maximized for the normal data and minimized for the labeled anomalous data. A straight-forward loss definition for the anomalous data is the negative squared norm loss.”, wherein the examiner interprets “Lrec-semi = 1/n ∑ ||x̂ u i - x u i||² + 1/m ∑ ||x̂ l j - Φ(x l j)||² where, Φ(x l j) = φ(x l j), if yj = -1 (Equation 5)” and “reconstruction likelihood … minimized for the labeled anomalous data” to be the same as “generating … the anomaly prediction loss parameter based on the plurality of anomaly training data objects and the plurality of reconstructed anomaly training data objects” because they are both describing the generation of a reconstruction-based loss specifically computed over the anomalous training data and their reconstructions, namely, both calculate a computing loss (Lrec-semi term for labeled samples) using the reconstructed outputs (xˆ^l_j) and the labeled anomalous training inputs (x^l_j with yj = -1, via Φ(·) selecting φ(x^l_j) for anomalies). Lübbering, Zhang, Huang, and the instant application are analogous art because they are all directed to semi-supervised anomaly detection systems that process labeled anomalous training samples using machine learning architectures. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 6 disclosed by Lübbering, Zhang, and Huang to include the end to end encoding technique disclosed by Huang. One would be motivated to do so to effectively enhance anomaly discrimination by leveraging structured latent representations and reconstruction-based supervision, as suggested by Huang ([Huang, page 5, sec. 3] “the encoder-decoder-encoder architecture can be expressed as: z = Enc1(x), xˆ = Dec(z) …”). Claim 14 is analogous to claim 7, aside from claim type and mild differences, and thus will face the same rejection as set forth above. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVAN KAPOOR/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Mar 23, 2023
Application Filed
Feb 26, 2026
Non-Final Rejection mailed — §103
May 06, 2026
Applicant Interview (Telephonic)
May 06, 2026
Examiner Interview Summary
May 21, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
7%
Grant Probability
18%
With Interview (+11.1%)
4y 4m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month