Prosecution Insights
Last updated: October 02, 2026
Application No. 17/825,788

Unsupervised Anomaly Detection With Self-Trained Classification

Non-Final OA §103
Filed
May 26, 2022
Priority
May 27, 2021 — provisional 63/193,875
Examiner
KAPOOR, DEVAN
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
3 (Non-Final)
7%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
18%
With Interview

Examiner Intelligence

Grants only 7% of cases
7%
Career Allowance Rate
1 granted / 14 resolved
-47.9% vs TC avg
Moderate +11% lift
Without
With
+11.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
29 currently pending
Career history
47
Total Applications
across all art units

Statute-Specific Performance

§101
34.0%
-6.0% vs TC avg
§103
57.4%
+17.4% vs TC avg
§102
5.8%
-34.2% vs TC avg
§112
2.2%
-37.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the application filed on 04/27/2026. Claims 1, 4-11, and 13-22 are pending and have been examined. This action is Non-final. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/27/2026 has been entered. Response to Arguments Argument 1: The applicant argues that the 101 rejection of claims 1, 4-11, and 13-22 is moot in view of the amended claims, because the claims have been amended to specify how self-training is performed. In particular, the applicant points to the newly added limitation of generating, using the STOC, a refined set of training data including the training examples categorized as non-anomalous training examples and excluding the training examples categorized as anomalous, wherein the refined set of training data is generated by dividing the set of training data into subsets of training data, providing each subset to one of the ensemble of OCCs, and each OCC training a respective model component. Relying on paragraph [0061] of the specification, the applicant contends that the use of multiple one-class classifiers accounts for variation or potential inaccuracy over using a single model, and that the resulting ensemble is more robust than a single model in generating pseudo-labels for categorizing training examples because the risk of a false positive or false negative is reduced through the consensus of multiple models. The applicant asserts that the above wherein clause captures this technological improvement, and therefore requests reconsideration and withdrawal of the 101 rejection. Response to Argument 1: The examiner has considered the argument set forth above and finds it persuasive. The rejection of claims 1, 4-11, and 13-22 under 35 U.S.C. 101 is withdrawn. Argument 2: The applicant argues that the 103 rejections over Pang in view of Xiao in view of Glassman (claims 1, 3-6, 11, 13-15, and 20-22), further in view of Li (claims 7 and 16), and further in view of Rudolph (claims 8-10 and 17-19) are moot in view of the amended claims. The applicant contends that the cited prior art does not disclose generating, using the STOC, a refined set of training data including the training examples categorized as non-anomalous training examples and excluding the training examples categorized as anomalous, wherein the refined set of training data is generated by dividing the set of training data into subsets of training data, providing each of the subsets to one of the ensemble of OCCs, and each OCC training a respective model component. Based on this limitation, the applicant asserts that independent claims 1, 11, and 20 are patentable, and that the dependent claims are patentable at least by virtue of their dependency on the independent claims as well as on their own merits. The applicant therefore requests that the 103 rejections be withdrawn and that the application be passed to issue. Response to Argument 2: The applicant’s arguments have been considered but are not persuasive. Although the claims have been amended to recite the STOC architecture, the refined training set, the subdivision of training data among respective OCCs, and the requirement that an example be treated as anomalous when at least one OCC indicates anomalous and non-anomalous when every OCC indicates non-anomalous, the presently applied combination addresses those limitations. Pang teaches the underlying self-trained anomaly-detection framework, including use of unlabeled data, categorization into anomalous and non-anomalous candidate sets, iterative updating of those sets, and retraining. Krawczyk teaches an ensemble of one-class classifiers in which the feature space is partitioned into subsets and each subset is used to train a respective one-class classifier, thereby supplying the claimed subset-to-OCC training arrangement. Tax teaches that the product combination rule operates as an AND combination, such that an object is accepted only when all one-class classifiers agree and the output changes when one classifier disagrees, which corresponds to the claimed every-OCC non-anomalous/at-least-one-OCC anomalous decision rule. Xiao further teaches identifying and removing outliers from the training set, retaining the remaining target samples, and training the final one-class classifier using that refined data, thereby addressing the claimed exclusion of anomalous examples and subsequent classifier training using the refined set. Li supplies the processor, network-interface, and receipt of unlabeled network data limitations. Accordingly, the amended limitations do not distinguish the claims from the combined teachings of the prior art, and the 103 rejection of claims 1, 11, and 20 is maintained, and the dependent claims likewise remain rejected because their additional limitations are separately addressed below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 4-7, 11, 13-16, and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over the NPL reference “Self-trained Deep Ordinal Regression for End-to-End Video Anomaly Detection” by Pang et al. (referred herein as Pang) in view of the NPL reference “Selecting locally specialised classifiers for one-class classification ensembles” by Krawczyk et al. (referred herein as Krawczyk) in view of the NPL reference “One-class classification; Concept-learning in the absence of counter-examples” by Tax (referred herein as Tax) in view of the NPL reference “Ramp Loss based robust one-class SVM” by Xiao et al. (referred herein as Xiao) further in view of US 20200195667A1 by Li et al. (referred herein as Li). Regarding claim 1, Pang teaches: A system for anomaly detection, comprising: ([Pang, Abstract] “Video anomaly detection is of critical practical importance to a variety of real applications” AND [Pang, page 12175, sec. 4] “Our formulation is implemented as a self-training deep neural network for ordinal regression.”, wherein the examiner interprets Pang’s implemented self-training anomaly-detection framework to be the same as a system for anomaly detection because Pang implements a trained neural-network framework directed to detecting anomalies, hence it is also a system for anomaly detection.) self-trained ([Pang, page 12177, sec. 4.3] “We further perform iterative learning using self-training [47] to iteratively improve our anomaly detector…we use the newly obtained pseudo labels, A and N, to replace the previous ones and then retrain the end-to-end anomaly learner φ.”, wherein the examiner interprets Pang’s repeated updating of pseudo labels and retraining of the anomaly detector to teach the self-trained aspect of the claimed classifier because the detector trains itself iteratively using its newly generated pseudo labels.) unlabeled training data comprising a plurality of training examples, wherein the unlabeled training data comprises one or more anomalous training examples and one or more non-anomalous training examples and there are fewer anomalous training examples than non-anomalous training examples; ([Pang, page 12175, sec. 3] “given a set of K video frames X = {x1, x2, · · · , xK} with no class label information” AND [Pang, page 12175, sec. 3] “A ⊂ X be a set of anomalous frame candidates… N ⊂ X (N∩A = ∅)” AND [Pang, page 12173, sec. 1] “anomalous events are rare” AND [Pang, page 12177, sec. 5.1] “the overwhelming presence of normal frames in the real-word datasets.”, wherein the examiner interprets Pang’s unlabeled video frames containing anomalous and normal frame candidates, with anomalies described as rare and normal frames overwhelmingly present, to be the same as unlabeled training data containing anomalous and non-anomalous examples with fewer anomalous examples than non-anomalous examples because both datasets contain a minority of anomalous examples and a majority of normal examples.) each of the training examples as an anomalous training example or as a non-anomalous training example ([Pang, page 12175, sec. 3] “we first initialize A and N using anomaly scores generated by some existing unsupervised anomaly detection methods” AND [Pang, page 12176, sec. 4.1] “We then use these anomaly scores to include the most likely anomalous frames into the pseudo anomaly set A and the most likely normal frames into the pseudo normal set N”, wherein the examiner interprets Pang’s assignment of training examples to pseudo-anomalous set A or pseudo-normal set N to be the same as categorizing each training example as anomalous or non-anomalous because A and N are expressly the anomalous and normal candidate sets used in Pang’s training procedure.) Pang does not teach one class classifier, STOC, having an ensemble of one class classifiers, OCCs… wherein the refined set of training data is generated by dividing the set of training data into subsets of training data and providing each of the subsets of training data to one of the ensemble of OCCs and in which each one of the ensemble of OCCs trains a respective model component...each OCC being configured to receive an input and output an indication whether the input is anomalous or non-anomalous…when said training example is predicted, respectively, as anomalous by at least one OCC or as non-anomalous by every OCC…one or more processors, wherein the one or more processors are configured to: receive, from one or more devices over a network…generate, using the STOC, a refined set of training data including the training examples categorized as non-anomalous training examples and excluding the training examples categorized as anomalous…retrain each OCC, using the STOC, according to the refined set of training data…an output classifier configured to receive input data and predict a final classification of the input data as anomalous or non-anomalous…train the output classifier, using the STOC, according to the refined set of training data…receive input data and predict by the output classifier a final classification whether the input data is anomalous or non-anomalous. Krawczyk teaches: an ensemble of one class classifiers, OCCs ([Krawczyk, page 430, sec. 3] “OCClustE [24, 25] was proposed as an efficient ensemble system for one-class learning.” AND [Krawczyk, page 430, sec. 3] “This leads to the formation of a pool of K classifiers assigned to the target class” AND [Krawczyk, page 430, sec. 3] “This allows us to easily create a pool of several one-class learners, dedicated to the target class.”, wherein the examiner interprets Krawczyk’s pool of K one-class learners forming an ensemble system to be the same as an ensemble of OCCs because each pool member is expressly a one-class classifier. When Krawczyk’s OCC ensemble is incorporated into Pang’s self-training framework, the resulting classifier is the claimed self-trained one class classifier, STOC.) wherein the refined set of training data is generated by dividing the set of training data into subsets of training data and providing each of the subsets of training data to one of the ensemble of OCCs and in which each one of the ensemble of OCCs trains a respective model component ([Krawczyk, page 428] “ while homogeneous ones use the same classifier model but each fed with a diverse input (e.g., different subsets of objects or features)” AND [Krawczyk, page 430, sec. 3] “This method uses a clustering algorithm to partition the feature space into atomic subsets. In the next step each of these clusters is used to train a one-class classifier.” AND [Krawczyk, page 430, sec. 3] “This leads to the formation of a pool of K classifiers assigned to the target class”, wherein the examiner interprets Krawczyk’s partitioning of training objects into atomic subsets and using each respective subset to train a respective one-class classifier that becomes a member of the K-classifier pool to be the same as dividing training data into subsets, providing the subsets to respective OCCs, and training respective model components because each trained OCC is a component of the disclosed ensemble.) Pang and Krawczyk do not teach each OCC being configured to receive an input and output an indication whether the input is anomalous or non-anomalous…when said training example is predicted, respectively, as anomalous by at least one OCC or as non-anomalous by every OCC… when said training example is predicted, respectively, as anomalous by at least one OCC or as non-anomalous by every OCC…generate, using the STOC, a refined set of training data including the training examples categorized as non-anomalous training examples and excluding the training examples categorized as anomalous…retrain each OCC, using the STOC, according to the refined set of training data…an output classifier configured to receive input data and predict a final classification of the input data as anomalous or non-anomalous…train the output classifier, using the STOC, according to the refined set of training data…receive input data and predict by the output classifier a final classification whether the input data is anomalous or non-anomalous. Tax teaches: each OCC being configured to receive an input and output an indication whether the input is anomalous or non-anomalous ([Tax, page 123, sec. 5.2] “the one-class classifiers do not model the complete p(x|ωT), but they only provide a yes-no output: object x is either accepted or rejected.” AND [Tax, page 109, sec 4.12] “a one-class classification method trained on the normal working situation has to detect anomalous signals and raise an alarm.”, wherein the examiner interprets Tax’s accepted/rejected yes-no output of each one-class classifier to be the same as an anomalous/non-anomalous indication because acceptance denotes membership in the target class while rejection denotes an outlier relative to the target class.) when said training example is predicted, respectively, as anomalous by at least one OCC or as non-anomalous by every OCC ([Tax, page 131, sec. 5.3] “The mean and product combination rules can be interpreted as being an OR combination and an AND combination, respectively.” AND [Tax, page 131, sec. 5.3] “When it is required that the new object is accepted by all one-class classifiers, a product combination rule gives high output when all classifiers agree that the object is acceptable, but immediately gives low output when one classifier disagrees…If the user wants to reject objects like that, then he should apply the product combination rule.” AND [Tax, page 120, sec. 5.1.1] “Combining the predictions of the classifiers in the two extremes, perfect correlation and complete independence of the data, will indicate where one combination rule can be preferred over the other.”, wherein the examiner interprets Tax’s product combination rule to be the same as predicting an example as non-anomalous only when every OCC indicates non-anomalous, and as anomalous when at least one OCC indicates anomalous, because Tax expressly characterizes the product rule as an AND combination, requires acceptance by all one-class classifiers for the acceptable result, and states that the output immediately becomes low when one classifier disagrees.) Pang, Krawczyk, and Tax do not teach generate, using the STOC, a refined set of training data including the training examples categorized as non-anomalous training examples and excluding the training examples categorized as anomalous…retrain each OCC, using the STOC, according to the refined set of training data…an output classifier configured to receive input data and predict a final classification of the input data as anomalous or non-anomalous…train the output classifier, using the STOC, according to the refined set of training data…receive input data and predict by the output classifier a final classification whether the input data is anomalous or non-anomalous. Xiao teaches: generate, using the STOC, a refined set of training data including the training examples categorized as non-anomalous training examples and excluding the training examples categorized as anomalous ([Xiao, Abstract] “Then the outliers are identified and removed from the training set. The final classification surface is obtained on the remaining training samples.” AND [Xiao, page 18, sec. 4.2] “The samples with the smallest values are identified as outliers. They are removed from the training data set, and the remaining samples are all target samples.”, wherein the examiner interprets Xiao’s identification and removal of outliers and retention of target samples to be the same as excluding examples categorized as anomalous while retaining examples categorized as non-anomalous in a refined training set because Xiao expressly treats the removed samples as outliers and the remaining samples as target samples.) an output classifier configured to receive input data and predict a final classification of the input data as anomalous or non-anomalous ([Xiao, page 18, sec. 4.1] “A sample x is classified as target if fR(x) ≥ 0; otherwise, it is classified as non-target.” AND [Xiao, page 16, sec. 2] “identify outliers by this surface and to exclude them from the training set, before training the final OCSVM model.”, wherein the examiner interprets Xiao’s final OCSVM model and its target/non-target decision to be the same as an output classifier configured to provide a final non-anomalous/anomalous classification because Xiao uses the final one-class model to classify an input sample as target or non-target.) train the output classifier, using the STOC, according to the refined set of training data ([Xiao, page 16, sec. 2] “this surface is used to identify outliers and remove them from the training set, so as to train the final OCSVM.” AND [Xiao, page 18, Algorithm 2] “Take X2 = {x|x ∈ X1, x ∉ Xout} as the training set…8: Train conventional OCSVM, and obtain the hyper-plane f 2 ( x ) = 0. Output: The classification hyper-plane f 2 ( x ) = 0.”, wherein the examiner interprets Xiao’s training of the final OCSVM on the training set remaining after identified outliers are removed to be the same as training the output classifier according to the refined set of training data because the final OCSVM is trained only on the cleaned remainder. In the combined STOC architecture, that final OCSVM functions as the STOC output classifier.) receive input data and predict by the output classifier a final classification whether the input data is anomalous or non-anomalous ([Xiao, page 18, sec. 4.1] “A sample x is classified as target if fR(x) ≥ 0; otherwise, it is classified as non-target.”, wherein the examiner interprets classification of an input sample as target versus non-target by the final OCSVM to be the same as a final non-anomalous versus anomalous classification because target samples are the normal class and non-target samples are treated as outliers.) retrain each OCC, using the STOC, according to the refined set of training data; ([Xiao, page 16, sec. 2] “Next, this surface is used to identify outliers and remove them from the training set, so as to train the final OCSVM.” AND [Xiao, page 18, Algorithm 2] “Take X2 = {x|x ∈ X1, x ∉ Xout} as the training set.” AND [Xiao, page 18, Algorithm 2] “Train conventional OCSVM, and obtain the hyper-plane f2(x) = 0.”, wherein the examiner interprets Xiao’s training of an OCSVM after outliers have been identified and removed, using the remaining cleaned training set, to be the same as retraining an OCC according to the refined set of training data because the one-class classifier is trained again after refinement of the training data; when this training procedure is applied to the OCCs of the STOC ensemble, each OCC is retrained according to the refined set of training data.) Pang, Krawczyk, Tax, and Xiao do not teach one or more processors, wherein the one or more processors are configured to: receive, from one or more devices over a network,. Li teaches one or more processors, wherein the one or more processors are configured to: receive, from one or more devices over a network, ([Li, [0096]] “In addition to a processor, a memory, a network interface, and a non-volatile memory that are shown in FIG. 3” AND [Li, [0110]] “The electronic device includes a processor and a memory configured to store a machine executable instruction.” AND [Li, [0110]] “the device may further include an external interface, so that the device can communicate with other devices or components.” AND [Li, [0034]] “a modeler can collect a large quantity of unlabeled URL access requests as unlabeled samples in advance”, wherein the examiner interprets Li’s processor-based electronic device having a network interface/external interface for communicating with other devices or components and collecting unlabeled URL access requests to be the same as one or more processors configured to receive unlabeled training data from one or more devices over a network because Li expressly provides network communications hardware and processor-executed anomaly-detection operations on unlabeled network requests.) Pang, Krawczyk, Tax, Xiao, Li, and the instant application are analogous art because they are all directed to machine-learning anomaly detection or one-class classification, including training anomaly-detection models, combining one-class classifier outputs, identifying outliers, refining training data, and implementing anomaly detection on computing systems. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the self-trained anomaly-detection framework disclosed by Pang to include the “pool of several one-class learners” disclosed by Krawczyk. One would be motivated to do so to effectively increase model diversity and complementarity while reducing the complexity handled by each classifier, as suggested by Krawczyk ([Krawczyk, page 430, sec. 3] “It assures the initial diversity (as a result of using different inputs in their training) and complementarity (as classifiers together cover all the decision space), which leads to better performance of the ensemble.”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the product-based binary combining rule disclosed by Tax. One would be motivated to do so to effectively obtain a consensus-based one-class decision and improve combined classification performance over individual classifiers, as suggested by Tax ([Tax, pages 126-127, sec. 5.2.2] “Only the product combination rule on the estimated probabilities improves over all the individual classifiers and produces an ROC curve which exceeds the other classifiers”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the outlier-removal and final-OCSVM training procedure disclosed by Xiao. One would be motivated to do so to effectively reduce the influence of anomalous or outlier contamination on one-class training and improve the resulting classification performance, as suggested by Xiao ([Xiao, page 20, sec. 6] “the proposed method reduces the outliers’ influence effectively, and achieves better classification performance.”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the resulting anomaly-detection system using the processor, network interface, and external communication interface disclosed by Li. One would be motivated to do so to efficiently deploy the anomaly detector in a networked electronic device capable of processing unlabeled network requests and identifying abnormal requests, as suggested by Li ([Li, [0094]] “a potential URL attack can be found in advance, thereby helping perform security protection in time for a potential abnormal URL access.”). Claims 11 and 20 are analogous to claim 1, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 4, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 1 (see rejection of claim 1). Pang further teaches: wherein the one or more processors are configured to perform additional iterations of: categorizing each of the training examples using the plurality of first machine learning models; ([Pang, page 12175] “…we first initialize A and N using anomaly scores generated by some existing unsupervised anomaly detection methods…and we iteratively update A and N and retrain φ until the best φ is achieved.”, wherein the examiner interprets “initialize A and N using anomaly scores” to be the same as “categorizing each of the training examples” because both describe dividing examples into anomalous (A) and non-anomalous (N) categories using model outputs; and interprets “iteratively update A and N and retrain φ” to be the same as “perform additional iterations” because both describe repeating the categorization process multiple times.) and updating, based on the additional iterations, the refined set of training data. ([Pang, page 12175] “A corresponding new set of anomaly scores are then generated, which are used to update the membership of A and N.”, wherein the examiner interprets “generate new anomaly scores…update the membership of A and N” to be the same as “updating the refined set of training data based on additional iterations” because both describe refining the anomalous and non-anomalous subsets after each iteration.) Claim 14 is analogous to claim 4, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 5, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 4 (see rejection of claim 4). Pang further teaches: wherein the one or more processors are further configured to train a machine learning model using the refined set of training data, wherein the machine learning model is trained to receive training examples and to generate one or more respective feature values for each of the received training examples; ([Pang, page 12176, sec. 4.2] “The end-to-end anomaly score learner takes A and N as inputs and learns to optimize the anomaly scores…The score learner can be defined as a function φ(·; Θ) : X ↦ R, which is a sequential stack of a feature representation learner ψ(·; Θr) : X ↦ Q and an anomaly score learner η(·; Θs) : Q ↦ R, where Q ∈ RM is an intermediate feature representation space and Θ = {Θr, Θs} contains all the parameters to be learned.” AND [Pang, page 12176, Eq. (5)] “q = ψ(x; Θr)”, wherein the examiner interprets Pang’s feature representation learner ψ, whose parameters are learned as part of the end-to-end learner trained using A and N, to be the same as the claimed machine learning model trained using the refined training data, and interprets q = ψ(x; Θr) to be the same as generating one or more respective feature values for each received training example because q is expressly the intermediate feature representation generated from input x.) Krawczyk further teaches wherein to categorize the unlabeled training data using the OCCs, the one or more processors are configured to process, using the OCCs, respective one or more feature values for each training example of the unlabeled training data ([Krawczyk, page 430, sec. 3] “This method uses a clustering algorithm to partition the feature space into atomic subsets. In the next step each of these clusters is used to train a one-class classifier…Finally, we need to combine the individual outputs of base classifiers at our disposal…For a given kth WOCSVM classifier, we propose to use the following heuristic function: [Eq. 2]”, wherein the examiner interprets Krawczyk’s pool of one-class classifiers operating on objects represented in feature space and producing individual classifier outputs to be the same as the OCCs processing respective feature values for each training example, because when Krawczyk’s OCC processing is applied to the feature representations q generated by Pang’s feature representation learner mapped above, the claimed OCC-based processing of feature values is obtained.) Pang, Krawczyk, Tax, Xiao, Li, and the instant application are analogous art because they are all directed to machine-learning anomaly detection or one-class classification using trained models to process data representations. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 4 disclosed by Pang, Krawczyk, Tax, Xiao, and Li to include the feature-representation anomaly learner disclosed by Pang. One would be motivated to do so to use Pang’s anomaly score learner to generate feature values and train them for the received values as to-be-learned parameters as suggested by Pang ([Pang, page 12176, sec. 4.2] “The end-to-end anomaly score learner takes A and N as inputs and learns to optimize the anomaly scores…The score learner can be defined as a function φ(·; Θ) : X ↦ R, which is a sequential stack of a feature representation learner ψ(·; Θr) : X ↦ Q and an anomaly score learner η(·; Θs) : Q ↦ R, where Q ∈ RM is an intermediate feature representation space and Θ = {Θr, Θs} contains all the parameters to be learned.” AND [Pang, page 12176, Eq. (5)] “q = ψ(x; Θr)”). It would also have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the pool of one-class learners operating on feature-space inputs disclosed by Krawczyk. One would be motivated to do so to effectively increase classifier diversity and complementarity while processing learned feature representations, as suggested by Krawczyk ([Krawczyk, page 430, sec. 3] “It assures the initial diversity (as a result of using different inputs in their training) and complementarity (as classifiers together cover all the decision space), which leads to better performance of the ensemble.”). Regarding claim 6, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 5 (see rejection of claim 5). Pang further teaches: wherein the one or more processors are configured to perform additional iterations of training the machine learning model using the refined set of training data. ([Pang, page 12175] “we first initialize A and N using anomaly scores generated by some existing unsupervised anomaly detection methods (see Sec. 4.1), and we iteratively update A and N and retrain φ until the best φ is achieved (see Section 4.3).” AND [Pang, page 12180] “Figure 6 shows the AUC results of our method at each iteration during self-training. Our performance gets larger improvement with increasing iterations in the first few iterations on most datasets and then becomes stable at the 4th or 5th iteration…We found empirically that five iterations are often sufficient to reach the possibly best performance on different datasets.”, wherein the examiner interprets “iteratively update A and N and retrain φ” to be the same as “perform additional iterations of training the machine learning model using the refined set of training data” because both are directed to repeatedly retraining a model on progressively refined anomalous/non-anomalous subsets.) Regarding claim 7, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 1 (see rejection of claim 1). Pang further teaches: determine that at least one first score does not meet one or more thresholds; and in response to the determination that the at least one first score does not meet one or more thresholds, exclude the first training example from the unlabeled training data. ([Pang, page 12177] “Particularly, we include the 10% most anomalous frames into [A] according to their anomaly scores, because anomaly scores often follow a Gaussian distribution and this decision threshold can provide an approximate 90% confidence level of making false positive errors in such cases …To generate the pseudo normal frame set N, we select the 20% most normal frames based on the anomaly scores…These two cutoff thresholds are used by default as they consistently obtain substantially improved performance on datasets with diverse anomaly rates.”, wherein the examiner interprets Pang’s use of cutoff thresholds to retain only examples falling within the selected most-anomalous or most-normal score ranges to be the same as determining whether a score meets a threshold and excluding examples that fail the selected threshold criteria from the corresponding retained training subset.) Krawczyk further teaches train a plurality of first machine learning models using a respective subset of the unlabeled training data; ([Krawczyk, page 428] “homogeneous ones use the same classifier model but each fed with a diverse input (e.g., different subsets of objects or features)” AND [Krawczyk, page 430, sec. 3] “This method uses a clustering algorithm to partition the feature space into atomic subsets. In the next step each of these clusters is used to train a one-class classifier.” AND [Krawczyk, page 430, sec. 3] “This leads to the formation of a pool of K classifiers assigned to the target class”, wherein the examiner interprets Krawczyk’s pool of K classifiers, each trained using a different subset of training objects or features, to be the same as a plurality of first machine learning models trained using respective subsets of the unlabeled training data because each model receives a respective portion of the training data.) Tax further teaches process a first training example of the unlabeled training data through each of the plurality of first machine learning models to generate a plurality of first scores corresponding to respective probabilities that the first training example is non-anomalous or anomalous; ([Tax, page 125, sec. 5.2.1] “For R one-class classifiers, with estimated probabilities P(xk|ωT) and threshold θk, this results in the following set of combining rules:” AND [Tax, page 125, Eq. (5.20)], wherein the examiner interprets Tax’s R one-class classifiers each producing an estimated target-class probability Pk(x|ωT) for the same object x, as reflected in Eq. (5.20), to be the same as processing a first training example through each of a plurality of machine learning models to generate a plurality of first scores corresponding to respective probabilities that the example is non-anomalous or anomalous.) Pang, Krawczyk, Tax, Xiao, Li, and the instant application are analogous art because they are all directed to anomaly detection or one-class classification using multiple trained models, model scores, thresholds, and filtered training data. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 1 disclosed by Pang, Krawczyk, Tax, Xiao, and Li to include the threshold-based training-sample selection disclosed by Pang. One would be motivated to do so to effectively cutoff thresholds to retain only examples falling within the selected most-anomalous or most-normal score ranges, as suggested by Pang ([Pang, page 12177] “Particularly, we include the 10% most anomalous frames into [A] according to their anomaly scores, because anomaly scores often follow a Gaussian distribution and this decision threshold can provide an approximate 90% confidence level of making false positive errors in such cases…To generate the pseudo normal frame set N, we select the 20% most normal frames based on the anomaly scores…These two cutoff thresholds are used by default as they consistently obtain substantially improved performance on datasets with diverse anomaly rates.”) It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the plurality of classifiers trained on respective subsets disclosed by Krawczyk. One would be motivated to do so to effectively increase model diversity and complementarity, as suggested by Krawczyk ([Krawczyk, page 430, sec. 3] “It assures the initial diversity (as a result of using different inputs in their training) and complementarity…which leads to better performance of the ensemble.”). It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the respective probability outputs of the one-class classifiers disclosed by Tax. One would be motivated to do so to efficiently place the outputs of the different one-class classifiers on a probabilistic basis suitable for combination and thresholding, as suggested by Tax ([Tax, page 124, sec. 5.2.1] “When one-class classifiers are to be combined based on posterior probabilities, an estimate for p(ωT|x) has to be used.”). Claim 16 is analogous to claim 7, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 13, Pang, Krawczyk, Tax, Xiao, and Li teaches The method of claim 11 (see rejection of claim 11). Pang further teaches wherein the method further comprises training a plurality of first machine learning models using the refined set of training data. ([Pang, page 12175] “…we first initialize A and N using anomaly scores generated by some existing unsupervised anomaly detection methods (see Sec. 4.1), and we iteratively update A and N and retrain φ until the best φ is achieved.”, wherein the examiner interprets “iteratively update A and N” to be the same as “refined set of training data” because both are directed to progressively filtering anomalous (A) and non-anomalous (N) subsets, and the examiner interprets “retrain φ” to be the same as “train the plurality of first machine learning models” because both describe training machine learning models using the refined subsets of data to improve anomaly detection accuracy.) Regarding claim 15, Pang, Krawczyk, Tax, Xiao, and Li teaches The method of claim 14 (see rejection of claim 14). Pang further teaches: wherein the method further comprises training a third machine learning model using the refined set of training data, wherein the third machine learning model is trained to receive training examples and to generate one or more respective feature values for each of the received training examples; ([Pang, page 12176, sec. 4.2] “The end-to-end anomaly score learner takes A and N as inputs and learns to optimize the anomaly scores…The score learner can be defined as a function φ(·; Θ) : X ↦ R, which is a sequential stack of a feature representation learner ψ(·; Θr) : X ↦ Q and an anomaly score learner η(·; Θs) : Q ↦ R” AND [Pang, page 12176, Eq. (5)] “q = ψ(x; Θr)”, wherein the examiner interprets Pang’s feature representation learner ψ is the same as the claimed third machine learning model because ψ is a separately identifiable learned component of the end-to-end model that receives input x and generates the intermediate feature representation q, and Pang trains the end-to-end learner using A and N, the refined anomalous/non-anomalous training sets.) Krawczyk further teaches when categorizing the unlabeled training the unlabeled training data using the plurality of first machine learning models comprises processing, using the plurality of first machine learning models, respective one or more feature values for each training example of the unlabeled training data ([Krawczyk, page 430, sec. 3] “This leads to the formation of a pool of K classifiers assigned to the target class…For a given kth WOCSVM classifier, we propose to use the following heuristic function…OCClustE uses kernel fuzzy c-means…that operates in an artificial feature space created by a kernel function”, wherein the examiner interprets Krawczyk’s plurality of K one-class classifiers operating on feature-space representations of an object to be the same as a plurality of first machine learning models processing respective feature values for each training example; when those feature-space inputs are supplied by Pang’s learned feature representation q mapped above, the feature values are generated using the third machine learning model as claimed.) Pang, Krawczyk, Tax, Xiao, Li, and the instant application are analogous art because they are all directed to machine-learning anomaly detection or one-class classification using learned data representations and trained classifiers. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 14 disclosed by Pang, Krawczyk, Tax, Xiao, and Li to include the learned feature-representation pipeline disclosed by Pang. One would be motivated to do so to effectively use a feature representation learner as a [third] ML model, as suggested by Pang ([Pang, page 12176, sec. 4.2] “The end-to-end anomaly score learner takes A and N as inputs and learns to optimize the anomaly scores…The score learner can be defined as a function φ(·; Θ) : X ↦ R, which is a sequential stack of a feature representation learner ψ(·; Θr) : X ↦ Q and an anomaly score learner η(·; Θs) : Q ↦ R” AND [Pang, page 12176, Eq. (5)] “q = ψ(x; Θr)”.) It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to provide those feature representations to the plurality of one-class classifiers disclosed by Krawczyk. One would be motivated to do so to effectively obtain complementary classifier decisions from learned feature-space inputs and thereby improve ensemble performance, as suggested by Krawczyk ([Krawczyk, page 430, sec. 3] “It assures the initial diversity…and complementarity…which leads to better performance of the ensemble.”). Regarding claim 21, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 1 (see rejection of claim 1). Pang further teaches: wherein the input data comprises one or more of images, video, audio, or data and the output data comprises one or more anomalous images, video, audio or data; ([Pang, page 12173] “unsupervised video anomaly detection...which requires identifying abnormal frames from a large volume of video frames with no manually labeled normal/abnormal training data.”, wherein the examiner interprets “video frames” and “abnormal frames” to be the same as “video” and “anomalous video” because they are both directed to video data in which certain frames are identified as anomalous, and because the claim uses “one or more of”, Pang meeting the “video” modality is sufficient even if Pang does not discuss audio.) Regarding claim 22, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 1 (see rejection of claim 1). Li further teaches wherein the output data comprises data that identifies one or more of anomalous parts as part of a first manufacturing process, an anomalous process as part of a second manufacturing process, fraudulent activity as part of a credit card transaction, a security breach on a monitored network, anomalous patterns associated with patient data, and improper usage of a cloud computing platform. ([Li, Abstract] “It is determined, based on the risk score, that the URL access request is a URL attack request.” AND [Li, [0017]] “a potential URL attack can be found in advance, thereby helping perform security protection in time for a potential abnormal URL access.”, wherein the examiner interprets Li’s output determination that a network URL access request is a URL attack request to be the same as output data identifying a security breach on a monitored network because a detected URL attack is an anomalous security event occurring in network access traffic, and the claim requires only one or more of the listed alternative anomaly types.) Pang, Krawczyk, Tax, Xiao, Li, and the instant application are analogous art because they are all directed to anomaly detection systems that classify input data and identify anomalous conditions. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 1 disclosed by Pang, Krawczyk, Tax, Xiao, and Li to identify the network-security anomaly disclosed by Li. One would be motivated to do so to effectively identify potentially abnormal network access and provide timely security protection, as suggested by Li ([Li, [0017]] “a potential URL attack can be found in advance, thereby helping perform security protection in time for a potential abnormal URL access.”). Claims 8-10, and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Pang in view of Krawczyk in view of Tax in view of Xiao in view of Li further in view of NPL reference “Same But DifferNet: Semi-Supervised Defect Detection with Normalizing Flows” by Rudolph et. al. (referred herein as Rudolph). Regarding claim 8, Pang, Krawczyk, Tax, Xiao, and Li teaches The system of claim 7 (see rejection of claim 7). Pang, Krawczyk, Tax, Xiao, and Li do not teach wherein the one or more thresholds are based on a predetermined percentile value of a distribution of scores corresponding to respective probabilities that training examples in the unlabeled training data are non-anomalous or anomalous. Rudolph teaches wherein the one or more thresholds are based on a predetermined percentile value of a distribution of scores corresponding to respective probabilities that training examples in the unlabeled training data are non-anomalous or anomalous. ([Rudolph, page 1907-1908] “The feature distribution of normal samples is captured by utilizing the latent space of a normalizing flow…each vector is assigned to a likelihood. This enables DifferNet to calculate a likelihood for each image. From this likelihood we derive a scoring function to decide if an image contains an anomaly.” AND [Rudolph, page 1908] “The most common samples are assigned to a high likelihood whereas uncommon images are assigned to a lower likelihood.”, wherein the examiner interprets “deriving a scoring function from likelihoods to decide if an image contains an anomaly” to be the same as “one or more thresholds based on a predetermined percentile value of a distribution of scores” because both describe setting a decision threshold on the statistical distribution of anomaly scores. The examiner further interprets “most common samples assigned to a high likelihood and uncommon images to a low likelihood” to be the same as “scores corresponding to respective probabilities that training examples are non-anomalous or anomalous” because both describe probabilistic values used to assign training examples as belonging to normal (high likelihood) or anomalous (low likelihood) categories.) Pang, Krawczyk, Tax, Xiao, Li, Rudolph, and the instant application are analogous art because they are all directed to anomaly detection using the aid of statistical thresholds. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim 7 disclosed by Pang, Krawczyk, Tax, Xiao, and Li to include the likelihood-based thresholding disclosed by Rudolph. One would be motivated to do so to effectively distinguish anomalous from non-anomalous training examples by leveraging probability-based decision boundaries, as suggested by Rudolph ([Rudolph, page 1908] “From this likelihood we derive a scoring function to decide if an image contains an anomaly. The most common samples are assigned to a high likelihood whereas uncommon images are assigned to a lower likelihood.”). Claim 17 is analogous to claim 8, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 9, Pang, Krawczyk, Tax, Xiao, Li, and Rudolph teaches The system of claim 8 (see rejection of claim 8). Pang further teaches Pang further teaches wherein the one or more thresholds comprise a plurality of thresholds, each threshold based on the predetermined percentile value of a respective distribution of scores generated from training examples processed by a respective first machine learning model of the plurality of first machine learning models. ([Pang, Sec. 5.1] “…we include the 10% most anomalous frames into [A] according to their anomaly scores, because anomaly scores often follow a Gaussian distribution [19] and this decision threshold can provide an approximate 90% confidence level of making false positive errors in such cases… To generate the pseudo normal frame set N, we select the 20% most normal frames based on the anomaly scores. These two cutoff thresholds are used by default as they consistently obtain substantially improved performance on datasets with diverse anomaly rates.” wherein the examiner interprets “10% most anomalous frames” and “20% most normal frames” to be the same as “a plurality of thresholds, each threshold based on the predetermined percentile value” because both are directed to selecting cutoff values based on percentile ranks of the score distributions; “anomaly scores often follow a Gaussian distribution” to be the same as “respective distribution of scores generated from training examples” because both describe the statistical distribution of scores output by trained models; and “cutoff thresholds…consistently obtain improved performance” to be the same as “thresholds… of a respective first machine learning model” because both apply percentile-based thresholds to the distributions of anomaly scores produced by the models). Claim 18 is analogous to claim 9, aside from claim type and minute differences, and thus the same rejection applies as above. Regarding claim 10, Pang, Krawczyk, Tax, Xiao, Li, and teaches The system of claim 9 (see rejection of claim 9). Pang further teaches wherein the one or more processors are further configured to: generate the one or more thresholds based on minimizing, over one or more iterations of an optimization process, respective intra-class variances among anomalous and non-anomalous training examples in the training data. ([Pang, page 12175] “we first initialize A and N using anomaly scores generated by some existing unsupervised anomaly detection methods (see Sec. 4.1), and we iteratively update A and N and retrain φ until the best φ is achieved (see Section 4.3).” AND [Pang, page 12175] “optimizing the objective in Eqn. (1) will identify Θ∗ corresponding to a version of φ(x; Θ∗) that assigns scores to X such that suspicious abnormal and normal samples have anomaly scores as close to respective c1 and c2 as possible, yielding an optimal anomaly ranking.”, wherein the examiner interprets “iteratively update A and N and retrain φ until the best φ is achieved” to be the same as “minimizing, over one or more iterations of an optimization process” because both describe repeatedly refining and optimizing parameters to improve separation between anomalous and non-anomalous classes. The examiner further interprets “assign scores…as close to respective c1 and c2 as possible” to be the same as “minimizing intra-class variances among anomalous and non-anomalous training examples” because both are directed to reducing the spread of scores within each class, thereby generating thresholds that separate anomalous and non-anomalous examples.) Claim 19 is analogous to claim 10, aside from claim type and minute differences, and thus the same rejection applies as above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVAN KAPOOR/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Show 2 earlier events
Dec 08, 2025
Response Filed
Feb 27, 2026
Final Rejection mailed — §103
Apr 27, 2026
Examiner Interview Summary
Apr 27, 2026
Response after Non-Final Action
Apr 27, 2026
Applicant Interview (Telephonic)
May 27, 2026
Request for Continued Examination
May 31, 2026
Response after Non-Final Action
Sep 21, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
7%
Grant Probability
18%
With Interview (+11.1%)
4y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month