Prosecution Insights
Last updated: October 01, 2026
Application No. 18/105,739

ANOMALY DETECTION METHOD, ELECTRONIC DEVICE, NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM, AND COMPUTER PROGRAM

Non-Final OA §101§103
Filed
Feb 03, 2023
Priority
Oct 31, 2022 — RE 10-2022-0143202
Examiner
ADMASU, MAHLIET TASEW
Art Unit
2123
Tech Center
2100 — Computer Architecture & Software
Assignee
Semes Co., Ltd.
OA Round
3 (Non-Final)
0%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
15 currently pending
Career history
12
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
62.0%
+22.0% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103
DETAILED ACTION This communication is in response to the Application No. 18/105,739 filed on August 06, 2026 in which Claims 1-11 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 08/ has been entered. Response to Arguments The amendments filed on August 06, 2026 have been considered. Claims 1-11 have been amended. Thus, Claims 1-11 are pending and presented for examination. Applicant’s arguments filled on August 06, 2026 with respect to the 35 U.S.C. 112(b) rejection have been fully considered and are persuasive. Thus, the previous 35 U.S.C. 112(b) rejection has been withdrawn. Applicant’s arguments filled on August 06, 2026 with respect to the 35 U.S.C. 101 rejection have been fully considered and they are not persuasive. Applicant’s argument on pg. 8 of Argument/Remarks state: PNG media_image1.png 840 982 media_image1.png Greyscale Examiner respectfully disagrees. The solution Applicant describes is achieved by performing the abstract idea, not by any additional element. The recited consolidation of two or more first subsets into a smaller number of second subsets is carried out by the clustering step, and that clustering step (setting initial centroids, allocating each feature to a closest initial centroid, calculating a first mean distance, and updating the centroids and calculating a second mean distance, and thereby grouping the training data into fewer subsets) is itself the mental process and mathematical concept that the claim recites as the abstract idea. The asserted benefit that fewer diagnostic models are trained and applied, and that the computing resources and learning time are correspondingly reduced, is a direct result of performing that abstract clustering, it is not an improvement to the operation of the computer effected by an additional element. The claims do not recite any improvement to how the encoder, decoder, or neural networks are structured, trained, or executed, nor any more efficient manner of performing the model training itself, they recite generically that the models are "trained" and "used," and the computer performs the same conventional clustering, training, and classification operations, merely on a smaller number of subsets. Reducing the number of conventional models that are trained, by grouping the data differently, does not improve the functioning of the computer, it is a consequence of applying the abstract idea. Applicant’s argument on pgs. 8-9 of Argument/Remarks state: PNG media_image2.png 778 967 media_image2.png Greyscale Examiner respectfully disagrees. Reciting the abstract idea with particularity does not make it eligible, because the recited mechanism is the abstract idea. The feature based clustering that merges two or more first subsets into fewer second subsets is carried out by the recited substeps (setting initial centroids, allocating each feature to a closest centroid, calculating a first mean distance, and updating the centroids and calculating a second mean distance) which recite mathematical concepts and a mental process. A detailed clustering procedure is still a clustering procedure, and the reduction in the number of models that follows is a consequence of performing it. The claims recite no improved model architecture, training algorithm, or parameter-update technique. The encoder, decoder, and neural networks are recited generically as trained and used and remain tools that apply the exception. Applicant’s argument on pg. 9 of Argument/Remarks state: PNG media_image3.png 343 1014 media_image3.png Greyscale Examiner respectfully disagrees. The "how" Applicant relies upon, the feature based clustering that consolidates subsets, is the abstract idea itself, not an additional element that ties the exception to the manufacturing technology. The recitation "manufacturing of semiconducting equipment" does no more than identify the environment from which the input data originates and the field in which the anomaly detection idea is applied. Claim 1 does not positively recite any manufacturing operation, any control or adjustment of equipment, or any corrective action taken in response to the detected abnormality. The claim ends at detecting, and the manufacturing environment merely provides the context and source of the data being analyzed. A limitation that merely indicates the field of use in which to apply a judicial exception, however particular that field, does not integrate the exception into a practical application. The "concrete technical benefit" Applicant identifies (fewer models and reduced computational cost) is the same computing benefit that flows from abstract clustering, not an effect on the manufacturing process itself. Applicant’s argument on pg. 9 of Argument/Remarks state: PNG media_image4.png 643 976 media_image4.png Greyscale Examiner respectfully disagrees. The claim in Example 39 was found not to recite any judicial exception, and so the analysis ended without reaching integration or significantly more. The present claims are different in kind. Beyond the training steps, they recite mathematical concepts (setting initial centroids, allocating each feature to a closest centroid, calculating a first mean distance, and updating the centroids and calculating a second mean) together with clustering and detecting steps that recite a mental process. The added feature based clustering that Applicant relies upon is not additional detail on an otherwise-eligible process. It is the recitation of the judicial exception itself. Because the present claims recite a judicial exception that the Example 39 claim did not, they do not resolve in Applicant's favor at the threshold and have been evaluated under Prong Two and Step 2B, where they fail for the reasons above. Accordingly, this argument is not persuasive. Thus, the 35 U.S.C. 101 rejection is maintained. Applicant’s arguments filled on August 06, 2026 with respect to the 35 U.S.C. 103 rejection have been fully considered but are moot because of the new ground of rejection. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-11 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an abstract idea without significantly more. Regarding Claim 1: Step 1: Claim 1 is a method type claim. Therefore, Claims 1-9 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. extracting features from the plurality of training data by computing the plurality of training data […](mental process - extracting features may be performed mentally by a user observing/analyzing and computing the training data by hand and accordingly using judgement/evaluation to extract relevant features based on said analysis); reconstructing the plurality of training data into a plurality of second subsets by clustering the plurality of training data based on the features extracted from the plurality of training data, the clustering comprising at least: setting initial centroids for the features distributed in a feature space; allocating each of the features to a closest initial centroid among the initial centroids; and calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated; (mental process/mathematical concept – reconstructing the training data into subsets can be performed mentally by observing and analyzing the data, and by manually grouping or clustering the data using human judgment and evaluation based on the extracted features. The clustering steps also recite mathematical concepts because setting centroids, assigning features to the closest centroid, and calculating mean distance require numerical comparison, distance measurement, and mathematical calculation in a feature space.); updating a location of the initial centroids based on the calculation of the first mean distance and calculating a second mean distance based on the updated location of the initial centroids (mathematical concept - updating a location of the initial centroids based on the calculation of the first mean distance and calculating a second mean distance recites mathematical calculations) wherein the clustering groups training data from two or more different ones of the plurality of first subsets into a same one of the plurality of second subsets based on the extracted features, […](mental process - grouping training data from two or more different ones of the plurality of training data subsets into a same one of the plurality of first subsets may be performed mentally or using pen and paper by a user observing/analyzing the clustered features and accordingly using judgement/evaluation to group the training data based on said analysis) and detecting an abnormality in […] input data based on one or more determinations […] (mental process – detecting abnormality may be performed mentally by a user observing/analyzing the input data and accordingly using judgement/evaluation to detect any abnormality based on said analysis). Step 2A Prong 2: This judicial exception is not integrated into a practical application. training a first classifier, […] using a plurality of training data to obtain a trained first classifier, the plurality of training data being classified into a plurality of first subsets and the first classifier comprising a first neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) ) – Examiner’s note: high level recitation of training a machine learning model by using a training data without significantly more) […] that includes an encoder and a decoder […] (recited at a high-level of generality (i.e., as an encoder, decoder) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with the encoder of the trained first classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more) training a plurality of second classifiers, that correspond to the plurality of second subsets using the plurality of second subsets to obtain a plurality of trained second classifiers, each of the second classifiers comprising a second neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more. The claim merely states that each second classifier includes a second neural network, but it does not describe any particular neural network structure or technical improvement to how the neural network operates. The neural network is recited only as a generic tool used to train classifiers from the second subsets.) […] and a number of the reconstructed training data subsets is less than a number of the plurality of training data subsets(Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that number of the plurality of second classifiers is less than the number of first subsets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) […] made by the plurality of trained second classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying machine learning model without significantly more) […detecting an abnormality…] in manufacturing of semiconducting equipment from […] (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that abnormality is detected in manufacturing of semiconducting equipment does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. training a first classifier, […] using a plurality of training data to obtain a trained first classifier, the plurality of training data being classified into a plurality of first subsets and the first classifier comprising a first neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) ) – Examiner’s note: high level recitation of training a machine learning model by using a training data without significantly more) […] that includes an encoder and a decoder […] (recited at a high-level of generality (i.e., as an encoder, decoder) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with the encoder of the trained first classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more) training a plurality of second classifiers, that correspond to the plurality of second subsets using the plurality of second subsets to obtain a plurality of trained second classifiers, each of the second classifiers comprising a second neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more. The claim merely states that each second classifier includes a second neural network, but it does not describe any particular neural network structure or technical improvement to how the neural network operates. The neural network is recited only as a generic tool used to train classifiers from the second subsets.) […] and a number of the reconstructed training data subsets is less than a number of the plurality of training data subsets(Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that number of the plurality of second classifiers is less than the number of first subsets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) […] made by the plurality of trained second classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying machine learning model without significantly more) […detecting an abnormality…] in manufacturing of semiconducting equipment from […] (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that abnormality is detected in manufacturing of semiconducting equipment does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 1-9. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. wherein the reconstructing of the plurality of training data into the plurality of second subsets comprises: creating a plurality of feature clusters by clustering the extracted features (mental process - creating feature clusters may be performed mentally by a user by observing/analyzing the extracted features and grouping similar features together based on human judgment, comparison, and evaluation) and creating the plurality of second subsets based on the plurality of feature clusters, wherein the plurality of second subsets correspond to the plurality of feature clusters such that training data corresponding to features included in each of the plurality of feature clusters is allocated to a corresponding second subset of the plurality of second subsets(mental process – creating the second subsets based on the feature clusters may be performed mentally by a person by reviewing which training data corresponds to each clustered feature and assigning that training data to a matching subset.) Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 3 depends on. wherein the clustering of the extracted features comprises: clustering the extracted features based on locations of the extracted features in the feature space (mental process –clustering extracted features may be performed mentally by a user observing/analyzing the locations of the extracted features and accordingly using judgement/evaluation to cluster the extracted features based on said analysis). Step 2A Prong 2 & Step 2B: Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. Step 2A Prong 2 & Step 2B: wherein the training of the plurality of second classifiers comprises: using a final weight of the learned first classifier to train the plurality of second classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model/classifier using a final weight, without significantly more) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 5 depends on. Step 2A Prong 2 & Step 2B: wherein the trained the plurality of second classifiers further comprises: setting the final weight of the learned first classifier as an initial weight of the plurality of second classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model by using the final weight of the learned first classifier as an initial weight of the plurality of second classifiers, without significantly more) Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 4. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 6 depends on. Step 2A Prong 2 & Step 2B: wherein the training of the plurality of second classifiers further comprises: training one of the plurality of second subsets first and then training another one of the plurality of second subsets after while setting the final weight of the learned first classifier as an initial weight of the one of the plurality of second subset that is trained first and setting a final weight of the one of the plurality of second subsets that is trained first as an initial weight of the another one of the plurality of the second subsets that is trained after the one of the plurality of second subsets (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model by sequentially applying weights between subsets without significantly more). Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 4. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 7 depends on. Step 2A Prong 2 & Step 2B: wherein the detecting of the abnormality in the input data, comprises: determining the input data as being normal if any one of the plurality of trained second classifiers determines that the input data is normal((Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of classifier decision logic without significantly more) and determining the input data as being abnormal if the plurality of trained second classifiers all determine that the input data is abnormal (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model to determine abnormality without significantly more). Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 8 depends on. Step 2A Prong 2 & Step 2B: wherein the number of the plurality of second classifiers is less than the number of first subsets (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that number of the plurality of second classifiers is less than the number of first subsets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)). Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 9 depends on. Step 2A Prong 2 & Step 2B: wherein the first classifier and the plurality of second classifiers are configured as auto encoders (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that first classifier and the plurality of second classifiers are configured as auto encoders does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)). Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception Regarding Claim 10: Step 1: Claim 10 is a method type claim. Therefore, Claims 10 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. extracting features for each training data of the training data set by computing each training data of the training data set […](mental process - extracting features may be performed mentally by a user observing/analyzing or computing the training data by hand and accordingly using judgement/evaluation to extract features based on said analysis); reconstructing the plurality of training data subsets to obtain reconstructed training data subsets by: clustering the features extracted for each of the training data of the training data based on locations of the features in a feature space to obtain clustered features by at least: setting initial centroids for features distributed in a feature space; allocating each of the features to a closest initial centroid among the initial centroids; and calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated, and clustering each of the training data of the training data set based on the clustered features (mental process – reconstructing the training data subsets may be performed mentally by a user observing/analyzing the training data subsets, identifying the extracted features, and grouping or clustering the training data based on similarities or relationships among those features. A person could manually assign data into groups using judgment/evaluation based on the observed features. The centroid steps also recite mathematical concepts because setting centroids in a feature space, allocating features to the closest centroid, and calculating mean distance involve numerical comparison, distance measurement, and mathematical calculation); updating a location of the initial centroids based on the calculation of the first mean distance and calculating a second mean distance based on the updated location of the initial centroids (mathematical concept - updating a location of the initial centroids based on the calculation of the first mean distance and calculating a second mean distance recites mathematical calculations) herein the reconstructing groups training data from two or more different ones of the plurality of training data subsets into a same one of the reconstructed training data subsets based on the clustered features […](mental process - grouping training data from two or more different ones of the plurality of training data subsets into a same one of the reconstructed training data subsets may be performed mentally or using pen and paper by a user observing/analyzing the clustered features and accordingly using judgement/evaluation to group the training data based on said analysis) generating a plurality of retrained classifiers by […] with the use of the reconstructed training data subsets (mental process –generating a plurality of retrained classifiers may be performed manually by a user observing/analyzing the reconstructed training data subsets and accordingly using judgement/evaluation to create/generate a plurality of retrained classifiers based on said evaluation); and detecting an abnormality in […] input data based on one or more determinations […](mental process – detecting abnormality may be performed mentally by a user observing/analyzing the input data and accordingly using judgement/evaluation to detect any abnormality based on said analysis). Step 2A Prong 2: This judicial exception is not integrated into a practical application. training a classifier […] using a training data set that includes a plurality of training data subsets to obtain a trained classifier, the classifier comprising a neural network(Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model/classifier by using a training data without significantly more) […] that includes an encoder and a decoder […] (recited at a high-level of generality (i.e., as an encoder, decoder) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with the encoder of the trained classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more) […] retraining the trained classifier […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of retraining/relearning a machine learning model without significantly more) […] and a number of the reconstructed training data subsets is less than a number of the plurality of training data subsets(Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that number of the plurality of second classifiers is less than the number of first subsets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) […] made by the plurality of retrained classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying a machine learning model without significantly more) […detecting an abnormality…] in manufacturing of semiconducting equipment from […] (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that abnormality is detected in manufacturing of semiconducting equipment does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. training a classifier […] using a training data set that includes a plurality of training data subsets to obtain a trained classifier, the classifier comprising a neural network(Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model/classifier by using a training data without significantly more) […] that includes an encoder and a decoder […] (recited at a high-level of generality (i.e., as an encoder, decoder) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with the encoder of the trained classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more) […] retraining the trained classifier […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of retraining/relearning a machine learning model without significantly more) […] and a number of the reconstructed training data subsets is less than a number of the plurality of training data subsets(Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that number of the plurality of second classifiers is less than the number of first subsets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) […] made by the plurality of retrained classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying a machine learning model without significantly more) […detecting an abnormality…] in manufacturing of semiconducting equipment from […] (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that abnormality is detected in manufacturing of semiconducting equipment does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) For the reasons above, Claim 10 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 11: Step 1: Claim 1 is a device/machine type claim. Therefore, Claim 11 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas. Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 11 depends on. extracting features from the plurality of training data by computing the plurality of training data […](mental process - extracting features may be performed mentally by a user observing/analyzing and computing the training data by hand and accordingly using judgement/evaluation to extract relevant features based on said analysis); reconstructing the plurality of training data into a plurality of second subsets by clustering the plurality of training data based on the features extracted from the plurality of training data, the clustering comprising at least: setting initial centroids for the features distributed in a feature space; allocating each of the features to a closest initial centroid among the initial centroids; and calculating a mean distance between each of the features and the closest initial centroid to which each of the features are allocated; (mental process/mathematical concept – reconstructing the training data into subsets can be performed mentally by observing and analyzing the data, and by manually grouping or clustering the data using human judgment and evaluation based on the extracted features. The clustering steps also recite mathematical concepts because setting centroids, assigning features to the closest centroid, and calculating mean distance require numerical comparison, distance measurement, and mathematical calculation in a feature space.); updating a location of the initial centroids based on the calculation of the first mean distance and calculating a second mean distance based on the updated location of the initial centroids (mathematical concept - updating a location of the initial centroids based on the calculation of the first mean distance and calculating a second mean distance recites mathematical calculations) wherein the clustering groups training data from two or more different ones of the plurality of first subsets into a same one of the plurality of second subsets based on the extracted features, […](mental process - grouping training data from two or more different ones of the plurality of training data subsets into a same one of the plurality of first subsets may be performed mentally or using pen and paper by a user observing/analyzing the clustered features and accordingly using judgement/evaluation to group the training data based on said analysis) and detecting an abnormality in […] input data based on one or more determinations […] (mental process – detecting abnormality may be performed mentally by a user observing/analyzing the input data and accordingly using judgement/evaluation to detect any abnormality based on said analysis). Step 2A Prong 2 & Step 2B: The judicial exception is not integrated into a practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. An electronic device comprising: a processor; and a memory connected to the processor (recited at a high-level of generality (i.e., as an electronic device, generic processor, and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components) wherein the memory stores instructions that can be executed by the processor, to cause the processor to perform operations that comprise: (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that memory stores instructions that can be executed by the processor, to cause the processor to perform operations does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)). training a first classifier, […] using a plurality of training data to obtain a trained first classifier, the plurality of training data being classified into a plurality of first subsets and the first classifier comprising a first neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) ) – Examiner’s note: high level recitation of training a machine learning model by using a training data without significantly more) […] that includes an encoder and a decoder […] (recited at a high-level of generality (i.e., as an encoder, decoder) such that it amounts to no more than mere instructions to apply the exception using generic computer components) […] with the encoder of the trained first classifier (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more) training a plurality of second classifiers, that correspond to the plurality of second subsets using the plurality of second subsets to obtain a plurality of trained second classifiers, each of the second classifiers comprising a second neural network (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more. The claim merely states that each second classifier includes a second neural network, but it does not describe any particular neural network structure or technical improvement to how the neural network operates. The neural network is recited only as a generic tool used to train classifiers from the second subsets.) […] and a number of the reconstructed training data subsets is less than a number of the plurality of training data subsets(Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that number of the plurality of second classifiers is less than the number of first subsets does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) […] made by the plurality of trained second classifiers (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of applying machine learning model without significantly more) […detecting an abnormality…] in manufacturing of semiconducting equipment from […] (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that abnormality is detected in manufacturing of semiconducting equipment does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h)) For the reasons above, Claim 11 is rejected as being directed to an abstract idea without significantly more. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 - 11 are rejected under 35 U.S.C. 103 as being unpatentable over Yoon et al. (hereinafter Yoon) (WO 2020013494), in view of Sohn et al. (hereinafter Sohn) (KR 102363737), and in further view of Baradaran et al. (hereinafter Baradaran) (US 20170124478) Regarding Claim 1, Yoon teaches a method (Yoon, Pg. 1 – Abstract - line 3, “a method for detecting an anomaly in data”, thus a method is disclosed) comprising: training a first classifier, that includes an encoder and a decoder using a plurality of training data to obtain a trained first classifier, the plurality of training data being classified into a plurality of first subsets and the first classifier comprising a first neural network (Yoon, Pg. 6 – lines 9-13, “The processor 110 may generate an anomaly sensing model for sensing anomaly of data by learning a network function using the learning data set. The training data set can include a plurality of training data subsets. The plurality of training data subsets may comprise different training data grouped by predetermined criteria”, &Pg. 12 – lines 46 – 52 & Pg. 13 – lines 1 - 3, “In one embodiment of the present disclosure, the network function 200 may include an autoencoder. The auto encoder may be a kind of artificial neural network for outputting output data similar to the input data. The auto encoder may include at least one hidden layer, and an odd number of hidden layers may be disposed between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding) and then expanded symmetrically from the bottleneck layer to the output layer (symmetrical with the input layer). In this case, in the example of FIG. 2, the dimensional reduction layer and the dimensional reconstruction layer are illustrated to be symmetrical, but the present disclosure is not limited thereto, and nodes of the dimensional reduction layer and the dimensional reconstruction layer may or may not be symmetrical.”, thus Yoon discloses training a first classifier/network function using a plurality of training data/learning data set to obtain a trained anomaly sensing model. Yoon further discloses that the training data set includes a plurality of training data subsets grouped by predetermined criteria. Yoon also teaches that the network function may include an autoencoder, which is an artificial neural network having an encoding/dimensional reduction portion and a decoding/dimensional reconstruction portion. Therefore, Yoon discloses a first classifier comprising a first neural network with an encoder and decoder, trained using a plurality of training data classified into a plurality of first subsets.); extracting features from the plurality of training data by computing the plurality of training data with the encoder of the trained first classifier (Yoon, Pg. 5 – lines 6 - 9, “The processor 110 may process neural networks, such as processing input data for learning in deep learning (DN), extracting features from the input data, calculating errors, and weighting neural networks using backpropagation. Can perform calculations for learning.” & Pg. 9 – lines 34 - 37, “Pre-learned network functions in the present disclosure can be learned to reduce and reconstruct the dimension of the training data. The network function of the present disclosure may include an auto encoder capable of dimensional reduction and dimensional reconstruction of input data.”, therefore extracting features from the plurality of training data by computing the plurality of training data with the encoder of the trained first classifier/pre-learned network functions is disclosed); reconstructing the plurality of training data into a plurality of second subsets […] based on the features extracted from the plurality of training data, […] (Yoon, Pg. 10 – lines 7 - 9, “In the input data of the present disclosure, in the production process, the production equipment may perform an operation in which the processor 110 transforms the input data into data different from the input data and restores the data using the anomaly sensing submodel.” &Pg. 10 – lines 11 - 16, “The processor 110 may extract a feature from the input data using the anomaly detection submodel, and restore the input data based on the feature. As described above, since the network function included in the anomaly sensing submodel of the present disclosure may include a network function capable of restoring the input data, the processor 110 may calculate the input data using the anomaly sensing submodel. Can restore the input data.”, therefore reconstructing the plurality of training data into a plurality of second subsets based on the extracted features is disclosed); wherein […] groups training data from two or more different ones of the plurality of first subsets into a same one of the plurality of second subsets based on the extracted features, and a number of the plurality of second subsets is less than a number of the plurality of first subsets(Yoon, Page 6, “For example, there is a recipe change in the process, but the input data generated after the recipe change (that is, the sensor data obtained in the process generated after the recipe change) is determined as a new pattern by the existing anomaly detection submodel. If not (i.e., having no novelty above the threshold), the input data generated before and after the recipe change may be grouped into one subset of training data”, & Page 13, “For example, when the training data generated for 6 months is configured into one subset, the training data generated by two or more recipes (ie, a plurality of normal patterns) may be included in one subset. In this case, one anomaly detection submodel may learn a plurality of normal patterns and may be used for anomaly detection (eg, novelty detection) for the plurality of normal patterns”, thus wherein […] groups training data from two or more different ones of the plurality of first subsets into a same one of the plurality of second subsets based on the extracted features, and a number of the plurality of second subsets is less than a number of the plurality of first subsets is disclosed because Yoon teaches that training data generated before and after a recipe change may be grouped into one subset of training data, and further teaches that training data generated by two or more recipes may be included in one subset, thereby consolidating multiple recipe-based subsets into fewer resulting subsets); training a plurality of second classifiers, that correspond to the plurality of second subsets using the plurality of second subsets to obtain a plurality of trained second classifiers, each of the second classifiers comprising a second neural network (Yoon, Pg. 17 – lines 3 - 6, “The computing device 100 generates a first anomaly sensing submodel that includes the first network function trained with the first training data subset, and then, if there is a change in the recipe, trained with the second training data subset. A second anomaly sensing submodel that includes a second network function may be generated.”, thus Yoon discloses training a plurality of second classifiers/anomaly sensing submodels corresponding to the plurality of second subsets/training data subsets. Yoon teaches generating a first anomaly sensing submodel including a first network function trained with a first training data subset, and generating a second anomaly sensing submodel including a second network function trained with a second training data subset. Because each submodel includes a network function trained using a corresponding training data subset, Yoon discloses obtaining a plurality of trained second classifiers, each comprising a second neural network); and detecting an abnormality in manufacturing of semiconducting equipment from input data based on one or more determinations made by the plurality of trained second classifiers (Yoon, Pg. 3 – lines 9 - 10, “In an alternative embodiment, determining whether anomaly exists in the input data comprises: determining whether anomaly exists in the input data using the second anomaly sensing submodel”, &Pg. 7 – lines 8 - 10, “And a second anomaly sensing submodel comprising a second network function pre-learned with a second subset of learning data composed of learning data generated during the second time interval”, & Page 5, “As a more specific example, work-in-progress information including about 120,000 items per lot acquired in a semiconductor fab, raw processing tool data, equipment interface information, process metrology information information) (e.g., contains more than 1000 items per lot), defect information accessible to yield engineers, operational test information, sort information (including datalogs and bitmaps), The present disclosure is not limited thereto”, thus, detecting an abnormality in manufacturing of semiconducting equipment from input data based on one or more determinations made by the plurality of trained second classifiers is disclosed) Yoon does not explicitly disclose clustering the plurality of training data based on the extracted features […] the clustering comprising at least: setting initial centroids for features distributed in a feature space; allocating each of the features to a closest initial centroid among the initial centroids; calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated, updating a location of […] and calculating a second mean distance based on the updated location of […], and […] the initial centroids based on the calculation of the first mean distance […] the initial centroids. However, Sohn teaches: reconstructing the plurality of training data into a plurality of second subsets by clustering the plurality of training data based on the extracted features (Sohn, Pg. 3 – lines 20 - 22, “The encoder 102a may extract a latent feature vector z from the input normal data. In this case, normal data of a data set having various classes may be input to the encoder 102a. The decoder 102b may reconstruct normal data based on the latent feature vector output from the encoder 102a.”, &Pg. 3 – lines 27 - 28, “The labeling module 104 may perform clustering on the latent feature vectors z output from the encoder 102a, and may give a pseudo label to each clustered cluster.”, therefore the plurality of training data is reconstructed into a plurality of second subsets(classes) by clustering the plurality of training data based on the extracted features) setting initial centroids for features distributed in a feature space (Sohn, Page – 3, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, & “ Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below”, thus establishing the centers of the initialized clusters i.e., initial centroids for the latent features distributed in the latent feature space) allocating each of the features to a closest initial centroid among the initial centroids (Sohn, Page 3, “Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below”, & Page 6, “the anomaly detection apparatus 100 calculates the similarity between each latent feature vector and the centers of the primary clustered clusters, and learns the second artificial neural network model so that the probability distribution for the calculated similarity matches the preset target probability distribution”, thus Sohn discloses allocating each latent feature vector to a cluster center based on the measured similarity/proximity between the position of that feature and the cluster centers i.e., the K-means assignment of each feature to its nearest centroid) calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated (Sohn, Page – 3, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, thus Sohn discloses calculating distance based clustering of extracted features because Sohn teaches performing initial clustering on latent feature vectors using a K-means algorithm. In K-means clustering, feature vectors are assigned to clusters based on proximity to cluster centroids, which requires calculating distances between the feature vectors and the centroids. Therefore, Sohn suggests allocating extracted features to the closest centroid and using distance calculations between the features and the corresponding centroid as part of the clustering process) […] the initial centroids based on the calculation of the first mean distance […] the initial centroids (Sohn, Page 3, “For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, & “the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster”, thus […] the initial centroids based on the calculation of the first mean distance […] the initial centroids is disclosed, because Sohn teaches performing K-means clustering on the latent feature vectors and measuring the similarity between each latent feature vector and the center of an initialized cluster, thereby determining the relationship of each feature to the initial centroids for clustering) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine Yoon’s approach of reconstructing the plurality of training data into a plurality of second subsets based on extracted features with Sohn’s approach of clustering the plurality of training data based on the extracted features to reconstruct the plurality of training data into a plurality of second subsets/classes, thereby improving the accuracy and efficiency of classifying an abnormal situation and anomaly detection (Sohn, Pg. 2 – lines 38 - 41, “2 is a diagram schematically illustrating multi-class classification in anomaly detection technology according to an embodiment of the present invention. Referring to FIG. 2 , the disclosed embodiment is to more accurately classify an abnormal situation by generating a strict decision boundary for each class when a normal sample includes multiple classes”, & Pg. 5 – lines 16 - 19, “According to the disclosed embodiment, even when multiple latent classes are included in the normal data set, by clustering and pseudo-labeling using latent characteristics extracted from normal data, it is more efficient than the single decision boundary-based anomaly detection technique. Anomaly detection can be performed with excellent performance.”) Yoon combined with Sohn does not explicitly disclose updating a location of […] and calculating a second mean distance based on the updated location of […]. However, Baradaran teaches: updating a location of […] and calculating a second mean distance based on the updated location of […] (Baradaran, Par. [0274], “The clustering engine 720 may calculate a new mean or centroid of the data points 731, 731′, and 733 for each cluster or partition. The clustering engine 720 may then repeat and iterate the process of calculating distances and new means or centroids for each cluster until the K means clustering algorithm converges in results”, & Par. [0257], “In some embodiments, the clustering engine 720 may be configured to determine, for each partition, a mean or centroid point of the data points for the respective partition, e.g., via an averaging and/or weighting process on characteristics or attributes of the data points. In some embodiments, the clustering engine 720 may be configured to determine or calculate distances (e.g., Euclidean distance) between the mean or centroid point to all data points of the training data. In some embodiments, the clustering engine 720 may be configured to iterate a process of calculating distances using each (e.g., new) mean or centroid point, partitioning the dataset into clusters, and/or assigning the data points to the respective partition based on the nearest mean or centroid point until convergence is reached”, thus updating a location of […] and calculating a second mean distance based on the updated location of […] is disclosed, because Baradaran first assigns the data points to a cluster based on their distance from a centroid, then recalculates a new mean or centroid from the assigned data points, thereby updating the centroid location, and subsequently recalculates distances using the new centroid location during the next iteration of the K-means process) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Yoon and Sohn with Baradaran by incorporating Baradaran’s iterative centroid update and distance recalculation technique into Sohn’s K-means clustering of the extracted latent feature vectors. Sohn teaches extracting latent feature vectors and clustering them using K-means based on their similarity to initialized cluster centers. Baradaran teaches updating the cluster centroids based on assigned data points and recalculating distances using the updated centroids until convergence. Therefore, a POSITA would have been motivated to apply Baradaran’s iterative K-means procedure to Sohn’s clustering of latent feature vectors so that the cluster centers are updated based on the assigned features and the distances are recalculated using the updated centroids, thereby refining the cluster assignments and improving the clustering used for anomaly detection (Baradaran, Par. [0271], “a method 703 for improving anomaly detection using injected outliers is depicted. In brief overview, a device may include a set of outliers into a training dataset of data points (706). The device may identify, using a K-means clustering algorithm applied on the training dataset, at least a first cluster of data points (709). The device may determine a center and an outer radius of a region that covers at least a spatial extent of the first cluster of data points (712). The device may determine a first normalcy radius for the first cluster by adjusting the region around the center until a point at which all artificial outliers are excluded from a region defined by the first normalcy radius (715)”) Regarding Claim 2, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 1 as cited above and Sohn further teaches: wherein the reconstructing of the plurality of training data into the plurality of second subsets comprises: creating a plurality of feature clusters by clustering the extracted features (Sohn, Pg. 3 – lines 29 - 36, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm. In addition, the labeling module 104 may perform secondary clustering on the initially clustered latent feature vectors (z) when the latent feature vectors (z) are sufficiently initialized and there is no further cluster change. Here, the secondary clustering may be performed through an artificial neural network model such as a deep neural network (DNN), but the example of the artificial neural network model is not limited thereto”, thus Sohn discloses creating a plurality of feature clusters by clustering the extracted features/latent feature vectors. Sohn teaches that the labeling module performs initial clustering on the latent feature vectors using a K-means algorithm, and further performs secondary clustering on the initially clustered latent feature vectors after the vectors are sufficiently initialized and there is no further cluster change. Therefore, Sohn discloses forming feature clusters from the extracted latent feature vectors) and creating the plurality of second subsets based on the plurality of feature clusters, wherein the plurality of the second subsets correspond to the plurality of feature clusters such that training data corresponding to features included in each of the plurality of feature clusters is allocated to a corresponding second subset of the plurality of second subsets (Sohn, Pg. 3 – lines 29 - 36, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm. In addition, the labeling module 104 may perform secondary clustering on the initially clustered latent feature vectors (z) when the latent feature vectors (z) are sufficiently initialized and there is no further cluster change. Here, the secondary clustering may be performed through an artificial neural network model such as a deep neural network (DNN), but the example of the artificial neural network model is not limited thereto.”, thus Sohn discloses creating a plurality of feature clusters by clustering extracted features/latent feature vectors using K-means clustering. Sohn further discloses performing secondary clustering on the initially clustered latent feature vectors after the latent feature vectors are initialized and there is no further cluster change. Because the latent feature vectors correspond to the data being clustered, the clustered feature vectors form feature-based groups, and the data associated with those clustered features is allocated into corresponding groups. Therefore, Sohn discloses creating the plurality of second subsets based on the plurality of feature clusters, wherein each second subset corresponds to a feature cluster and includes training data corresponding to the features included in that feature cluster). The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 3, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 2 as cited above and Sohn further teaches: wherein the clustering of the extracted features comprises: clustering the extracted features based on locations of the extracted features in the feature space (Sohn, Pg. 3 – lines 27 - 28, “The labeling module 104 may perform clustering on the latent feature vectors z output from the encoder 102a, and may give a pseudo label to each clustered cluster.”, & Pg. 3 – line 37 - 38, “Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters.”, therefore Sohn teaches clustering the extracted features comprising clustering the extracted features based on locations of the extracted features(positions of the latent feature vectors) in feature space). The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 4, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 1 as cited above and Yoon further teaches: wherein the training of the plurality of second classifiers comprises: using a final weight of the trained first classifier to train the plurality of second classifier (Yoon, Pg. 7 – line 52 & Pg. 8 – lines 1-3, “A predetermined number of layers from the output layers of the second network function can be learned using the weight of the latest network function as the initial weight. The processor 110 may train the second network function by using the initial weights of some layers proximate to the output layer of the newly learned second network function as the weights of the already learned first network functions”, & Pg. 8 – lines 2 - 7, “This weight sharing can reduce the amount of computation required to learn the second network function. That is, by using the initial weight of a predetermined number of layers proximate to the output layer of the second network function as the weight of the pre-learned network function rather than random, the knowledge of the pre-learned network function can be utilized in the dimensional reconstruction of the input data.”, thus Yoon discloses using a final weight of the learned first classifier/pre-learned first network function to train the plurality of second classifiers/second network functions. Yoon teaches using the weight of the latest or already learned first network function as the initial weight for layers of the second network function.) Regarding Claim 5, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 4 as cited above and Yoon further teaches: wherein the training of the plurality of second classifiers, further comprises: setting the final weight of the trained first classifier as an initial weight of the plurality of second classifiers (Yoon, Pg. 7 – line 52 & Pg. 8 – lines 1-3, “A predetermined number of layers from the output layers of the second network function can be learned using the weight of the latest network function as the initial weight. The processor 110 may train the second network function by using the initial weights of some layers proximate to the output layer of the newly learned second network function as the weights of the already learned first network functions.”, therefore setting the final weight of the trained first classifier as an initial weight of the plurality of second classifiers is disclosed) Regarding Claim 6, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 4 as cited above and Yoon further teaches: wherein the training of the plurality of second classifiers further comprises: training one of the plurality of second subsets first and then training another one of the plurality of second subsets after while setting the final weight of the trained first classifier as an initial weight of the one of the plurality of second subsets that is trained first and setting a final weight of the one of the plurality of second subset that is trained first as an initial weight of the another one of the plurality of the second subsets that is trained after the one of the plurality of second subsets (Yoon, Pg. 7 – lines 19 - 23, “A second anomaly detection submodel that includes two network functions may be generated. The first training data subset and the second data subset may include data obtained in a production process produced by different recipes. Here, the initial weight of the second network function may share at least a portion of the weight of the pre-learned first network function.”, & Pg. 7 – lines 36 - 40, “Dimension Reduction of the Second Network Function A predetermined number of layers from the layers closest to the input layer of the layers of the network may use initial weights as weights of corresponding layers of the first learned network function. A predetermined number of layers from the input layer of the second network function can be learned using the weight of the latest network function as the initial weight.”, thus training one of the plurality of second subsets first and then training another one of the plurality of second subsets after while setting the final weight of the trained first classifier as an initial weight of the one of the plurality of second subsets that is trained first and setting a final weight of the one of the plurality of second subset that is trained first as an initial weight of the another one of the plurality of the second subsets that is trained after the one of the plurality of second subsets is disclosed) Regarding Claim 7, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 1 as cited above and Yoon further teaches: wherein the detecting of the abnormality in the input data, comprises: determining the input data as being normal if any one of the plurality of trained second classifiers determines that the input data is normal (Yoon, Pg. 11 – lines 2 - 8, “The input data may be determined as anomaly in the latest anomaly detection submodel, but, for example, when there is a change of a recipe in a process, when the input data is sensor data obtained in a process produced by a previous recipe, The input data may be normal in the previous recipe. In this case, the processor 110 may determine the input data as normal data. When all of the plurality of anomaly sensing submodels included in the anomaly sensing model determine that anomaly exists in the input data, the processor 110 may determine that the anomaly exists as the input data.”, thus determining the input data as being normal if any one of the plurality of trained second classifiers determines that the input data is normal is disclosed) and determining the input data as being abnormal if the plurality of trained second classifiers all determine that the input data is abnormal (Yoon, Pg. 11 – lines 2 - 8, “The input data may be determined as anomaly in the latest anomaly detection submodel, but, for example, when there is a change of a recipe in a process, when the input data is sensor data obtained in a process produced by a previous recipe, The input data may be normal in the previous recipe. In this case, the processor 110 may determine the input data as normal data. When all of the plurality of anomaly sensing submodels included in the anomaly sensing model determine that anomaly exists in the input data, the processor 110 may determine that the anomaly exists as the input data.”, therefore determining the input data as being abnormal if the plurality of second classifiers all determine that the input data is abnormal is disclosed) Regarding Claim 8, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 1 as cited above and Yoon further teaches: Wherein the number of the plurality of second classifiers is less than the number of the plurality of first subsets (Yoon, Pg. 8 – line 50 – 51 & Pg. 9 – lines 1 - 11, “Since the second anomaly sensing submodel is trained with the training data including the training data submodel for which the previous anomaly sensing submodel is trained, the second anomaly sensing submodel is the knowledge learned in the first anomaly sensing submodel. Can succeed. In this case, in the training data for training the second anomaly sensing submodel, the sampling rate of the second training data subset and the first training data subset (that is, the training data subset for which the previous submodel was trained) is different. can do. The first training data subset consisting of the training data generated during the first time interval for generating the first anomaly sensing submodel is such that only a portion of the training data included in the first training data subset is used for training. The sample rate may be sampled at a sampling rate lower than that of the second training data subset for generating the Mali sense submodel. The first training data subset may be used for training the second anomaly sensing submodel, but in this case may be sampled at a lower sampling rate than the second training data subset.”, therefore, this shows that the second anomaly detection submodel is trained using fewer training samples from the first training data subset than from the second training data subset, which is analogous to a situation where the number of second classifiers is less than the number of first subsets) Regarding Claim 9, Yoon and Sohn combined with Baradaran teaches all of the limitations of claim 1 as cited above and Yoon further teaches: wherein the first classifier and the plurality of second classifiers are configured as auto encoders (Yoon, Pg. 12 – lines 46 - 48, “In one embodiment of the present disclosure, the network function 200 may include an autoencoder. The auto encoder may be a kind of artificial neural network for outputting output data similar to the input data.”, thus wherein the first classifier and the plurality of second classifiers are configured as auto encoders is disclosed) Regarding Claim 10, Yoon teaches a method (Yoon, Pg. 1 – Abstract – line 3, “a method for detecting an anomaly in data”, thus a method is disclosed) comprising: training a classifier, that includes an encoder and a decoder using a training data set that includes a plurality of training data subsets to obtain a trained classifier, the classifier comprising a neural network(Yoon, Pg. 6 – lines 9-13, “The processor 110 may generate an anomaly sensing model for sensing anomaly of data by learning a network function using the learning data set. The training data set can include a plurality of training data subsets. The plurality of training data subsets may comprise different training data grouped by predetermined criteria”, &Pg. 12 – lines 46 – 52 & Pg. 13 – lines 1 - 3, “In one embodiment of the present disclosure, the network function 200 may include an autoencoder. The auto encoder may be a kind of artificial neural network for outputting output data similar to the input data. The auto encoder may include at least one hidden layer, and an odd number of hidden layers may be disposed between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding) and then expanded symmetrically from the bottleneck layer to the output layer (symmetrical with the input layer). In this case, in the example of FIG. 2, the dimensional reduction layer and the dimensional reconstruction layer are illustrated to be symmetrical, but the present disclosure is not limited thereto, and nodes of the dimensional reduction layer and the dimensional reconstruction layer may or may not be symmetrical.”, thus Yoon discloses training a classifier/network function using a training data set/learning data set to obtain a trained anomaly sensing model. Yoon further discloses that the training data set includes a plurality of training data subsets grouped by predetermined criteria. Yoon also discloses that the network function may include an autoencoder, which is an artificial neural network having an encoding/dimensional reduction portion and a decoding/dimensional reconstruction portion. Therefore, Yoon discloses a classifier comprising a neural network with an encoder and a decoder, trained using a training data set that includes a plurality of training data subsets.); extracting features for each training data of the training data set by computing each training data of the training data set with the encoder of the trained first classifier (Yoon, Pg. 5 – lines 6 - 9, “The processor 110 may process neural networks, such as processing input data for learning in deep learning (DN), extracting features from the input data, calculating errors, and weighting neural networks using backpropagation. Can perform calculations for learning.” & Pg. 9 – lines 34 - 37, “Pre-learned network functions in the present disclosure can be learned to reduce and reconstruct the dimension of the training data. The network function of the present disclosure may include an auto encoder capable of dimensional reduction and dimensional reconstruction of input data.”, thus Yoon discloses extracting features for each training data of the training data set by computing each training data with the encoder of the trained classifier. Yoon teaches that the processor performs neural network learning operations, including extracting features from input data. Yoon also discloses that the trained network function may include an autoencoder capable of dimensional reduction and dimensional reconstruction of input data. The dimensional reduction portion of the autoencoder corresponds to the encoder. Accordingly, Yoon discloses using the encoder portion of the trained autoencoder to process each training data and extract corresponding features); reconstructing the plurality of training data subsets to obtain reconstructed training data subsets by: […](Yoon, Pg. 10 – lines 7 - 9, “In the input data of the present disclosure, in the production process, the production equipment may perform an operation in which the processor 110 transforms the input data into data different from the input data and restores the data using the anomaly sensing submodel.” &Pg. 10 – lines 11 - 16, “The processor 110 may extract a feature from the input data using the anomaly detection submodel, and restore the input data based on the feature. As described above, since the network function included in the anomaly sensing submodel of the present disclosure may include a network function capable of restoring the input data, the processor 110 may calculate the input data using the anomaly sensing submodel. Can restore the input data.”, therefore reconstructing the plurality of training data subsets to obtain reconstructed training data subsets is disclosed); wherein […] groups training data from two or more different ones of the plurality of training data subsets into a same one of […] training data subsets based on the […features…], and a number of […] training data subsets is less than a number of the plurality of training data subsets (Yoon, Page 6, “For example, there is a recipe change in the process, but the input data generated after the recipe change (that is, the sensor data obtained in the process generated after the recipe change) is determined as a new pattern by the existing anomaly detection submodel. If not (i.e., having no novelty above the threshold), the input data generated before and after the recipe change may be grouped into one subset of training data”, & Page 13, “For example, when the training data generated for 6 months is configured into one subset, the training data generated by two or more recipes (ie, a plurality of normal patterns) may be included in one subset. In this case, one anomaly detection submodel may learn a plurality of normal patterns and may be used for anomaly detection (eg, novelty detection) for the plurality of normal patterns”, thus wherein […] groups training data from two or more different ones of the plurality of training data subsets into a same one of […] training data subsets based on the […features…], and a number of […] training data subsets is less than a number of the plurality of training data subsets is disclosed because Yoon teaches that training data generated before and after a recipe change may be grouped into one subset of training data, and further teaches that training data generated by two or more recipes may be included in one subset, thereby consolidating multiple recipe-based subsets into fewer resulting subsets) generating a plurality of retrained classifiers by retraining the trained classifier with the use of the reconstructed training data subsets (Yoon, Pg. 17 – lines 3 - 6, “The computing device 100 generates a first anomaly sensing submodel that includes the first network function trained with the first training data subset, and then, if there is a change in the recipe, trained with the second training data subset. A second anomaly sensing submodel that includes a second network function may be generated.”, thus generating a plurality of retrained classifiers by retraining the trained classifier with the use of the reconstructed training data subsets is disclosed); and detecting an abnormality in manufacturing of semiconducting equipment from input data based on one or more determinations made by the plurality of retrained classifiers (Yoon, Pg. 3 – lines 9 - 10, “In an alternative embodiment, determining whether anomaly exists in the input data comprises: determining whether anomaly exists in the input data using the second anomaly sensing submodel”, &Pg. 7 – lines 8 - 10, “And a second anomaly sensing submodel comprising a second network function pre-learned with a second subset of learning data composed of learning data generated during the second time interval”, & Page 5, “As a more specific example, work-in-progress information including about 120,000 items per lot acquired in a semiconductor fab, raw processing tool data, equipment interface information, process metrology information information) (e.g., contains more than 1000 items per lot), defect information accessible to yield engineers, operational test information, sort information (including datalogs and bitmaps), The present disclosure is not limited thereto”, thus detecting an abnormality in manufacturing of semiconducting equipment from input data based on one or more determinations made by the plurality of retrained classifiers is disclosed). Yoon does not explicitly disclose clustering features extracted for each of the training data of the training data set based on locations of the features in a feature space to obtain clustered features by at least: setting initial centroids for features distributed in a feature space; allocating each of the features to a closest initial centroid among the initial centroids; and calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated, and clustering each of the training data of the training data set based on the clustered features, […] the initial centroids based on the calculation of the first mean distance […] the initial centroids, and updating a location of […] and calculating a second mean distance based on the updated location of […]. However, Sohn teaches: reconstructing the training data subsets by clustering the extracted features based on locations of the extracted features in feature space and clustering each training data of the training data set based on the clustered features (Sohn, Pg. 3 – lines 20 - 22, “The encoder 102a may extract a latent feature vector z from the input normal data. In this case, normal data of a data set having various classes may be input to the encoder 102a. The decoder 102b may reconstruct normal data based on the latent feature vector output from the encoder 102a.”, Pg. 3 – lines 27 - 28, “The labeling module 104 may perform clustering on the latent feature vectors z output from the encoder 102a, and may give a pseudo label to each clustered cluster.” & Pg. 3 – lines 37 - 40, “Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below.”, therefore reconstructing the training data subsets by clustering the extracted features based on locations of the extracted features in feature space and clustering each training data of the training data set based on the clustered features is disclosed) setting initial centroids for features distributed in a feature space (Sohn, Page – 3, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, & “ Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below”, thus establishing the centers of the initialized clusters i.e., initial centroids for the latent features distributed in the latent feature space) allocating each of the features to a closest initial centroid among the initial centroids (Sohn, Page 3, “Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below”, & Page 6, “the anomaly detection apparatus 100 calculates the similarity between each latent feature vector and the centers of the primary clustered clusters, and learns the second artificial neural network model so that the probability distribution for the calculated similarity matches the preset target probability distribution”, thus Sohn discloses allocating each latent feature vector to a cluster center based on the measured similarity/proximity between the position of that feature and the cluster centers i.e., the K-means assignment of each feature to its nearest centroid) calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated (Sohn, Page – 3, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, thus Sohn discloses calculating distance based clustering of extracted features because Sohn teaches performing initial clustering on latent feature vectors using a K-means algorithm. In K-means clustering, feature vectors are assigned to clusters based on proximity to cluster centroids, which requires calculating distances between the feature vectors and the centroids. Therefore, Sohn suggests allocating extracted features to the closest centroid and using distance calculations between the features and the corresponding centroid as part of the clustering process) […] the initial centroids based on the calculation of the first mean distance […] the initial centroids (Sohn, Page 3, “For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, & “the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster”, thus […] the initial centroids based on the calculation of the first mean distance […] the initial centroids is disclosed, because Sohn teaches performing K-means clustering on the latent feature vectors and measuring the similarity between each latent feature vector and the center of an initialized cluster, thereby determining the relationship of each feature to the initial centroids for clustering) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine Yoon’s approach of reconstructing the training data subsets by clustering the extracted features based on locations of the extracted features in feature space and clustering each training data of the training data set based on the clustered features, thereby improving the accuracy and efficiency of classifying an abnormal situation and anomaly detection (Sohn, Pg. 2 – lines 38 - 41, “2 is a diagram schematically illustrating multi-class classification in anomaly detection technology according to an embodiment of the present invention. Referring to FIG. 2 , the disclosed embodiment is to more accurately classify an abnormal situation by generating a strict decision boundary for each class when a normal sample includes multiple classes”, & Pg. 5 – lines 16 - 19, “According to the disclosed embodiment, even when multiple latent classes are included in the normal data set, by clustering and pseudo-labeling using latent characteristics extracted from normal data, it is more efficient than the single decision boundary-based anomaly detection technique. Anomaly detection can be performed with excellent performance.”) Yoon combined with Sohn does not explicitly disclose updating a location of […] and calculating a second mean distance based on the updated location of […]. However, Baradaran teaches: updating a location of […] and calculating a second mean distance based on the updated location of […] (Baradaran, Par. [0274], “The clustering engine 720 may calculate a new mean or centroid of the data points 731, 731′, and 733 for each cluster or partition. The clustering engine 720 may then repeat and iterate the process of calculating distances and new means or centroids for each cluster until the K means clustering algorithm converges in results”, & Par. [0257], “In some embodiments, the clustering engine 720 may be configured to determine, for each partition, a mean or centroid point of the data points for the respective partition, e.g., via an averaging and/or weighting process on characteristics or attributes of the data points. In some embodiments, the clustering engine 720 may be configured to determine or calculate distances (e.g., Euclidean distance) between the mean or centroid point to all data points of the training data. In some embodiments, the clustering engine 720 may be configured to iterate a process of calculating distances using each (e.g., new) mean or centroid point, partitioning the dataset into clusters, and/or assigning the data points to the respective partition based on the nearest mean or centroid point until convergence is reached”, thus updating a location of […] and calculating a second mean distance based on the updated location of […] is disclosed, because Baradaran first assigns the data points to a cluster based on their distance from a centroid, then recalculates a new mean or centroid from the assigned data points, thereby updating the centroid location, and subsequently recalculates distances using the new centroid location during the next iteration of the K-means process) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Yoon and Sohn with Baradaran by incorporating Baradaran’s iterative centroid update and distance recalculation technique into Sohn’s K-means clustering of the extracted latent feature vectors. Sohn teaches extracting latent feature vectors and clustering them using K-means based on their similarity to initialized cluster centers. Baradaran teaches updating the cluster centroids based on assigned data points and recalculating distances using the updated centroids until convergence. Therefore, a POSITA would have been motivated to apply Baradaran’s iterative K-means procedure to Sohn’s clustering of latent feature vectors so that the cluster centers are updated based on the assigned features and the distances are recalculated using the updated centroids, thereby refining the cluster assignments and improving the clustering used for anomaly detection (Baradaran, Par. [0271], “a method 703 for improving anomaly detection using injected outliers is depicted. In brief overview, a device may include a set of outliers into a training dataset of data points (706). The device may identify, using a K-means clustering algorithm applied on the training dataset, at least a first cluster of data points (709). The device may determine a center and an outer radius of a region that covers at least a spatial extent of the first cluster of data points (712). The device may determine a first normalcy radius for the first cluster by adjusting the region around the center until a point at which all artificial outliers are excluded from a region defined by the first normalcy radius (715)”) Regarding Claim 11, Yoon combined with Sohn teaches all of the limitations of claim 1 as cited above and Yoon further teaches: An electronic device (Yoon, Pg. 4 – line 5, “ a computing device”, thus an electronic device is disclosed) comprising: a processor (Yoon, Pg. 4 – line 48, “a processor”, thus a processor is disclosed); and a memory (Yoon, Pg. 4 – line 48, “a memory”, thus a memory is disclosed) connected to the processor, wherein the memory stores instructions that can be executed by the processor, and to cause the processor to perform operations that comprise: training a first classifier, that includes an encoder and a decoder using a plurality of training data to obtain a trained first classifier, the plurality of training data being classified into a plurality of first subsets and the first classifier comprising a first neural network (Yoon, Pg. 6 – lines 9-13, “The processor 110 may generate an anomaly sensing model for sensing anomaly of data by learning a network function using the learning data set. The training data set can include a plurality of training data subsets. The plurality of training data subsets may comprise different training data grouped by predetermined criteria”, &Pg. 12 – lines 46 – 52 & Pg. 13 – lines 1 - 3, “In one embodiment of the present disclosure, the network function 200 may include an autoencoder. The auto encoder may be a kind of artificial neural network for outputting output data similar to the input data. The auto encoder may include at least one hidden layer, and an odd number of hidden layers may be disposed between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding) and then expanded symmetrically from the bottleneck layer to the output layer (symmetrical with the input layer). In this case, in the example of FIG. 2, the dimensional reduction layer and the dimensional reconstruction layer are illustrated to be symmetrical, but the present disclosure is not limited thereto, and nodes of the dimensional reduction layer and the dimensional reconstruction layer may or may not be symmetrical.”, thus Yoon discloses training a first classifier/network function using a plurality of training data/learning data set to obtain a trained anomaly sensing model. Yoon further discloses that the training data set includes a plurality of training data subsets grouped by predetermined criteria. Yoon also teaches that the network function may include an autoencoder, which is an artificial neural network having an encoding/dimensional reduction portion and a decoding/dimensional reconstruction portion. Therefore, Yoon discloses a first classifier comprising a first neural network with an encoder and decoder, trained using a plurality of training data classified into a plurality of first subsets.); extracting features from the plurality of training data by computing the plurality of training data with the encoder of the trained first classifier (Yoon, Pg. 5 – lines 6 - 9, “The processor 110 may process neural networks, such as processing input data for learning in deep learning (DN), extracting features from the input data, calculating errors, and weighting neural networks using backpropagation. Can perform calculations for learning.” & Pg. 9 – lines 34 - 37, “Pre-learned network functions in the present disclosure can be learned to reduce and reconstruct the dimension of the training data. The network function of the present disclosure may include an auto encoder capable of dimensional reduction and dimensional reconstruction of input data.”, therefore extracting features from the plurality of training data by computing the plurality of training data with the encoder of the trained first classifier/pre-learned network functions is disclosed); reconstructing the plurality of training data into a plurality of second subsets […] based on the features extracted from the plurality of training data, […] (Yoon, Pg. 10 – lines 7 - 9, “In the input data of the present disclosure, in the production process, the production equipment may perform an operation in which the processor 110 transforms the input data into data different from the input data and restores the data using the anomaly sensing submodel.” &Pg. 10 – lines 11 - 16, “The processor 110 may extract a feature from the input data using the anomaly detection submodel, and restore the input data based on the feature. As described above, since the network function included in the anomaly sensing submodel of the present disclosure may include a network function capable of restoring the input data, the processor 110 may calculate the input data using the anomaly sensing submodel. Can restore the input data.”, therefore reconstructing the plurality of training data into a plurality of second subsets based on the extracted features is disclosed); wherein […] groups training data from two or more different ones of the plurality of first subsets into a same one of the plurality of second subsets based on the extracted features, and a number of the plurality of second subsets is less than a number of the plurality of first subsets(Yoon, Page 6, “For example, there is a recipe change in the process, but the input data generated after the recipe change (that is, the sensor data obtained in the process generated after the recipe change) is determined as a new pattern by the existing anomaly detection submodel. If not (i.e., having no novelty above the threshold), the input data generated before and after the recipe change may be grouped into one subset of training data”, & Page 13, “For example, when the training data generated for 6 months is configured into one subset, the training data generated by two or more recipes (ie, a plurality of normal patterns) may be included in one subset. In this case, one anomaly detection submodel may learn a plurality of normal patterns and may be used for anomaly detection (eg, novelty detection) for the plurality of normal patterns”, thus wherein […] groups training data from two or more different ones of the plurality of first subsets into a same one of the plurality of second subsets based on the extracted features, and a number of the plurality of second subsets is less than a number of the plurality of first subsets is disclosed because Yoon teaches that training data generated before and after a recipe change may be grouped into one subset of training data, and further teaches that training data generated by two or more recipes may be included in one subset, thereby consolidating multiple recipe-based subsets into fewer resulting subsets); training a plurality of second classifiers, that correspond to the plurality of second subsets using the plurality of second subsets to obtain a plurality of trained second classifiers, each of the second classifiers comprising a second neural network (Yoon, Pg. 17 – lines 3 - 6, “The computing device 100 generates a first anomaly sensing submodel that includes the first network function trained with the first training data subset, and then, if there is a change in the recipe, trained with the second training data subset. A second anomaly sensing submodel that includes a second network function may be generated.”, thus Yoon discloses training a plurality of second classifiers/anomaly sensing submodels corresponding to the plurality of second subsets/training data subsets. Yoon teaches generating a first anomaly sensing submodel including a first network function trained with a first training data subset, and generating a second anomaly sensing submodel including a second network function trained with a second training data subset. Because each submodel includes a network function trained using a corresponding training data subset, Yoon discloses obtaining a plurality of trained second classifiers, each comprising a second neural network); and detecting an abnormality in manufacturing of semiconducting equipment from input data based on one or more determinations made by the plurality of trained second classifiers (Yoon, Pg. 3 – lines 9 - 10, “In an alternative embodiment, determining whether anomaly exists in the input data comprises: determining whether anomaly exists in the input data using the second anomaly sensing submodel”, &Pg. 7 – lines 8 - 10, “And a second anomaly sensing submodel comprising a second network function pre-learned with a second subset of learning data composed of learning data generated during the second time interval”, & Page 5, “As a more specific example, work-in-progress information including about 120,000 items per lot acquired in a semiconductor fab, raw processing tool data, equipment interface information, process metrology information information) (e.g., contains more than 1000 items per lot), defect information accessible to yield engineers, operational test information, sort information (including datalogs and bitmaps), The present disclosure is not limited thereto”, thus detecting an abnormality in manufacturing of semiconducting equipment from input data based on one or more determinations made by the plurality of trained second classifiers is disclosed). Yoon does not explicitly disclose clustering the plurality of training data based on the extracted features […] the clustering comprising at least: setting initial centroids for features distributed in a feature space; allocating each of the features to a closest initial centroid among the initial centroids; and calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated, […] the initial centroids based on the calculation of the first mean distance […] the initial centroids, and updating a location of […] and calculating a second mean distance based on the updated location of […]. However, Sohn teaches: reconstructing the plurality of training data into a plurality of second subsets by clustering the plurality of training data based on the extracted features (Sohn, Pg. 3 – lines 20 - 22, “The encoder 102a may extract a latent feature vector z from the input normal data. In this case, normal data of a data set having various classes may be input to the encoder 102a. The decoder 102b may reconstruct normal data based on the latent feature vector output from the encoder 102a.”, &Pg. 3 – lines 27 - 28, “The labeling module 104 may perform clustering on the latent feature vectors z output from the encoder 102a, and may give a pseudo label to each clustered cluster.”, therefore the plurality of training data is reconstructed into a plurality of second subsets(classes) by clustering the plurality of training data based on the extracted features) setting initial centroids for features distributed in a feature space (Sohn, Page – 3, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, & “ Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below”, thus establishing the centers of the initialized clusters i.e., initial centroids for the latent features distributed in the latent feature space) allocating each of the features to a closest initial centroid among the initial centroids (Sohn, Page 3, “Specifically, the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster can be expressed by Equation 1 below”, & Page 6, “the anomaly detection apparatus 100 calculates the similarity between each latent feature vector and the centers of the primary clustered clusters, and learns the second artificial neural network model so that the probability distribution for the calculated similarity matches the preset target probability distribution”, thus Sohn discloses allocating each latent feature vector to a cluster center based on the measured similarity/proximity between the position of that feature and the cluster centers i.e., the K-means assignment of each feature to its nearest centroid) calculating a first mean distance between each of the features and the closest initial centroid to which each of the features are allocated (Sohn, Page – 3, “Specifically, the labeling module 104 may perform initial clustering (primary clustering) on the latent feature vectors z. For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, thus Sohn discloses calculating distance based clustering of extracted features because Sohn teaches performing initial clustering on latent feature vectors using a K-means algorithm. In K-means clustering, feature vectors are assigned to clusters based on proximity to cluster centroids, which requires calculating distances between the feature vectors and the centroids. Therefore, Sohn suggests allocating extracted features to the closest centroid and using distance calculations between the features and the corresponding centroid as part of the clustering process) […] the initial centroids based on the calculation of the first mean distance […] the initial centroids (Sohn, Page 3, “For example, the labeling module 104 may perform initial clustering on the latent feature vectors (z) using a K-means algorithm”, & “the labeling module 104 may measure the similarity between the positions of the latent feature vectors (z) and the centers of the initialized clusters. Here, the similarity (q .sub.ij ) between the i-th latent feature vector (z .sub.i ) and the center (μ .sub.j ) of the j-th cluster”, thus […] the initial centroids based on the calculation of the first mean distance […] the initial centroids is disclosed, because Sohn teaches performing K-means clustering on the latent feature vectors and measuring the similarity between each latent feature vector and the center of an initialized cluster, thereby determining the relationship of each feature to the initial centroids for clustering) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine Yoon’s approach of reconstructing the plurality of training data into a plurality of second subsets based on extracted features with Sohn’s approach of clustering the plurality of training data based on the extracted features to reconstruct the plurality of training data into a plurality of second subsets/classes, thereby improving the accuracy and efficiency of classifying an abnormal situation and anomaly detection (Sohn, Pg. 2 – lines 38 - 41, “2 is a diagram schematically illustrating multi-class classification in anomaly detection technology according to an embodiment of the present invention. Referring to FIG. 2 , the disclosed embodiment is to more accurately classify an abnormal situation by generating a strict decision boundary for each class when a normal sample includes multiple classes”, & Pg. 5 – lines 16 - 19, “According to the disclosed embodiment, even when multiple latent classes are included in the normal data set, by clustering and pseudo-labeling using latent characteristics extracted from normal data, it is more efficient than the single decision boundary-based anomaly detection technique. Anomaly detection can be performed with excellent performance.”) Yoon combined with Sohn does not explicitly disclose updating a location of […] and calculating a second mean distance based on the updated location of […]. However, Baradaran teaches: updating a location of […] and calculating a second mean distance based on the updated location of […] (Baradaran, Par. [0274], “The clustering engine 720 may calculate a new mean or centroid of the data points 731, 731′, and 733 for each cluster or partition. The clustering engine 720 may then repeat and iterate the process of calculating distances and new means or centroids for each cluster until the K means clustering algorithm converges in results”, & Par. [0257], “In some embodiments, the clustering engine 720 may be configured to determine, for each partition, a mean or centroid point of the data points for the respective partition, e.g., via an averaging and/or weighting process on characteristics or attributes of the data points. In some embodiments, the clustering engine 720 may be configured to determine or calculate distances (e.g., Euclidean distance) between the mean or centroid point to all data points of the training data. In some embodiments, the clustering engine 720 may be configured to iterate a process of calculating distances using each (e.g., new) mean or centroid point, partitioning the dataset into clusters, and/or assigning the data points to the respective partition based on the nearest mean or centroid point until convergence is reached”, thus updating a location of […] and calculating a second mean distance based on the updated location of […] is disclosed, because Baradaran first assigns the data points to a cluster based on their distance from a centroid, then recalculates a new mean or centroid from the assigned data points, thereby updating the centroid location, and subsequently recalculates distances using the new centroid location during the next iteration of the K-means process) It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Yoon and Sohn with Baradaran by incorporating Baradaran’s iterative centroid update and distance recalculation technique into Sohn’s K-means clustering of the extracted latent feature vectors. Sohn teaches extracting latent feature vectors and clustering them using K-means based on their similarity to initialized cluster centers. Baradaran teaches updating the cluster centroids based on assigned data points and recalculating distances using the updated centroids until convergence. Therefore, a POSITA would have been motivated to apply Baradaran’s iterative K-means procedure to Sohn’s clustering of latent feature vectors so that the cluster centers are updated based on the assigned features and the distances are recalculated using the updated centroids, thereby refining the cluster assignments and improving the clustering used for anomaly detection (Baradaran, Par. [0271], “a method 703 for improving anomaly detection using injected outliers is depicted. In brief overview, a device may include a set of outliers into a training dataset of data points (706). The device may identify, using a K-means clustering algorithm applied on the training dataset, at least a first cluster of data points (709). The device may determine a center and an outer radius of a region that covers at least a spatial extent of the first cluster of data points (712). The device may determine a first normalcy radius for the first cluster by adjusting the region around the center until a point at which all artificial outliers are excluded from a region defined by the first normalcy radius (715)”) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.T.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Feb 03, 2023
Application Filed
Dec 03, 2025
Non-Final Rejection mailed — §101, §103
Feb 24, 2026
Response Filed
Jun 10, 2026
Final Rejection mailed — §101, §103
Aug 06, 2026
Request for Continued Examination
Aug 09, 2026
Response after Non-Final Action
Sep 08, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month