Prosecution Insights
Last updated: October 02, 2026
Application No. 18/193,781

METHOD AND DEVICE WITH ENSEMBLE MODEL FOR DATA LABELING

Final Rejection §102§103§112
Filed
Mar 31, 2023
Priority
Nov 01, 2022 — RE 10-2022-0143771
Examiner
ROHD, BENJAMIN MATTHEW
Art Unit
2147
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
2 (Final)
20%
Grant Probability
At Risk
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 20% of cases
20%
Career Allowance Rate
1 granted / 5 resolved
-35.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
19 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
22.3%
-17.7% vs TC avg
§103
52.1%
+12.1% vs TC avg
§102
10.0%
-30.0% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 5 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION This office action is in response to amendments filed on 05/18/2026. Claims 1, 3, 5, 7-11, 13, 15-17, and 19-20 have been amended. Claim 6 has been canceled. Claims 1-5 and 7-20 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Rejections Under 35 USC § 112(b): In light of applicant’s amendments to the claims, the previous rejections under 35 USC § 112(b) have been withdrawn. However, a new rejection under 35 USC § 112(b) has been introduced by the amendment to claim 7. Prior Art Rejections: Applicant's arguments regarding the prior art rejections have been fully considered but they are not persuasive. Applicant argues (pg. 12-13) that Sturlaugson’s ensemble model is configured to infer a positive or negative performance status, while the claimed ensemble model is configured to infer probabilities of each of N classes. Examiner respectfully notes that while Sturlaugson’s positive and negative performance status are probabilities of the same performance status, they each represent a different state (i.e. class) of said performance status. Further, while the positive and negative scores are complementary and can be derived from one another, Sturlaugson makes clear that both scores are directly output by the models: “Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score… Primary models 66 also produce a negative classification score…” (Sturlaugson, 0050). Thus, the positive and negative states inferred by Sturlaugson’s models fall within the broadest reasonable interpretation of the claimed N classes. Applicant argues (pg. 13) that neither Sturlaugson nor Dogan teach determining the inference performance values by integrating confidence values with consistency values, as recited in claim 1. Examiner respectfully notes that, as can be seen in the rejection below, this limitation is taught by Ren. Applicant argues (pg. 13) that neither Sturlaugson nor Dogan teach the additional features of claim 11. However, examiner respectfully notes that, as can be seen in the rejection below, each limitation of claim 11 is taught by the combination of Sturlaugson and Dogan. The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 7-9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 7 recites the limitation "the M sets of score data form an MxN first individual score matrix, and M sets of score data form an MxN score matrix to which the MxN first individual score matrix is applied.” Both the first individual score matrix and the score matrix are MxN matrices formed by the M sets of score data. It is unclear whether these are intended to refer to a single matrix or two different matrices, and it is unclear what it would mean to apply the first individual score matrix to the score matrix as the contents of these matrices are identical. For examination purposes, this limitation will be interpreted as referring to a single first individual score matrix, which is formed by the M sets of score data. Claims 8-9 are additionally rejected due to their dependence on the rejected claims for the reasons outlined above. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 19-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Sturlaugson et al. (hereinafter Sturlaugson), U.S. Patent Application Publication US-20180346151-A1 (published 12/06/2018). Regarding Claim 19, Sturlaugson teaches A method comprising: storing constituent neural networks (NNs) of an ensemble model, the constituent NNs each trained to infer a same set of class labels from input data items inputted thereto; (0008: “An ensemble of related machine learning models is applied to the feature data. Each model is characterized by a false positive rate and a false negative rate…” 0061: “Candidate models and primary models 66 may be the result of supervised machine learning and/or guided machine learning… The underlying function may be… a classification algorithm… Examples of classification algorithms include… neural networks.” An ensemble of neural networks is stored and applied to feature data (i.e. input data items) to infer positive and negative labels (i.e. class labels).) wherein each of the NNs is configured to recognize and infer probabilities for each of N classes, and wherein when a data item is inputted to the ensemble model, each constituent NN infers therefrom a corresponding set of probabilities of the respective classes; (0049-0051: “Each of the primary models 66 (also referred to as models and machine learning models) of the ensemble 64 are machine learning algorithms that identify (e.g., classify) the category (one of at least two states) to which a new observation (set of extracted features) belongs… Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66… For example, the positive classification score 76 and the negative classification score 78 may be likelihood estimates of respective positive and negative outcomes…” For extracted feature data (i.e. a data item) inputted to the ensemble model, each of the primary models (i.e. constituent neural networks) infers positive and negative classification likelihood scores (i.e. a corresponding set of probabilities for each of the N classes).) for each constituent NN, providing a respective score set comprising scores of the respective classes, wherein each constituent NN's score set comprises scores that are specific thereto, and wherein each score comprises a measure of inference performance of a corresponding constituent NN with respect to a corresponding class label; (0008: “An ensemble of related machine learning models is applied to the feature data. Each model is characterized by a false positive rate and a false negative rate…” 0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results).” For each constituent model, false positive and false negative rates (i.e. scores measuring inference performance on each class label) are determined.) inputting an input data item to the constituent NNs and based thereon, the constituent NNs generate respective sets of prediction values, each set of prediction values comprising prediction values of the respective classes for the corresponding constituent NN; (0050: “Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66.” The primary models (i.e. constituent neural network models) generate positive and negative scores (i.e. prediction values for each class label) based on extracted feature data of an input data item.) assigning a class, among the classes, to the input data item by applying the score sets of the constituent NNs to the respective sets of prediction values of the constituent NNs. (0066: “The positive weighting function and the negative weighting function may take the form of an inverse exponential function with an exponent proportional to the respective false positive rate (positive weighting function) or the false negative rate (negative weighting function). For example, the weighted positive score (the product of the positive weighting function and the positive classification score of a primary model) may be given by: W p = P e - α F P R where W p is the weighted positive score, P is the positive classification score, F P R is the false positive rate, and α is a constant… The weighted negative score (the product of the negative weighting function and the negative classification score of a primary model) may be given by: W N = N e - β F N R where W N is the weighted negative score, N is the negative classification score, F N R is the false negative rate, and β is a constant.” 0069: “The performance indicator 28 may be determined according to the combined weighted positive score and/or the combined weighted negative score… As a specific method of determining the performance indicator 28, the performance indicator 28 may be determined to be a positive category if the combined weighted positive score is greater than a threshold or a negative category if the combined weighted positive score is less than the same threshold.” The false positive and false negative rates (i.e. score sets) are applied to the positive and negative classification scores (i.e. prediction values) to generate weighted positive and negative scores, which are then combined and compared to a threshold to assign a performance indicator (i.e. class label) indicating a positive or negative category.) Regarding Claim 20, Sturlaugson teaches The method of claim 19, as shown above. Sturlaugson also teaches wherein the scores are generated based on measures of model confidence and/or model consistency of the constituent NNs with respect to the class labels. (0008: “An ensemble of related machine learning models is applied to the feature data. Each model is characterized by a false positive rate and a false negative rate…” 0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results).” False positive and false negative rates are scores generated based on model classification consistency for each class.) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 and 3-5 are rejected under 35 U.S.C. 103 as being unpatentable over Sturlaugson in view of Ren et al. (hereinafter Ren), “Multi-classifier ensemble based on dynamic weights” (published 12/30/2017) and Dogan et al. (hereinafter Dogan), “A Weighted Majority Voting Ensemble Approach for Classification” (published 11/21/2019). Regarding Claim 1, Sturlaugson teaches A labeling method performed by one or more processors and comprising: (0026: “As illustrated in FIG. 1, a predictive maintenance system 10 includes a computerized system 200 (as further discussed with respect to FIG. 7). The predictive maintenance system 10 may be programmed to perform, and/or may store instructions to perform, the methods described herein.” 0097: “The computerized system 200 includes a processing unit 202 operatively coupled to a computer-readable memory 206 by a communications infrastructure 210. The processing unit 202 may include one or more computer processors 204…”) N classes for which the neural network models are trained to recognize and infer probabilities of; (0049-0051: “Each of the primary models 66 (also referred to as models and machine learning models) of the ensemble 64 are machine learning algorithms that identify (e.g., classify) the category (one of at least two states) to which a new observation (set of extracted features) belongs… Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66… For example, the positive classification score 76 and the negative classification score 78 may be likelihood estimates of respective positive and negative outcomes…” 0061: “Candidate models and primary models 66 may be the result of supervised machine learning and/or guided machine learning… The underlying function may be… a classification algorithm… Examples of classification algorithms include… neural networks.” Neural network models of an ensemble are trained to recognize and infer likelihood estimates (i.e. probabilities) of positive and negative outcomes (i.e. N classes).) applying the M neural network models to second validation data to determine M sets of consistency values of the M respective neural network models with respect to the N classes, wherein each set of consistency values comprises N consistency values of the N respective classes for the corresponding neural network model, and wherein MxN consistency values are determined; (0008: “An ensemble of related machine learning models is applied to the feature data. Each model is characterized by a false positive rate and a false negative rate…” 0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results).” An ensemble of M neural network models is applied to training or verification inputs (i.e. validation data) to determine a false positive rate and a false negative rate for each model. Positive and negative outcomes represent N classes, and thus the false positive and false negative rates represent N consistency values with respect to the N classes. Since each of the M models is characterized by a false positive and false negative rate, there are a total of MxN consistency values determined.) determining M sets of inference performance values of the M respective neural network models by [integrating the M confidence values with] the M sets of consistency values, wherein each set of inference performance values comprises N consistency values of the N respective classes for the corresponding neural network model, and wherein MxN inference performance values are determined; (See the portions of 0008, 0055, and 0061 cited above. For each of the M neural networks, false positive and false negative rates (i.e. a set of inference performance values comprising N consistency values of the N classes) are determined, for a total of M sets of N inference performance values, or MxN.) based on the M sets of inference performance values, determining M sets of weights for the M respective neural network models, wherein each set of weights comprises N weights of the respective N classes for the corresponding neural network model, wherein the M sets of weights data are not weights of nodes of the neural network models; (0066: “The positive weighting function and the negative weighting function may take the form of an inverse exponential function with an exponent proportional to the respective false positive rate (positive weighting function) or the false negative rate (negative weighting function). For example, the weighted positive score (the product of the positive weighting function and the positive classification score of a primary model) may be given by: W p = P e - α F P R where W p is the weighted positive score, P is the positive classification score, F P R is the false positive rate, and α is a constant… The weighted negative score (the product of the negative weighting function and the negative classification score of a primary model) may be given by: W N = N e - β F N R where W N is the weighted negative score, N is the negative classification score, F N R is the false negative rate, and β is a constant.” For each of the M neural network models, a positive weight e - α F P R is determined based on the false positive rate, and a negative weight e - β F N R is determined based on the false negative rate (i.e. M sets of N weights are determined for the N classes based on the M sets of inference performance features).) generating M sets of classification result data by performing a classification inference operation on labeling target inputs by the neural network models, each set of classification result data comprising N respective predictions for the N respective classes; (0050: “Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66.” Each of the M primary models (i.e. neural network models) perform a classification inference operation to transform extracted feature data (i.e. labeling target inputs) into positive and negative classification scores (i.e. a set of classification result data comprising N predictions for the N classes).) determining M sets of score data, each set of score data representing N confidences for the respective N classes for the labeling target inputs, wherein the M sets of score data are determined by applying the M sets of weights to the respectively corresponding M sets of classification result data; and (See the portion of 0066 cited above. For each of the M neural network models, weighted positive and negative scores W p and W N (i.e. a set of score data representing N confidences for the N classes) are determined by the product of the positive and negative weighting functions and the positive and negative classification scores (i.e. by applying the M sets of weight data to the M sets of classification result data).) Sturlaugson does not appear to explicitly disclose the remaining features of claim 1. However, Ren teaches applying M neural network models comprised in an ensemble model to first validation data to determine M confidence values of the M respective neural network models with respect to N classes for which the neural network models are trained to recognize and infer probabilities of; (Pg. 21085. Section 1: “A dynamic weighted multi-classifier ensemble method that defines reliability and credibility is proposed… The reliability of the classifier is defined on the basis of its recognition capability, which is determined by the prior knowledge gained during the training process… The credibility of the classifier is obtained by calculating the posterior probability distribution, which is used to determine the separability feature of the classifier.” For each of the M classifier models of a multi-classifier ensemble, a reliability (i.e. confidence value) is determined by its class recognition capability (i.e. confidence with respect to the N classes) based on knowledge gained in the training process (i.e. based on applying the model to validation data).) Determining the inference performance values by integrating the M confidence values with the M sets of consistency values (Pg. 21083, Abstract: “The algorithm defines decision credibility to describe the real-time importance of the classifier to the current target, combines this credibility with the reliability calculated by the classifier on the training data set and dynamically assigns the fusion weight to the classifier.” Fusion weights (i.e. inference performance values) are determined by combining reliability and credibility (i.e. integrating confidence values and consistency values).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson and Ren. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Ren teaches a model ensemble which calculates weights based on an integration of two model metrics: reliability and credibility. One of ordinary skill would have motivation to combine Sturlaugson and Ren because “[c]ompared with single classifier and current popular fusion methods, our method can more effectively reduce the effect of unreliable instance information in the training phase of the classifier and can therefore fuse the decisions in classification efficiently and improve the overall performance of the integrated method” (Ren, pg. 21085, section 1). Sturlaugson and Ren do not appear to explicitly disclose measuring classification accuracy of the classification operation for the labeling target inputs based on the M sets of score data. However, Dogan teaches measuring classification accuracy of the classification operation for the labeling target inputs based on the M sets of score data. (Pg. 369, section IV.B: “In this study, classification accuracy was used as evaluation metric to compare the performances of the methods.”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson, Ren, and Dogan. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Ren teaches a model ensemble which calculates weights based on an integration of two model metrics: reliability and credibility. Dogan teaches a model ensemble with weighted voting based on base model performance, including measuring ensemble classification accuracy. One of ordinary skill would have motivation to combine Sturlaugson, Ren, and Dogan in order to evaluate the performance of the ensemble model. Regarding Claim 3, Sturlaugson, Ren, and Dogan teach The labeling method of claim 1, as shown above. Sturlaugson also teaches wherein the determining of the M sets of weight data comprises: determining […] consistency data indicating classification consistency for each class of the neural network models, based on a comparison result between labels of validation inputs and validation result data; (0008: “An ensemble of related machine learning models is applied to the feature data. Each model is characterized by a false positive rate and a false negative rate…” 0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results).” False positive and false negative rates (i.e. consistency data indicating classification consistency for each class) are determined based on performance on verification inputs with known results (i.e. based on a comparison between labels of validation inputs and validation result data).) Ren teaches determining model confidence data indicating model confidence of the neural network models and consistency data indicating classification consistency for each class of the neural network models, based on a comparison result between labels of validation inputs and validation result data; and (Pg. 21085. Section 1: “The reliability of the classifier is defined on the basis of its recognition capability, which is determined by the prior knowledge gained during the training process… The credibility of the classifier is obtained by calculating the posterior probability distribution, which is used to determine the separability feature of the classifier.” For each classifier model, reliability (i.e. confidence data) is determined based on knowledge gained in the training process (i.e. based on validation data), and credibility (i.e. consistency data) is determined based on class separability (i.e. indicating classification consistency for each class).) determining the M sets of weight data by integrating the model confidence data and the consistency data. (Pg. 21083, Abstract: “The algorithm defines decision credibility to describe the real-time importance of the classifier to the current target, combines this credibility with the reliability calculated by the classifier on the training data set and dynamically assigns the fusion weight to the classifier.” Fusion weights for each of the M classifiers (i.e. the M sets of weight data) are determined by combining reliability and credibility (i.e. integrating confidence data and consistency data).) Regarding Claim 4, Sturlaugson, Ren, and Dogan teach The labeling method of claim 1, as shown above. Sturlaugson also teaches wherein the measuring of the classification accuracy comprises: determining representative score values based on the labeling target inputs based on the score data; and (0068: “The weighted positive scores of the primary models 66 may be combined by summing (and/or averaging, etc.) the individual weighted positive scores for all (or a subset, e.g., a filtered subset) of the primary models 66… The weighted negative scores of the primary models 66 may be combined in a manner analogous to the weighted positive scores.” Combined weighted positive and negative scores (i.e. representative score values) are determined based on the weighted positive and negative scores (i.e. score data).) classifying each of the labeling target inputs into a first group or a second group based on the representative score values. (0069: “The performance indicator 28 may be determined according to the combined weighted positive score and/or the combined weighted negative score… As a specific method of determining the performance indicator 28, the performance indicator 28 may be determined to be a positive category if the combined weighted positive score is greater than a threshold or a negative category if the combined weighted positive score is less than the same threshold.” Each input is assigned a performance indicator representing a positive or negative classification (i.e. classification into a first or second group) based on its combined weighted positive and negative scores (i.e. representative score values).) Regarding Claim 5, Sturlaugson, Ren, and Dogan teach The labeling method of claim 4, as shown above. Sturlaugson also teaches wherein the M sets of score data comprises individual score data of each of the labeling target inputs, and (0008-0009: “Feature data is extracted from the flight data. The feature data relates to performance of one or more components of the aircraft. An ensemble of related machine learning models is applied to the feature data… Each model produces a positive score and a complementary negative score related to performance of the selected component.” The positive and negative scores (i.e. the M sets of score data) are produced individually based on the feature data associated with each selected component (i.e. labeling target input).) Dogan teaches the determining of the representative score values comprises determining a maximum value of each piece of individual score data of the labeling target inputs of the score data to be the representative score value of each of the labeling target inputs. (Pg. 366, section I: “Each classifier is associated with a coefficient (weight), usually proportional to its classification performance on a validation set. The final decision is made by summing up all weighted votes and by selecting the class with the highest aggregate.” For an input, the final decision (i.e. representative score value) of the classifier is the class with the highest aggregate (i.e. the maximum value of the individual score data).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson, Ren, and Dogan. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Ren teaches a model ensemble which calculates weights based on an integration of two model metrics: reliability and credibility. Dogan teaches a voting procedure for a model ensemble with weighted voting based on base model performance. One of ordinary skill would have motivation to combine Sturlaugson, Ren, and Dogan in order to extend the binary classification capabilities of Sturlaugson’s ensemble to work with an arbitrary number of classes by providing a weighted voting procedure which “generally produce[s] better classification results than both simple majority voting ensemble (SMVE) approach and individual standard classification algorithms in terms of accuracy” (Dogan, pg. 366, section I). Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Sturlaugson in view of Ren and Dogan, and further in view of Zhang et al. (hereinafter Zhang), “MEMO: Test Time Robustness via Adaptation and Augmentation” (published 10/18/2021). Regarding Claim 2, Sturlaugson, Ren, and Dogan teach The labeling method of claim 1, as shown above. Sturlaugson also teaches further comprising: generating validation result data by performing a classification operation on validation inputs by the neural network models; (0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results). The model outcomes may be categorized as true positive outcomes (the predicted value of the model is positive and the known result is also positive), true negative outcomes (the predicted value is negative and the known result is negative), false positive outcomes (the predicted value is positive and the known result is negative), or false negative outcomes (the predicted value is negative and the known result is positive).” Each model performs classification on verification inputs (i.e. validation inputs) to generate predicted values (i.e. validation result data).) Sturlaugson, Ren, and Dogan do not appear to explicitly disclose the remaining features of claim 2. However, Zhang teaches generating first partial data of the validation result data by performing a first classification operation on the validation inputs by the neural network models; (Pg. 4, section 3.1: “After this step, we use the adapted model f θ ' to predict on the original test input x (line 4).” The model makes a prediction (i.e. generates first partial validation result data) based on the test input (i.e. validation input).) generating additional validation inputs by transforming the validation inputs; and (pg. 4, section 3.1: “Given a test point x and set of augmentation functions A , we sample B augmentations from A and apply them to x in order to produce a batch of augmented data x ~ 1 ,   .   .   .   , x ~ B .” The test point (i.e. validation input) is augmented (i.e. transformed) to produce a batch of augmented data (i.e. additional validation inputs).) generating second partial data of the validation result data by performing a second classification operation on the additional validation inputs by the neural network models. (Pg. 4, section 3.1: “The model’s average, or marginal, output distribution with respect to the augmented points is given…” The model makes predictions (i.e. generates second partial validation result data) based on the augmented test input (i.e. additional validation inputs).) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson, Ren, Dogan, and Zhang. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Ren teaches a model ensemble which calculates weights based on an integration of two model metrics: reliability and credibility. Dogan teaches a voting procedure for a model ensemble with weighted voting based on base model performance. Zhang teaches test-time data augmentation for model adaptation and robustification. One of ordinary skill would have motivation to combine Sturlaugson, Ren, Dogan, and Zhang because Zhang’s ‘MEMO’ method “does not require access or changes to the model training procedure and is thus broadly applicable for a wide range of model architectures pretrained in a number of different ways. Furthermore, MEMO adapts at test time using single test inputs, thus it does not assume access to multiple test points as in several recent methods for test time adaptation [40, 46]. On a range of CIFAR-10 and ImageNet distribution shift benchmarks, and for ResNet, vision transformer, and, to an extent, ResNext models, MEMO consistently improves performance at test time and achieves several new state-of-the-art results for these models in the single test point setting” (Zhang, pg. 10, section 5). Claims 7-9 are rejected under 35 U.S.C. 103 as being unpatentable over Sturlaugson in view of Ren and Dogan, and further in view of Saha et al. (hereinafter Saha), “Combining multiple classifiers using vote based classifier ensemble technique for named entity recognition” (published 07/20/2012). Regarding Claim 7, Sturlaugson, Ren, and Dogan, teach The labeling method of claim 1, as shown above. Sturlaugson also teaches wherein each set of score data comprises N scores of the corresponding neural network model for the N respective classes, and (0066: “The positive weighting function and the negative weighting function may take the form of an inverse exponential function with an exponent proportional to the respective false positive rate (positive weighting function) or the false negative rate (negative weighting function). For example, the weighted positive score (the product of the positive weighting function and the positive classification score of a primary model) may be given by: W p = P e - α F P R where W p is the weighted positive score, P is the positive classification score, F P R is the false positive rate, and α is a constant… The weighted negative score (the product of the negative weighting function and the negative classification score of a primary model) may be given by: W N = N e - β F N R where W N is the weighted negative score, N is the negative classification score, F N R is the false negative rate, and β is a constant.” For each of the M neural network models, weighted positive and negative scores W p and W N (i.e. a set of score data representing N scores for the N classes) are determined.) Sturlaugson, Ren, and Dogan do not appear to explicitly disclose the remaining features of claim 7. However, Saha teaches wherein the M sets of score data form an MxN first individual score matrix, the M sets of score data form an MxN score matrix to which the MxN first individual score matrix is applied. (Pg. 19, section 3.1: “Let, the N number of available classifiers be denoted by C_1,…,C_N and A={C_i:i=1;N}. Suppose, there are M number of output classes. The classifier ensemble problem is then stated as follows: Find the combination of votes V per classifier C_i which will optimize a function F(V). Here, V can be either a Boolean array (binary vote based ensemble) of size N×M or a real array (real/weighted vote based ensemble) of size N×M… In the case of real array: V(i,j) denotes the weight of vote of the i^th classifier for the j^th class.” Vote array V (i.e. a first individual score matrix) is an N×M weight matrix, where N is the number of classifiers (i.e. neural network models) and M is the number of classes.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson, Ren, Dogan, and Saha. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Ren teaches a model ensemble which calculates weights based on an integration of two model metrics: reliability and credibility. Dogan teaches a voting procedure for a model ensemble with weighted voting based on base model performance. Saha teaches a model ensemble with per-class per-model confidence weighting, including storing the vote values in an array. One of ordinary skill would have motivation to combine Sturlaugson, Ren, Dogan, and Saha in order to efficiently store and access weight and score values. Regarding Claim 8, Sturlaugson, Ren, Dogan, and Saha teach The labeling method of claim 7, as shown above. Sturlaugson also teaches wherein the measuring of the classification accuracy comprises: determining a first representative score value for the first labeling target input from the MxN first individual score matrix; and (0068: “The weighted positive scores of the primary models 66 may be combined by summing (and/or averaging, etc.) the individual weighted positive scores for all (or a subset, e.g., a filtered subset) of the primary models 66.” A combined weighted positive score (i.e. a first representative score value) is determined for the input (i.e. first labeling target input) based on the weighted positive scores (i.e. first individual score matrix).) classifying the first labeling target input into a certain label group or a review group based on the first representative score value. (0069: “The performance indicator 28 may be determined according to the combined weighted positive score… the positive category may be assigned to values greater than a first threshold and the negative category may be assigned to values less than a second threshold. Values between the first and second threshold may be assigned an unclassified or undetermined category.” Based on the combined weighted positive score (i.e. first representative score value), the input is classified as positive or negative (i.e. a certain label group), or as undetermined (i.e. a review group).) Regarding Claim 9, Sturlaugson, Ren, Dogan, and Saha teach The labeling method of claim 8, as shown above. Sturlaugson also teaches wherein the determining of the first representative score value comprises: determining a 1xN first individual score vector by integrating elements of the first individual score matrix for each class; and (0068: “The weighted positive scores of the primary models 66 may be combined by summing (and/or averaging, etc.) the individual weighted positive scores for all (or a subset, e.g., a filtered subset) of the primary models 66… The weighted negative scores of the primary models 66 may be combined in a manner analogous to the weighted positive scores.” The weighted positive and negative scores are combined (i.e. elements of the score matrix are integrated for each class), resulting in a combined weighted positive and negative score (i.e. a 1*n first individual score vector).) Dogan teaches determining a maximum value of the elements of the first individual score vector to be the first representative score value. (Pg. 366, section I: “Each classifier is associated with a coefficient (weight), usually proportional to its classification performance on a validation set. The final decision is made by summing up all weighted votes and by selecting the class with the highest aggregate.” For an input, the final decision (i.e. first representative score value) of the classifier is the class with the highest aggregate (i.e. the maximum value of the first individual score vector).) Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Sturlaugson in view of Ren and Dogan, and further in view of Saqlain et al. (hereinafter Saqlain), “A Voting Ensemble Classifier for Wafer Map Defect Patterns Identification in Semiconductor Manufacturing” (published 03/11/2019). Regarding Claim 10, Sturlaugson, Ren, and Dogan teach The labeling method of claim 1, as shown above. Sturlaugson also teaches further comprising generating validation result data by performing a classification operation on validation inputs by the neural network models, and (0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results). The model outcomes may be categorized as true positive outcomes (the predicted value of the model is positive and the known result is also positive), true negative outcomes (the predicted value is negative and the known result is negative), false positive outcomes (the predicted value is positive and the known result is negative), or false negative outcomes (the predicted value is negative and the known result is positive).” Each model performs classification on verification inputs (i.e. validation inputs) to generate predicted values (i.e. validation result data).) Sturlaugson, Ren, and Dogan do not appear to explicitly disclose the remaining features of claim 10. However, Saqlain teaches wherein the validation inputs and the labeling target inputs correspond to semiconductor images based on a semiconductor manufacturing process, and (Pg. 171-172, section I: “A wafer map (WM) is a collection of visual data about the physical parameter that are collected from semiconductor wafers… this paper proposes a voting ensemble classifier with multi-types features to identify wafer map defect patterns in semiconductor manufacturing.” Pg. 179, section V.A: “We used a real-world dataset WM-811K collected from 46,293 lots and consist of 811,457 original wafer images.”) the classes of the neural network models correspond to types of manufacturing defects based on the semiconductor manufacturing process. (Pg. 174, section III.B: “The labeled dataset also has two major classes such as pattern class and no-pattern class. The no-pattern class has no specific defect pattern of WM and labeled as None. It contains 147,431 wafer entities (85.2%) of the whole labeled data set. In addition, pattern class has only 25,519 wafer entities (14.8%) and consists of eight actual defect classes, labeled as Center, Donut, Edge-local, Edge-ring, Local, Random, Scratch, and Near-full.”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson, Ren, Dogan, and Saqlain. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Ren teaches a model ensemble which calculates weights based on an integration of two model metrics: reliability and credibility. Dogan teaches a voting procedure for a model ensemble with weighted voting based on base model performance. Saqlain teaches an ensemble model for detecting and classifying defects in semiconductor wafer manufacturing. One of ordinary skill would have motivation to combine Sturlaugson, Ren, Dogan, and Saqlain because “the research methodologies about WM defects detection and recognition of spatial patterns are in high demand. These methods can be further used for early prevention of defects by diagnosing their root causes and to enhance the product quality by improving the reliability of the manufacturing system” (Saqlain, pg. 171, section I). Sturlaugson and Dogan teach improvements to ensemble learning, and Saqlain shows that ensemble learning is well-suited for wafer defect detection: “the SVE [soft voting ensemble] classifier outperformed all individual classification algorithms as well as previously proposed wafer map failure pattern recognition (WMFPR) [4] method and CNN model” (Saqlain, pg. 172, section I). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Sturlaugson in view of Dogan. Regarding Claim 11, Sturlaugson teaches A labeling device comprising: one or more processors; and a memory storing instructions configured to, when executed by the one or more processors, cause the one or more processors to: (0026: “As illustrated in FIG. 1, a predictive maintenance system 10 includes a computerized system 200 (as further discussed with respect to FIG. 7). The predictive maintenance system 10 may be programmed to perform, and/or may store instructions to perform, the methods described herein.” 0097: “The computerized system 200 includes a processing unit 202 operatively coupled to a computer-readable memory 206 by a communications infrastructure 210. The processing unit 202 may include one or more computer processors 204…”) apply M neural network models comprised in an ensemble model to validation data to generate M validation result data of the M respective neural network models with respect to N classes for which the neural network models are trained to recognize and infer probabilities of, wherein each validation result data comprises N validation values of the N respective classes for the corresponding neural network model, the validation data comprising validation inputs inputted to the M neural network models (0049-0051: “Each of the primary models 66 (also referred to as models and machine learning models) of the ensemble 64 are machine learning algorithms that identify (e.g., classify) the category (one of at least two states) to which a new observation (set of extracted features) belongs… Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66… For example, the positive classification score 76 and the negative classification score 78 may be likelihood estimates of respective positive and negative outcomes…” 0061: “Candidate models and primary models 66 may be the result of supervised machine learning and/or guided machine learning… The underlying function may be… a classification algorithm… Examples of classification algorithms include… neural networks.” 0008: “An ensemble of related machine learning models is applied to the feature data. Each model is characterized by a false positive rate and a false negative rate…” 0055: “Each primary model 66 may be characterized by performance with training or verification inputs (e.g., using inputs with known results).” An ensemble of neural network models is trained to recognize and infer likelihood estimates (i.e. probabilities) of positive and negative outcomes (i.e. N classes). Each of the M neural network models of the ensemble is applied to training or verification inputs with known results (i.e. validation data comprising validation inputs) to generate positive and negative classification scores (i.e. validation result data of the M neural network models comprising N validation values of the N classes).) determine M sets of weight data indicating weights of the respective M neural network models, based on a comparison result between the M sets of validation result data and labels of the validation inputs, wherein the M sets of weight data are not weights of nodes of the neural network models; (0066: “The positive weighting function and the negative weighting function may take the form of an inverse exponential function with an exponent proportional to the respective false positive rate (positive weighting function) or the false negative rate (negative weighting function). For example, the weighted positive score (the product of the positive weighting function and the positive classification score of a primary model) may be given by: W p = P e - α F P R where W p is the weighted positive score, P is the positive classification score, F P R is the false positive rate, and α is a constant… The weighted negative score (the product of the negative weighting function and the negative classification score of a primary model) may be given by: W N = N e - β F N R where W N is the weighted negative score, N is the negative classification score, F N R is the false negative rate, and β is a constant.” For each of the M neural network models, a positive weight e - α F P R is determined based on the false positive rate, and a negative weight e - β F N R is determined based on the false negative rate (i.e. M sets of weight data are determined for the M neural networks based on a comparison between the validation result data and the labels of the validation inputs).) generate M sets of classification result data of the M respective neural network models by performing a classification operation for the N classes on labeling target inputs by the M neural network models, wherein each set of classification result data comprises N classification values of the N respective classes for the corresponding neural network model, wherein MxN classification values are determined, and wherein each of the neural network models is configured to recognize and infer probabilities for each of the N classes; (0050: “Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66.” Each of the M primary models (i.e. neural network models) perform a classification inference operation to transform extracted feature data (i.e. labeling target inputs) into positive and negative classification scores (i.e. a set of classification result data comprising N classification values for the N classes). Since each of the M models generates a positive and negative classification score, there are a total of MxN classification values determined.) determine M sets of score data representing confidences of the M respective neural network models for the N classes for the labeling target inputs by applying the M sets of weight data to the respective M sets of classification result data, and (See the portion of 0066 cited above. For each of the M neural network models, weighted positive and negative scores W p and W N (i.e. a set of score data representing confidences for the N classes) are determined by the product of the positive and negative weighting functions and the positive and negative classification scores (i.e. by applying the M sets of weight data to the M sets of classification result data).) Sturlaugson does not appear to explicitly disclose measure classification accuracy of the classification operation for the labeling target inputs, based on the M sets of score data. However, Dogan teaches measure classification accuracy of the classification operation for the labeling target inputs, based on the M sets of score data. (Pg. 369, section IV.B: “In this study, classification accuracy was used as evaluation metric to compare the performances of the methods.”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson and Dogan. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Dogan teaches a model ensemble with weighted voting based on base model performance, including measuring ensemble classification accuracy. One of ordinary skill would have motivation to combine Sturlaugson and Dogan in order to evaluate the performance of the ensemble model. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Sturlaugson in view of Dogan, and further in view of Saha. Regarding Claim 16, Sturlaugson and Dogan teach The labeling device of claim 11, as shown above. Sturlaugson also teaches wherein each set of weight data comprises N weights of the corresponding neural network model for the N respective classes, wherein the M sets of weight data form an MxN weight [matrix], (0066: “The positive weighting function and the negative weighting function may take the form of an inverse exponential function with an exponent proportional to the respective false positive rate (positive weighting function) or the false negative rate (negative weighting function). For example, the weighted positive score (the product of the positive weighting function and the positive classification score of a primary model) may be given by: W p = P e - α F P R where W p is the weighted positive score, P is the positive classification score, F P R is the false positive rate, and α is a constant… The weighted negative score (the product of the negative weighting function and the negative classification score of a primary model) may be given by: W N = N e - β F N R where W N is the weighted negative score, N is the negative classification score, F N R is the false negative rate, and β is a constant.” For each of the M neural network models, a positive weight e - α F P R is determined based on the false positive rate, and a negative weight e - β F N R is determined based on the false negative rate (i.e. M sets of N weights are determined for the N classes, for a total of MxN weights).) wherein each set of the M sets of classification result data comprises N classification values for the N respective classes, wherein the M sets of classification result data form an MxN first classification result [matrix], and (0050: “Primary models 66 are selected and/or configured to transform extracted feature data 26 into an estimate of performance status by producing a positive classification score 76. The positive classification score 76, which may be referred to as the positive score, is a numeric value that relates to the positive outcome or result of the primary model 66. Primary models 66 also produce a negative classification score 78 that is complementary to the positive classification score 76. The negative classification score 78, which may be referred to as a negative score, is a numeric value that relates to the negative outcome or result of the primary model 66.” For each of the M neural network models a positive and a negative classification score are determined (i.e. M sets of N classification values for the N classes, for a total of MxN classification result values).) wherein the instructions are further configured to cause the one or more processors to: determine the M sets of score data by applying the MxN weight [matrix] to the MxN first classification result [matrix]. (See the portion of 0066 cited above. For each of the M neural network models, weighted positive and negative scores W p and W N (i.e. a set of score data) are determined by the product of the positive and negative weighting functions and the positive and negative classification scores (i.e. by applying the MxN weights to the MxN classification result values).) Sturlaugson and Dogan do not appear to explicitly disclose representing the MxN weights and the MxN classification values as a matrix. However, Saha teaches representing the MxN weights and the MxN classification values as a matrix. (Pg. 19, section 3.1: “Let, the N number of available classifiers be denoted by C_1,…,C_N and A={C_i:i=1;N}. Suppose, there are M number of output classes. The classifier ensemble problem is then stated as follows: Find the combination of votes V per classifier C_i which will optimize a function F(V). Here, V can be either a Boolean array (binary vote based ensemble) of size N×M or a real array (real/weighted vote based ensemble) of size N×M… In the case of real array: V(i,j) denotes the weight of vote of the i^th classifier for the j^th class.” Vote array V is an N×M matrix, where N is the number of classifiers (i.e. neural network models) and M is the number of classes.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Sturlaugson, Dogan, and Saha. Sturlaugson teaches a model ensemble with per-class per-model confidence weighting based on base model performance. Dogan teaches a voting procedure for a model ensemble with weighted voting based on base model performance. Saha teaches a model ensemble with per-class per-model confidence weighting, including storing the vote values in an array. One of ordinary skill would have motivation to combine Sturlaugson, Dogan, and Saha in order to efficiently store and access weight and score values. Claims 12-15 and 17-18 are device claims containing substantially the same elements as method claims 2-5, 8, and 10. Sturlaugson, Dogan, Zhang, Ren, Saha, and Saqlain teach the elements of claims 2-5, 8, and 10, as shown above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN M ROHD whose telephone number is (571)272-6445. The examiner can normally be reached Mon-Thurs 8:00-6:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /B.M.R./Examiner, Art Unit 2147 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Mar 31, 2023
Application Filed
Feb 18, 2026
Non-Final Rejection mailed — §102, §103, §112
May 18, 2026
Response Filed
May 18, 2026
Applicant Interview (Telephonic)
May 18, 2026
Examiner Interview Summary
Aug 20, 2026
Final Rejection mailed — §102, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
20%
Grant Probability
99%
With Interview (+100.0%)
4y 2m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 5 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month