Prosecution Insights
Last updated: September 18, 2026
Application No. 18/173,661

Apparatus for Training and Method Thereof

Final Rejection §101§103
Filed
Feb 23, 2023
Priority
Mar 07, 2022 — RE 10-2022-0028825 +1 more
Examiner
BENOURAIDA, AMINA MORENO
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
Hyperconnect LLC
OA Round
2 (Final)
0%
Grant Probability
At Risk
3-4
OA Rounds
9m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 4 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
11 currently pending
Career history
23
Total Applications
across all art units

Statute-Specific Performance

§101
22.4%
-17.6% vs TC avg
§103
63.5%
+23.5% vs TC avg
§102
9.4%
-30.6% vs TC avg
§112
4.7%
-35.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d) based on an application filed in REPUBLIC OF KOREA on March 3, 2022 and on August 23, 2022. The certified copy has been filed in parent Application No. 18/173,661, filed on February 23, 2023. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Response to Amendment The amendment filed on April 6th, 2026, has been entered and Claim(s) 1-20 are pending. Claim(s) 2, 4-5, 11-18 have been withdrawn from consideration. Response to Arguments Applicant's arguments filed on April 6th, 2026, have been fully considered but they are not persuasive. Applicant’s arguments starting on page 7 regarding Claim 2, para 2. This is not persuasive. In the Office Action claim 2 was rejected as further modifying the abstract idea of claim 1 and therefore inherited the analysis applicable to claim 1. The applicant has amended claim 1 to now recite the limitations of claim 2. However, the claim limitation acquiring loss information based on the target training dynamics information and the predictive training dynamics information merely recites a mathematical concept, as it involves mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP 2106.04(a)(2), (I)). And the amended claim 1 recites collecting information, analyzing information, generating results, and using the results to train a model. The claim does not integrate the judicial exception into a practical application therefore, the rejection under 101 is maintained. Please refer to the updated rejection below. Applicant’s arguments starting on page 7 regarding Claim 2, para 4. Applicant argues and points to the specification and asserts claim 1 reflects an improvement in machine learning training by reducing computation. This is not persuasive. According to the MPEP 2106.05(a), states that "if it is asserted that the invention improves upon conventional functioning of a computer, or upon conventional technology or technological processes...the claim must include the components or steps of the invention that provide the improvement described in the specification." Although the specification discusses those improvements the claim language of claim 1 does not recite an improvement, rather the claim recites collecting information, analyzing information, generating results, and using the results to train a model. Applicant’s arguments starting on page 8. Applicant argues that Zhang and Yoo do not disclose or suggest the amended features. This is not persuasive. The Office Action sets forth the reliance of Zhang in view of Yoo and not individually to teach. Zhang teaches acquiring classification information and target training dynamics information within active learning framework. Yoo teaches training dynamic prediction model and the loss information. Please see updated rejection below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1, 3, 6-10, 19-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., an abstract idea) without significantly more. Regarding claim 1 and analogous claim 19, 20: Step 1 (whether a claim is to a statutory category): Yes, the claim is within the four statutory categories (a process, machine, manufacture or composition of matter). Claim 1 recites a method, therefore, falls within a process category. Claim 19 recites a non-transitory computer-readable medium, therefore, falls within a manufacture category. Claim 20 recites an apparatus, therefore, falls within a machine category. Step 2A Prong 1 (whether a claim is directed to a judicial exception): Yes, “acquiring, based on a classification model, classification information on training data included in a first dataset;” describes a mental process (observation, evaluation) wherein receiving a dataset and its labels recites concepts that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). And “acquiring, based on the classification information and previous classification information acquired based on the classification model in one or more previous epochs, target training dynamics information;” describes a mental process (observation, evaluation, judgement) wherein receiving information and past model information to compare and judge a target value based on previous data recites concepts that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). And “acquiring, based on the training dynamics prediction model, predictive training dynamics information on the training data; and” describes a mental process (observation, evaluation) wherein receiving information about data recites concepts that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). And, wherein, the predictive training dynamics information includes a result of calculating probability for each of a plurality of classes that the data belongs to that class” recites a mathematical concept, as it involves mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP 2106.04(a)(2), (I)). And, “acquiring loss information based on the target training dynamics information and the predictive training dynamics information” describes a mental process (observation) wherein receiving data information recites concepts that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Step 2A Prong 2 (evaluate whether the claim recites additional elements that integrate the exception into a practical application): No, “wherein the first dataset includes one or more data pre-labeled with a class.” the claim does not recite additional elements that integrate the judicial exception into a practical application with the words "apply it" (or an equivalent), such as mere (i.e., selecting a particular data source or type of data to be manipulated) to implement an abstract idea on a computer (see MPEP 2106.05(f)). And, “training the classification model based on the classification information and the pre-labeled class” the claim does not recite additional elements that integrate the judicial exception into a practical application with the words "apply it" (or an equivalent), such as mere (i.e., selecting a particular data source or type of data to be manipulated) to implement an abstract idea on a computer (see MPEP 2106.05(f)). And, “training, based on the loss information, the training dynamics prediction model” the claim does not recite additional elements that integrate the judicial exception into a practical application with the words "apply it" (or an equivalent), such as merely applies the mathematical calculations to update a machine learning model. The claim does not recite any particular technological improvement rather, applies the concept using generic computer functionality (see MPEP 2106.05(f)). Step 2B (Inventive concept): No, it does not add significantly more since the intended practical application is well-understood, routine, and conventional and stated at a generic level (i.e. “apply it”, see MPEP 2106.05(f)). Therefore, claim 1 is ineligible. Regarding claim 3: Further modifies the abstract idea of claim 1. Step 2A Prong 1 (whether a claim is directed to a judicial exception): Yes, “wherein the acquiring the loss information includes acquiring a Kullback-Leibler divergence value based on the target training dynamics information and the predictive training dynamics information.” wherein acquiring loss information using recites a mathematical concept, as it involves mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP 2106.04(a)(2), (I)). Therefore, claim 3 is ineligible. Regarding claim 6: Further modifies the abstract idea of claim 1. Step 2A Prong 1 (whether a claim is directed to a judicial exception): Yes, “wherein the training the training dynamics prediction model includes acquiring first loss information based on the classification information and the pre-labeled class on the training data included in the first dataset.” wherein acquiring loss information using recites a mathematical concept, as it involves mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP 2106.04(a)(2), (I)). Therefore, claim 6 is ineligible. Regarding claim 7: Further modifies the abstract idea of claim 6. Step 2A Prong 1 (whether a claim is directed to a judicial exception): Yes, “wherein the acquiring the first loss information includes acquiring a cross-entropy loss value based on the classification information and the pre- labeled class.” wherein acquiring loss information using recites a mathematical concept, as it involves mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP 2106.04(a)(2), (I)). Therefore, claim 7 is ineligible. Regarding claim 8: Further modifies the abstract idea of claim 6. Step 2A Prong 1 (whether a claim is directed to a judicial exception): Yes, “determining, based on the classification information, a class to which the training data included in the first dataset is most likely to belong; and checking whether the determined class matches the pre-labeled class.” describes a mental process (evaluation, judgement) wherein evaluating data and comparing values, then determining if it matches recites concepts that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Therefore, claim 8 is ineligible. Regarding claim 9: Further modifies the abstract idea of claim 1. Step 2A Prong 1 (whether a claim is directed to a judicial exception): Yes, “wherein the acquiring the target training dynamics information includes calculating, based on the classification information and previous classification information, average values of probability of data belonging to each class.” recites a mathematical concept, as it involves mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP 2106.04(a)(2), (I)), and describes a mental process (evaluation, judgement) wherein evaluating data and comparing values, then determining if it matches recites concepts that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III). Therefore claim 9 is ineligible. Regarding claim 10: Further modifies the abstract idea of claim 1. Step 2A Prong 2 (evaluate whether the claim recites additional elements that integrate the exception into a practical application): No, “wherein the acquiring the predictive training dynamics information includes: acquiring, based on the classification model, hidden feature information on the training data; and acquiring, based on the hidden feature information, the predictive training dynamics information.” does not recite additional elements that integrate the judicial exception into a practical application with the words "apply it" (or an equivalent), such as mere (i.e., selecting a particular data source or type of data to be manipulated) to implement an abstract idea on a computer (see MPEP 2106.05(f)). Step 2B (Inventive concept): No, it does not add significantly more since the intended practical application is well-understood, routine, and conventional and stated at a generic level (i.e. “apply it”, see MPEP 2106.05(f)). Therefore, claim 10 is ineligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 6-7, 9 and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al., Non-Patent Literature “Cartography Active Learning” in view of Yoo et al., Non-Patent Literature “Learning Loss for Active Learning” further in view of Swayamdipta et al., Non-Patent Literature “Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics” Regarding claim 1 analogous claim 19 and 20: Zhang teaches: acquiring, based on a classification model, classification information on training data included in a first dataset; (Section 3, paragraph 9, “To select presumably informative instances, we train a binary classifier [classification model] on the seed set L [first dataset] and apply it to U to select the instances that are the closest to the decision boundary between ambiguous and hard-to-learn instances (i.e., wherein classification information is interpreted as the measures of distance to the decision)”) wherein the first dataset includes one or more data pre-labeled with a class (Section 3, paragraph 2, “Then we introduce CAL, which pro poses to learn a data map from the seed labeled data L [first dataset] and identifying regions of instances with a binary classifier…To identify these regions, we require: (1) a data map can be learned from limited data, and (2) a classifier to identify informative instances (i.e., wherein seed labeled data is interpreted as the first dataset that is pre-labeled, hence, with a class)”) training the classification model based on the classification information and the pre-labeled class (Section 3, paragraph 2, “Then we introduce CAL, which pro poses to learn a data map from the seed labeled data L and identifying regions of instances with a binary classifier…To identify these regions, we require: (1) a data map can be learned from limited data, and (2) a classifier to identify informative instances (i.e., wherein seed labeled data is interpreted as the first dataset that is pre-labeled, hence, with a class)”…“We start with a seed set size of 1,000 for AGNews and 500 for TREC, this means after the AL iterations we will have 2,500 labeled instances for AGNews and 2,000 for TREC. Our motivation here is to keep the AL simulation realistic. We assume enough annotation budget initially annotate 500-1,000 samples. Then, in every AL iteration annotate an additional 50 samples, which seems manageable for an annotator. Finally, we run 30 AL iterations to give a good overview of the performance of the acquisition functions over the iterations towards convergence”) acquiring, based on the classification information and previous classification information acquired based on the classification model in one or more previous epochs, target training dynamics information; (Section 3, paragraph 1, “to use model-independent measures, from fitting the model on the seed data L [acquiring, based on the classification information], by using data maps (Swayamdipta et al., 2020) for AL. Data maps help identify characteristics [classification information] of instances within the broader trends of a dataset by leveraging their training dynamics [target training dynamics information] (i.e., the behavior of a model during training, such as mean and standard deviation of confidence and correctness with respect to the gold label) (i.e., wherein behavior of a model during training under the broadest reasonable interpretation (BRI) is interpreted as observations over time, hence, ‘previous epochs’). These model-dependent measures reveal distinct regions in a data map, by and large, reflecting instance properties (see Figure 1 and details below on easy-to-learn, ambiguous, and hard-to-learn instances). Training dynamics encapsulate information of data quality that has been largely ignored in AL: the sweet spot of instances at the boundary of hard-to-learn and ambiguous instances, which are quick to label while providing informative samples, as shown in full data training") acquiring, based on the training dynamics prediction model, predictive training dynamics information on the training data; and (Algorithm 1, “Algorithm 1: Cartography Active Learning 1 input: Labeled seed set L, Unlabeled set U, Total budget K, Number of queries n, Correctness threshold tcor = 0.2; for i = 1, ..., n do 3 Ψ(L), Ψ(U) ← train main classifier θ on L, get representations of L and U; 4 µˆ, σˆ, φˆ ← get data map statistics of L with θ; [training dynamics information] 5 Pθ 0 ← train binary classifier θ 0 on Ψ(L) with yΨ(xˆi) = ( 1, if φˆ i > tcor 0, else [training dynamics prediction model] (i.e., wherein the classifier predicts, hence, ‘prediction model’) 6 for j=1, ..., K n do 7 xˆ ← argmin x∈Ψ(U) |0.5−Pθ 0(ˆy = 1 | x)|;”) PNG media_image1.png 648 463 media_image1.png Greyscale Zhang does not explicitly teach: A method for training a training dynamics prediction model, the method comprising: training, based on the target training dynamics information and the predictive training dynamics information, the training dynamics prediction model. Yoo teaches: A method for training a training dynamics prediction model, the method comprising: (Abstract, “active learning method that is simple but task-agnostic, and works efficiently with the deep networks. We attach a small parametric module, named “loss prediction module,” to a target network, and learn it to predict target losses of unlabeled inputs.”) Acquiring loss information based on the target training dynamics information and the predictive training dynamics information (Section 3.3, paragraph 2, “Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss [loss information] through the loss prediction module as ˆl = Θloss(h) [predictive training dynamics information]. With the target annotation y of x, the target loss can be computed as l = Ltarget(ˆy, y) to learn the target model . Since this loss l is a ground-truth target of h [target training dynamics information] for the loss prediction module, we can also compute the loss for the loss prediction module as Lloss( ˆl, l). Then, the final loss function to jointly learn both of the target model and the loss prediction module is defined as Ltarget(ˆy, y) + λ · Lloss( ˆl, l) (1) where λ is a scaling constant”) training, based on the loss information, the training dynamics prediction model (Section 3.3, “We have a labeled dataset L s K·(s+1) and a model set composed of a target model Θtarget and a loss prediction module [training dynamics prediction model] Θloss. Our objective is to learn the model set for this stage s to obtain {Θs target, Θs loss}. Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss [loss information] through the loss prediction module as ˆl = Θloss(h). With the target annotation y of x, the target loss can be computed as l = Ltarget(ˆy, y) to learn the target model. Since this loss l is a ground-truth target of h for the loss prediction module, we can also compute the loss for the loss prediction module as Lloss( ˆl, l). Then, the final loss function to jointly learn both of the target model and the loss prediction module is defined as Ltarget(ˆy, y) + λ · Lloss( ˆl, l) (1)”) Yoo and Zhang are both related to the same field of endeavor (i.e., active learning). In view of the teachings of Yoo it would have been obvious for a person of ordinary skill in the art to apply the teachings of Yoo to Zhang before the effective filing date of the claim invention in order to improve the efficiency of classification models (Yoo, Abstract, “The performance of deep neural networks improves with more annotated data. The problem is that the budget for annotation is limited. One solution to this is active learning, where a model asks human to annotate data that it perceived as uncertain. A variety of recent methods have been proposed to apply active learning to deep networks but most of them are either designed specific for their target tasks or computationally inefficient for large networks. In this paper, we propose a novel active learning method that is simple but task-agnostic, and works efficiently with the deep networks. We attach a small parametric module, named “loss prediction module,” to a target network, and learn it to predict target losses of unlabeled inputs. Then, this module can suggest data that the target model is likely to produce a wrong prediction.”) Swayamdipta teaches: wherein, the predictive training dynamics information includes a result of calculating probability for each of a plurality of classes that the data belongs to that class (Introduction, “We construct coordinates for data maps by leveraging training dynamics—the behavior of a model as training progresses [training dynamics information]. We consider the mean and standard deviation of the gold label probabilities [a result of calculating probability], predicted for each example across training epochs; these are referred to as confidence and variability. The map reveals three distinct regions in the dataset: a region with instances whose true class probabilities fluctuate frequently during training (high variability), and are hence ambiguous for the model; a region with easy-to-learn instances that the model predicts correctly and consistently (high confidence, low variability); and a region with hard-to-learn instances with low confidence, low variability, many of which we find are mislabeled during annotation”) Swayamdipta and Zhang are both related to the same field of endeavor (i.e., training dynamics). In view of the teachings of Swayamdipta it would have been obvious for a person of ordinary skill in the art to apply the teachings of Swayamdipta to Zhang before the effective filing date of the claim invention in order to improve the efficiency of training models by measuring the models’ confidence per epoch (Swayamdipta, Abstract, “We introduce Data Maps—a model-based tool to characterize and diagnose datasets. We leverage a largely ignored source of information: the behavior of the model on individual instances during training (training dynamics) for building data maps. This yields two intuitive measures for each example—the model’s confidence in the true class, and the variability of this confidence across epochs—obtained in a single run of training.”) Regarding claim 6: Zhang, as modified by Yoo, teaches the method of claim 1. Zhang further teaches: based on the classification information and the pre-labeled class on the training data included in the first dataset (Section 3, paragraph 2, “Then we introduce CAL, which pro poses to learn a data map from the seed labeled data L [first dataset] and identifying regions of instances with a binary classifier…To identify these regions, we require: (1) a data map can be learned from limited data, and (2) a classifier to identify informative instances (i.e., wherein seed labeled data is interpreted as the first dataset that is pre-labeled, hence, with a class)”) Zhang does not explicitly teach: wherein the training the training dynamics prediction model includes acquiring loss information Yoo further teaches: wherein the training the training dynamics prediction model (Section 3.3, paragraph 2, “Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss through the loss prediction module as ˆl = Θloss(h). With the target annotation y of x, the target loss can be computed as l = Ltarget(ˆy, y) to learn the target model . Since this loss l is a ground-truth target of h for the loss prediction module, we can also compute the loss for the loss prediction module as Lloss( ˆl, l). Then, the final loss function to jointly learn both of the target model and the loss prediction module is defined as Ltarget(ˆy, y) + λ · Lloss( ˆl, l) (1) where λ is a scaling constant. This procedure to define the final loss [training dynamics prediction model]”) includes acquiring loss information (Section 3.3, paragraph 2, “Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss [acquiring loss information] through the loss prediction module as ˆl = Θloss(h)) The motivation for claim 6 is the same motivation for claim 1. Regarding claim 7: Zhang, as modified by Yoo, teaches the method of claim 6. Zhang further teaches: includes acquiring a cross-entropy loss value based on the classification information and the pre- labeled class (Section 4.3, paragraph 3, “This model is suited for the binary classification task of DAL and CAL. In this case it is a single demb = 300 ReLu layer. We minimize the weighted cross-entropy as well. We use the Adam optimizer with the same parameters as above (i.e., wherein under the broadest reasonable interpretation (BRI) classification model use pre-labeled class data)”) Zhang does not explicitly teach: wherein acquiring the loss information Yoo further teaches: wherein the acquiring the first loss information (Section 3.3, paragraph 2, “Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss [acquiring loss information] through the loss prediction module as ˆl = Θloss(h)) The motivation for claim 7 is the same motivation for claim 1. Regarding claim 9: Zhang, as modified by Yoo, teaches the method of claim 1. Zhang, as modified by Yoo, does not explicitly teach: wherein the acquiring the target training dynamics information includes calculating, based on the classification information and previous classification information, average values of probability of data belonging to each class Swayamdipta teaches: wherein acquiring the target training dynamics information includes calculating, based on the classification information and previous classification information, average values of probability of data belonging to each class (Section 2.1, “Consider a training dataset of size N, D = {(x, y∗ )i} N i=1 where the ith instance consists of the observation, xi and its true label under the task, y ∗ i . Our method assumes a particular model (family) whose parameters are selected to minimize empirical risk using a particular algorithm.4 We assume the model defines a probability distribution over labels given an observation. We assume a stochastic gradient-based optimization procedure is used, with training instances randomly ordered at each epoch, across E epochs. The training dynamics [target training dynamics information] of instance i are defined as statistics calculated across the E epochs (i.e., wherein calculating based on classification information ‘statistics’ is interpreted as calculating over 'epoch'). The values of these measures then serve as coordinates in our map. The first measure aims to capture how confidently the learner assigns the true label to the observation, based on its probability distribution. We define confidence as the mean [average values] model probability of the true label [belonging to each class] (y ∗ i ) across epochs: µˆi = 1 E X E e=1 pθ (e) (y ∗ i | xi) where pθ (e) denotes the model’s probability with parameters θ (e) at the end of the eth epoch”) Swayamdipta and Zhang are both related to the same field of endeavor (i.e., training dynamics). In view of the teachings of Swayamdipta it would have been obvious for a person of ordinary skill in the art to apply the teachings of Swayamdipta to Zhang before the effective filing date of the claim invention in order to improve the efficiency of training models by measuring the models’ confidence per epoch (Swayamdipta, Abstract, “We introduce Data Maps—a model-based tool to characterize and diagnose datasets. We leverage a largely ignored source of information: the behavior of the model on individual instances during training (training dynamics) for building data maps. This yields two intuitive measures for each example—the model’s confidence in the true class, and the variability of this confidence across epochs—obtained in a single run of training.”) Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as modified by Yoo and Swayamdipta et al., further in view of Wu et al., Non-Patent Literature “Learning Kullback-Leibler Divergence-Based Gaussian Model for Multivariate Time Series Classification” Regarding claim 3: Zhang, as modified by Yoo and Swayamdipta, teaches the method of claim 1. Zhang does not explicitly teach: wherein the acquiring the loss information includes acquiring a Kullback-Leibler divergence value based on the target training dynamics information and the predictive training dynamics information Yoo further teaches: wherein the acquiring the loss information includes ((Section 3.3, paragraph 2, “Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss [acquiring loss information] through the loss prediction module as ˆl = Θloss(h)) Wu teaches: acquiring a Kullback-Leibler divergence value based on the target training dynamics information and the predictive training dynamics information (Section III, Part A, paragraph 2, “Assuming that the statistical models P1 and P2 represent two N-dimensional probability distribution functions, respectively, the Kullback-Leibler divergence between those two models is defined as Equations 1 and 2 in the case of discrete and continuous random variables, respectively: KL (P1||P2) = X x∈X P1 (x)log P1 (x) P2 (x) (1) KL (P1||P2) = Z x∈X P1 (x)log P1 (x) P2 (x) dx (2) The physical meaning of the above equations is to calculate the degree of difference between the statistical model and the given statistical model. The Kullback-Leibler divergence is non-negative and asymmetrical (i.e., wherein the target training dynamics information and the predictive training dynamics information is interpreted as P1 and P2 of the equation.)”) Wu and Zhang are both related to the same field of endeavor (i.e., classification models). In view of the teachings of Wu it would have been obvious for a person of ordinary skill in the art to apply the teachings of Wu to Zhang before the effective filing date of the claim invention in order to improve the efficiency of classification models by measuring the probability distribution during training (Wu, Abstract, “Furthermore, the Kullback-Leibler divergence is used as the similarity measurement to implement the classification of unlabeled subsequences, because it can effectively measure the similarity between different distributions.”) Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as modified by Yoo and Swayamdipta et al., further in view of Raschka et al., Non-Patent Literature “Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning.” Regarding claim 8: Zhang, as modified by Yoo and Swayamdipta, teaches the method of claim 6. Zhang does not explicitly teach: Wherein the acquiring the first loss information includes: determining, based on the classification information, a class to which the training data included in the first dataset is most likely to belong; and checking whether the determined class matches the pre-labeled class Yoo further teaches: wherein the acquiring the first loss information includes (Section 3.3, paragraph 2, “Given a training data point x, we obtain a target prediction through the target model as yˆ = Θtarget(x), and also a predicted loss [acquiring loss information] through the loss prediction module as ˆl = Θloss(h)) Raschka teaches: determining, based on the classification information, a class to which the training data included in the first dataset is most likely to belong; and checking whether the determined class matches the pre-labeled class (“After the learning algorithm fit a model in the previous step, the next question is: How "good" is the performance of the resulting model? This is where the independent test set comes into play. Since the learning algorithm has not "seen" this test set before, it should provide a relatively unbiased estimate of its performance on new, unseen data. Now, we take this test set and use the model to predict the class labels (i.e., wherein using the previously trained model to make predictions, hence, ‘first dataset used in the trained model’). Then, we take the predicted class labels and compare them to the "ground truth," the correct class labels, to estimate the models generalization accuracy or error (i.e., wherein the predicted class is compared to the ground truth is interpreted as ‘checking whether the determined class matches the pre-labeled class’)”) Raschka and Zhang are both related to the same field of endeavor (i.e., classification models). In view of the teachings of Raschka it would have been obvious for a person of ordinary skill in the art to apply the teachings of Raschka to Zhang before the effective filing date of the claim invention in order to improve the efficiency of classification models by verifying the accuracy of the model’s predictions (Raschka, Section 1.1, “"How do we estimate the performance of a machine learning model?" A typical answer to this question might be as follows: "First, we feed the training data to our learning algorithm to learn a model. Second, we predict the labels of our test set. Third, we count the number of wrong predictions on the test dataset to compute the model’s prediction accuracy."”) Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as modified by Yoo and Swayamdipta et al., further in view of Raj et al., Non-Patent Literature “Understanding Learning Dynamics of Binary Neural Networks via Information Bottleneck.” Regarding claim 10: Zhang, as modified by Yoo and Swayamdipta, teaches the method of claim 1. Zhang, as modified by Yoo and Swayamdipta, does not explicitly teach: wherein the acquiring the predictive training dynamics information includes: acquiring, based on the classification model, hidden feature information on the training data; and acquiring, based on the hidden feature information, the predictive training dynamics information Raj teaches: wherein the acquiring the predictive training dynamics information includes: acquiring, based on the classification model, hidden feature information on the training data; and (Introduction, paragraph 3, “as a framework to extract [acquiring] relevant information [hidden feature information] from random variable X (features) about another random variable Y (labels) (i.e., wherein predicting labels is interpreted as ‘based on the classification model’). By viewing each step of processing the input feature as a trade-off between compression of information and keeping relevant statistics helpful in prediction, IB principle presents an information-theoretic interpretation for explaining the learning dynamics in a supervised learning system [predictive training dynamics information]”) acquiring, based on the hidden feature information, the predictive training dynamics information (Introduction, paragraph 3 “By viewing each step of processing the input feature as a trade-off between compression of information and keeping relevant statistics helpful in prediction [hidden feature information], IB principle presents an information-theoretic interpretation for explaining the learning dynamics [predictive training dynamics information] in a supervised learning system. [13] laid the groundwork for applying IB principle to analyze deep learning by formulating the information flow from the input layer to the output layer as successive Markov chains of intermediate representations. In a follow-up work, [15] analyzed the information plane dynamics of each layer during training to obtain insights into the DNN training process”) A person of ordinary skill in the art would reasonably find the teachings of Raj to be helpful in solving the problem of understanding a models’ hidden features for training dynamics present in Zhang. In view of the teachings of Raj it would have been obvious for a person of ordinary skill in the art to apply the teachings of Raj to Zhang before the effective filing date of the claim invention in order to improve the efficiency of training models by the use of hidden features for predictive training dynamics (Raj, Abstract, “We analyze BNNs through the Information Bottleneck principle and observe that the training dynamics of BNNs is considerably different from that of Deep Neural Networks (DNNs). While DNNs have a separate empirical risk minimization and representation compression phases, our numerical experiments show that in BNNs, both these phases are simultaneous. Since BNNs have a less expressive capacity, they tend to find efficient hidden representations concurrently with label fitting. Experiments in multiple datasets support these observations, and we see a consistent behavior across different activation functions in BNNs.”) Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMINA BENOURAIDA whose telephone number is (571)272-4340. The examiner can normally be reached Monday-Friday 8:30am-5pm ET.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMINA MORENO BENOURAIDA/Examiner, Art Unit 2129 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Feb 23, 2023
Application Filed
Dec 05, 2025
Non-Final Rejection mailed — §101, §103
Apr 06, 2026
Response Filed
Jul 27, 2026
Final Rejection mailed — §101, §103
Sep 14, 2026
Request for Continued Examination
Sep 17, 2026
Response after Non-Final Action

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
4y 3m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month