Prosecution Insights
Last updated: October 02, 2026
Application No. 17/317,421

METHOD AND APPARATUS FOR INCREMENTAL LEARNING

Non-Final OA §101§103
Filed
May 11, 2021
Priority
Nov 05, 2020 — provisional 63/110,063
Examiner
STORK, KYLE R
Art Unit
2128
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
5 (Non-Final)
63%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
559 granted / 884 resolved
+8.2% vs TC avg
Strong +29% interview lift
Without
With
+28.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
45 currently pending
Career history
931
Total Applications
across all art units

Statute-Specific Performance

§101
15.5%
-24.5% vs TC avg
§103
61.3%
+21.3% vs TC avg
§102
10.5%
-29.5% vs TC avg
§112
5.7%
-34.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 884 resolved cases

Office Action

§101 §103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This non-final office action is in response to the amendment filed 4 May 2026 and the RCE filed 27 May 2026. Claims 1-18 are pending. Claims 1 and 10 are independent claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 4, 8-10, 13, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Meng et al. (US 2020/0334538, published 22 October 2020, hereafter Meng) and further in view of Sridhar et al. (US 11900260, filed 5 March 2020, hereafter Sridhar) and further in view of Sainz de Cea et al. (US 2021/0035285, filed 19 August 2019, hereafter Sainz de Cea) and further in view of Sarpatwar et al. (US 2021/0397988, filed 22 June 2020, hereafter Sarpatwar) and further in view of Huang et al. (WO 2018/206504, published 15 November 2018, hereafter Huang). As per independent claim 1, Meng discloses a method of performing class-incremental learning, the method comprising: designating a pre-trained first model for at least one past data class as a first teacher (Figure 1A, item S110; paragraphs 0019-0020: Here, a teacher model is used to produce a teacher posterior representing training data. When the teacher posterior matches a ground truth label, the teacher posterior is used to train the student) training a second model using at least one new data class (Figure 1A, item S120; paragraphs 0019-0020: Here, a student model is trained using at least one of the generated teacher posterior or the ground truth data (new data class)) designating the trained second model as a second teacher (Figure 1A, item S130; paragraphs 0019-0020 0058-0059: Here, the teach model and student model are used for training via forward and/or backward propagation) performing dual-teacher information distillation (Figure 2C; paragraphs 0058-0059: Here, a plurality of posteriors are used to train the student model. The divergence is minimized (maximizing mutual information) between the posteriors and the student model) determining a dual-teach information distillation loss based on the first set of samples and the second set of samples (paragraphs 0025-0030) Meng fails to specifically disclose: computing an information distillation loss based on mutual information between intermediate feature maps of the first teach, the second teacher, and a combined student model wherein the combined student model is configured to perform classification on the at least one past data class and the at least one new data class However, Sridhar, which is analogous to the claimed invention because it is directed toward calculating loss in a teach-student system, discloses: computing an information distillation loss based on mutual information between intermediate feature maps of the first teach, the second teacher, and a combined student model (column 2, lines 20-35: Here, one or more student neural sub-networks are trained as part of the teacher network. The teacher network may include multiple teachers (Figure 1B; column 8, line 55-63). An interference map is generated and a function loss is calculated (column 3, line 60- column 4, line 16) wherein the combined student model is configured to perform classification on the at least one past data class and the at least one new data class (column 6, lines 34-63: Here, a neural network is used to perform inference tasks relating to input data, such as classification. Additionally, loss is calculated by comparing the classification of previous data and newly classified data) It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sridhar with Meng, with a reasonable expectation of success, as it would have allowed for integrating a plurality of teach models for training student networks (Sridhar: Figure 1B; column 2, lines 20-45). Meng fails to specifically disclose wherein the combined student model is configured to incrementally collect a feature extractor output, and generate a classification weight of the at least one new data class based on the feature extractor output. However, Sainz de Cea, which is analogous to the claimed invention because it is directed toward an adaptive classifier, discloses performing classification on the past data class and the at least one new data class, wherein the combined student model is configured to incrementally collect a feature extractor output, and generate a classification weight of the at least one new data class based on the feature extractor output (paragraph 0047: Here, a first classification is performed based on a prior (first) set of examination data and a second classification is performed based on a current (second) set of examination data to generate compressed multi-dimensional data. These data from these two sets of multi-dimensional data are then compared using quality measures (weights) to detect and classify outliers in the data). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sainz de Cea with Meng-Sridhar, with a reasonable expectation of success, as it would have allowed for detecting outliers within a set of combined classifications (Sainz de Cea: paragraph 0047). This would have facilitated an improved classification, as only data that is an outlier to the combined sets would be marked as an outlier. Meng fails to specifically disclose: applying data-free generative replay to generate a first set of synthetic samples with a first conditional generator for a first class at a first time applying data-free generative replay to generate a second set of synthetic samples with a second conditional generator for a second class at a second time, wherein the second time is after the first time However, Sarpatwar, which is analogous to the claimed invention because it is directed toward training teacher model using synthetic data, discloses: applying data-free generative replay to generate a first set of synthetic samples with a first conditional generator for a first class at a first time (Figure 5; paragraphs 0006 and 0079-0080) applying data-free generative replay to generate a second set of synthetic samples with a second conditional generator for a second class at a second time, wherein the second time is after the first time (Figure 5; paragraphs 0006 and 0079-0080) It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng- Sridhar-Sainz de Cea, with a reasonable expectation of success, as it would have facilitated customizing synthetic data to account for problems with the original dataset. This would have provide the benefit of creating a model which satisfies domain constraints. Meng fails to specifically disclose wherein the dual-teach information distillation further comprises generating a first set of synthetic samples for a first class at a first time and a second set of synthetic samples for a second class at a second time. However, Huang, which is analogous to the claimed invention because it is directed toward training a reinforcement learning model using multiple sets of synthetic data, discloses wherein the dual-teach information distillation further comprises generating a first set of synthetic samples for a first class at a first time and a second set of synthetic samples for a second class at a second time (page 2, lines 13-27: Here, a Generative Adversarial Network (GAN) is trained using data from the real environment. This training data includes a data slice corresponding to a state-reward pair and a state-action pair. A data generator trained with the first state-reward pair is used to generate a first set of synthetic data. The examiner considers this data be generated at the claimed “first time” and the state as being analogous to the claimed “first class.” A portion of the first synthetic data, generated at a first time, is processed to generate a resulting data slice. The first synthetic data corresponding to a second state-action pair and the result data slice corresponding to a second state-reward pair. The second state-action pair portion of the first synthetic data is merged with the second state-reward pair from the relations network to generate second synthetic data. This second synthetic data is generated at a “second time, wherein the second time is after the first time” because this second synthetic data is at least partially generated based on the first synthetic data generated at the “first time.” Further, the examiner considers the second state as analogous to the claimed “second class.” Further, a third synthetic data set may be generated from the first synthetic data and the second synthetic data (page 3, lines 3-4)). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Huang with Meng- Sridhar-Sainz de Cea- Sarpatwar, with a reasonable expectation of success, as it would have allowed improving training of a reinforcement learning model by improving the quality of synthetic data used for training (Huang: page 7, lines 16-24). As per dependent claim 2, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Meng fails to specifically disclose training at least one of the first conditional generator or a second conditional generator or generate synthetic data, given the first model or the second model, without using any stored data, wherein the synthetic data is configured to mimic training data used to train the first teacher or the second teacher. However, Sarpatwar, which is analogous to the claimed invention because it is directed toward training teacher model using synthetic data, discloses training at least one of the first conditional generator or a second conditional generator or generate synthetic data, given the first model or the second model, without using any stored data, wherein the synthetic data is configured to mimic training data used to train the first teacher or the second teacher (paragraph 0006). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng-Tao, with a reasonable expectation of success, as it would have facilitated training using synthetic data. This would have provide the benefit of creating a model which satisfies domain constraints). As per dependent claim 3, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 2, and the same rejection is incorporated herein. Meng discloses: determining a cross-entropy loss between a label input into the conditional generator and a value output from the first teacher or the second teacher (paragraphs 0020-0022) determining a batch-normalization statistics loss by matching mean and variance variables stored in the batch-normalization layers of the first teacher or the second teach with mean and variance variables computed at the same batch-normalization layers of the first teacher or the second teacher for information output from the conditional generator Meng fails to specifically disclose adjusting the conditional generator to account for the change in parameters. However, Sarpatwar, which is analogous to the claimed invention because it is directed toward training teacher model using synthetic data, discloses adjusting the conditional generator to account for the change in parameters (paragraph 0006). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng-Tao, with a reasonable expectation of success, as it would have facilitated customizing synthetic data to account for problems with the original dataset. This would have provide the benefit of creating a model which satisfies domain constraints. As per dependent claim 4, Meng discloses wherein the first model designated as the first teacher is updated using weight imprinting by accessing stored training data (paragraphs 0022-0023). As per dependent claim 6, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Meng discloses: accounting for the dual-teacher information distillation loss when performing dual-teacher information distillation (paragraphs 0025-0030) As per dependent claim 7, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 2, and the same rejection is incorporated herein. Sarpatwar discloses wherein training the first conditional generator or the second conditional generator further comprises a pre-trained model to generate the synthetic data that is used to train the first conditional generator or the second conditional generator without using any stored training data (Figure 5; paragraph 0006). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng-Tao, with a reasonable expectation of success, as it would have facilitated customizing synthetic data to account for problems with the original dataset. This would have provide the benefit of creating a model which satisfies domain constraints. As per dependent claim 8, Meng discloses wherein the second model designated as the second teacher is trained with new data for each new class that is introduced (Figures 2A-2B; paragraphs 0034-0035: Here, the model is adapted to each new domain). As per dependent claim 9, Meng discloses wherein data output from the second teacher and data output from the first teacher are applied to the combined student model to perform dual-teach information distillation (Figure 2B; paragraph 0035). With respect to claim 10, the applicant discloses the limitations similar to those in claim 1. Additionally, Meng discloses an electronic device comprising a non-transitory computer readable memory and a process (Figure 4). Claim 10 is similarly rejected. With respect to claims 11-13 and 15-18, the applicant discloses the limitations substantially similar to those in claims 2-4 and 6-9, respectively. Claims 11-13 and 15-18 are similarly rejected. Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang, and further in view of Stone et al. (US 11640447, filed 18 April 2018, hereafter Stone). As per dependent claim 5, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Meng fails to specifically disclose wherein the trained second model designated as the second teacher is trained by using a “none” class in response to training data not being accessible. However, Stone, which is analogous to the claimed invention because it is directed toward training a model classifier, discloses wherein the trained second model is trained by using a “none” class in response to training data not being accessible (column 11, line 52- column 12, line 2). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Stone with Wang, with a reasonable expectation of success, as it would have allowed for assigning values to each feature. This would have insured that all fields in the model are filled in order to improve classification. With respect to claim 14, the applicant discloses the limitations substantially similar to those in claim 5. Claim 14 is similarly rejected. Response to Arguments Applicant’s arguments with respect to the rejection of claims under 35 USC 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground of rejection is made in view of Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang. Applicant’s arguments with respect to the rejection of claims under 35 USC 101 have been fully considered and are persuasive. The rejection has been withdrawn. Specifically, the applicant’s arguments beginning on page 8 culminate in the assessment that the claims as a whole, “define a specific arrangement of multiple neural networks, intermediate feature extracting, an information distillation loss based on mutual information, feature-extractor-based classification-weight generation, and temporally distinct data-free generative reply” that “constitutes a concrete improvement to neural network training technology and to the functioning of the model itself (page 10).” The applicant’s arguments are persuasive. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Liu et al. (US 2021/0150340): Discloses training a BERT model to a smaller student model using the same inputs as the teacher model (Abstract) Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLE R STORK whose telephone number is (571)272-4130. The examiner can normally be reached 8am - 2pm; 4pm - 6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571/272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KYLE R STORK/Primary Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Show 12 earlier events
Oct 29, 2025
Response Filed
Feb 02, 2026
Final Rejection mailed — §101, §103
Apr 15, 2026
Examiner Interview Summary
Apr 15, 2026
Applicant Interview (Telephonic)
May 04, 2026
Response after Non-Final Action
May 27, 2026
Request for Continued Examination
May 31, 2026
Response after Non-Final Action
Aug 25, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749004
HYPER-PERSONALIZED QUALIFIED APPLICANT MODELS
5y 7m to grant Granted Sep 29, 2026
Patent 12731020
NEUROMORPHIC CIRCUIT, NEUROMORPHIC ARRAY LEARNING METHOD, AND PROGRAM
5y 4m to grant Granted Sep 08, 2026
Patent 12675682
NEURAL NETWORK ACCELERATOR OUTPUT RANKING
5y 8m to grant Granted Jul 07, 2026
Patent 12645924
HARDWARE CIRCUIT FOR ACCELERATING NEURAL NETWORK COMPUTATIONS
5y 5m to grant Granted Jun 02, 2026
Patent 12585935
EXECUTION BEHAVIOR ANALYSIS TEXT-BASED ENSEMBLE MALWARE DETECTOR
5y 1m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
63%
Grant Probability
92%
With Interview (+28.7%)
3y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 884 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month