DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This non-final office action is in response to the amendment filed 4 May 2026 and the RCE filed 27 May 2026.
Claims 1-18 are pending. Claims 1 and 10 are independent claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4, 8-10, 13, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Meng et al. (US 2020/0334538, published 22 October 2020, hereafter Meng) and further in view of Sridhar et al. (US 11900260, filed 5 March 2020, hereafter Sridhar) and further in view of Sainz de Cea et al. (US 2021/0035285, filed 19 August 2019, hereafter Sainz de Cea) and further in view of Sarpatwar et al. (US 2021/0397988, filed 22 June 2020, hereafter Sarpatwar) and further in view of Huang et al. (WO 2018/206504, published 15 November 2018, hereafter Huang).
As per independent claim 1, Meng discloses a method of performing class-incremental learning, the method comprising:
designating a pre-trained first model for at least one past data class as a first teacher (Figure 1A, item S110; paragraphs 0019-0020: Here, a teacher model is used to produce a teacher posterior representing training data. When the teacher posterior matches a ground truth label, the teacher posterior is used to train the student)
training a second model using at least one new data class (Figure 1A, item S120; paragraphs 0019-0020: Here, a student model is trained using at least one of the generated teacher posterior or the ground truth data (new data class))
designating the trained second model as a second teacher (Figure 1A, item S130; paragraphs 0019-0020 0058-0059: Here, the teach model and student model are used for training via forward and/or backward propagation)
performing dual-teacher information distillation (Figure 2C; paragraphs 0058-0059: Here, a plurality of posteriors are used to train the student model. The divergence is minimized (maximizing mutual information) between the posteriors and the student model)
determining a dual-teach information distillation loss based on the first set of samples and the second set of samples (paragraphs 0025-0030)
Meng fails to specifically disclose:
computing an information distillation loss based on mutual information between intermediate feature maps of the first teach, the second teacher, and a combined student model
wherein the combined student model is configured to perform classification on the at least one past data class and the at least one new data class
However, Sridhar, which is analogous to the claimed invention because it is directed toward calculating loss in a teach-student system, discloses:
computing an information distillation loss based on mutual information between intermediate feature maps of the first teach, the second teacher, and a combined student model (column 2, lines 20-35: Here, one or more student neural sub-networks are trained as part of the teacher network. The teacher network may include multiple teachers (Figure 1B; column 8, line 55-63). An interference map is generated and a function loss is calculated (column 3, line 60- column 4, line 16)
wherein the combined student model is configured to perform classification on the at least one past data class and the at least one new data class (column 6, lines 34-63: Here, a neural network is used to perform inference tasks relating to input data, such as classification. Additionally, loss is calculated by comparing the classification of previous data and newly classified data)
It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sridhar with Meng, with a reasonable expectation of success, as it would have allowed for integrating a plurality of teach models for training student networks (Sridhar: Figure 1B; column 2, lines 20-45).
Meng fails to specifically disclose wherein the combined student model is configured to incrementally collect a feature extractor output, and generate a classification weight of the at least one new data class based on the feature extractor output.
However, Sainz de Cea, which is analogous to the claimed invention because it is directed toward an adaptive classifier, discloses performing classification on the past data class and the at least one new data class, wherein the combined student model is configured to incrementally collect a feature extractor output, and generate a classification weight of the at least one new data class based on the feature extractor output (paragraph 0047: Here, a first classification is performed based on a prior (first) set of examination data and a second classification is performed based on a current (second) set of examination data to generate compressed multi-dimensional data. These data from these two sets of multi-dimensional data are then compared using quality measures (weights) to detect and classify outliers in the data). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sainz de Cea with Meng-Sridhar, with a reasonable expectation of success, as it would have allowed for detecting outliers within a set of combined classifications (Sainz de Cea: paragraph 0047). This would have facilitated an improved classification, as only data that is an outlier to the combined sets would be marked as an outlier.
Meng fails to specifically disclose:
applying data-free generative replay to generate a first set of synthetic samples with a first conditional generator for a first class at a first time
applying data-free generative replay to generate a second set of synthetic samples with a second conditional generator for a second class at a second time, wherein the second time is after the first time
However, Sarpatwar, which is analogous to the claimed invention because it is directed toward training teacher model using synthetic data, discloses:
applying data-free generative replay to generate a first set of synthetic samples with a first conditional generator for a first class at a first time (Figure 5; paragraphs 0006 and 0079-0080)
applying data-free generative replay to generate a second set of synthetic samples with a second conditional generator for a second class at a second time, wherein the second time is after the first time (Figure 5; paragraphs 0006 and 0079-0080)
It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng- Sridhar-Sainz de Cea, with a reasonable expectation of success, as it would have facilitated customizing synthetic data to account for problems with the original dataset. This would have provide the benefit of creating a model which satisfies domain constraints.
Meng fails to specifically disclose wherein the dual-teach information distillation further comprises generating a first set of synthetic samples for a first class at a first time and a second set of synthetic samples for a second class at a second time. However, Huang, which is analogous to the claimed invention because it is directed toward training a reinforcement learning model using multiple sets of synthetic data, discloses wherein the dual-teach information distillation further comprises generating a first set of synthetic samples for a first class at a first time and a second set of synthetic samples for a second class at a second time (page 2, lines 13-27: Here, a Generative Adversarial Network (GAN) is trained using data from the real environment. This training data includes a data slice corresponding to a state-reward pair and a state-action pair. A data generator trained with the first state-reward pair is used to generate a first set of synthetic data. The examiner considers this data be generated at the claimed “first time” and the state as being analogous to the claimed “first class.” A portion of the first synthetic data, generated at a first time, is processed to generate a resulting data slice. The first synthetic data corresponding to a second state-action pair and the result data slice corresponding to a second state-reward pair. The second state-action pair portion of the first synthetic data is merged with the second state-reward pair from the relations network to generate second synthetic data. This second synthetic data is generated at a “second time, wherein the second time is after the first time” because this second synthetic data is at least partially generated based on the first synthetic data generated at the “first time.” Further, the examiner considers the second state as analogous to the claimed “second class.” Further, a third synthetic data set may be generated from the first synthetic data and the second synthetic data (page 3, lines 3-4)). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Huang with Meng- Sridhar-Sainz de Cea- Sarpatwar, with a reasonable expectation of success, as it would have allowed improving training of a reinforcement learning model by improving the quality of synthetic data used for training (Huang: page 7, lines 16-24).
As per dependent claim 2, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Meng fails to specifically disclose training at least one of the first conditional generator or a second conditional generator or generate synthetic data, given the first model or the second model, without using any stored data, wherein the synthetic data is configured to mimic training data used to train the first teacher or the second teacher.
However, Sarpatwar, which is analogous to the claimed invention because it is directed toward training teacher model using synthetic data, discloses training at least one of the first conditional generator or a second conditional generator or generate synthetic data, given the first model or the second model, without using any stored data, wherein the synthetic data is configured to mimic training data used to train the first teacher or the second teacher (paragraph 0006).
It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng-Tao, with a reasonable expectation of success, as it would have facilitated training using synthetic data. This would have provide the benefit of creating a model which satisfies domain constraints).
As per dependent claim 3, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 2, and the same rejection is incorporated herein. Meng discloses:
determining a cross-entropy loss between a label input into the conditional generator and a value output from the first teacher or the second teacher (paragraphs 0020-0022)
determining a batch-normalization statistics loss by matching mean and variance variables stored in the batch-normalization layers of the first teacher or the second teach with mean and variance variables computed at the same batch-normalization layers of the first teacher or the second teacher for information output from the conditional generator
Meng fails to specifically disclose adjusting the conditional generator to account for the change in parameters. However, Sarpatwar, which is analogous to the claimed invention because it is directed toward training teacher model using synthetic data, discloses adjusting the conditional generator to account for the change in parameters (paragraph 0006). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng-Tao, with a reasonable expectation of success, as it would have facilitated customizing synthetic data to account for problems with the original dataset. This would have provide the benefit of creating a model which satisfies domain constraints.
As per dependent claim 4, Meng discloses wherein the first model designated as the first teacher is updated using weight imprinting by accessing stored training data (paragraphs 0022-0023).
As per dependent claim 6, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Meng discloses:
accounting for the dual-teacher information distillation loss when performing dual-teacher information distillation (paragraphs 0025-0030)
As per dependent claim 7, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 2, and the same rejection is incorporated herein. Sarpatwar discloses wherein training the first conditional generator or the second conditional generator further comprises a pre-trained model to generate the synthetic data that is used to train the first conditional generator or the second conditional generator without using any stored training data (Figure 5; paragraph 0006). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Sarpatwar with Meng-Tao, with a reasonable expectation of success, as it would have facilitated customizing synthetic data to account for problems with the original dataset. This would have provide the benefit of creating a model which satisfies domain constraints.
As per dependent claim 8, Meng discloses wherein the second model designated as the second teacher is trained with new data for each new class that is introduced (Figures 2A-2B; paragraphs 0034-0035: Here, the model is adapted to each new domain).
As per dependent claim 9, Meng discloses wherein data output from the second teacher and data output from the first teacher are applied to the combined student model to perform dual-teach information distillation (Figure 2B; paragraph 0035).
With respect to claim 10, the applicant discloses the limitations similar to those in claim 1. Additionally, Meng discloses an electronic device comprising a non-transitory computer readable memory and a process (Figure 4). Claim 10 is similarly rejected.
With respect to claims 11-13 and 15-18, the applicant discloses the limitations substantially similar to those in claims 2-4 and 6-9, respectively. Claims 11-13 and 15-18 are similarly rejected.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang, and further in view of Stone et al. (US 11640447, filed 18 April 2018, hereafter Stone).
As per dependent claim 5, Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang disclose the limitations similar to those in claim 1, and the same rejection is incorporated herein. Meng fails to specifically disclose wherein the trained second model designated as the second teacher is trained by using a “none” class in response to training data not being accessible.
However, Stone, which is analogous to the claimed invention because it is directed toward training a model classifier, discloses wherein the trained second model is trained by using a “none” class in response to training data not being accessible (column 11, line 52- column 12, line 2). It would have been obvious to one of ordinary skill in the art at the time of the applicant’s effective filing date to have combined Stone with Wang, with a reasonable expectation of success, as it would have allowed for assigning values to each feature. This would have insured that all fields in the model are filled in order to improve classification.
With respect to claim 14, the applicant discloses the limitations substantially similar to those in claim 5. Claim 14 is similarly rejected.
Response to Arguments
Applicant’s arguments with respect to the rejection of claims under 35 USC 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground of rejection is made in view of Meng, Sridhar, Sainz de Cea, Sarpatwar, and Huang.
Applicant’s arguments with respect to the rejection of claims under 35 USC 101 have been fully considered and are persuasive. The rejection has been withdrawn.
Specifically, the applicant’s arguments beginning on page 8 culminate in the assessment that the claims as a whole, “define a specific arrangement of multiple neural networks, intermediate feature extracting, an information distillation loss based on mutual information, feature-extractor-based classification-weight generation, and temporally distinct data-free generative reply” that “constitutes a concrete improvement to neural network training technology and to the functioning of the model itself (page 10).” The applicant’s arguments are persuasive.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Liu et al. (US 2021/0150340): Discloses training a BERT model to a smaller student model using the same inputs as the teacher model (Abstract)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLE R STORK whose telephone number is (571)272-4130. The examiner can normally be reached 8am - 2pm; 4pm - 6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571/272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KYLE R STORK/Primary Examiner, Art Unit 2128