Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al (20220343205) in view of Merler (20230259716).
As per claim 1, Yu et al (20220343205) teaches a system comprising:
a processor; and a computer-readable medium storing instructions that are operative upon execution by the processor to (as processor and computer readable medium storing the executable instructions – para 0074 – 0076):
identify training samples from a dataset via active learning using a teacher model (as teacher version of the machine learning model – para 0015); generate a first group of soft labels for the training samples using a large language machine learning model (LLM – using the teacher model as a larger heavier network – para 0060), the LLM being independent of
training the student model using the training samples, with the first group of soft labels, the student model being configured to output class membership probabilities for input samples (as, using probabilities to compare for results of the language model – para 0086, as applied to training the student model – para 0070, 0068) ;
evaluate a performance metric of the trained student model based on human-annotated ground truth samples (as, conventional systems perform ground-truth labels – para 0064); upon determining that, that the performance metric is below a threshold level, identify additional training samples from the dataset using the teacher model (and using the human annotated ground truth samples in training the teacher model – para 0064);
using the LLM, generating a first group of soft labels and generate a second group of soft labels for the identified additional training samples (as, using a smaller subgroup to re-prompt the model – para 0085); and retrain the student model using at least the training samples and the additional training samples with the second group of labels (as using user input – para 0096, and using human supervision to updates the teacher version as well as the distillation of the student model – para 0020 – examiner notes, that the “distillation process” includes labeling and human input into the training of the student model – see para 0020 – “incrementally finetuning the student version” – and the student version is trained using labels – see para 0015, 0025). Regarding the claim language toward using an embedding model independent of the teacher model, and to train the student model, and to identify additional training examples to improve the student model; Yu et al (20220343205) teaches a direct relationship between the teacher/student models, as evidenced in Figure 5 of Yu et al (20220343205). Merler (20230259716) teaches a knowledge-distillation process that uses previous knowledge, distilled to a student model to train the student model (see para 0042, para 0043; examiner notes that the “NAS” is similar to the technique in Yu et al (20220343205), with “knowledge distillation” – ‘KD’ to train the student model not directly from the teacher model – e.g., in para 0042, wherein the knowledge-distillation model uses filtered/gleaned information from the teacher model), and then using additional training examples to improve the student model – see para 0006. Therefore, it would have been obvious to one of ordinary skill in the art of distributed language models, to modify the model structure of Yu et al (20220343205) with knowledge-distillation guided student models embedded with a traditional teacher-student model architecture, as taught by Merler (20230259716) because it would advantageously improve the accuracy and the latency scores of the student model, compared to randomly trained student models (see Merler (20230259716), para 0038).
As per claim 2, the combination of Yu et al (20220343205) in view of Merler (20230259716) teaches the system of claim 1, wherein the instructions are further operative to: cause a user interface (UI) to be displayed on a display device (Yu et al (20220343205), as GUI presenting a visual representation of the data – para 0095),
the UI including a graph comprising data points, each of the data points representing a training sample from the training samples (Yu et al (20220343205), as, showing the change in data –para 0095; wherein the data refers to monitoring of certain data to identify potential bias, errors, and unintended outcomes – para 0094);
receive second user input indicating selection of a first data point; cause to be displayed sample data associated with the first data point; and receive third user input identifying a label for the first data point, thereby causing the first data point to become a human-annotated training sample of the additional training sample used to retrain the student model (Yu et al (20220343205), as, using human supervision to updates the teacher version as well as the distillation of the student model – para 0020 – examiner notes, that the “distillation process” includes labeling and human input into the training of the student model – see para 0020 – “incrementally finetuning the student version” – and the student version is trained using labels – see para 0015, 0025).
As per claim 3, the combination of Yu et al (20220343205) in view of Merler (20230259716) teaches the system of claim 2, wherein the instructions are further operative to: in response to receiving the second user input indicating selection of the first data point, prompt the LLM to generate a label recommendation for the first data point (Yu et al (20220343205), LLM – using the teacher model as a larger heavier network – para 0060) (Yu et al (20220343205), deriving labels for the training samples – as, the student version generates programmed labels for data snippets – para 0015), wherein causing to be displayed sample data associated with the first data point includes causing the label recommendation to be displayed (Yu et al (20220343205), as GUI presenting a visual representation of the data – para 0095) .
As per claim 4, the combination of Yu et al (20220343205) in view of Merler (20230259716) teaches the system of claim 1, wherein the instructions are further operative to: cause a user interface (UI) to be displayed on a display device, the UI including a graph comprising data points, each of the data points representing a training sample from the training samples (Yu et al (20220343205), as, showing the change in data –para 0095; wherein the data refers to monitoring of certain data to identify potential bias, errors, and unintended outcomes – para 0094);
receive second user input indicating selection of a region of the graph; identify data points occurring within the region; cause the UI to display sample data for each of the data points occurring within the region; and receive additional user input identifying a label for each of the data points (Yu et al (20220343205), as, when the user constantly watches the display – para 0095, monitoring the constant updated display – para 0094; and choosing which sections, based on the display, that needs user/human input to increase the accuracy of the models – see para 0100 – human intervention, and human updating of the labels on the student models – para 0015).
As per claim 5, the combination of Yu et al (20220343205) in view of Merler (20230259716) teaches the system of claim 1, wherein the instructions are further operative to: perform iterations of student model retraining;
at each of the iterations of student model retraining: compare a current performance metric of a current student model to a previous performance metric of a prior student model, thereby identifying a performance differential (Yu et al (20220343205), as, measuring performance and improving performance by increasing accuracy in the programmed labels compare with pseudo labels – para 0038);
and based on the comparison, add an additional soft labeled training sample to the training samples when the performance differential is above a threshold and add an additional human-labeled training sample to the training samples when the performance differential is below the threshold (Yu et al (20220343205), as, using human supervision to updates the teacher version as well as the distillation of the student model – para 0020 – examiner notes, that the “distillation process” includes labeling and human input into the training of the student model – see para 0020 – “incrementally finetuning the student version” – and the student version is trained using labels – see para 0015, 0025).
As per claim 6, the combination of Yu et al (20220343205) in view of Merler (20230259716) teaches the system of claim 1, wherein the instructions are further operative to: determine, using the student model, a class membership probability for a first sample belonging to a first class (Yu et al (20220343205), as, using probabilities to compare for results of the language model – para 0086, as applied to training the student model – para 0070, 0068; wherein the classification can be based on classes – para 0085); and assign the first class as a soft label to the first sample when the class membership probability is above a threshold (Yu et al (20220343205), as labelling, and additional labeling, according to threshold confidence levels – para 0040).
As per claim 7, the combination of Yu et al (20220343205) in view of Merler (20230259716) teaches the system of claim 1, wherein the LLM generates semantic embeddings for the additional training samples, and the student model is retrained using at least the additional training samples and the generated semantic embeddings (see Yu et al (20220343205), as using language models – para 0088, 0089, and para 0041 showing similarity in content features to the label; Merler (20230259716) teaching deep contextual word processing – para 0021, and common context meanings – para 0043) ).
Claims 8-14 are method claims whose steps are performed by the system claims 1-7 above and as such, claims 8-14 are similar in scope and content to claims 1-7 above; therefore, claims 8-14 are rejected under similar rationale as presented against claims 1-7 above. Furthermore, to claims 9,10, (Examiner notes the list of categories are in the alternative form, “or”, and Yu et al (20220343205) teaches operating on image information – see para 0099).
Claims 15-20 are device claims contain steps that are performed by the systems claims 1-7 above and as such, claims 15-20 are similar in scope and content to claims 1-7 above; therefore, claims 15-20 are rejected under similar rationale as presented against claims 1-7 above. Furthermore, Yu et al (20220343205) teaches processor/memories performing stored steps (para 0074).
Response to Arguments
Applicant's arguments filed 5/20/2026 have been fully considered but are moot in view of the new grounds of rejection. Examiner notes the introduction of the Merler (20230259716) to teach knowledge-distillation of student models as an additional sidestep of traditional teacher-student direct modeling.
Lastly, as to soft labeling, examiner notes the Fukuda et al (20220414448) teaching separate teacher model, student model, language model, -- see paragraphs 0048, 0017, with soft labeling – para 0003, 0027.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Please see related art listed on the PTO-892 form.
Furthermore, the following references were found:
Fukuda et al (20220414448) teaches soft labeling by the teacher model see para 0003, 0027.
Bui et al (20220114476) teaches pseudo labeling in addition to an initial round of labeling (para 0020).
Kim et al (20230143721) teaches natural language processing (NLP) with few-shot processing (para 0021) with label sequencing (para 0030-0032), differentiating between teacher and student models – para 0039)
Luong et al (20220383206) teaches pretrained language models, improving accuracy with few-shot benchmarks (para 0034), using ground truth labels (para 0043), operating on teacher/student models (para 0056-0060).
Balasubramanian et al (20230368786) teaches student/teacher models operating on speech (para 0032), with displays to monitor/modify the models – para 0036, using modified labeling (para 0061).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Michael N Opsasnick/Primary Examiner, Art Unit 2658 08/20/2026