Prosecution Insights
Last updated: October 02, 2026
Application No. 18/243,259

MODEL ASSEMBLY WITH KNOWLEDGE DISTILLATION

Final Rejection §103
Filed
Sep 07, 2023
Examiner
SHOLEMAN, ABU S
Art Unit
2496
Tech Center
2400 — Computer Networks
Assignee
Cisco Technology Inc.
OA Round
2 (Final)
78%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
623 granted / 796 resolved
+20.3% vs TC avg
Strong +28% interview lift
Without
With
+27.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
26 currently pending
Career history
832
Total Applications
across all art units

Statute-Specific Performance

§101
14.2%
-25.8% vs TC avg
§103
54.6%
+14.6% vs TC avg
§102
4.4%
-35.6% vs TC avg
§112
18.9%
-21.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 796 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim(s) are rejected under 103(a), have been considered but are moot because the new ground of rejection based on the IDs filed on 09/07/2023, does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant argued in the remark that Prior arts do not disclose the generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models. Examiner respectfully disagrees. Fukuda discloses [0008] FIG. 1 illustrates a block diagram of a knowledge distillation system for speech recognition according to an exemplary embodiment of the present invention; 0023] FIG. 1 illustrates a block diagram of a knowledge distillation system 100 for the speech recognition. As shown in FIG. 1, the knowledge distillation system 100 includes a training data pool 102 that stores a collection of training data; a training module 110 that performs a training process based on knowledge distillation technique; a teacher model 120, i.e. one model, that produces data for generating soft labels for the knowledge distillation from training data stored in the training data pool 102; and a student model 130, i.e. second model, under training by the training module 110. Thus, Fukuda discloses knowledge distillation system as a unified model using the two different model 1) a teacher model and 2) a student model. However, Yang US 11,487,944(IDs submitted 09/07/2023), fig.2, 240, discloses system computes a distillation loss(i.e. a unified model) between that student model and each teacher model based on the tag predictions from each model. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3-11 and 13-22 are rejected under 35 U.S.C. 103 as being unpatentable over Fukuda et al US 2019/0205748 in view of Yang et al US 11,487,944. As per claim 1, Fukuda discloses a method comprising: receiving, at a device and via a user interface, one or more constraint parameters for each of a plurality of machine learning models that perform different analytics tasks( 0064 preparing the teacher model 120 means making the teacher model 120 available by reading the teacher model 120 onto a memory space of the local computer system; or establishing a connection with the teacher model 120 that operates on a remote computer system such that the training input can be fed into the teacher model 120 and a result for the training input can be received from the teacher model 120. ); computing, by the device and based on the one or more constraint parameters, a set of weights for the plurality of machine learning models ([0069] At step S104, the processing circuitry may create a confusion matrix 106 based on the alignments of the phoneme classes for the respective frames. Since it has been described with reference to FIGS. 2-4, a detailed description about the way of creating the confusion matrix 106 is omitted here. 0080, the soft labels q.sub.i after the softmax computation 142 are compared with the output p.sub.i after the softmax computation 138 to encourage the posterior probabilities of the student model 130 close to those of the teacher model 120. Comparing value after the softmax computation is preferable. However, in other embodiment, comparing the soft labels before the softmax computation 142 with the output before the softmax computation 138 may not be excluded. 0077] In one embodiment, all of the soft labels calculated for each training input are used to train the student model 130. Alternatively, in other embodiment, merely at least a part of the set of the soft labels calculated for each training input is used to train the student model 130. For example, posterior probabilities of top K most likely class labels in q.sub.i are used to train the student model 130 after the top K class labels from the teacher model 120 are normalized so that the sum of the top K equals to 1. This normalization may be performed after the softmax computation. ); generating, by the device, a unified model by performing knowledge distillation on the plurality of machine learning models using the set of weights ([0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130. FIG. 6 further describes a way of comparing the soft labels q.sub.i with the outputs p.sub.i of the student model 130 during the training of the knowledge distillation. [0081] After performing the training process at S108, the process may proceed to step S109 and end at the step S109. The parameters of the student model 130, which may include weights between each units and biases of each unit, are optimized during the training of the knowledge distillation process so as to classify the input correctly. ); and generating including performing knowledge distillation on the plurality of machine learning models using the set of weights(0023] FIG. 1 illustrates a block diagram of a knowledge distillation system 100 for the speech recognition. As shown in FIG. 1, the knowledge distillation system 100 includes a training data pool 102 that stores a collection of training data; a training module 110 that performs a training process based on knowledge distillation technique; a teacher model 120 that produces data for generating soft labels for the knowledge distillation from training data stored in the training data pool 102; and a student model 130 under training by the training module 110. And [0051]/0060/0081 In a preferable embodiment, the soft label convertor 140 uses an output obtained for a teacher side class that is frequently observed in the collection together with the corresponding student side member. In a further preferable embodiment, a most frequently observed class is mapped to the corresponding student side member, and the output for this teacher side class is used for calculating a soft label for the corresponding student side member by using softmax function. However, in other embodiments, multiple outputs corresponding to multiple teacher side classes that are frequently observed in the collection together with the corresponding student side member may be used for calculating the soft label by weighted or unweighted average. ) Fukuda does not disclose generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models, the generating including performing knowledge distillation on the plurality of machine learning models using the set of weights (emphasis added ); and deploying, by the device, the unified model for execution by a particular node in a network. However, Yang discloses generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models, the generating including performing knowledge distillation on the plurality of machine learning models using the set of weights ( fig.2, 240, discloses system computes a distillation loss( i.e. a unified model) between that student model and each teacher model based on the tag predictions from each model and col 2, lines 55-62 The result is a unified named-entity recognition model); and deploying, by the device, the unified model for execution by a particular node in a network(col 7, lines 58-60 The methods described herein are embodied in software and performed by a computer system (comprising one or more computing devices) executing the software. So those method can be deployed to oner or more computers by performed by the one or more computers ). Fukuda and Yang are both considered to be analogous to the claimed invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fukuda to incorporate the teachings of Yang and provided a NER model on datasets with multiple tag sets. Doing so would provide datasets with different tags typically originate from different domains, thereby increasing overcome domain mismatch and language variations (col 2, lines 17-18). As per claim 3. Fukuda and Yang disclose the method as in claim 1, Yang discloses wherein the different analytics tasks comprise video analytics tasks( col 7, lines 15-32 obtaining the student NER model from a plurality of teacher NER models. Steps 410-460 are the same as steps 310-360. However, in addition to computing an aggregate distillation loss and a student loss, the system also computes a contrastive representation distillation (CRD) loss. A contrastive representation distillation loss between the student model and a teacher model is a measure of the differences between vector representations generated by the two models for the input data sequences as part of the prediction process. Minimizing this loss enables the student to distill domain-invariant knowledge from the teacher models and enables the student model to produce vector representations of input data sequences that are domain insensitive (or at least less domain sensitive than would otherwise be). This in turn enables the student model to adapt to and to perform better on another domain (i.e., on a domain that is different from the one on which it was trained ). As per claim 4. Fukuda and Yang disclose the method as in claim 3, Fukuda discloses wherein the video analytics tasks comprise one or more of: pose estimation, object detection, semantic segmentation, or instance segmentation ([0111] In the image recognition system, the components of the output layers may also be different depending on environments. However, according to one or more embodiments of the present invention, it is possible to train a student model having different image class set from a teacher model, thereby leading that the process time to build the student model is expected to be largely cut down). As per clam 5 Fukuda and Yang disclose the method as in claim 1, Fukuda discloses wherein the set of weights are associated with distillation losses of the plurality of machine learning models during the knowledge distillation to control how much each of the plurality of machine learning models contribute to the unified model ( [0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. ). As per claim 6. Fukuda and Yang disclose the method as in claim 1,Fukuda discloses wherein the one or more constraint parameters comprise a compression ratio( 075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130.). As per claim 7. Fukuda and Yang disclose the method as in claim 1,Fukuda discloses wherein the one or more constraint parameters comprise an accuracy threshold (075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. [0076] In a particular embodiment, a cost function used for training the student model 130 is represented as follow: [00002]ℒ(θ)=-.Math.i.Math.qi.Math..Math.log.Math..Math.pi where q.sub.i represents the soft label determined by the confusion matrix 106 for each student side class i, which works as a pseudo label, and p.sub.i represents output probability for each student side class i. In a particular embodiment, the hard label and the soft labels are used alternately to update the parameters of the student model 130 during the training process). As per claim 8. Fukuda and Yang disclose the method as in claim 1, Fukuda discloses wherein the one or more constraint parameters comprise a resource constraint of the particular node (075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130. ). As per claim 9. Fukuda and Yang disclose the method as in claim 1, Fukuda discloses further comprising; providing, to the user interface, size metrics for the plurality of machine learning models and for the unified model ( [0078] Also, it is described that a feature vector 102a that is same as that fed into the teacher model 120 is fed into the student model 130 during the training process. However, the input feature vector to be fed into the student model 130 may not be necessary to be same as that fed into the teacher model 120. In a particular embodiment, the input layer 122 of the teacher model 120 may be different from the input layer 132 of the student model 130 in sizes (i.e., the number of the frames) and acoustic features. Thus, a feature vector that shares the same central frame with a feature vector for the teacher model 120 and that originates from the same speech data as that generates this feature vector for the teacher model 120 may be fed into the student model 130 during the training process.). As per claim 10. Fukuda and Yang disclose the method as in claim 1, Yang discloses further comprising: providing, to the user interface, a performance metric for the unified model ( fig.2, 240, discloses system computes a distillation loss( i.e. a unified model) between that student model and each teacher model based on the tag predictions from each model and col 2, lines 55-62 The result is a unified named-entity recognition model and col 5, lines 43-60 The system aggregates the distillation losses of each of the student-teacher model pairs to compute an aggregate distillation loss (step 250). The system computes an overall loss as function of the aggregate distillation loss (step 260). In certain embodiments, the overall loss may be equal to the aggregate distillation loss. In other embodiments, it may also include other losses, such as a student loss or a contrastive representation distillation (CRD) loss, as described below with respect to FIGS. 3A-3B and 4A-4B. The system repeats steps 230-260 for a number of iterations, adjusting the parameters of the student model with each iteration to reduce the overall loss (step 270). The steps may be repeated for a fixed number of iterations or until convergence is achieved. The result is a unified named-entity recognition model (i.e., the student) with the collective predictive capabilities of the teacher models without the need for the annotated training data used to train the teacher models. ). As per claim 11. Fukuda discloses an apparatus, comprising: a network interface to communicate with a computer network; a processor coupled to the network interface and configured to execute one or more processes; and a memory configured to store a process that is executed by the processor (FIG. 10, the computer system 10 is shown in the form of a general-purpose computing device. The components of the computer system 10 may include, but are not limited to, a processor (or processing circuitry) 12 and a memory 16 coupled to the processor 12 by a bus including a memory bus or memory controller, and a processor or local bus using any of a variety of bus architectures. ), the process when executed configured to: receiving, at a device and via a user interface, one or more constraint parameters for each of a plurality of machine learning models that perform different analytics tasks( 0064 preparing the teacher model 120 means making the teacher model 120 available by reading the teacher model 120 onto a memory space of the local computer system; or establishing a connection with the teacher model 120 that operates on a remote computer system such that the training input can be fed into the teacher model 120 and a result for the training input can be received from the teacher model 120. ); computing, by the device and based on the one or more constraint parameters, a set of weights for the plurality of machine learning models ([0069] At step S104, the processing circuitry may create a confusion matrix 106 based on the alignments of the phoneme classes for the respective frames. Since it has been described with reference to FIGS. 2-4, a detailed description about the way of creating the confusion matrix 106 is omitted here. 0080, the soft labels q.sub.i after the softmax computation 142 are compared with the output p.sub.i after the softmax computation 138 to encourage the posterior probabilities of the student model 130 close to those of the teacher model 120. Comparing value after the softmax computation is preferable. However, in other embodiment, comparing the soft labels before the softmax computation 142 with the output before the softmax computation 138 may not be excluded. 0077] In one embodiment, all of the soft labels calculated for each training input are used to train the student model 130. Alternatively, in other embodiment, merely at least a part of the set of the soft labels calculated for each training input is used to train the student model 130. For example, posterior probabilities of top K most likely class labels in q.sub.i are used to train the student model 130 after the top K class labels from the teacher model 120 are normalized so that the sum of the top K equals to 1. This normalization may be performed after the softmax computation. ); generating, by the device, a unified model by performing knowledge distillation on the plurality of machine learning models using the set of weights ([0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130. FIG. 6 further describes a way of comparing the soft labels q.sub.i with the outputs p.sub.i of the student model 130 during the training of the knowledge distillation. [0081] After performing the training process at S108, the process may proceed to step S109 and end at the step S109. The parameters of the student model 130, which may include weights between each units and biases of each unit, are optimized during the training of the knowledge distillation process so as to classify the input correctly. ); and generating including performing knowledge distillation on the plurality of machine learning models using the set of weights(0023] FIG. 1 illustrates a block diagram of a knowledge distillation system 100 for the speech recognition. As shown in FIG. 1, the knowledge distillation system 100 includes a training data pool 102 that stores a collection of training data; a training module 110 that performs a training process based on knowledge distillation technique; a teacher model 120 that produces data for generating soft labels for the knowledge distillation from training data stored in the training data pool 102; and a student model 130 under training by the training module 110. And [0051]/0060/0081 In a preferable embodiment, the soft label convertor 140 uses an output obtained for a teacher side class that is frequently observed in the collection together with the corresponding student side member. In a further preferable embodiment, a most frequently observed class is mapped to the corresponding student side member, and the output for this teacher side class is used for calculating a soft label for the corresponding student side member by using softmax function. However, in other embodiments, multiple outputs corresponding to multiple teacher side classes that are frequently observed in the collection together with the corresponding student side member may be used for calculating the soft label by weighted or unweighted average. ) Fukuda does not disclose generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models, the generating including performing knowledge distillation on the plurality of machine learning models using the set of weights; and deploying, by the device, the unified model for execution by a particular node in a network. However, Yang discloses generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models, the generating including performing knowledge distillation on the plurality of machine learning models using the set of weights ( fig.2, 240, discloses system computes a distillation loss( i.e. a unified model) between that student model and each teacher model based on the tag predictions from each model and col 2, 55-62 The result is a unified named-entity recognition model); and deploying, by the device, the unified model for execution by a particular node in a network(col 7, lines 58-60 The methods described herein are embodied in software and performed by a computer system (comprising one or more computing devices) executing the software. So those method can be deployed to oner or more computers by performed by the one or more computers ). Fukuda and Yang are both considered to be analogous to the claimed invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fukuda to incorporate the teachings of Yang and provided a NER model on datasets with multiple tag sets. Doing so would provide datasets with different tags typically originate from different domains, thereby increasing overcome domain mismatch and language variations (col 2, lines 17-18). As per claim 13. Fukuda and Yang discloses The apparatus as in claim 11, Yang discloses wherein the different analytics tasks comprise video analytics tasks ( col 7, lines 15-32 obtaining the student NER model from a plurality of teacher NER models. Steps 410-460 are the same as steps 310-360. However, in addition to computing an aggregate distillation loss and a student loss, the system also computes a contrastive representation distillation (CRD) loss. A contrastive representation distillation loss between the student model and a teacher model is a measure of the differences between vector representations generated by the two models for the input data sequences as part of the prediction process. Minimizing this loss enables the student to distill domain-invariant knowledge from the teacher models and enables the student model to produce vector representations of input data sequences that are domain insensitive (or at least less domain sensitive than would otherwise be). This in turn enables the student model to adapt to and to perform better on another domain (i.e., on a domain that is different from the one on which it was trained)). As per claim 14. Fukuda and Yang disclose The apparatus as in claim 13, Fukuda discloses wherein the video analytics tasks comprise one or more of: pose estimation, object detection, semantic segmentation, or instance segmentation (([0111] In the image recognition system, the components of the output layers may also be different depending on environments. However, according to one or more embodiments of the present invention, it is possible to train a student model having different image class set from a teacher model, thereby leading that the process time to build the student model is expected to be largely cut down ). As per claim 15. Fukuda and Yang discloses The apparatus as in claim 11, Fukuda wherein the set of weights are associated with distillation losses of the plurality of machine learning models during the knowledge distillation to control how much each of the plurality of machine learning models contribute to the unified model ([0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set ). As per claim 16. Fukuda and Yang discloses The apparatus as in claim 11, Fukuda wherein the one or more constraint parameters comprise a compression ratio(0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130 ). As per claim 17. Fukuda and Yang discloses The apparatus as in claim 11, Fukuda wherein the one or more constraint parameters comprise an accuracy threshold (0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. [0076] In a particular embodiment, a cost function used for training the student model 130 is represented as follow: [00002]ℒ(θ)=-.Math.i.Math.qi.Math..Math.log.Math..Math.pi where q.sub.i represents the soft label determined by the confusion matrix 106 for each student side class i, which works as a pseudo label, and p.sub.i represents output probability for each student side class). As per claim 18. Fukuda and Yang discloses The apparatus as in claim 11,Fukuda discloses wherein the one or more constraint parameters comprise a resource constraint of the particular node ( 0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130). As per claim 19. Fukuda and Yang discloses The apparatus as in claim 11,Fukuda wherein the process when executed is further configured to: provide, to the user interface, size metrics for the plurality of machine learning models and for the unified model( [0078] Also, it is described that a feature vector 102a that is same as that fed into the teacher model 120 is fed into the student model 130 during the training process. However, the input feature vector to be fed into the student model 130 may not be necessary to be same as that fed into the teacher model 120. In a particular embodiment, the input layer 122 of the teacher model 120 may be different from the input layer 132 of the student model 130 in sizes (i.e., the number of the frames) and acoustic features. Thus, a feature vector that shares the same central frame with a feature vector for the teacher model 120 and that originates from the same speech data as that generates this feature vector for the teacher model 120 may be fed into the student model 130 during the training process). As per claim 20. Fukuda discloses a tangible, non-transitory, computer-readable medium storing program instructions that cause a device to execute a process (FIG. 10, the computer system 10 is shown in the form of a general-purpose computing device. The components of the computer system 10 may include, but are not limited to, a processor (or processing circuitry) 12 and a memory 16 coupled to the processor 12 by a bus including a memory bus or memory controller, and a processor or local bus using any of a variety of bus architectures.)comprising : receiving, at a device and via a user interface, one or more constraint parameters for each of a plurality of machine learning models that perform different analytics tasks( 0064 preparing the teacher model 120 means making the teacher model 120 available by reading the teacher model 120 onto a memory space of the local computer system; or establishing a connection with the teacher model 120 that operates on a remote computer system such that the training input can be fed into the teacher model 120 and a result for the training input can be received from the teacher model 120. ); computing, by the device and based on the one or more constraint parameters, a set of weights for the plurality of machine learning models ([0069] At step S104, the processing circuitry may create a confusion matrix 106 based on the alignments of the phoneme classes for the respective frames. Since it has been described with reference to FIGS. 2-4, a detailed description about the way of creating the confusion matrix 106 is omitted here. 0080, the soft labels q.sub.i after the softmax computation 142 are compared with the output p.sub.i after the softmax computation 138 to encourage the posterior probabilities of the student model 130 close to those of the teacher model 120. Comparing value after the softmax computation is preferable. However, in other embodiment, comparing the soft labels before the softmax computation 142 with the output before the softmax computation 138 may not be excluded. 0077] In one embodiment, all of the soft labels calculated for each training input are used to train the student model 130. Alternatively, in other embodiment, merely at least a part of the set of the soft labels calculated for each training input is used to train the student model 130. For example, posterior probabilities of top K most likely class labels in q.sub.i are used to train the student model 130 after the top K class labels from the teacher model 120 are normalized so that the sum of the top K equals to 1. This normalization may be performed after the softmax computation); generating, by the device, a unified model by performing knowledge distillation on the plurality of machine learning models using the set of weights ([0075] At step S108, the processing circuitry may train the student model 130 by the knowledge distillation technique using the soft labels and optionally hard labels for each input feature vector. During the training, the processing circuitry may pick a feature vector 102a from the training data pool 102 and feed the vector 102a into the student model 130 to obtain a set of outputs p.sub.i for the student side class set. The outputs p.sub.i (i=1, . . . , M) obtained at step S108 are probabilities after the softmax computation, as illustrated in FIG. 6. The soft labels q.sub.i (1, . . . , M) obtained from the teacher model 120 are compared with the output p.sub.i (1, . . . , M) obtained from the student model 130. FIG. 6 further describes a way of comparing the soft labels q.sub.i with the outputs p.sub.i of the student model 130 during the training of the knowledge distillation. [0081] After performing the training process at S108, the process may proceed to step S109 and end at the step S109. The parameters of the student model 130, which may include weights between each units and biases of each unit, are optimized during the training of the knowledge distillation process so as to classify the input correctly. ); and Fukuda does not disclose deploying, by the device, the unified model for execution by a particular node in a network. generating including performing knowledge distillation on the plurality of machine learning models using the set of weights(0023] FIG. 1 illustrates a block diagram of a knowledge distillation system 100 for the speech recognition. As shown in FIG. 1, the knowledge distillation system 100 includes a training data pool 102 that stores a collection of training data; a training module 110 that performs a training process based on knowledge distillation technique; a teacher model 120 that produces data for generating soft labels for the knowledge distillation from training data stored in the training data pool 102; and a student model 130 under training by the training module 110. And [0051]/0060/0081 In a preferable embodiment, the soft label convertor 140 uses an output obtained for a teacher side class that is frequently observed in the collection together with the corresponding student side member. In a further preferable embodiment, a most frequently observed class is mapped to the corresponding student side member, and the output for this teacher side class is used for calculating a soft label for the corresponding student side member by using softmax function. However, in other embodiments, multiple outputs corresponding to multiple teacher side classes that are frequently observed in the collection together with the corresponding student side member may be used for calculating the soft label by weighted or unweighted average. ) Fukuda does not disclose generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models, the generating including performing knowledge distillation on the plurality of machine learning models using the set of weights; and deploying, by the device, the unified model for execution by a particular node in a network. However, Yang discloses generating, by the device, a unified model configured to perform the different analytics tasks of the plurality of machine learning models, the generating including performing knowledge distillation on the plurality of machine learning models using the set of weights ( fig.2, 240, discloses system computes a distillation loss( i.e. a unified model) between that student model and each teacher model based on the tag predictions from each model and col 2, 55-62 The result is a unified named-entity recognition model); and deploying, by the device, the unified model for execution by a particular node in a network(col 7, lines 58-60 The methods described herein are embodied in software and performed by a computer system (comprising one or more computing devices) executing the software. So, those methods can be deployed to oner or more computers by performed by the one or more computers). Fukuda and Yang are both considered to be analogous to the claimed invention because they are in the same field of machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fukuda to incorporate the teachings of Yang and provided a NER model on datasets with multiple tag sets. Doing so would provide datasets with different tags typically originate from different domains, thereby increasing overcome domain mismatch and language variations (col 2, lines 17-18). As per claim 21. Fukuda and Yang discloses The tangible, non-transitory, computer-readable medium as in claim 20, wherein the set of weights are associated with distillation losses of the plurality of machine learning models during the knowledge distillation to control how much each of the plurality of machine learning models contribute to the unified model(Yang col 7, lines 14-32 FIGS. 4A-4B illustrate a further embodiment of the method for obtaining the student NER model from a plurality of teacher NER models. Steps 410-460 are the same as steps 310-360. However, in addition to computing an aggregate distillation loss and a student loss, the system also computes a contrastive representation distillation (CRD) loss. A contrastive representation distillation loss between the student model and a teacher model is a measure of the differences between vector representations generated by the two models for the input data sequences as part of the prediction process. Minimizing this loss enables the student to distill domain-invariant knowledge from the teacher models and enables the student model to produce vector representations of input data sequences that are domain insensitive (or at least less domain sensitive than would otherwise be). This in turn enables the student model to adapt to and to perform better on another domain (i.e., on a domain that is different from the one on which it was trained)). As per claim 22. Fukuda and Yang discloses The tangible, non-transitory, computer-readable medium as in claim 20, wherein the one or more constraint parameters comprise an accuracy threshold( Yang col 5, lines 43-60 aggregates the distillation losses of each of the student-teacher model pairs to compute an aggregate distillation loss (step 250). The system computes an overall loss as function of the aggregate distillation loss (step 260). In certain embodiments, the overall loss may be equal to the aggregate distillation loss. In other embodiments, it may also include other losses, such as a student loss or a contrastive representation distillation (CRD) loss, as described below with respect to FIGS. 3A-3B and 4A-4B. The system repeats steps 230-260 for a number of iterations, adjusting the parameters of the student model with each iteration to reduce the overall loss (step 270). The steps may be repeated for a fixed number of iterations or until convergence is achieved. The result is a unified named-entity recognition model (i.e., the student) with the collective predictive capabilities of the teacher models without the need for the annotated training data used to train the teacher models. And col 3, lines 3-7 the overall loss is a function of the aggregate distillation loss and a student loss. The student loss is computed based on the student model's tag predictions and ground truth hard labels for data sequences in the input set. This increases the accuracy of the student model. And col 6, lines 50-67 FIGS. 3A-3B illustrate a further embodiment of the method for obtaining the student NER model from a plurality of teacher NER models. Steps 310-350 are the same 210-250 in FIG. 2. However, in addition to calculating an aggregate distillation loss, which is a measure of the difference in predictions between the student model and each of the teacher models, the system also calculates a “student loss,” which is a measure of the difference in the student model's prediction and ground truth hard labels for the input data sequences (step 360). Including a student loss in the overall loss increases the accuracy of the student model. In one embodiment, the student loss, notated as [AltContent: rect].sub.NLL herein, is calculated by replacing the soft target label q with the ground truth label in Equation 1. The system computes an overall loss as a function of the aggregate distillation loss across all student-teacher pairs and the student loss (step 370)). Conclusion Applicant's submission of an information disclosure statement under 37 CFR 1.97(c) with the timing fee set forth in 37 CFR 1.17(p) on 09/07/2023 prompted the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 609.04(b). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABU S SHOLEMAN whose telephone number is (571)270-7314. The examiner can normally be reached EST: 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JORGE ORTIZ CRIADO can be reached at 571-272-7624. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ABU S SHOLEMAN/Primary Examiner, Art Unit 2496
Read full office action

Prosecution Timeline

Sep 07, 2023
Application Filed
Apr 08, 2026
Non-Final Rejection mailed — §103
Jul 01, 2026
Interview Requested
Jul 07, 2026
Applicant Interview (Telephonic)
Jul 07, 2026
Examiner Interview Summary
Jul 07, 2026
Response Filed
Aug 27, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12739122
USER AUTHENTICATION FOR A RESOURCE USING CONTEXT BASED ENCRYPTION OF AUTHENTICATION TOKENS
2y 3m to grant Granted Sep 15, 2026
Patent 12730875
SYSTEMS AND METHODS FOR SOFTWARE AUTHENTICATION WITHOUT PROVIDING FULL SOFTWARE DISCLOSURE TO A SOFTWARE NOTARIZER
2y 9m to grant Granted Sep 08, 2026
Patent 12689513
LEVERAGING USER'S VIRTUAL INTERACTIONS TO INFLUENCE PREFERRED MESSAGE COMMUNICATION TIMING
2y 7m to grant Granted Jul 21, 2026
Patent 12683784
DATA ANALYSIS SYSTEMS AND METHODS FOR DETECTING ANOMALIES IN TOKENIZED DATASETS
2y 11m to grant Granted Jul 14, 2026
Patent 12659742
ENSURING SECURE ATTACHMENT IN SIZE CONSTRAINED AUTHENTICATION PROTOCOLS
5y 0m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+27.6%)
3y 0m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 796 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month