DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-6, 8-9, 11, 16, 18-22, and 26-35 are presented for examination.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 11, 2026, has been entered.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on April 3, 2026, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 26, 27, 33, and 34 are objected to because of the following informalities:
Claim 26: “outputting the second ML model will operate on the target hardware without the pruned one or more parameters” is grammatically incorrect; Examiner suggests “outputting the second ML model, such that the second ML model will operate on the target hardware without the pruned one or more parameters”
Claim 33: “the subnet ML model” should read “the second ML model”
Claim 34: “the subnet ML model” should read “the second ML model”
Claim 27 is objected to due to dependency on objected-to claim 26.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 20 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “similar” in claim 20 is a relative term which renders the claim indefinite. The term “similar” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The limitation “a spatial attention map that is similar to a spatial attention map of the first ML model” is rendered indefinite by the use of the term “similar”.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-6, 8-9, 11, 16, 18-22, and 26-35 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance (“2019 PEG”).
Claim 1
Step 1: The claim recites an apparatus comprising interface circuitry, instructions, and at least one programmable circuit, and thus is directed to the statutory category of machines.
Step 2A Prong 1: The claim recites:
“distill knowledge of a supernet ML model to generate a subnet ML model during a single ML training epoch, the supernet ML model having one or more feature maps describing activation patterns indicative of an output of the supernet ML model”; This limitation encompasses mentally distilling knowledge of a supernet ML model, such as by mentally selecting weights/parameters for a subnet ML model based on the weights/parameters of the supernet ML model.
“during the same single ML training epoch as the distilling: calculate a training loss using a loss function, the loss function including a distillation loss, a task loss, and an attention transfer loss, wherein: the distillation loss describes a divergence between predictions of the supernet ML model and predictions of the subnet ML model, the task loss describes a difference between predictions of the subnet ML model and labels of a training dataset, and the attention transfer loss describes a difference between the one or more feature maps of the supernet ML model and one or more corresponding feature maps of the subnet ML model”; This limitation encompasses a mathematical concept of calculating a training loss using a loss function.
“prune one or more parameters from the subnet ML model based on the training loss”; This limitation encompasses mentally pruning one or more parameters from the subnet ML model based on the training loss, such as by mentally using the training loss to determine which model parameters to remove.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “An apparatus for sparse distillation of machine learning (ML) models, the apparatus comprising: interface circuitry; instructions; and at least one programmable circuit to be programmed by the instructions,” however this limitation amounts to mere instructions to apply the exception on a generic computer (MPEP 2106.05(f)). The claim further recites “after pruning the one or more parameters from the subnet ML model, output the subnet ML model,” however this limitation amounts to the insignificant extra-solution activity of mere data outputting (MPEP 2106.05(g)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “output the subnet ML model” limitation, in addition to reciting insignificant extra-solution activity, is also directed to the well understood, routine, and conventional activity of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i) OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). Otherwise, the analysis at this step mirrors that of Step 2A, Prong 2. As an ordered whole, the claim is directed to an abstract idea of distilling knowledge of a supernet ML model to generate a subnet model, calculating a training loss using a loss function, and pruning one or more parameters from the subnet ML model based on the training loss. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim 2
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites the same judicial exception as claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein one or more of the at least one programmable circuit is to distill the knowledge and prune the one or more parameters simultaneously,” however this limitation amounts to mere instructions to apply a judicial exception on a generic computer (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 3
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“perform a single pass over a training dataset during the single ML training epoch”; This limitation encompasses mentally performing a single pass over a training dataset during a single ML training epoch, such as by mentally perusing the training dataset.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Claim 4
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“…extract the knowledge to be distilled into the subnet ML model”; This limitation encompasses mentally extracting the knowledge to be distilled into the subnet ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “train the supernet ML model using a training dataset” and “train the subnet using the training datset and using the extracted knowledge to guide the training of the subnet,” however this amounts to generally linking the judicial exception to the technological environment of model training (MPEP 2106.05(h)). The claim further recites “operate the supernet…” however, this limitation amounts to mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 5
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“wherein the knowledge includes both logits and feature maps extracted from the supernet ML model”; This limitation encompasses mentally extracting logits and feature maps from the supernet ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 4.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 4.
Claim 6
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“transfer the knowledge from the supernet ML model to the subnet ML model”; This limitation encompasses mentally transferring the knowledge from the supernet ML model to the subnet ML model, such as by mentally selecting weights/parameters for the subnet ML model based on the logits and feature maps extracted from the supernet ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 4.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 4.
Claim 8
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“generate, based on input data, a queries matrix, a values matrix, and a keys matrix, wherein the queries matrix, the values matrix, and the keys matrix include parameters to be pruned”; This limitation encompasses mentally generating queries, values, and keys matrices based on input data.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “provide the queries matrix, the values matrix, and the keys matrix to the pruning,” however this limitation amounts to the insignificant extra-solution activity of mere data gathering and outputting (MPEP 2106.05(g)).
Step 2B: The claim does not contain significantly more than the judicial exception. The provide the matrices to the pruning limitation, in addition to reciting insignificant extra solution activity, is also directed to the well-understood, routine, and conventional activity of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i) OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)).
Claim 9
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites the same judicial exception as claim 8.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “apply the input data to a parameterized learnable transformation (PLT) to generate the queries matrix, the values matrix, and the keys matrix,” however this limitation amounts to mere instructions to apply the judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 11
Step 1: A process, as above.
Step 2A Prong 1: The claim recites:
“perform an operation on the queries matrix and the keys matrix, wherein the operation is a matrix multiplication operation or a 1 X 1 convolution on the queries matrix and the keys matrix”; This limitation encompasses the mathematical calculation of matrix multiplication or 1 x 1 convolution.
“apply a softmax function to an output of the operation”; This limitation encompasses the mathematical calculation of applying a softmax function.
“generate a self-attention output based on a combination of the values matrix and an output of the softmax function”; This limitation encompasses the mathematical concept of combining values to generate an output.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 8.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 8.
Claim 16
Step 1: The claim recites a non-transitory machine readable storage medium comprising instructions, and thus is directed to the statutory category of articles of manufacture.
Step 2A Prong 1: The claim recites:
“during training of the first and second ML models, distill knowledge of the first ML model into the second ML model during one pass over the training dataset”; This limitation encompasses mentally distilling knowledge of the first ML model into the second ML model, such as by mentally selecting weights/parameters for a second ML model based on the weights/parameters of the first ML model.
“during the same one pass over the training dataset as the distill: calculate a training loss using a loss function, the loss function including a distillation loss, a task loss, and an attention transfer loss, wherein: the distillation loss describes a divergence between predictions of the first ML model and predictions of the second ML model, the task loss describes a difference between predictions of the second ML model and labels of the training dataset, and the attention transfer loss describes a difference between the one or more feature maps of the first ML model and one or more corresponding feature maps of the second ML model”; This limitation encompasses a mathematical concept of calculating a training loss using a loss function.
“prune one or more parameters from the second ML model based on the training loss”; This limitation encompasses mentally pruning one or more parameters from the second ML model based on the training loss, such as by mentally using the training loss to determine which model parameters to remove.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “A non-transitory machine readable storage medium comprising instructions,” however this limitation amounts to mere instructions to apply the exception on a generic computer (MPEP 2106.05(f)). The claim also further recites “provide a training dataset to a first machine learning (ML) model and a second ML model, the second ML model having fewer parameters than the first ML model, the first ML model having one or more feature maps describing activation patterns indicative of an output of the first ML model,” and “output the second ML model without the pruned one or more parameters,” however, these limitations amount to the insignificant extra solution activity of mere data gathering and outputting (MPEP 2106.05(g)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “provide a training dataset” and “output the second ML model” limitations, in addition to reciting insignificant extra-solution activity, are also directed to the well understood, routine, and conventional activity of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i) OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). Otherwise, the analysis at this step mirrors that of Step 2A, Prong 2. As an ordered whole, the claim is directed to an abstract idea of distilling knowledge of a first ML model into a second ML model, calculating a training loss using a loss function, and pruning one or more parameters from the second ML model based on the training loss. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim 18
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“…extract the knowledge to be distilled into the second ML model”; This limitation encompasses mentally extracting the knowledge to be distilled into the second ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “train the first ML model using the training dataset” and “train the second ML model using the training dataset and using the extracted knowledge to guide the training of the second ML model,” however this amounts to generally linking the judicial exception to the technological environment of model training (MPEP 2106.05(h)). The claim further recites “operate the first ML model…” however, this limitation amounts to mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 19
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“extract logits and feature maps from the first ML model, wherein the knowledge includes both the extracted logits and the extracted feature maps”; This limitation encompasses mentally extracting logits and feature maps from the first ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 18.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 18.
Claim 20
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“…transfer the knowledge from the first ML model to the second ML model such that the second ML model includes a spatial attention map that is similar to a spatial attention map of the first ML model”; This limitation encompasses mentally transferring the knowledge from the first ML model to the second ML model, such as by mentally selecting weights/parameters for the second ML model based on the logits and feature maps extracted from the first ML model, such that the second ML model includes a spatial attention map that is similar to the first ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “operate an attention transfer distillation algorithm to…” however this limitation amounts to mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 21
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“generate, based on the training dataset, a queries matrix, a values matrix, and a keys matrix, wherein the queries matrix, the values matrix, and the keys matrix include parameters to be pruned”; This limitation encompasses mentally generating queries, values, and keys matrices based on input data.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “provide the queries matrix, the values matrix, and the keys matrix to the pruning,” however this limitation amounts to the insignificant extra-solution activity of mere data gathering and outputting (MPEP 2106.05(g)).
Step 2B: The claim does not contain significantly more than the judicial exception. The provide the matrices to the pruning limitation, in addition to reciting insignificant extra solution activity, is also directed to the well-understood, routine, and conventional activity of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i) OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)).
Claim 22
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“perform an operation on the queries matrix and the keys matrix, wherein the operation is a matrix multiplication operation or a 1 X 1 convolution on the queries matrix and the keys matrix”; This limitation encompasses the mathematical calculation of matrix multiplication or 1 x 1 convolution.
“apply a softmax function to an output of the operation”; This limitation encompasses the mathematical calculation of applying a softmax function.
“generate self-attention output based on a combination of the values matrix and an output of the softmax function”; This limitation encompasses the mathematical concept of combining values to generate an output.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “apply parts of the training dataset to a parameterized learnable transformation (PLT) to generate the queries matrix, the values matrix, and the keys matrix,” however this limitation amounts to mere instructions to apply the judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 26
Step 1: The claim recites a method and thus is directed to the statutory category of processes.
Step 2A Prong 1: The claim recites:
“during training of the first and second ML models, distilling knowledge of the first ML model into the second ML model during one pass over the training dataset”; This limitation encompasses mentally distilling knowledge of the first ML model into the second ML model, such as by mentally selecting weights/parameters for a second ML model based on the weights/parameters of the first ML model.
“during the same one pass over the training dataset as the distilling: calculating a training loss using a loss function, the loss function including a distillation loss, a task loss, and an attention transfer loss, wherein: the distillation loss describes a divergence between predictions of the first ML model and predictions of the second ML model, the task loss describes a difference between predictions of the second ML model and labels of the training dataset, and the attention transfer loss describes a difference between the one or more feature maps of the first ML model and one or more corresponding feature maps of the second ML model”; This limitation encompasses a mathematical concept of calculating a training loss using a loss function.
“pruning one or more parameters from the second ML model based on an identified parameter of target hardware, the training loss”; This limitation encompasses mentally pruning one or more parameters from the second ML model based on an identified parameter of target hardware and the training loss, such as by mentally using the training loss and the hardware parameter to determine which model parameters to remove.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “providing a training dataset to a first machine learning (ML) model and a second ML model, the second ML model having fewer parameters than the first ML model, the first ML model having one or more feature maps describing activation patterns indicative of an output of the first ML model,” and “outputting the second ML model will operate on the target hardware without the pruned one or more parameters,” however, these limitations amount to the insignificant extra solution activity of mere data gathering and outputting (MPEP 2106.05(g)).
Step 2B: The claim does not contain significantly more than the judicial exception. The “providing a training dataset” and “outputting the second ML model” limitations, in addition to reciting insignificant extra-solution activity, are also directed to the well understood, routine, and conventional activity of receiving or transmitting data over a network (MPEP 2106.05(d)(II)(i) OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network)). Otherwise, the analysis at this step mirrors that of Step 2A, Prong 2. As an ordered whole, the claim is directed to an abstract idea of distilling knowledge of a first ML model into a second ML model, calculating a training loss using a loss function, and pruning the second ML model based on an identified parameter of target hardware and the training loss. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim 27
Step 1: A process, as above.
Step 2A Prong 1: The claim recites:
“…extract the knowledge to be distilled into the second ML model”; This limitation encompasses mentally extracting the knowledge to be distilled into the second ML model.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “training the first ML model using the training dataset” and “training the second ML model using the training dataset and using the extracted knowledge to guide the training of the second ML model,” however this amounts to generally linking the judicial exception to the technological environment of model training (MPEP 2106.05(h)). The claim further recites “operating the first ML model…” however, this limitation amounts to mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 28
Step 1: A machine, as claim 1 above.
Step 2A Prong 1: The claim recites:
“wherein the subnet ML model has a plurality of layers, each layer having a plurality of weights, wherein pruning one or more parameters includes setting one or more weights of a layer of the subnet ML model to zero”; This limitation encompasses mentally setting one or more weights of a layer of the subnet ML model to zero.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Claim 29
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“update, for a plurality of batches within the single ML training epoch and based on the training loss, weights of the subnet ML model and a momentum contribution associated with each layer of a plurality of layers”; This limitation encompasses mentally updating weights of the subnet ML model and a momentum contribution associated with each layer of a plurality of layers based on the training loss.
“the prune of the one or more parameters from the subnet ML model includes: after processing the plurality of batches, deactivating a first fraction of active weights in each layer of the plurality of layers; and reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer”; This limitation encompasses mentally deactivating a first fraction of active weights in each layer of the plurality of layers, such as by mentally setting their values to zero, and mentally reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer, such as by mentally setting their values to a number greater than zero.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Claim 30
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“wherein the loss function combines the distillation loss, the task loss, and the attention transfer loss as a weighted sum, wherein: the task loss is weighted by a first hyperparameter; the distillation loss is weighted by a complement of the first hyperparameter; and the attention transfer loss is weighted by a second hyperparameter”; This limitation encompasses the mathematical concept of a weighted sum.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Claim 31
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“wherein the weighted sum is calculated by: applying the first hyperparameter as a weight to the task loss; applying a complement of the first hyperparameter as a weight to the distillation loss; and applying the second hyperparameter as a weight to the attention transfer loss”; This limitation encompasses the mathematical concept of a weighted sum.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 1.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 1.
Claim 32
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
“generate a queries matrix, a keys matrix, and a values matrix based on input data”; This limitation encompasses mentally generating a queries matrix, a keys matrix, and a values matrix based on input data.
“produce a self-attention output based on the queries matrix, the keys matrix, and the values matrix”; This limitation encompasses mentally producing a self-attention output based on the queries matrix, the keys matrix, and the values matrix.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites “wherein the subnet ML model comprises a transformer architecture that includes a self-attention mechanism, the self-attention mechanism configured to…” however this limitation amounts to mere instructions to apply a judicial exception on a generic computer programmed with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong 2 above.
Claim 33
Step 1: An article of manufacture, as claim 16 above.
Step 2A Prong 1: The claim recites:
“wherein the subnet ML model has a plurality of layers, each layer having a plurality of weights, wherein pruning one or more parameters includes setting one or more weights of a layer of the subnet ML model to zero”; This limitation encompasses mentally setting one or more weights of a layer of the subnet ML model to zero.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 16.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 16.
Claim 34
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“update, for a plurality of batches within the single ML training epoch and based on the training loss, weights of the subnet ML model and a momentum contribution associated with each layer of a plurality of layers”; This limitation encompasses mentally updating weights of the subnet ML model and a momentum contribution associated with each layer of a plurality of layers based on the training loss.
“the prune of the one or more parameters from the subnet ML model includes: after processing the plurality of batches, deactivating a first fraction of active weights in each layer of the plurality of layers; and reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer”; This limitation encompasses mentally deactivating a first fraction of active weights in each layer of the plurality of layers, such as by mentally setting their values to zero, and mentally reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer, such as by mentally setting their values to a number greater than zero.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 16.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 16.
Claim 35
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
“wherein the loss function combines the distillation loss, the task loss, and the attention transfer loss as a weighted sum, wherein: the task loss is weighted by a first hyperparameter; the distillation loss is weighted by a complement of the first hyperparameter; and the attention transfer loss is weighted by a second hyperparameter”; This limitation encompasses the mathematical concept of a weighted sum.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. No further additional elements are recited, see analysis of claim 16.
Step 2B: The claim does not contain significantly more than the judicial exception. No further additional elements are recited, see analysis of claim 16.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-6, 8-9, 11, 16, 18-22, 28, and 32-33 are rejected under 35 U.S.C. 103 as being unpatentable over Cui et al. (“Joint structured pruning and dense knowledge distillation for efficient transformer model compression”) (“Cui”) in view of Zagoruyko et al. (“Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer”) (“Zagoruyko”), further in view of Zhang et al. (“Student Network Learning via Evolutionary Knowledge Distillation”) (“Zhang”).
Regarding claim 1, Cui discloses “An apparatus for sparse distillation of machine learning (ML) models, the apparatus comprising:
interface circuitry;
instructions; and
at least one programmable circuit to be programmed by the instructions to (Cui, 6.2.1 Implementation details: “The experiments are conducted with PyTorch framework on GeForce GTX 1080Ti GPU”):
distill knowledge of a supernet ML model to generate a subnet ML model during a single ML training epoch… (Cui, 5.1 Dense knowledge distillation: “Knowledge distillation aims to transfer the useful knowledge of a large teacher network to a lightweight student network with fewer layers…We assume that the teacher model T has M Transformer layers T = {T1, … ,TM} and the student S contains N layers as S = {S1, …, SN}… a many-to-one layer mapping strategy is introduced to encourage each student layer to flexibly learn from multiple teacher layers to obtain the rich linguistic knowledge in different levels simultaneously (as illustrated in Fig.4b)” and 6.3.3 Results: “Our BERT6-JMC overcomes the limitation of them through performing the pruning and distillation in one training process. The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4”; Examiner notes that teacher model T corresponds to a supernet ML model and student model S corresponds to a subnet ML model); and
during the same single ML training epoch as the distilling (Cui, 6.3.3 Results: “Our BERT6-JMC overcomes the limitation of them through performing the pruning and distillation in one training process. The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4” and 5.2 Integrating pruning and distillation: “Different from the Two-stage compression (Fig. 1c) of continuing to apply pruning method on the learned distilled model, our JMC performs pruning and distillation simultaneously, which allows the joint optimization of the two compression techniques”):
calculate a training loss using a loss function, the loss function including a distillation loss, a task loss… (Cui, 5.2 Integrating pruning and distillation: “The final loss function of this JMC approach is formulated by the fusion of the original loss for Transformer Lc, the pruning objective function LP, and the dense knowledge distillation objective LKD: LJMC = Lc + αLp + LKD”; Examiner notes that LJMC corresponds to a training loss, LKD corresponds to a distillation loss, and Lc corresponds to a task loss) wherein:
the distillation loss describes a divergence between predictions of the supernet ML model and predictions of the subnet ML model (Cui, 5.1 Dense Knowledge Distillation: “Prediction Layer Distillation: Moreover, in the prediction layer, the student network is also required to learn the soft labels provided by the teacher for receiving more information about the prediction distribution of the training samples. We minimize the distance between the prediction of the teacher and that of the student as follows: Lpred = -softmax (
z
T
t
e
) ∙ logsoftmax (
z
S
t
e
), where zT and zS represent the probability logits predicted by the teacher and the student model respectively… Finally, we integrate dense Transformer layer distillation and prediction layer distillation together to build the overall objective function for knowledge distillation: LKD = n1Lhidden + n2Lpred”) …
prune one or more parameters from the subnet ML model based on the training loss (Cui, 5.2 Integrating pruning and distillation: “During the joint compression, the proposed structured pruning approach DISP is applied on the student to prune the redundant parameters in this shallow network including the heads in the multi-head attention and weight blocks of intermediate layer in FFN” and Cui, 4.2 Structured pruning for transformers: “Gradual Pruning Optimization In order to encourage the model to be pruned to the target sparsity, a sparsity-aware pruning objective function LP is introduced to impose an explicit control over the entire model. The pruning procedure is optimized with the original loss function Lc for the Transformer model (e.g., cross-entropy loss). In particular, the model is trained to reach the overall target sparsity pt by minimizing the new training objective L as follows: L = Lc + Lp”; Examiner notes that the pruning is based on Lc and Lp, which are part of the joint loss function LJMC); and
after pruning the one or more parameters from the subnet ML model, output the subnet ML model” (Cui, 6.3.1 Training setup: “In the joint compression procedure, we reduce the unimportant parameters in the student based on our pruning approach DISP, and perform the dense knowledge distillation to transfer the knowledge from the teacher to the pruned shallow network at the same time. Such joint compression strategy enables us to obtain a compressed shallow model BERT6-JMC”).
Cui does not appear to explicitly disclose the further limitations of the claim. However, Zagoruyko discloses “the supernet ML model having one or more feature maps describing activation patterns indicative of an output of the supernet ML model” (Zagoruyko, 3.1 Activation-Based Attention Transfer, Paragraph 1: “Let us consider a CNN layer and its corresponding activation tensor A ∈ RC×H×W, which consists of C feature planes with spatial dimensions H×W. An activation-based mapping function F (w.r.t. that layer) takes as input the above 3D tensor A and outputs a spatial attention map, i.e., a flattened 2D tensor defined over the spatial dimensions, or F : RC× H× W →RH× W . To define such a spatial attention mapping function, the implicit assumption that we make in this section is that the absolute value of a hidden neuron activation (that results when the network is evaluated on given input) can be used as an indication about the importance of that neuron w.r.t. the specific input” and 3.1 Activation-Based Attention Transfer, Paragraph 6: “In attention transfer, given the spatial attention maps of a teacher network (computed using any of the above attention mapping functions), the goal is to train a student network that will not only make correct predictions but will also have attentions maps that are similar to those of the teacher” and Figure 1: “(a) An input image and a corresponding spatial attention map of a convolutional network that shows where the network focuses in order to classify the given image”; Examiner notes that the teacher network corresponds to the supernet ML model, and the spatial attention maps correspond to “one or more features maps describing activation patterns indicative of an output of the supernet ML model” as they describe neuron activations in generating an output (e.g. image classification) for a given input of the network) and “the loss function including… an attention transfer loss, wherein: … the attention transfer loss describes a difference between the one or more feature maps of the supernet ML model and one or more corresponding feature maps of the subnet ML model” (3.1 Activation-Based Attention Transfer, Paragraph 8: “Without loss of generality, we assume that transfer losses are placed between student and teacher attention maps of same spatial resolution, but, if needed, attention maps can be interpolated to match their shapes. Let S, T and WS, WT denote student, teacher and their weights correspondingly, and let L(W, x) denote a standard cross entropy loss. Let also I denote the indices of all teacher-student activation layer pairs for which we want to transfer attention maps. Then we can define the following total loss:
PNG
media_image1.png
224
824
media_image1.png
Greyscale
”; Examiner notes that the second term in the loss function corresponds to an attention transfer loss, which describes a difference between the teacher (supernet) attention maps and the student (subnet) attention maps).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cui with the teachings of Zagoruyko such that the supernet ML model has one or more feature maps describing activation patterns indicative of an output of the supernet ML model, and such that the loss function further includes an attention transfer loss, wherein the attention transfer loss describes a difference between the one or more feature maps of the supernet ML model and one or more corresponding feature maps of the subnet ML model, and one would have been motivated to do so. Doing so would significantly improve the performance of a student CNN network by forcing it to mimic the attention maps of a powerful teacher network (see Zagoruyko, Abstract).
Neither Cui nor Zagoruyko appear to explicitly disclose the further limitations of the claim.
However, Zhang discloses “the task loss describes a difference between predictions of… [an] ML model and labels of a training dataset” (Zhang, III. A: “Denote the training set is D
=
{
(
x
i
,
y
i
)
i
=
1
D
, where xi is the ith sample with label yi ∈{1,2,...,m}” and III.B, paragraph 8: “For each classifier, we compute cross entropy loss lce between φ(x;w(c)) and y. In this way, the label y directs each classifier’s probability as possible”; Examiner notes that φ(x;w(c)) corresponds to predictions of an ML model and y corresponds to labels of a training dataset).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the task loss describes a difference between predictions of the subnet ML model and labels of the training dataset, and one would have been motivated to do so. Doing so would improve the student model’s classification performance by directing the model to match ground truth labels (see Zhang III.B, paragraph 8).
Regarding claim 2, the rejection of claim 1 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein one or more of the at least one programmable circuit is to distill the knowledge and prune the one or more parameters simultaneously” (Cui, 5.2 Integrating pruning and distillation: “Different from the Two-stage compression (Fig. 1c) of continuing to apply pruning method on the learned distilled model, our JMC performs pruning and distillation simultaneously, which allows the joint optimization of the two compression techniques”).
Regarding claim 3, the rejection of claim 1 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein one or more of the at least one programmable circuit is to perform a single pass over a training dataset during the single ML training epoch” (Cui, 6.3.3. Results: “The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4”; Examiner notes that one epoch is a single pass over a training dataset).
Regarding claim 4, the rejection of claim 1 is incorporated. Zhang further discloses “wherein one or more of the at least one programmable circuit is to:
train the supernet ML model using a training dataset (Zhang, Algorithm 1: “3: Get a batch of data; 4: Feed the data into teacher stream…9: Compute gradient to model parameters wt and update with the SGD optimizer”)
operate the supernet to extract the knowledge to be distilled into the subnet ML model (Zhang, Algorithm 1: “5: Get the feature map, soften probabilities of the teacher stream”); and
train the subnet using the training dataset and using the extracted knowledge to guide the training of the subnet” (Zhang, Algorithm 1: “3: Get a batch of data; 10: Feed the data into student stream; 13: Compute the cross-stream distillation loss with Eq. (6) and Eq. (7); 15: Compute the total loss of student with Eq. (11); 16: Compute gradient to model parameters ws and update with the SGD optimizer” and III.B.: “Cross-stream Distillation: For the cross stream knowledge transfer, the student will learn under the supervision of the corresponding guided modules of the teacher stream… The cross-stream distillation loss contains two types of losses: distillation loss and feature loss… The form of feature loss is as follows,
PNG
media_image2.png
30
594
media_image2.png
Greyscale
each pair of guided modules is to measure differences in the feature map, which can promote intermediate knowledge transfer cross stream”; Examiner notes that the feature maps of the teacher model (
F
g
,
t
(
b
)
) correspond to “knowledge” that guides the training of the student model).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the programmable circuit is to “train the supernet ML model using a training dataset; operate the supernet to extract knowledge to be distilled into the subnet ML model; and train the subnet using the training dataset and using the extracted knowledge to guide the training of the subnet,” and one would have been motivated to do so. Doing so would enable more efficient knowledge transfer by minimizing the capability gap between teacher and student (see Zhang, Fig. 1 Description).
Regarding claim 5, the rejection of claim 4 is incorporated. Cui further discloses “wherein the knowledge includes both logits and… [hidden states] extracted from the supernet ML model” (Cui, 5.1 Dense knowledge distillation, paragraphs 5-7: “Then, we minimize the mean-square error between the hidden states of the student and the fused hidden states from the teacher. The distillation objective is defined as follows: Lhidden= [see eq(16)]… Prediction Layer Distillation: Moreover, in the prediction layer, the student network is also required to learn the soft labels provided by the teacher for receiving more information about the prediction distribution of the training samples. We minimize the distance between the prediction of the teacher and that of the student as follows: Lpred = -softmax (
z
T
t
e
) ∙ logsoftmax (
z
S
t
e
), where zT and zS represent the probability logits predicted by the teacher and the student model respectively… Finally, we integrate dense Transformer layer distillation and prediction layer distillation together to build the overall objective function for knowledge distillation: LKD = n1Lhidden + n2Lpred”).
Neither Cui nor Zagoruyko appear to explicitly disclose the further limitations of the claim.
However, Zhang discloses “wherein the knowledge includes… feature maps extracted from the supernet ML model” (Zhang, III.B.: “Cross-stream Distillation: For the cross stream knowledge transfer, the student will learn under the supervision of the corresponding guided modules of the teacher stream… The cross-stream distillation loss contains two types of losses: distillation loss and feature loss… The form of feature loss is as follows,
PNG
media_image2.png
30
594
media_image2.png
Greyscale
each pair of guided modules is to measure differences in the feature map, which can promote intermediate knowledge transfer cross stream”; Examiner notes that the feature maps of the teacher model (
F
g
,
t
(
b
)
) correspond to “knowledge” that guides the training of the student model).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the knowledge includes feature maps extracted from the supernet ML model, and one would have been motivated to do so. Doing so would provide rich information about image intensity and spatial correlation to improve the learning of the student model (see Zhang, E. Influence of Loss Function).
Regarding claim 6, the rejection of claim 5 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein one or more of the at least one programmable circuit is to transfer the knowledge from the supernet ML model to the subnet ML model” (Cui, 6.3.1 Training setup: “In the joint compression procedure, we reduce the unimportant parameters in the student based on our pruning approach DISP, and perform the dense knowledge distillation to transfer the knowledge from the teacher to the pruned shallow network at the same time”).
Regarding claim 8, the rejection of claim 1 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “generate, based on input data, a queries matrix, a values matrix, and a keys matrix, wherein the queries matrix, the values matrix, and the keys matrix include parameters to be pruned (Cui, 3. Transformer overview: “Given an input sequence
x
=
x
1
,
…
,
x
L
∈
R
d
and a query vector q
∈
R
d
for an Ha heads model, the output of the h-th head can be computed as:
PNG
media_image3.png
70
694
media_image3.png
Greyscale
” and 4. Structured pruning for transformers: “In this section, we first propose a new Direct Importance-aware Structured Pruning method by allowing the model to make the pruning decision directly based on its corresponding parameter matrices. Further, this approach is applied to the pruning of Transformer models”; Examiner notes that
W
Q
h
corresponds to a queries matrix,
W
K
h
corresponds to a keys matrix, and
W
V
h
corresponds to a values matrix); and
provide the queries matrix, the values matrix, and the keys matrix to the pruning” (Cui, 4.2 Structured pruning for transformers: “In this section, the direct importance-aware structured pruning strategy is employed to encourage the Transformer model to automatically learn what to prune based on the related parameter matrices (as shown in Fig. 2). In the following, we start by pruning the heads in the multi-head attention to elaborate our pruning strategy, and then apply it to FFN layer for further model compression. Different from the previous approach of introducing gate variable gh to the headh to decide whether this head should be pruned or not [30,46,29,6,19], our DISP method allows the head to adaptively make a decision by itself based on its corresponding parameter matrices. Intuitively, these parameter matrices are more relevant to the head than the extra gate variable, and thus are more straightforward to reflect the importance of the head and more suitable to determine its existence”).
Regarding claim 9, the rejection of claim 8 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein one or more of the at least one programmable circuit is to: apply the input data to a parameterized learnable transformation (PLT) to generate the queries matrix, the values matrix, and the keys matrix” (Cui 3. Transformer overview: “Specifically, a standard Transformer model is composed of M Transformer layers. In each layer, there are two main sub-layers: multi-head attention and fully connected feed-forward network. Given an input sequence
x
=
x
1
,
…
,
x
L
∈
R
d
and a query vector q
∈
R
d
for an Ha heads model, the output of the h-th head can be computed as:
PNG
media_image3.png
70
694
media_image3.png
Greyscale
” and 4.2 Structured pruning for transformers: “The pruning procedure is optimized with the original loss function Lc for the Transformer model (e.g., cross-entropy loss)”; Examiner notes that the transformer model corresponds to a “parameterized learnable transformation (PLT)”).
Regarding claim 11, the rejection of claim 8 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein one or more of the at least one programmable circuit is to:
perform an operation on the queries matrix and the keys matrix, wherein the operation is a matrix multiplication operation or a 1 x 1 convolution on the queries matrix and the keys matrix (Cui 3. Transformer overview: the output of the h-th head can be computed as:
PNG
media_image3.png
70
694
media_image3.png
Greyscale
; Examiner notes that ai contains a matrix multiplication operation of the queries and keys matrices);
apply a softmax function to an output of the operation (Cui 3. Transformer overview:
see “softmax(ai)” in eq (1)); and
generate a self-attention output based on a combination of the values matrix and an output of the softmax function” (Cui 3. Transformer overview: see “softmax(ai)” multiplied by values matrix in eq (1) to generate “headh” output).
Regarding claim 16, Cui discloses “A non-transitory machine readable storage medium comprising instructions to cause at least one programmable circuit to at least (Cui, 6.2.1 Implementation details: “The experiments are conducted with PyTorch framework on GeForce GTX 1080Ti GPU”):
provide a training dataset to a first machine learning (ML) model and a second ML model, the second ML model having fewer parameters than the first ML model (Cui, 5.1 Dense knowledge distillation: “Knowledge distillation aims to transfer the useful knowledge of a large teacher network to a lightweight student network with fewer layers…We assume that the teacher model T has M Transformer layers T = {T1, … ,TM} and the student S contains N layers as S = {S1, …, SN}… a many-to-one layer mapping strategy is introduced to encourage each student layer to flexibly learn from multiple teacher layers to obtain the rich linguistic knowledge in different levels simultaneously (as illustrated in Fig.4b)” and 6.1. Datasets: “The experiments are conducted on four NLP tasks with seven public datasets” and Fig. 3; Examiner notes that teacher model T corresponds to a first ML model and student model S corresponds to a second ML model, and Fig. 3 depicts the input dataset being provided to the teacher and student models) …
during training of the… second ML model[[s]], distill knowledge of the first ML model into the second ML model during one pass over the training dataset (Cui, 5.1 Dense knowledge distillation: “Knowledge distillation aims to transfer the useful knowledge of a large teacher network to a lightweight student network with fewer layers…We assume that the teacher model T has M Transformer layers T = {T1, … ,TM} and the student S contains N layers as S = {S1, …, SN}… a many-to-one layer mapping strategy is introduced to encourage each student layer to flexibly learn from multiple teacher layers to obtain the rich linguistic knowledge in different levels simultaneously (as illustrated in Fig.4b)” and 6.3.3 Results: “Our BERT6-JMC overcomes the limitation of them through performing the pruning and distillation in one training process. The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4”; Examiner notes that one epoch corresponds to one pass over the training dataset); and
during the same one pass over the training dataset as the distill (Cui, 6.3.3 Results: “Our BERT6-JMC overcomes the limitation of them through performing the pruning and distillation in one training process. The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4” and 5.2 Integrating pruning and distillation: “Different from the Two-stage compression (Fig. 1c) of continuing to apply pruning method on the learned distilled model, our JMC performs pruning and distillation simultaneously, which allows the joint optimization of the two compression techniques”):
calculate a training loss using a loss function, the loss function including a distillation loss, a task loss… (Cui, 5.2 Integrating pruning and distillation: “The final loss function of this JMC approach is formulated by the fusion of the original loss for Transformer Lc, the pruning objective function LP, and the dense knowledge distillation objective LKD: LJMC = Lc + αLp + LKD”; Examiner notes that LJMC corresponds to a training loss, LKD corresponds to a distillation loss and Lc corresponds to a task loss) wherein:
the distillation loss describes a divergence between predictions of the first ML model and predictions of the second ML model (Cui, 5.1 Dense Knowledge Distillation: “Prediction Layer Distillation: Moreover, in the prediction layer, the student network is also required to learn the soft labels provided by the teacher for receiving more information about the prediction distribution of the training samples. We minimize the distance between the prediction of the teacher and that of the student as follows: Lpred = -softmax (
z
T
t
e
) ∙ logsoftmax (
z
S
t
e
), where zT and zS represent the probability logits predicted by the teacher and the student model respectively… Finally, we integrate dense Transformer layer distillation and prediction layer distillation together to build the overall objective function for knowledge distillation: LKD = n1Lhidden + n2Lpred”) …
prune one or more parameters from the second ML model based on the training loss (Cui, 5.2 Integrating pruning and distillation: “During the joint compression, the proposed structured pruning approach DISP is applied on the student to prune the redundant parameters in this shallow network including the heads in the multi-head attention and weight blocks of intermediate layer in FFN” and Cui, 4.2 Structured pruning for transformers: “Gradual Pruning Optimization In order to encourage the model to be pruned to the target sparsity, a sparsity-aware pruning objective function LP is introduced to impose an explicit control over the entire model. The pruning procedure is optimized with the original loss function Lc for the Transformer model (e.g., cross-entropy loss). In particular, the model is trained to reach the overall target sparsity pt by minimizing the new training objective L as follows: L = Lc + Lp”; Examiner notes that the pruning is based on Lc and Lp, which are part of the joint loss function LJMC); and
output the second ML model without the pruned one or more parameters” (Cui, 6.3.1 Training setup: “In the joint compression procedure, we reduce the unimportant parameters in the student based on our pruning approach DISP, and perform the dense knowledge distillation to transfer the knowledge from the teacher to the pruned shallow network at the same time. Such joint compression strategy enables us to obtain a compressed shallow model BERT6-JMC”).
Cui does not appear to explicitly disclose the further limitations of the claim. However, Zagoruyko discloses “the first ML model having one or more feature maps describing activation patterns indicative of an output of the first ML model” (Zagoruyko, 3.1 Activation-Based Attention Transfer, Paragraph 1: “Let us consider a CNN layer and its corresponding activation tensor A ∈ RC×H×W, which consists of C feature planes with spatial dimensions H×W. An activation-based mapping function F (w.r.t. that layer) takes as input the above 3D tensor A and outputs a spatial attention map, i.e., a flattened 2D tensor defined over the spatial dimensions, or F : RC× H× W →RH× W . To define such a spatial attention mapping function, the implicit assumption that we make in this section is that the absolute value of a hidden neuron activation (that results when the network is evaluated on given input) can be used as an indication about the importance of that neuron w.r.t. the specific input” and 3.1 Activation-Based Attention Transfer, Paragraph 6: “In attention transfer, given the spatial attention maps of a teacher network (computed using any of the above attention mapping functions), the goal is to train a student network that will not only make correct predictions but will also have attentions maps that are similar to those of the teacher” and Figure 1: “(a) An input image and a corresponding spatial attention map of a convolutional network that shows where the network focuses in order to classify the given image”; Examiner notes that the teacher network corresponds to the supernet ML model, and the spatial attention maps correspond to “one or more features maps describing activation patterns indicative of an output of the supernet ML model” as they describe neuron activations in generating an output (e.g. image classification) for a given input of the network) and “the loss function including… an attention transfer loss, wherein: … the attention transfer loss describes a difference between the one or more feature maps of the first ML model and one or more corresponding feature maps of the second ML model” (3.1 Activation-Based Attention Transfer, Paragraph 8: “Without loss of generality, we assume that transfer losses are placed between student and teacher attention maps of same spatial resolution, but, if needed, attention maps can be interpolated to match their shapes. Let S, T and WS, WT denote student, teacher and their weights correspondingly, and let L(W, x) denote a standard cross entropy loss. Let also I denote the indices of all teacher-student activation layer pairs for which we want to transfer attention maps. Then we can define the following total loss:
PNG
media_image1.png
224
824
media_image1.png
Greyscale
”; Examiner notes that the second term in the loss function corresponds to an attention transfer loss, which describes a difference between the teacher (first ML model) attention maps and the student (second ML model) attention maps).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cui with the teachings of Zagoruyko such that the first ML model has one or more feature maps describing activation patterns indicative of an output of the first ML model, and such that the loss function further includes an attention transfer loss, wherein the attention transfer loss describes a difference between the one or more feature maps of the first ML model and one or more corresponding feature maps of the second ML model, and one would have been motivated to do so. Doing so would significantly improve the performance of a student CNN network by forcing it to mimic the attention maps of a powerful teacher network (see Zagoruyko, Abstract).
Neither Cui nor Zagoruyko appear to explicitly disclose the further limitations of the claim.
However Zhang discloses “during training of the first and second ML models, distill knowledge of the first ML model into the second ML model during one pass over the training dataset” (Zhang, III.B: “The proposed evolutionary knowledge distillation adopts online training of teacher and student model synchronously, the parameters of student and teacher models are updated in each batch data process, which can reduce the teacher-student capability gap. When coupled with Fig. 2, we can see that teacher and student streams are divided into C blocks. Each block followed by a guided module with a fully connected layer constitutes multiple classifiers. We assume a stream with C classifiers. The training process includes two simultaneous stages, within-stream distillation and cross-stream distillation. For within-stream distillation, deeper classifiers provide super vision to help the learning of shallow classifiers, which can improve the ability of the stream itself to represent knowledge. Cross-stream distillation can improve knowledge transfer from the evolutionary teacher to student. The guided modules can help facilitate knowledge representation and transfer”) and “the task loss describes a difference between predictions of… [an] ML model and labels of a training dataset” (Zhang, III. A: “Denote the training set is D
=
{
(
x
i
,
y
i
)
i
=
1
D
, where xi is the ith sample with label yi ∈{1,2,...,m}” and III.B, paragraph 8: “For each classifier, we compute cross entropy loss lce between φ(x;w(c)) and y. In this way, the label y directs each classifier’s probability as possible”; Examiner notes that φ(x;w(c)) corresponds to predictions of an ML model and y corresponds to labels of a training dataset).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the knowledge distillation occurs during training of both the first and second ML models rather than only the second, and such that the task loss describes a difference between predictions of the second ML model and labels of the training dataset, and one would have been motivated to do so. Doing so would enable more efficient knowledge transfer by minimizing the capability gap between teacher and student (see Zhang, Fig. 1 Description) and improve the student model’s classification performance by directing the model to match ground truth labels (see Zhang III.B, paragraph 8).
Regarding claim 18, the rejection of claim 16 is incorporated. Zhang further discloses “wherein the instructions cause one or more of the at least one programmable circuit to:
train the first ML model using the training dataset (Zhang, Algorithm 1: “3: Get a batch of data; 4: Feed the data into teacher stream…9: Compute gradient to model parameters wt and update with the SGD optimizer”)
operate the first ML model to extract the knowledge to be distilled into the second ML model (Zhang, Algorithm 1: “5: Get the feature map, soften probabilities of the teacher stream”); and
train the second ML model using the training dataset and using the extracted knowledge to guide the training of the second ML model” (Zhang, Algorithm 1: “3: Get a batch of data; 10: Feed the data into student stream; 13: Compute the cross-stream distillation loss with Eq. (6) and Eq. (7); 15: Compute the total loss of student with Eq. (11); 16: Compute gradient to model parameters ws and update with the SGD optimizer” and III.B.: “Cross-stream Distillation: For the cross stream knowledge transfer, the student will learn under the supervision of the corresponding guided modules of the teacher stream… The cross-stream distillation loss contains two types of losses: distillation loss and feature loss… The form of feature loss is as follows,
PNG
media_image2.png
30
594
media_image2.png
Greyscale
”; Examiner notes that the feature maps of the teacher model (
F
g
,
t
(
b
)
) correspond to “knowledge” that guides the training of the student model).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the instructions cause the programmable circuit to “train the first ML model using a training dataset; operate the first ML model to extract knowledge to be distilled into the second ML model; and train the second ML model using the training dataset and using the extracted knowledge to guide the training of the second ML model,” and one would have been motivated to do so. Doing so would enable more efficient knowledge transfer by minimizing the capability gap between teacher and student (see Zhang, Fig. 1 Description).
Regarding claim 19, the rejection of claim 18 is incorporated. Cui further discloses “wherein the instructions cause one or more of the at least one programmable circuit to: extract logits and… [hidden states] from the first ML model, wherein the knowledge includes both the extracted logits and the extracted… [hidden states]” (Cui, 5.1 Dense knowledge distillation: “Prediction Layer Distillation: Moreover, in the prediction layer, the student network is also required to learn the soft labels provided by the teacher for receiving more information about the prediction distribution of the training samples. We minimize the distance between the prediction of the teacher and that of the student as follows: Lpred = -softmax (
z
T
t
e
) ∙ logsoftmax (
z
S
t
e
), where zT and zS represent the probability logits predicted by the teacher and the student model respectively… Finally, we integrate dense Transformer layer distillation and prediction layer distillation together to build the overall objective function for knowledge distillation: LKD = n1Lhidden + n2Lpred”).
Neither Cui nor Zagoruyko appear to explicitly disclose the further limitations of the claim.
However, Zhang discloses “extract… feature maps from the first ML model, wherein the knowledge includes… the extracted feature maps” (Zhang, III.B.: “Cross-stream Distillation: For the cross stream knowledge transfer, the student will learn under the supervision of the corresponding guided modules of the teacher stream… The cross-stream distillation loss contains two types of losses: distillation loss and feature loss… The form of feature loss is as follows,
PNG
media_image2.png
30
594
media_image2.png
Greyscale
each pair of guided modules is to measure differences in the feature map, which can promote intermediate knowledge transfer cross stream”; Examiner notes that the feature maps of the teacher model (
F
g
,
t
(
b
)
) correspond to “knowledge” that guides the training of the student model).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the instructions cause the programmable circuit to extract feature maps from the first ML model and such that the knowledge includes the extracted feature maps, and one would have been motivated to do so. Doing so would provide rich information about image intensity and spatial correlation to improve the learning of the student model (see Zhang, E. Influence of Loss Function).
Regarding claim 20, the rejection of claim 16 is incorporated. Zagoruyko further discloses “wherein the instructions cause one or more of the at least one programmable circuit to: operate an attention transfer distillation algorithm to transfer the knowledge from the first ML model to the second ML model such that the second ML model includes a spatial attention map that is similar to a spatial attention map of the first ML model” (Zagoruyko, Figure 1: “(b) Schematic representation of attention transfer: a student CNN is trained so as, not only to make good predictions, but to also have similar spatial attention maps to those of an already trained teacher CNN” and Zagoruyko, 4.1.1: “Results of attention transfer (using
F
s
u
m
2
attention maps) for various networks on CIFAR-10 can befound in table 1. We experimented with teacher/student having the same depth (WRN-16-2/WRN 16-1), as well as different depth (WRN-40-1/WRN-16-1, WRN-40-2/WRN-16-2). In all combinations, attention transfer (AT) shows significant improvements, which are also higher when it is combined with knowledge distillation (AT+KD)”).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cui/Zhang with the teachings of Zagoruyko such that the instructions cause the programmable circuit to operate an attention transfer distillation algorithm to transfer the knowledge from the first ML model to the second ML model such that the second ML model includes a spatial attention map that is similar to a spatial attention map of the first ML model, and one would have been motivated to do so. Doing so would significantly improve the performance of a student CNN network by forcing it to mimic the attention maps of a powerful teacher network (see Zagoruyko, Abstract).
Regarding claim 21, the rejection of claim 16 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein one or more of the at least one programmable circuit is to: generate, based on the training dataset, a queries matrix, a values matrix, and a keys matrix, wherein the queries matrix, the values matrix, and the keys matrix include parameters to be pruned (Cui, 3. Transformer overview: “Given an input sequence
x
=
x
1
,
…
,
x
L
∈
R
d
and a query vector q
∈
R
d
for an Ha heads model, the output of the h-th head can be computed as:
PNG
media_image3.png
70
694
media_image3.png
Greyscale
” and 4. Structured pruning for transformers: “In this section, we first propose a new Direct Importance-aware Structured Pruning method by allowing the model to make the pruning decision directly based on its corresponding parameter matrices. Further, this approach is applied to the pruning of Transformer models”; Examiner notes that
W
Q
h
corresponds to a queries matrix,
W
K
h
corresponds to a keys matrix, and
W
V
h
corresponds to a values matrix); and
provide the queries matrix, the values matrix, and the keys matrix to the pruning” (4.2 Structured pruning for transformers: “In this section, the direct importance-aware structured pruning strategy is employed to encourage the Transformer model to automatically learn what to prune based on the related parameter matrices (as shown in Fig. 2). In the following, we start by pruning the heads in the multi-head attention to elaborate our pruning strategy, and then apply it to FFN layer for further model compression. Different from the previous approach of introducing gate variable gh to the headh to decide whether this head should be pruned or not [30,46,29,6,19], our DISP method allows the head to adaptively make a decision by itself based on its corresponding parameter matrices. Intuitively, these parameter matrices are more relevant to the head than the extra gate variable, and thus are more straightforward to reflect the importance of the head and more suitable to determine its existence”).
Regarding claim 22, the rejection of claim 21 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein the instructions cause one or more of the at least one programmable circuit to: apply parts of the training dataset to a parameterized learnable transformation (PLT) to generate the queries matrix, the values matrix, and the keys matrix (Cui 3. Transformer overview: “Specifically, a standard Transformer model is composed of M Transformer layers. In each layer, there are two main sub-layers: multi-head attention and fully connected feed-forward network. Given an input sequence
x
=
x
1
,
…
,
x
L
∈
R
d
and a query vector q
∈
R
d
for an Ha heads model, the output of the h-th head can be computed as:
PNG
media_image3.png
70
694
media_image3.png
Greyscale
”; Examiner notes that the transformer model corresponds to a “parameterized learnable transformation (PLT)”);
perform an operation on the queries matrix and the keys matrix, wherein the operation is a matrix multiplication operation or a 1 x 1 convolution on the queries matrix and the keys matrix (Cui 3. Transformer overview: see ai containing a matrix multiplication operation of the queries and keys matrices in eq (1));
apply a softmax function to an output of the operation (Cui 3. Transformer overview:
see “softmax(ai)” in eq (1)); and
generate self-attention output based on a combination of the values matrix and an output of the softmax function” (Cui 3. Transformer overview: see “softmax(ai)” multiplied by values matrix in eq (1) to generate “headh” output).
Regarding claim 28, the rejection of claim 1 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein the subnet ML model has a plurality of layers, each layer having a plurality of weights, wherein pruning one or more parameters includes setting one or more weights of a layer of the subnet ML model to zero” (Cui, 4.2 Structured pruning for transformers: “In this section, the direct importance-aware structured pruning strategy is employed to encourage the Transformer model to automatically learn what to prune based on the related parameter matrices (as shown in Fig. 2). In the following, we start by pruning the heads in the multi-head attention to elaborate our pruning strategy, and then apply it to FFN layer for further model compression… the corresponding representation headh is automatically pruned in the output state by “zeroing” out the subset weights
W
O
h
.
This observation allows us to view the pruning problem as finding the important head weights to be retained in each multi-head attention layer”).
Regarding claim 32, the rejection of claim 1 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein the subnet ML model comprises a transformer architecture that includes a self-attention mechanism, the self-attention mechanism configured to generate a queries matrix, a keys matrix, and a values matrix based on input data, and to produce a self-attention output based on the queries matrix, the keys matrix, and the values matrix” (Cui 3. Transformer overview: “Specifically, a standard Transformer model is composed of M Transformer layers. In each layer, there are two main sub-layers: multi-head attention and fully connected feed-forward network. Given an input sequence
x
=
x
1
,
…
,
x
L
∈
R
d
and a query vector q
∈
R
d
for an Ha heads model, the output of the h-th head can be computed as:
PNG
media_image3.png
70
694
media_image3.png
Greyscale
” and 5.2 Integrating pruning and distillation: “Specifically, the original full network is regarded as the teacher model and the student model is initialized with a shallow network with fewer Transformer layers (as shown in Fig. 1d)”; Examiner notes that
W
Q
h
corresponds to a queries matrix,
W
K
h
corresponds to a keys matrix, and
W
V
h
corresponds to a values matrix).
Regarding claim 33, the rejection of claim 16 is incorporated. Cui as modified by Zagoruyko and Zhang further discloses “wherein the subnet ML model has a plurality of layers, each layer having a plurality of weights, wherein pruning one or more parameters includes setting one or more weights of a layer of the subnet ML model to zero” (Cui, 4.2 Structured pruning for transformers: “In this section, the direct importance-aware structured pruning strategy is employed to encourage the Transformer model to automatically learn what to prune based on the related parameter matrices (as shown in Fig. 2). In the following, we start by pruning the heads in the multi-head attention to elaborate our pruning strategy, and then apply it to FFN layer for further model compression… the corresponding representation headh is automatically pruned in the output state by “zeroing” out the subset weights
W
O
h
.
This observation allows us to view the pruning problem as finding the important head weights to be retained in each multi-head attention layer”).
Claims 26 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Cui in view of Zagoruyko and Zhang, and further in view of Liu et al. (US20210264278) (hereinafter “Liu”).
Regarding claim 26, Cui discloses “A method comprising:
providing a training dataset to a first machine learning (ML) model and a second ML model, the second ML model having fewer parameters than the first ML model (Cui, 5.1 Dense knowledge distillation: “Knowledge distillation aims to transfer the useful knowledge of a large teacher network to a lightweight student network with fewer layers…We assume that the teacher model T has M Transformer layers T = {T1, … ,TM} and the student S contains N layers as S = {S1, …, SN}… a many-to-one layer mapping strategy is introduced to encourage each student layer to flexibly learn from multiple teacher layers to obtain the rich linguistic knowledge in different levels simultaneously (as illustrated in Fig.4b)” and 6.1. Datasets: “The experiments are conducted on four NLP tasks with seven public datasets” and Fig. 3; Examiner notes that teacher model T corresponds to a first ML model and student model S corresponds to a second ML model, and Fig. 3 depicts the input dataset being provided to the teacher and student model) …
during training of the… second ML model[[s]], distilling knowledge of the first ML model into the second ML model during one pass over the training dataset (Cui, 5.1 Dense knowledge distillation: “Knowledge distillation aims to transfer the useful knowledge of a large teacher network to a lightweight student network with fewer layers…We assume that the teacher model T has M Transformer layers T = {T1, … ,TM} and the student S contains N layers as S = {S1, …, SN}… a many-to-one layer mapping strategy is introduced to encourage each student layer to flexibly learn from multiple teacher layers to obtain the rich linguistic knowledge in different levels simultaneously (as illustrated in Fig.4b)” and 6.3.3 Results: “Our BERT6-JMC overcomes the limitation of them through performing the pruning and distillation in one training process. The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4”; Examiner notes that one epoch corresponds to one pass over the training dataset); and
during the same one pass over the training dataset as the distilling (Cui, 6.3.3 Results: “Our BERT6-JMC overcomes the limitation of them through performing the pruning and distillation in one training process. The training time spent in BERT6-JMC and the two-stage compression models on each dataset (for one epoch) are reported in Table 4” and 5.2 Integrating pruning and distillation: “Different from the Two-stage compression (Fig. 1c) of continuing to apply pruning method on the learned distilled model, our JMC performs pruning and distillation simultaneously, which allows the joint optimization of the two compression techniques”):
calculating a training loss using a loss function, the loss function including a distillation loss, a task loss… (Cui, 5.2 Integrating pruning and distillation: “The final loss function of this JMC approach is formulated by the fusion of the original loss for Transformer Lc, the pruning objective function LP, and the dense knowledge distillation objective LKD: LJMC = Lc + αLp + LKD”; Examiner notes that LJMC corresponds to a training loss, LKD corresponds to a distillation loss and Lc corresponds to a task loss) wherein:
the distillation loss describes a divergence between predictions of the first ML model and predictions of the second ML model (Cui, 5.1 Dense Knowledge Distillation: “Prediction Layer Distillation: Moreover, in the prediction layer, the student network is also required to learn the soft labels provided by the teacher for receiving more information about the prediction distribution of the training samples. We minimize the distance between the prediction of the teacher and that of the student as follows: Lpred = -softmax (
z
T
t
e
) ∙ logsoftmax (
z
S
t
e
), where zT and zS represent the probability logits predicted by the teacher and the student model respectively… Finally, we integrate dense Transformer layer distillation and prediction layer distillation together to build the overall objective function for knowledge distillation: LKD = n1Lhidden + n2Lpred”) …
pruning one or more parameters from the second ML model based on… the training loss (Cui, 5.2 Integrating pruning and distillation: “During the joint compression, the proposed structured pruning approach DISP is applied on the student to prune the redundant parameters in this shallow network including the heads in the multi-head attention and weight blocks of intermediate layer in FFN” and Cui, 4.2 Structured pruning for transformers: “Gradual Pruning Optimization In order to encourage the model to be pruned to the target sparsity, a sparsity-aware pruning objective function LP is introduced to impose an explicit control over the entire model. The pruning procedure is optimized with the original loss function Lc for the Transformer model (e.g., cross-entropy loss). In particular, the model is trained to reach the overall target sparsity pt by minimizing the new training objective L as follows: L = Lc + Lp”; Examiner notes that the pruning is based on Lc and Lp, which are part of the joint loss function LJMC); and
outputting the second ML model… without the pruned one or more parameters” (Cui, 6.3.1 Training setup: “In the joint compression procedure, we reduce the unimportant parameters in the student based on our pruning approach DISP, and perform the dense knowledge distillation to transfer the knowledge from the teacher to the pruned shallow network at the same time. Such joint compression strategy enables us to obtain a compressed shallow model BERT6-JMC”).
Cui does not appear to explicitly disclose the further limitations of the claim. However, Zagoruyko discloses “the first ML model having one or more feature maps describing activation patterns indicative of an output of the first ML model” (Zagoruyko, 3.1 Activation-Based Attention Transfer, Paragraph 1: “Let us consider a CNN layer and its corresponding activation tensor A ∈ RC×H×W, which consists of C feature planes with spatial dimensions H×W. An activation-based mapping function F (w.r.t. that layer) takes as input the above 3D tensor A and outputs a spatial attention map, i.e., a flattened 2D tensor defined over the spatial dimensions, or F : RC× H× W →RH× W . To define such a spatial attention mapping function, the implicit assumption that we make in this section is that the absolute value of a hidden neuron activation (that results when the network is evaluated on given input) can be used as an indication about the importance of that neuron w.r.t. the specific input” and 3.1 Activation-Based Attention Transfer, Paragraph 6: “In attention transfer, given the spatial attention maps of a teacher network (computed using any of the above attention mapping functions), the goal is to train a student network that will not only make correct predictions but will also have attentions maps that are similar to those of the teacher” and Figure 1: “(a) An input image and a corresponding spatial attention map of a convolutional network that shows where the network focuses in order to classify the given image”; Examiner notes that the teacher network corresponds to the supernet ML model, and the spatial attention maps correspond to “one or more features maps describing activation patterns indicative of an output of the supernet ML model” as they describe neuron activations in generating an output (e.g. image classification) for a given input of the network) and “the loss function including… an attention transfer loss, wherein: … the attention transfer loss describes a difference between the one or more feature maps of the first ML model and one or more corresponding feature maps of the second ML model” (3.1 Activation-Based Attention Transfer, Paragraph 8: “Without loss of generality, we assume that transfer losses are placed between student and teacher attention maps of same spatial resolution, but, if needed, attention maps can be interpolated to match their shapes. Let S, T and WS, WT denote student, teacher and their weights correspondingly, and let L(W, x) denote a standard cross entropy loss. Let also I denote the indices of all teacher-student activation layer pairs for which we want to transfer attention maps. Then we can define the following total loss:
PNG
media_image1.png
224
824
media_image1.png
Greyscale
”; Examiner notes that the second term in the loss function corresponds to an attention transfer loss, which describes a difference between the teacher (first ML model) attention maps and the student (second ML model) attention maps).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Cui with the teachings of Zagoruyko such that the first ML model has one or more feature maps describing activation patterns indicative of an output of the supernet ML model, and such that the loss function further includes an attention transfer loss, wherein the attention transfer loss describes a difference between the one or more feature maps of the first ML model and one or more corresponding feature maps of the second ML model, and one would have been motivated to do so. Doing so would significantly improve the performance of a student CNN network by forcing it to mimic the attention maps of a powerful teacher network (see Zagoruyko, Abstract).
Neither Cui nor Zagoruyko appear to explicitly disclose the further limitations of the claim.
However Zhang discloses “during training of the first and second ML models, distilling knowledge of the first ML model into the second ML model during one pass over the training dataset” (Zhang, III.B: “The proposed evolutionary knowledge distillation adopts online training of teacher and student model synchronously, the parameters of student and teacher models are updated in each batch data process, which can reduce the teacher-student capability gap. When coupled with Fig. 2, we can see that teacher and student streams are divided into C blocks. Each block followed by a guided module with a fully connected layer constitutes multiple classifiers. We assume a stream with C classifiers. The training process includes two simultaneous stages, within-stream distillation and cross-stream distillation. For within-stream distillation, deeper classifiers provide super vision to help the learning of shallow classifiers, which can improve the ability of the stream itself to represent knowledge. Cross-stream distillation can improve knowledge transfer from the evolutionary teacher to student. The guided modules can help facilitate knowledge representation and transfer”) and “the task loss describes a difference between predictions of… [an] ML model and labels of a training dataset” (Zhang, III. A: “Denote the training set is D
=
{
(
x
i
,
y
i
)
i
=
1
D
, where xi is the ith sample with label yi ∈{1,2,...,m}” and III.B, paragraph 8: “For each classifier, we compute cross entropy loss lce between φ(x;w(c)) and y. In this way, the label y directs each classifier’s probability as possible”; Examiner notes that φ(x;w(c)) corresponds to predictions of an ML model and y corresponds to labels of a training dataset).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko with the teachings of Zhang such that the knowledge distillation occurs during training of both the first and second ML models rather than only the second, and such that the task loss describes a difference between predictions of the second ML model and labels of the training dataset, and one would have been motivated to do so. Doing so would enable more efficient knowledge transfer by minimizing the capability gap between teacher and student (see Zhang, Fig. 1 Description) and improve the student model’s classification performance by directing the model to match ground truth labels (see Zhang III.B, paragraph 8).
Neither Cui, Zagoruyko, nor Zhang appear to explicitly disclose the further limitations of the claim.
However, Liu discloses “pruning one or more parameters from… [an] ML model based on an identified parameter of target hardware (Liu, [0054]: “As mentioned above, the neural network pruning system can prune the neural network based on a pruning parameter, which can include a network size pruning parameter (or simply “size parameter”)” and [0056]: “In some implementations, the size parameter is based on the hardware constraints of a computing device. For instance, when pruning a neural network to be implemented on a particular computing device (e.g., a mobile client device), the neural network pruning system can identify the size parameter based on the memory capacity and/or availability of the computing device”) …and outputting the…ML model will operate on the target hardware without the pruned one or more parameters” (Liu, [0074]: “Upon progressively pruning the neural network based on the identified pruning parameter, the neural network pruning system can utilize the pruned neural network, as shown as the act 212 within the series of acts 200. In some implementations, the act 212 can include the neural network pruning system providing a query input to the pruned neural network and predicting an output. In one or more implementations, the act 212 can include providing the pruned neural network to another computing device (e.g., a hardware-constrained mobile client device)”).
Liu and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Zhang with the teachings of Liu such that the pruning is based on an identified parameter of target hardware and such that the outputted second ML model will operate on the target hardware, and one would have been motivated to do so. Doing so would allow for optimizing the student model to be able to fit on a computing device with hardware restrictions (see Liu, [0006]).
Regarding claim 27, the rejection of claim 26 is incorporated. Zhang further discloses “training the first ML model using the training dataset (Zhang, Algorithm 1: “3: Get a batch of data; 4: Feed the data into teacher stream…9: Compute gradient to model parameters wt and update with the SGD optimizer”);
operating the first ML model to extract the knowledge to be distilled into the subnet ML model (Zhang, Algorithm 1: “5: Get the feature map, soften probabilities of the teacher stream”); and
training the second ML model using the training dataset and using the extracted knowledge to guide the training of the second ML model” (Zhang, Algorithm 1: “3: Get a batch of data; 10: Feed the data into student stream; 13: Compute the cross-stream distillation loss with Eq. (6) and Eq. (7); 15: Compute the total loss of student with Eq. (11); 16: Compute gradient to model parameters ws and update with the SGD optimizer” and III.B.: “Cross-stream Distillation: For the cross stream knowledge transfer, the student will learn under the supervision of the corresponding guided modules of the teacher stream… The cross-stream distillation loss contains two types of losses: distillation loss and feature loss… The form of feature loss is as follows,
PNG
media_image2.png
30
594
media_image2.png
Greyscale
each pair of guided modules is to measure differences in the feature map, which can promote intermediate knowledge transfer cross stream”; Examiner notes that the feature maps of the teacher model (
F
g
,
t
(
b
)
) correspond to “knowledge” that guides the training of the student model).
Zhang and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Liu with the teachings of Zhang such that the method further includes “training the first ML model using the training dataset; operating the first ML model to extract the knowledge to be distilled into the second ML model; and training the second ML model using the training dataset and using the extracted knowledge to guide the training of the second ML model,” and one would have been motivated to do so. Doing so would enable more efficient knowledge transfer by minimizing the capability gap between teacher and student (see Zhang, Fig. 1 Description).
Claims 29 and 34 are rejected under 35 U.S.C. 103 as being unpatentable over Cui in view of Zagoruyko and Zhang, and further in view of Dettmers et al. (“Sparse Networks from Scratch: Faster Training without Losing Performance”) (“Dettmers”).
Regarding claim 29, the rejection of claim 1 is incorporated. Neither Cui, Zagoruyko, nor Zhang appear to explicitly disclose the further limitations of the claim.
However, Dettmers discloses “wherein: the at least one programmable circuit is to update, for a plurality of batches within the single ML training epoch and based on the training loss, weights of… [a] model and a momentum contribution associated with each layer of a plurality of layers (Dettmers, 3.2 Sparse Momentum, paragraph 5: “We train the network normally and mask the weights after each gradient update to enforce sparsity. We apply sparse momentum after each epoch. We can break the sparse momentum into three major parts: (a) redistribution of weights, (b) pruning weights, (c) regrowing weights. In step (a), we we take the mean of the element-wise momentum magnitude mi that belongs to all nonzero weights for each layer i and normalize the value by the total momentum magnitude of all layers
∑
i
=
o
k
m
i
” and Algorithm 1; Examiner notes that lines 6-15 of Algorithm 1 show computeGradients, UpdateMomentum (using the computed gradients/training loss,) and UpdateWeights occurring for each batch in a single epoch); and
the prune of the one or more parameters from the… model includes: after processing the plurality of batches, deactivating a first fraction of active weights in each layer of the plurality of layers (Dettmers, 3.2 Sparse Momentum, paragraph 5: “In step (b), we prune a proportion of p (prune rate) of the weights with the lowest magnitude for each layer”); and
reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer” (Dettmers, 3.2 Sparse Momentum, paragraph 5: “The number of weights to be regrow in each layer is the total number of removed weights multiplied by each layers momentum contribution: Regrowi = Total Removed · mi… In step (c), we regrow weights by enabling the gradient flow of zero-valued (missing) weights which have the largest momentum magnitude”).
Dettmers and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Zhang with the teachings of Dettmers such that the programmable circuit is further programmed to “update, for a plurality of batches within the single ML training epoch and based on the training loss, weights of the subnet ML model and a momentum contribution associated with each layer of a plurality of layers; and the prune of the one or more parameters from the subnet ML model includes: after processing the plurality of batches, deactivating a first fraction of active weights in each layer of the plurality of layers; and reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer,” and one would have been motivated to do so. Doing so would allow for efficiently identifying layers and weights which reduce the error, which would increase performance levels while providing faster training (see Dettmers, Abstract).
Regarding claim 34, the rejection of claim 16 is incorporated. Neither Cui, Zagoruyko, nor Zhang appear to explicitly disclose the further limitations of the claim.
However, Dettmers discloses “wherein the instructions cause one or more of the at least one programmable circuit to update, for a plurality of batches within the single ML training epoch and based on the training loss, weights of… [a] model and a momentum contribution associated with each layer of a plurality of layers (Dettmers, 3.2 Sparse Momentum, paragraph 5: “We train the network normally and mask the weights after each gradient update to enforce sparsity. We apply sparse momentum after each epoch. We can break the sparse momentum into three major parts: (a) redistribution of weights, (b) pruning weights, (c) regrowing weights. In step (a), we we take the mean of the element-wise momentum magnitude mi that belongs to all nonzero weights for each layer i and normalize the value by the total momentum magnitude of all layers
∑
i
=
o
k
m
i
” and Algorithm 1; Examiner notes that lines 6-15 of Algorithm 1 show computeGradients, UpdateMomentum (using the computed gradients/training loss,) and UpdateWeights occurring for each batch in a single epoch); and
the prune of the one or more parameters from the… model includes: after processing the plurality of batches, deactivating a first fraction of active weights in each layer of the plurality of layers (Dettmers, 3.2 Sparse Momentum, paragraph 5: “In step (b), we prune a proportion of p (prune rate) of the weights with the lowest magnitude for each layer”); and
reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer” (Dettmers, 3.2 Sparse Momentum, paragraph 5: “The number of weights to be regrow in each layer is the total number of removed weights multiplied by each layers momentum contribution: Regrowi = Total Removed · mi… In step (c), we regrow weights by enabling the gradient flow of zero-valued (missing) weights which have the largest momentum magnitude”).
Dettmers and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Zhang with the teachings of Dettmers such that the instructions further cause the programmable circuit to “update, for a plurality of batches within the single ML training epoch and based on the training loss, weights of the subnet ML model and a momentum contribution associated with each layer of a plurality of layers; and the prune of the one or more parameters from the subnet ML model includes: after processing the plurality of batches, deactivating a first fraction of active weights in each layer of the plurality of layers; and reactivating a second fraction of inactive weights in each layer of the plurality of layers based on the momentum contribution associated with the respective layer,” and one would have been motivated to do so. Doing so would allow for efficiently identifying layers and weights which reduce the error, which would increase performance levels while providing faster training (see Dettmers, Abstract).
Claims 30, 31, and 35 are rejected under 35 U.S.C. 103 as being unpatentable over Cui in view of Zagoruyko and Zhang, and further in view of Tian et al. (“Contrastive Representation Distillation”) (“Tian”).
Regarding claim 30, the rejection of claim 1 is incorporated. Zagoruyko further discloses “wherein the loss function combines the distillation loss, the task loss, and the attention transfer loss as a weighted sum, wherein… the attention transfer loss is weighted by a second hyperparameter” (Zagoruyko, 3.1 Activation-Based Attention Transfer, Paragraph 8: “Let S, T and WS, WT denote student, teacher and their weights correspondingly, and let L(W, x) denote a standard cross entropy loss. Let also I denote the indices of all teacher-student activation layer pairs for which we want to transfer attention maps. Then we can define the following total loss:
PNG
media_image4.png
204
806
media_image4.png
Greyscale
… Attention transfer can also be combined with knowledge distillation Hinton et al. (2015), in which case an additional term (corresponding to the cross entropy between softened distributions over labels of teacher and student) simply needs to be included to the above loss”; Examiner notes that L(WS, x) corresponds to a task loss, the “additional term” corresponds to a distillation loss, and
β
2
corresponds to a “second hyperparameter” which the attention transfer loss is weighted by).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zhang with Zagoruyko such that the loss function combines the distillation loss, the task loss, and the attention transfer loss as a weighted sum, wherein the attention transfer loss is weighted by a second hyperparameter, and one would have been motivated to do so. Doing so would significantly improve the performance of a student model (see Zagoruyko, Abstract).
Neither Cui, Zagoruyko, nor Zhang appear to explicitly disclose the further limitations of the claim.
However, Tian discloses “wherein the loss function combines the distillation loss, the task loss… as a weighted sum, wherein: the task loss is weighted by a first hyperparameter; the distillation loss is weighted by a complement of the first hyperparameter” (Tian, 3.2 KNOWLEDGE DISTILLATION OBJECTIVE: “The knowledge distillation loss was proposed in Hinton et al. (2015). In addition to the regular cross-entropy loss between the student output yS and one-hot label y, it asks the student network output to be as similar as possible to the teacher output by minimizing the cross-entropy between their output probabilities. The complete objective is:
PNG
media_image5.png
120
806
media_image5.png
Greyscale
”; Examiner notes that H(y,ys) corresponds to a task loss weighted by first hyperparameter (1
-
α
), and the second term of the loss function corresponds to a distillation loss weighted by a complement of the first hyperparameter,
α
)
.
Tian and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Zhang with the teachings of Tian such that the task loss is weighted by a first hyperparameter and the distillation loss is weighted by a complement of the first hyperparameter, and one would have been motivated to do so. Doing so would balance the influence that the task loss and the distillation loss have on the total loss (see Tian, 3.2).
Regarding claim 31, the rejection of claim 30 is incorporated. Zagoruyko further discloses “wherein the weighted sum is calculated by… applying the second hyperparameter as a weight to the attention transfer loss” (Zagoruyko, 3.1 Activation-Based Attention Transfer, Paragraph 8: “Let S, T and WS, WT denote student, teacher and their weights correspondingly, and let L(W, x) denote a standard cross entropy loss. Let also I denote the indices of all teacher-student activation layer pairs for which we want to transfer attention maps. Then we can define the following total loss:
PNG
media_image4.png
204
806
media_image4.png
Greyscale
… Attention transfer can also be combined with knowledge distillation Hinton et al. (2015), in which case an additional term (corresponding to the cross entropy between softened distributions over labels of teacher and student) simply needs to be included to the above loss”; Examiner notes that
β
2
corresponds to a “second hyperparameter” which the attention transfer loss is weighted by).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zhang with Zagoruyko such that the weighted sum is calculated by applying the second hyperparameter as a weight to the attention transfer loss, and one would have been motivated to do so. Doing so would significantly improve the performance of a student model (see Zagoruyko, Abstract).
Neither Cui, Zagoruyko, nor Zhang appear to explicitly disclose the further limitations of the claim.
However, Tian discloses “wherein the weighted sum is calculated by: applying the first hyperparameter as a weight to the task loss; applying a complement of the first hyperparameter as a weight to the distillation loss” (Tian, 3.2 KNOWLEDGE DISTILLATION OBJECTIVE: “The knowledge distillation loss was proposed in Hinton et al. (2015). In addition to the regular cross-entropy loss between the student output yS and one-hot label y, it asks the student network output to be as similar as possible to the teacher output by minimizing the cross-entropy between their output probabilities. The complete objective is:
PNG
media_image5.png
120
806
media_image5.png
Greyscale
”; Examiner notes that H(y,ys) corresponds to a task loss weighted by first hyperparameter (1
-
α
), and the second term of the loss function corresponds to a distillation loss weighted by a complement of the first hyperparameter,
α
)
.
Tian and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Zhang with the teachings of Tian such that the first hyperparameter is applied as a weight to the task loss and a complement of the first hyperparameter is applied as a weight to the distillation loss, and one would have been motivated to do so. Doing so would balance the influence that the task loss and the distillation loss have on the total loss (see Tian, 3.2).
Regarding claim 35, the rejection of claim 16 is incorporated. Zagoruyko further discloses “wherein the loss function combines the distillation loss, the task loss, and the attention transfer loss as a weighted sum, wherein… the attention transfer loss is weighted by a second hyperparameter” (Zagoruyko, 3.1 Activation-Based Attention Transfer, Paragraph 8: “Let S, T and WS, WT denote student, teacher and their weights correspondingly, and let L(W, x) denote a standard cross entropy loss. Let also I denote the indices of all teacher-student activation layer pairs for which we want to transfer attention maps. Then we can define the following total loss:
PNG
media_image4.png
204
806
media_image4.png
Greyscale
… Attention transfer can also be combined with knowledge distillation Hinton et al. (2015), in which case an additional term (corresponding to the cross entropy between softened distributions over labels of teacher and student) simply needs to be included to the above loss”; Examiner notes that L(WS, x) corresponds to a task loss, the “additional term” corresponds to a distillation loss, and
β
2
corresponds to a “second hyperparameter” which the attention transfer loss is weighted by).
Zagoruyko and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zhang with Zagoruyko such that the loss function combines the distillation loss, the task loss, and the attention transfer loss as a weighted sum, wherein the attention transfer loss is weighted by a second hyperparameter, and one would have been motivated to do so. Doing so would significantly improve the performance of a student model (see Zagoruyko, Abstract).
Neither Cui, Zagoruyko, nor Zhang appear to explicitly disclose the further limitations of the claim.
However, Tian discloses “wherein the loss function combines the distillation loss, the task loss as a weighted sum, wherein: the task loss is weighted by a first hyperparameter; the distillation loss is weighted by a complement of the first hyperparameter” (Tian, 3.2 KNOWLEDGE DISTILLATION OBJECTIVE: “The knowledge distillation loss was proposed in Hinton et al. (2015). In addition to the regular cross-entropy loss between the student output yS and one-hot label y, it asks the student network output to be as similar as possible to the teacher output by minimizing the cross-entropy between their output probabilities. The complete objective is:
PNG
media_image5.png
120
806
media_image5.png
Greyscale
”; Examiner notes that H(y,ys) corresponds to a task loss weighted by first hyperparameter (1
-
α
), and the second term of the loss function corresponds to a distillation loss weighted by a complement of the first hyperparameter,
α
)
.
Tian and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Cui/Zagoruyko/Zhang with the teachings of Tian such that the task loss is weighted by a first hyperparameter and the distillation loss is weighted by a complement of the first hyperparameter, and one would have been motivated to do so. Doing so would balance the influence that the task loss and the distillation loss have on the total loss (see Tian, 3.2).
Response to Arguments
Applicant’s arguments filed May 11, 2026 regarding the rejections under 35 U.S.C. 101 have been considered but are not persuasive.
Applicant argues on page 13 that the claimed operations are not mental processes, stating that “a human cannot mentally distill knowledge between neural network models comprising potentially billions of parameters, cannot mentally compute divergence between probability distributions over millions of predicted classes, and cannot mentally evaluate differences between feature maps that exist as multi-dimensional tensors.” Regarding knowledge distillation, the claims do not require that the models comprise “billions of parameters,” thus under the broadest reasonable interpretation of the claims, a human could mentally distill knowledge between neural network models given a reasonable number of parameters. Calculating a training loss is being analyzed as a mathematical concept rather than a mental process.
Applicant argues on page 14 that the abstract idea is integrated into a practical application. Applicant argues that the claimed three-component loss function “achieves training efficiency,” the distillation operating on feature maps describing activation patterns indicative of the supernet’s output “enables the smaller subnet to learn the supernet’s internal representations” and performing pruning based on the training loss within the same single training epoch as the distillation “avoids the resource-intensive iterative pruning-and-retraining cycles required by conventional approaches.” Examiner submits that these claim limitations are part of the judicial exception, and the judicial exception alone cannot provide the improvement (see MPEP 2106.05(a)). Further, the claims merely recite that the supernet has one or more feature maps describing activation patterns indicative of the supernet’s output, but do not recite that the distillation operates on them, or how the distillation operates on them.
Applicant’s arguments regarding the rejections under 35 U.S.C. 103 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GWYNEVERE A DETERDING whose telephone number is (571)272-7657. The examiner can normally be reached Mon-Fri. 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/G.A.D./Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125