Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This action is a responsive to the papers filed on 11/18/2025.
Claims 1-20 are pending.
Claims 1, 10, and 15 have been amended.
Response to Arguments
Applicant’s arguments, with respect to the specification objections have been fully considered and are persuasive. Therefore, the objections set forth in the previous office action have been withdrawn.
Applicant’s arguments, with respect to the rejection(s) of claim(s) 1-20 under 35 U.S.C. 101, have been considered but they are not persuasive. The applicant argues that the amendments overcome the 101 rejection, since they “integrate any such abstract idea into a practical application”. The examiner respectfully disagrees.
The recitations of time steps corresponding to different states are recited at a high level and do not integrate the judicial exceptions into a practical application since the amendments are deemed able to be performed in the human mind being performed on a computer; thus, are maintained as rejected under 101 abstract idea. See 35 U.S.C 101 section for full, updated analysis of claim limitations necessitated by applicant amendments.
Applicant’s arguments, with respect to the rejection(s) of claim(s) 1, 10, and 15 under 35 U.S.C. 103, have been considered but they are not persuasive. Applicant argues that no art of reference teaches the amended claim language of claims 1, 10, and 15, since Wang’s “trainings steps interval does not teach” the amended language. Upon review of the combination of prior art, the examiner respectfully disagrees due to the broadness of the claim language.
Wang, 3.5, Memory Replay Knowledge Distillation, pp. 10, paragraph 2, and Fig.1 teach; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8).” Here, the “copy step interval
κ
” encompasses a checkpoint-update frequency value used to determine each “t” or additional historical time step occur[ring] after a given historical time step. Said differently, each training step “t” corresponds to a given time step. Therefore, a first “t” corresponds to a first historical time step (historical time step that occurred before the current time step), and the next “t” determined using the “copy step interval
κ
” encompasses the additional historical time step (current step).
See 35 U.S.C 103 section for full mapping of claim limitations necessitated by applicant amendments.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2016 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”).
Claim 1
Step 1: The claim recites [a] non-transitory computer-readable medium; therefore, it is directed to the statutory category of an article of manufacture.
Step 2A Prong 1: The claim recites, inter alia:
determining a retrospective knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional elements, [a] non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional elements, generating output logits from a teacher machine learning model and generating, in a second state determined according to a current time step, output logits from a student machine learning model, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, utilizing the output logits of the student machine learning model and combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from a first state, wherein the first state is determined according to a historical time step that occurred before the current time step of the second state, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning parameters of the student machine learning model utilizing the retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites hardware elements as well as additional limitations directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 2
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites, inter alia:
determining the combined student-regularized teacher output logits utilizing an interpolation of the output logits of the teacher machine learning model and the past-state output logits of the student machine learning model: This limitation encompasses the mental process of interpolating logits from two sets of output, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, generated utilizing student machine learning model parameters from the first state, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites an additional element directed to generating an output using a machine learning model, which amounts to a mere invocation of a generic computing component. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 3
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites, inter alia:
determining a knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, utilizing outputs from the student machine learning model and outputs from the teacher machine learning model; and, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning prior parameters of the student machine learning model utilizing the knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional elements directed to generating output with machine learning models that amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 4
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites, inter alia:
identifying additional past-state output logits of the student machine learning model: This limitation encompasses the mental process of identifying logits from the output of a student model, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
determining an additional retrospective knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional elements, generated utilizing student machine learning model parameters from a third state, wherein the third state occurs after the first state and utilizing additional output logits of the student machine learning model and additional combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and the additional past-state output logits of the student machine learning model generated utilizing the student machine learning model parameters from the third state; and, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning additional parameters of the student machine learning model utilizing the additional retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the additional retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional elements directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 5
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
determining a time step of the third state utilizing a checkpoint-update frequency value: This limitation encompasses the mental process of determining a time step using an update frequency value, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject-matter ineligible.
Claim 6
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites:
determining the time step for the third state based on a remainder between a candidate time step and the checkpoint-update frequency value: This limitation encompasses the mental process of determining a time step using a remainder between a candidate time step and an update frequency value, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject-matter ineligible.
Claim 7
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim inherits the abstract idea of claim 1, from which it depends.
Step 2A Prong 2: The abstract idea inherited from claim 1 is not integrated into a practical application. Specifically, the additional element, retrieving the past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from the first state from stored memory corresponding to the student machine learning model, amounts to no more than mere data gathering and outputting, which is insignificant extra-solution activity that does not amount to an inventive concept (see MPEP § 2106.05(g) “Whether the limitation amounts to necessary data gathering and outputting, (i.e., all uses of the recited judicial exception require such data gathering or data output). See Mayo, 566 U.S. at 79, 101 USPQ2d at 1968; OIP Techs., Inc. v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1092-93 (Fed. Cir. 2015) (presenting offers and gathering statistics amounted to mere data gathering).”).
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional
elements directed to retrieving information in memory (see MPEP § 2106.05(d) “Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93”).
The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 8
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim recites, inter alia:
determining a student loss utilizing the output logits from the student machine learning model and ground truth data; and: This limitation encompasses the mathematical concept of calculating the loss between model output and ground truth data, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, learning the parameters of the student machine learning model utilizing a combination of the student loss and the retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using a combination of the student loss and the retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites an element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 9
Step 1: An article of manufacture, as above.
Step 2A Prong 1: The claim inherits the abstract idea of claim 1, from which it depends.
Step 2A Prong 2: The abstract idea inherited from claim 1 is not integrated into a practical application. Specifically, the additional element, wherein the student machine learning model comprises a smaller size than the teacher machine learning model, is a property inherent in knowledge distillation frameworks. As such, this limitation amounts to insignificant extra-solution activity (see MPEP § 2106.05(g)).
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites an additional element outlining an inherent property of knowledge distillation that therefore amounts to well-known, routine, conventional activity. The seminal work in the field of knowledge distillation, Hinton et al., which is cited by Applicant, establishes this property (Hinton et al., 1 Introduction, pp. 1, paragraph 1; “Once the cumbersome model has been trained, we can then use a different kind of training, which we call “distillation” to transfer the knowledge from the cumbersome model to a small model that is more suitable for deployment.”). The additional element listed above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 10
Step 1: The claim recites [a] system; therefore, it is directed to the statutory category of a machine.
Step 2A Prong 1: The claim recites, inter alia:
determining a retrospective knowledge distillation loss between the teacher
machine learning model and the student machine learning model by: identifying past-state output logits from the student machine learning model in a first state: This limitation encompasses the mental process of identifying logits from the output of a student model, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
comparing the student-regularized teacher output logits and output logits from the student machine learning model in the second state; and: This limitation encompasses the mental process of comparing two sets of output logits, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional elements, a memory component comprising a teacher machine learning model and a student machine learning model; and a processing device coupled to the memory component, the processing device to perform operations comprising, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, generating student-regularized teacher output logits utilizing a combination of output logits from the teacher machine learning model during a second state determined according to a current time step and the past-state output logits from the student machine learning model in the first state, wherein the first state is determined according to a historical time step that occurs prior to the current time step of to the second state; and, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning parameters of the student machine learning model utilizing the retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites hardware elements as well as additional limitations directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 11
Step 1: A machine, as above.
Step 2A Prong 1The claim recites:
generating the student-regularized teacher output logits utilizing an interpolation of the output logits from the teacher machine learning model and the past-state output logits from the student machine learning model in the first state: This limitation encompasses the mental process of interpolating logits from two sets of output, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject-matter ineligible.
Claim 12
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites, inter alia:
identifying additional past-state output logits of the student machine learning model: This limitation encompasses the mental process of identifying logits from the output of a student model, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
determining an additional retrospective knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional elements, generated utilizing student machine learning model parameters from a third state, wherein the third state occurs after the first state and utilizing additional output logits of the student machine learning model and additional combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and the additional past-state output logits of the student machine learning model generated utilizing the student machine learning model parameters from the third state; and, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning additional parameters of the student machine learning model utilizing the additional retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the additional retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional limitations directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 13
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites:
determining a time step of the third state utilizing a checkpoint-update frequency value: This limitation encompasses the mental process of determining a time step using an update frequency value, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject-matter ineligible.
Claim 14
Step 1: A machine, as above.
Step 2A Prong 1: The claim recites, inter alia:
determining a student loss utilizing the output logits from the student machine learning model and ground truth data; and: This limitation encompasses the mathematical concept of calculating the loss between model output and ground truth data, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, learning the parameters of the student machine learning model utilizing a combination of the student loss and the retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using a combination of the student loss and the retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites an element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 15
Step 1: The claim recites [a] computer-implemented method; therefore, it is directed to the statutory category of a process.
Step 2A Prong 1: The claim recites, inter alia:
identifying output logits from a teacher machine learning model: This limitation encompasses the mental process of identifying logits from the output of a teacher model, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
identifying output logits from a student machine learning model in a second state determined according to a current time step: This limitation encompasses the mental process of identifying logits from the output of a student model, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
determining a retrospective knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional element, from the output logits from the student machine learning model and combined student-regularized teacher output logits determined utilizing the output logits from the teacher machine learning model and historical output logits from the student machine learning model in a first state, wherein the first state is determined according to a historical time step that occurred before the current time step of the second state, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning parameters of the student machine learning model utilizing the retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional limitations directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 16
Step 1: A process, as above.
Step 2A Prong 1: The claim recites:
determining the combined student-regularized teacher output logits utilizing an interpolation of the output logits from the teacher machine learning model and the historical output logits from the student machine learning model: This limitation encompasses the mental process of interpolating logits from two sets of output, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject-matter ineligible.
Claim 17
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
determining a knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, utilizing the output logits from the teacher machine learning model and prior output logits from the student machine learning model; and, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning prior parameters of the student machine learning model utilizing the knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional limitations directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 18
Step 1: A process, as above.
Step 2A Prong 1: The claim recites:
identifying the historical output logits from the student machine learning model from a historical time step of the student machine learning model: This limitation encompasses the mental process of identifying logits from past output of a student model, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject-matter ineligible.
Claim 19
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
identifying additional historical output logits from the student machine learning model from an additional historical time step of the student machine learning model utilizing a checkpoint-update frequency value, wherein the additional historical time step occurs after the historical time step: This limitation encompasses the mental process of identifying logits from past output of a student model using an update frequency value, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
determining an additional retrospective knowledge distillation loss: This limitation encompasses the mathematical concept of calculating a retrospective knowledge distillation loss, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional element, from additional output logits from the student machine learning model and additional combined student-regularized teacher output logits determined utilizing the output logits from the teacher machine learning model and the additional historical output logits from the student machine learning model; and, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)).
The additional element, learning additional parameters of the student machine learning model utilizing the additional retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using the additional retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional limitations directed to generating output with machine learning models, which amount to mere invocations of generic computing components. The claim recites a further element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim 20
Step 1: A process, as above.
Step 2A Prong 1: The claim recites, inter alia:
determining a student loss utilizing the output logits from the student machine learning model and ground truth data; and: This limitation encompasses the mathematical concept of calculating the loss between model output and ground truth data, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper.
Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, learning the parameters of the student machine learning model utilizing a combination of the student loss and the retrospective knowledge distillation loss, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the claim language how the learning takes place, only that it is done using a combination of the student loss and the retrospective knowledge distillation loss.
Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites an element directed to model training which is recited at a high level of generality such that it amounts to a mere recitation of a solution or outcome. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Haidar et al. (US 2022/0335303 A1) hereinafter Haidar, in view of Wang et al. (“Memory-Replay Knowledge Distillation”, 2021) hereinafter Wang.
Regarding claim 1, Haidar teaches [a] non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: (Haidar, [0038]; “In some aspects, the present disclosure provides a non-transitory processor-readable medium containing instructions which, when executed by a processor of a device, cause the device to use knowledge distillation to train a student model, comprising a first number of student intermediate layers, to perform an inference task.”).
generating output logits from a teacher machine learning model; (Haidar, [0073]; “The method 500 then performs a first training epoch consisting of operations 504 through 514. At 504, the input data of each labeled training data sample of a batch of labeled training data samples (referred to herein as training batch) obtained from the training dataset 240 is forward propagated through the teacher model 234 to generate, for the input data of each labeled training data sample in the batch of training data, a teacher prediction (i.e. teacher inference data 24 from FIG. 1), shown in FIG. 3 as teacher predicted logits 310.”)
generating, , output logits from a student machine learning model; (Haidar, [0074]; “At 506, the student model 234 processes the training batch 302 (i.e. the input data of each labeled training data sample in the training batch) to generate, for the input data of each labeled training data sample in the training batch 302, a student prediction (i.e. student inference data 34 from FIG. 1), shown in FIG. 3 as student predicted logits 306,” wherein the output layer of “the student model” encompasses a second state.)
determining a retrospective knowledge distillation loss utilizing the combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from a first state, ; and (Haidar, Fig. 3;
PNG
media_image2.png
934
710
media_image2.png
Greyscale
Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate representation” encompasses past-state output logits…from a first state relative to the student model’s final output generated at a second state. Haidar, [0076]; “At 512, the intermediate representation loss module 222 processes the teacher intermediate representations and the student intermediate representations for the training batch 302 to compute an intermediate representation loss 316,” thereby determining a retrospective knowledge distillation loss from combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and past-state output logits of the student machine learning model. Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating” further specifies that the output logits of the teacher and student models are combined.).
Haidar does not explicitly teach in a second state determined according to a current time step; the output logits of the student machine learning model and; and wherein the first state is determined according to a historical time step that occurred before the current time step of the second state. However, Wang, in the area of knowledge distillation, teaches these limitations (Wang, Fig. 1;
PNG
media_image3.png
678
800
media_image3.png
Greyscale
As depicted in the figure above, the memory-replay knowledge distillation method of Wang compares ensembled prior-state backup outputs with the output logits of the “current” student machine learning model. Wang, 3.5. Memory Replay Knowledge Distillation, pp. 10, paragraph 2; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8).” Here, the “copy step interval
κ
” encompasses a checkpoint-update frequency value used to determine each “t” or additional historical time step occur[ring] after a given historical time step. Said differently, each training step “t” corresponds to a given time step. Therefore, a first “t” corresponds to a first historical time step (historical time step that occurred before the current time step), and the next “t” determined using the “copy step interval
κ
” encompasses the additional historical time step (current step).).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the current student model output of Wang into the calculation of the intermediate representation loss of Haidar. The motivation to do so is to stabilize the training process of the student model by regularizing it with its former output distributions (Wang, Abstract; “Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model’s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning.”).
Haidar further teaches learning parameters of the student machine learning model utilizing the retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 2, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 1 (and thus the rejection of claim 1 is incorporated).
Haidar further teaches wherein the operations further comprise determining the combined student-regularized teacher output logits utilizing an interpolation of the output logits of the teacher machine learning model and the past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from the first state (Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating the sets of intermediate representations” encompasses an interpolation of the output logits of the teacher machine learning model and the past-state output logits of the student machine learning model. Note that no explicit definition is given for the term interpolation in the specification of the claimed invention; rather, the method by which the two sets of output logits are combined is provided as an open-ended lest of possible embodiments, “the retrospective knowledge distillation learning system 106, in some cases, utilizes various output composition functions, such as, but not limited to, inverse distance weighted interpolation, spline interpolation, multiplication, and/or averaging.”).
Regarding claim 3, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 1 (and thus the rejection of claim 1 is incorporated).
Haidar further teaches wherein the operations further comprise, prior to utilizing the retrospective knowledge distillation loss: determining a knowledge distillation loss utilizing outputs from the student machine learning model and outputs from the teacher machine learning model; and (Haidar, Fig. 3; As depicted in the figure, outputs from the student machine learning model and outputs from the teacher machine learning model are fed into “KD loss module 216” for determining a knowledge distillation loss prior to utilizing “the intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss at the “weighted loss module 220.”).
learning prior parameters of the student machine learning model utilizing the knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “KD loss 314” corresponding to the knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning prior parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 4, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 1 (and thus the rejection of claim 1 is incorporated).
Haidar further teaches identifying additional past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from a third state, wherein the third state occurs after the first state; (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers,” thereby identifying additional past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from a third state. Said differently, the student machine learning model produces past-state output logits at each of its hidden layers or states. Therefore, there is necessarily a third state [that] occurs after the first state.).
determining an additional retrospective knowledge distillation loss utilizing… additional combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and the additional past-state output logits of the student machine learning model generated utilizing the student machine learning model parameters from the third state; and (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate representation” encompasses additional past-state output logits…from the third state. Haidar, [0076]; “At 512, the intermediate representation loss module 222 processes the teacher intermediate representations and the student intermediate representations for the training batch 302 to compute an intermediate representation loss 316,” thereby determining an additional retrospective knowledge distillation loss from additional combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and past-state output logits of the student machine learning model. Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating” further specifies that the output logits of the teacher and student models are combined.).
Haidar does not explicitly teach additional output logits of the student machine learning model and. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Fig. 1; As depicted in the figure, the memory-replay knowledge distillation method of Wang combines ensembled prior-state backup outputs with the output logits of the “current” student machine learning model. Wang, 3.5. Memory Replay Knowledge Distillation, pp. 10, paragraph 2; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8),” wherein “
p
θ
;
τ
” encompasses additional output logits of the student machine learning model relative to “
p
θ
;
τ
=
1
.
”).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the current student model output of Wang into the calculation of the intermediate representation loss of Haidar. The motivation to do so is to stabilize the training process of the student model by regularizing it with its former output distributions (Wang, Abstract; “Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model’s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning.”).
Haidar further teaches learning additional parameters of the student machine learning model utilizing the additional retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “intermediate representation loss 316” corresponding to the additional retrospective knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning additional parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 5, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 4 (and thus the rejection of claim 4 is incorporated).
Haidar does not explicitly teach wherein the operations further comprise determining a time step of the third state utilizing a checkpoint-update frequency value. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Algorithm 2;
PNG
media_image4.png
644
649
media_image4.png
Greyscale
Here, the “copy step interval
κ
” encompasses a checkpoint-update frequency value used to determine each “t” from the “total training steps T” encompassing a time step of the third state.).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the intermediate representation loss calculation method of Haidar to incorporate the copy step interval of Wang. The motivation to do so is to provide a further mechanism with which to control the tradeoff between stability and variance when updating the student model (Wang, 4.1. CIFAR-100, pp. 12, paragraph 3; “In Figure 7, we can see that if the step interval of the model backups was quite small, the error rate rose because the copy was too similar to the current model, then the regularization would not be helpful and may stumble the current model from learning. On the other hand, if the step was too large, the copies would be worse and lagging, then MrKD would also mislead and destabilize the learning.”).
Regarding claim 6, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 5 (and thus the rejection of claim 5 is incorporated).
Haidar does not explicitly teach wherein the operations further comprise determining the time step for the third state based on a remainder between a candidate time step and the checkpoint-update frequency value. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Algorithm 2; Here, the “copy step interval
κ
” encompasses a checkpoint-update frequency value used to determine each “t” from the “total training steps T” encompassing a time step of the third state. More specifically, the time step for the third state is calculated by taking the modulus or remainder between a candidate time step and the “copy step interval
κ
” corresponding to the checkpoint-update frequency value.).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the intermediate representation loss calculation method of Haidar to incorporate the copy step interval of Wang. The motivation to do so is to provide a further mechanism with which to control the tradeoff between stability and variance when updating the student model (Wang, 4.1. CIFAR-100, pp. 12, paragraph 3; “In Figure 7, we can see that if the step interval of the model backups was quite small, the error rate rose because the copy was too similar to the current model, then the regularization would not be helpful and may stumble the current model from learning. On the other hand, if the step was too large, the copies would be worse and lagging, then MrKD would also mislead and destabilize the learning.”).
Regarding claim 7, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 1 (and thus the rejection of claim 1 is incorporated).
Haidar further teaches wherein the operations further comprise retrieving the past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from the first state from stored memory corresponding to the student machine learning model (Haidar, Fig. 2;
PNG
media_image5.png
685
877
media_image5.png
Greyscale
As depicted in the figure above, the “intermediate layer KD system 300” housing the “student model 234” is located within “memory 208.” Therefore, the “intermediate representations” or past-state output logits of the student machine learning model [are] generated utilizing student machine learning model parameters from the first state from stored memory corresponding to the student machine learning model. Haidar, [0063]; “The memory 208 may also store the student model 234 and teacher model 232, each of which may include values for a plurality of learnable parameters (referred to herein as "learnable parameter values"), as well as values for a plurality of hyperparameters (referred to herein as ‘hyper-parameter values’) used to control the structure and operation of the student model 234 and teacher model 232,” further confirming that “the memory 208” stores the “learnable parameters” corresponding to the student machine learning model.).
Regarding claim 8, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 1 (and thus the rejection of claim 1 is incorporated).
Haidar further teaches wherein the operations further comprise: determining a student loss utilizing the output logits from the student machine learning model and ground truth data; and (Haidar, Fig. 3; As depicted in the figure, the “student predicted logits” are used to calculate “ground truth loss 312” corresponding to a student loss.)
learning the parameters of the student machine learning model utilizing a combination of the student loss and the retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “weighted loss 330” is calculated utilizing a combination of the “ground truth loss 312” corresponding to the student loss and the “intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss.).
Regarding claim 9, the combination of Haidar and Wang teaches [t]he non-transitory computer-readable medium of claim 1 (and thus the rejection of claim 1 is incorporated).
Haidar further teaches wherein the student machine learning model comprises a smaller size than the teacher machine learning model (Haidar, [0004]; “KD utilizes the generalization ability of the larger trained neural network model (referred to as the "teacher model" or "teacher") using the inference data output by the larger trained model as "soft targets", which are used as a supervision signal for training a smaller neural network model ( called the "student model" or "student").”).
Regarding claim 10, Haidar teaches [a] system comprising: a memory component comprising a teacher machine learning model and a student machine learning model; and a processing device coupled to the memory component, the processing device to perform operations comprising: (Haidar, Fig. 2; “Processor 202,” “memory 208,” “teacher model 232” and “student model 234.”).
determining a retrospective knowledge distillation loss between the teacher machine learning model and the student machine learning model by: identifying past-state output logits from the student machine learning model in a first state; (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers.”)
generating student-regularized teacher output logits utilizing a combination of output logits from the teacher machine learning model during a second state and the past-state output logits from the student machine learning model in the first state, wherein the first state occurs prior to the second state; and (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate representation” encompasses past-state output logits…in the first state relative to the student model’s final output generated at the second state. Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” thereby generating student-regularized teacher output logits utilizing a combination of output logits from the teacher machine learning model…and the past-state output logits from the student machine learning model.).
Haidar does not explicitly teach comparing the student-regularized teacher output logits and output logits from the student machine learning model in the second state; and. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Fig. 1;
As depicted in the figure, the memory-replay knowledge distillation method of Wang compar[es] ensembled prior-state backup outputs corresponding to the student-regularized teacher output logits with the output logits of the “current” student machine learning model. Wang, 3.5. Memory Replay Knowledge Distillation, pp. 10, paragraph 2; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8).”).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the current student model output of Wang into the calculation of the intermediate representation loss of Haidar. The motivation to do so is to stabilize the training process of the student model by regularizing it with its former output distributions (Wang, Abstract; “Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model’s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning.”).
Haidar further teaches learning parameters of the student machine learning model utilizing the retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 11, the combination of Haidar and Wang teaches [t]he system of claim 10 (and thus the rejection of claim 10 is incorporated).
Haidar further teaches wherein the operations further comprise generating the student-regularized teacher output logits utilizing an interpolation of the output logits from the teacher machine learning model and the past-state output logits from the student machine learning model in the first state (Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating the sets of intermediate representations” encompasses an interpolation of the output logits of the teacher machine learning model and the past-state output logits of the student machine learning model. Note that no explicit definition is given for the term interpolation in the specification of the claimed invention; rather, the method by which the two sets of output logits are combined is provided as an open-ended lest of possible embodiments, “the retrospective knowledge distillation learning system 106, in some cases, utilizes various output composition functions, such as, but not limited to, inverse distance weighted interpolation, spline interpolation, multiplication, and/or averaging.”).
Regarding claim 12, the combination of Haidar and Wang teaches [t]he system of claim 10 (and thus the rejection of claim 10 is incorporated).
Haidar further teaches identifying additional past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from a third state, wherein the third state occurs after the first state; (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers,” thereby identifying additional past-state output logits of the student machine learning model generated utilizing student machine learning model parameters from a third state. Said differently, the student machine learning model produces past-state output logits at each of its hidden layers or states. Therefore, there is necessarily a third state [that] occurs after the first state.).
determining an additional retrospective knowledge distillation loss utilizing… additional combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and the additional past-state output logits of the student machine learning model generated utilizing the student machine learning model parameters from the third state; and (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate representation” encompasses additional past-state output logits…from the third state. Haidar, [0076]; “At 512, the intermediate representation loss module 222 processes the teacher intermediate representations and the student intermediate representations for the training batch 302 to compute an intermediate representation loss 316,” thereby determining an additional retrospective knowledge distillation loss from additional combined student-regularized teacher output logits based on the output logits of the teacher machine learning model and past-state output logits of the student machine learning model. Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating” further specifies that the output logits of the teacher and student models are combined.).
Haidar does not explicitly teach additional output logits of the student machine learning model and. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Fig. 1; As depicted in the figure, the memory-replay knowledge distillation method of Wang combines ensembled prior-state backup outputs with the output logits of the “current” student machine learning model. Wang, 3.5. Memory Replay Knowledge Distillation, pp. 10, paragraph 2; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8),” wherein “
p
θ
;
τ
” encompasses additional output logits of the student machine learning model relative to “
p
θ
;
τ
=
1
.
”).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the current student model output of Wang into the calculation of the intermediate representation loss of Haidar. The motivation to do so is to stabilize the training process of the student model by regularizing it with its former output distributions (Wang, Abstract; “Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model’s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning.”).
Haidar further teaches learning additional parameters of the student machine learning model utilizing the additional retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “intermediate representation loss 316” corresponding to the additional retrospective knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning additional parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 13, the combination of Haidar and Wang teaches [t]he system of claim 12 (and thus the rejection of claim 12 is incorporated).
Haidar does not explicitly teach wherein the operations further comprise determining a time step of the third state utilizing a checkpoint-update frequency value. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Algorithm 2; Here, the “copy step interval
κ
” encompasses a checkpoint-update frequency value used to determine each “t” from the “total training steps T” encompassing a time step of the third state.).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the intermediate representation loss calculation method of Haidar to incorporate the copy step interval of Wang. The motivation to do so is to provide a further mechanism with which to control the tradeoff between stability and variance when updating the student model (Wang, 4.1. CIFAR-100, pp. 12, paragraph 3; “In Figure 7, we can see that if the step interval of the model backups was quite small, the error rate rose because the copy was too similar to the current model, then the regularization would not be helpful and may stumble the current model from learning. On the other hand, if the step was too large, the copies would be worse and lagging, then MrKD would also mislead and destabilize the learning.”).
Regarding claim 14, the combination of Haidar and Wang teaches [t]he system of claim 10 (and thus the rejection of claim 10 is incorporated).
Haidar further teaches wherein the operations further comprise: determining a student loss utilizing the output logits from the student machine learning model and ground truth data; and (Haidar, Fig. 3; As depicted in the figure, the “student predicted logits” are used to calculate “ground truth loss 312” corresponding to a student loss.)
learning the parameters of the student machine learning model utilizing a combination of the student loss and the retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “weighted loss 330” is calculated utilizing a combination of the “ground truth loss 312” corresponding to the student loss and the “intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss.).
Regarding claim 15, Haidar teaches [a] computer-implemented method comprising identifying output logits from a teacher machine learning model; (Haidar, [0073]; “The method 500 then performs a first training epoch consisting of operations 504 through 514. At 504, the input data of each labeled training data sample of a batch of labeled training data samples (referred to herein as training batch) obtained from the training dataset 240 is forward propagated through the teacher model 234 to generate, for the input data of each labeled training data sample in the batch of training data, a teacher prediction (i.e. teacher inference data 24 from FIG. 1), shown in FIG. 3 as teacher predicted logits 310.”)
identifying output logits from a student machine learning model; (Haidar, [0074]; “At 506, the student model 234 processes the training batch 302 (i.e. the input data of each labeled training data sample in the training batch) to generate, for the input data of each labeled training data sample in the training batch 302, a student prediction (i.e. student inference data 34 from FIG. 1), shown in FIG. 3 as student predicted logits 306.”)
determining a retrospective knowledge distillation loss from the…combined student-regularized teacher output logits determined utilizing the output logits from the teacher machine learning model and historical output logits from the student machine learning model generated utilizing student machine learning model; and (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate representation” encompasses historical output logits relative to the student model’s final output. Haidar, [0076]; “At 512, the intermediate representation loss module 222 processes the teacher intermediate representations and the student intermediate representations for the training batch 302 to compute an intermediate representation loss 316,” thereby determining a retrospective knowledge distillation loss from combined student-regularized teacher output logits determined utilizing the output logits of the teacher machine learning model and historical output logits of the student machine learning model. Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating” further specifies that the output logits of the teacher and student models are combined.).
Haidar does not explicitly teach the output logits from the student machine learning model and. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Fig. 1; As depicted in the figure, the memory-replay knowledge distillation method of Wang compares ensembled prior-state backup outputs with the output logits of the “current” student machine learning model. Wang, 3.5. Memory Replay Knowledge Distillation, pp. 10, paragraph 2; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8).”).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the current student model output of Wang into the calculation of the intermediate representation loss of Haidar. The motivation to do so is to stabilize the training process of the student model by regularizing it with its former output distributions (Wang, Abstract; “Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model’s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning.”).
Haidar further teaches learning parameters of the student machine learning model utilizing the retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 16, the combination of Haidar and Wang teaches [t]he computer-implemented method of claim 15 (and thus the rejection of claim 15 is incorporated).
Haidar further teaches further comprising determining the student-regularized teacher output logits utilizing an interpolation of the output logits from the teacher machine learning model and the historical output logits from the student machine learning model in the first state (Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating the sets of intermediate representations” encompasses an interpolation of the output logits of the teacher machine learning model and the historical output logits of the student machine learning model. Note that no explicit definition is given for the term interpolation in the specification of the claimed invention; rather, the method by which the two sets of output logits are combined is provided as an open-ended lest of possible embodiments, “the retrospective knowledge distillation learning system 106, in some cases, utilizes various output composition functions, such as, but not limited to, inverse distance weighted interpolation, spline interpolation, multiplication, and/or averaging.”).
claim 17, the combination of Haidar and Wang teaches [t]he computer-implemented method of claim 15 (and thus the rejection of claim 15 is incorporated).
Haidar further teaches further comprising, prior to utilizing the retrospective knowledge distillation loss: determining a knowledge distillation loss utilizing outputs from the student machine learning model and prior output logits from the teacher machine learning model; and (Haidar, Fig. 3; As depicted in the figure, outputs from the student machine learning model and prior outputs from the teacher machine learning model are fed into “KD loss module 216” for determining a knowledge distillation loss prior to utilizing “the intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss at the “weighted loss module 220.”).
learning prior parameters of the student machine learning model utilizing the knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “KD loss 314” corresponding to the knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning prior parameters of the student machine learning model via the “gradient descent module 218.”).
Regarding claim 18, the combination of Haidar and Wang teaches [t]he system of claim 15 (and thus the rejection of claim 15 is incorporated).
Haidar further teaches further comprising identifying the historical output logits from the student machine learning model from a historical time step of the student machine learning model (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate layers” encompass historical time step[s] of the student machine learning model relative to the final output.).
Regarding claim 19, the combination of Haidar and Wang teaches [t]he system of claim 18 (and thus the rejection of claim 18 is incorporated).
Haidar further teaches further comprising: identifying additional historical output logits from the student machine learning model from an additional historical time step of the student machine learning model (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers,” thereby identifying additional historical output logits from the student machine learning model from an additional historical time step of the student machine learning model. Said differently, the student machine learning model produces historical output logits at each of its hidden layers or states. Therefore, there is necessarily an additional historical time step after the initial historical time step.).
Haidar does not explicitly teach utilizing a checkpoint-update frequency value, wherein the additional historical time step occurs after the historical time step. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Algorithm 2; Here, the “copy step interval
κ
” encompasses a checkpoint-update frequency value used to determine each “t” or additional historical time step occur[ring] after a given historical time step. Said differently, each training step “t” corresponds to a given time step. Therefore, a first “t” corresponds to a first historical time step, and the next “t” determined using the “copy step interval
κ
” encompasses the additional historical time step.)
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the intermediate representation loss calculation method of Haidar to incorporate the copy step interval of Wang. The motivation to do so is to provide a further mechanism with which to control the tradeoff between stability and variance when updating the student model (Wang, 4.1. CIFAR-100, pp. 12, paragraph 3; “In Figure 7, we can see that if the step interval of the model backups was quite small, the error rate rose because the copy was too similar to the current model, then the regularization would not be helpful and may stumble the current model from learning. On the other hand, if the step was too large, the copies would be worse and lagging, then MrKD would also mislead and destabilize the learning.”).
Haidar further teaches determining an additional retrospective knowledge distillation loss utilizing…additional combined student-regularized teacher output logits determined utilizing the output logits from the teacher machine learning model and the additional historical output logits of the student machine learning model; and (Haidar, [0075]; “At 510, a student intermediate representation is obtained from each of the m student intermediate layers. Each student intermediate representation is generated by its respective student intermediate layer based on the student intermediate representation received from a previous student intermediate layer,” wherein a “student intermediate representation” encompasses additional historical output logits. Haidar, [0076]; “At 512, the intermediate representation loss module 222 processes the teacher intermediate representations and the student intermediate representations for the training batch 302 to compute an intermediate representation loss 316,” thereby determining an additional retrospective knowledge distillation loss from additional combined student-regularized teacher output logits determined utilizing the output logits of the teacher machine learning model and historical output logits of the student machine learning model. Haidar, [0089]; “The layer concatenating intermediate representation loss module 222a operates by concatenating the sets of intermediate representations from each of the teacher and student model 232, 234 and mapping them to a common information space in order to compare them to generate the intermediate representations loss 316,” wherein “concatenating” further specifies that the output logits of the teacher and student models are combined.).
Haidar does not explicitly teach additional output logits of the student machine learning model and. However, Wang, in the area of knowledge distillation, teaches this limitation (Wang, Fig. 1; As depicted in the figure, the memory-replay knowledge distillation method of Wang combines ensembled prior-state backup outputs with the output logits of the “current” student machine learning model. Wang, 3.5. Memory Replay Knowledge Distillation, pp. 10, paragraph 2; “Over all, the total loss of MrKD with FCN and KA is as follows:
PNG
media_image1.png
40
548
media_image1.png
Greyscale
Note that
p
c
o
r
r
e
c
t
(
ϕ
;
τ
)
is the soft target of
z
c
o
r
r
e
c
t
(
ϕ
)
obtain by Equation (1). Additionally, the parameter
ϕ
for FCN is also trained by the second term of Equation (8).” Wang, Algorithm 2; “Feed
z
(
θ
^
1
)
,
…
,
z
(
θ
^
n
)
to Fully Connected Network and get logits
z
e
n
s
e
m
b
l
e
(
ϕ
)
; Correct the value
z
e
n
s
e
m
b
l
e
(
ϕ
)
to
z
c
o
r
r
e
c
t
(
ϕ
)
refer to label
p
by Knowledge Adjustment method; Compute the predictions
p
θ
;
τ
=
1
,
p
θ
;
τ
,
p
e
n
s
e
m
b
l
e
(
ϕ
;
τ
)
by Equation (1); Compute loss
L
M
r
K
D
(
θ
)
by Equation (8),” wherein “
p
θ
;
τ
” encompasses additional output logits of the student machine learning model relative to “
p
θ
;
τ
=
1
.
”).
Wang is analogous to the claimed invention as both are from the same field of endeavor, that is, knowledge distillation utilizing output from prior-state models. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the current student model output of Wang into the calculation of the intermediate representation loss of Haidar. The motivation to do so is to stabilize the training process of the student model by regularizing it with its former output distributions (Wang, Abstract; “Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model’s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning.”).
Haidar further teaches learning additional parameters of the student machine learning model utilizing the additional retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “intermediate representation loss 316” corresponding to the additional retrospective knowledge distillation loss is used to calculate the “weighted loss 330” which is in turn used for learning additional
Regarding claim 20, the combination of Haidar and Wang teaches [t]he system of claim 15 (and thus the rejection of claim 15 is incorporated).
Haidar further teaches further comprising: determining a student loss utilizing the output logits from the student machine learning model and ground truth data; and (Haidar, Fig. 3; As depicted in the figure, the “student predicted logits” are used to calculate “ground truth loss 312” corresponding to a student loss.)
learning the parameters of the student machine learning model utilizing a combination of the student loss and the retrospective knowledge distillation loss (Haidar, Fig. 3; As depicted in the figure, the “weighted loss 330” is calculated utilizing a combination of the “ground truth loss 312” corresponding to the student loss and the “intermediate representation loss 316” corresponding to the retrospective knowledge distillation loss.).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CLINT MULLINAX whose telephone number is 571-272-3241. The examiner can normally be reached on Mon - Fri 8:00-4:30 PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached on 571-270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/C.M./Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123