Prosecution Insights
Last updated: October 01, 2026
Application No. 18/486,792

Efficient Knowledge Distillation Framework for Training Machine-Learned Models

Non-Final OA §101§103§112
Filed
Oct 13, 2023
Examiner
BLAIN, SHAWN MICHAEL
Art Unit
4100
Tech Center
4100
Assignee
Google LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
2 currently pending
Career history
2
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §112
DETAILED ACTION This action is in response to the original filing on 10/13/2023. Claims 1 – 20 are pending and have been considered below. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The disclosure is objected to because of the following informalities: Page 23, Paragraph 0104, Unclear when making reference to Fig. 2D as seen in previous references pages 18, 19, paragraph 0082, 0083. Page 27 Improper reference to Multiscale Refinement Objective 222 (Labelled in Fig. 2D as 210) Student distribution 222 Teacher distribution 224 Output-level refinement learning signal 220 (Labelled in Fig. 2D as 216) Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION. —The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 2, line 3, the phrase “the value” renders the claim indefinite because it is uncertain whether “the value” referred here is “the student value” or “the teacher value” or another value that is different from “the student value” and “the teacher value”. For the purpose of examination, examiner will interpret the limitation as “the student value”. Regarding Claim 3, line 7 recites the limitation “the portion” lacks antecedent basis. Regarding Claim 3, line 8, the phrase recites the limitation “the plurality of portion-specific divergence metrics” lacks antecedent basis. Regarding claim 4, line 2, the term “at least partially in parallel ” is a relative term which renders the claim indefinite. The term “at least partially in parallel ” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claim 1 - 20 rejected under 35 U.S.C. 101 because the claimed invention is directed to a an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised paten Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”). Claims 1 Step 1: The claim recites “a method for…, comprising:”; therefore, it is direct to the statutory category of a process. Step2A Prong 1: The claim recites, inter alia: generating a multiscale refinement objective configured to jointly distill knowledge from a teacher machine-learned sequence processing model and reinforce preferred behavior of the student machine-learned sequence processing model, wherein the multiscale refinement objective comprises: a first component based on a divergence metric characterizing, for the respective input, a comparison of a plurality of predictions of the student machine-learned sequence processing model to a plurality of predictions of the teacher machine-learned sequence processing model; and a second component based on a reinforcement learning signal associated with the respective output: Under its broadest reasonable interpretations in light of the specification, this limitation encompasses a mathematical relationship similar to organizing information and manipulating information, e.g. generating a multiscale refinement objective, a first component for the respective input, a comparison of a plurality of predictions of the student machine-learned sequence processing model to a plurality of predictions of the teacher machine-learned sequence processing model, and a second component based on a reinforcement learning signal associated with the respective output, through mathematical correlations, e.g. based on a divergence metric characterizing, described in MPEP 2106.04(a)(2)(I)(A)(iv). Thus, this claim recites a judicial exception. Step 2A Prong 2: The judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: obtaining a respective input; obtaining, from the student machine-learned sequence processing model, a respective output corresponding to the respective input: These additional elements are recited at a high level of generality and merely recites insignificant extra-solution activity of data gathering of obtaining a respective input and selecting a particular data source of a respective output from the student machine-learned sequence processing model to be manipulated. See MPEP 2106.05(g). and updating the machine-learned student sequence processing model based on the multiscale refinement objective. These additional elements are mere additional instructions to implement the judicial exception because the additional elements only recite the idea of a solution or outcome, e.g. updating the machine-learning student sequence processing model, but fails to recite details of how the updating is achieved based on the multiscale refinement objective, i.e. there is no details as to whether or not the updating is done by other components in additional than first and second components or using just one of the components/etc., to provide meaningful limitations to the claimed invention. See MPEP 2106.05(f). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional element from Step 2A Prong 2 include insignificant extra-solution activities of data gathering and selecting a particular data source recited by “obtaining a respective input; obtaining, from the student machine-learned sequence processing model, a respective output corresponding to the respective input:” These insignificant extra solution activity well understood routine and conventional activity similar to presenting offers and gathering statistics as described in MPEP 2106.05(d)(II). Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 2 Step 1: a process, as above. Step 2A Prong 1: The claim recites, inter alia: wherein the divergence metric is evaluated using: a student value generated by the machine-learned student sequence processing model for one or more portions of the respective output based on the respective input, the value corresponding to a student probability of the one or more portions of the respective output conditioned on the respective input; and a teacher value generated by the machine-learned teacher sequence processing model for the one or more portions of the respective output based on the respective input, the teacher value corresponding to a teacher probability of the one or more portions of the respective output conditioned on the respective input.: Under its broadest reasonable interpretations in light of the specification, this limitation encompasses a mathematical relationship similar to organizing information and manipulating information through mathematical correlations. described in MPEP 2106.04(a)(2)(I)(A)(iv). Thus, this claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 3 Step 1: a process, as above. Step 2A Prong 1: The claim recites, inter alia: comprising: for each portion of a plurality of portions of the respective output:determining a portion-specific divergence metric that characterizes a similarity between a student probability distribution over a set of candidate output portions and a teacher probability distribution over the set of candidate output portions, wherein each of the student probability distribution and the teacher probability distribution are conditioned on the respective input and one or more portions of the respective output that precede the portion; and aggregating the plurality of portion-specific divergence metrics for the respective output to obtain the first component. Under its broadest reasonable interpretations in light of the specification, this limitation encompasses a mathematical calculation similar to using an algorithm for calculating values then performing an operation to obtain a value, Additionally, Aggregating is summation operation, e.g. determining a portion-specific divergence metric, a student probability distribution over a set of candidate output, a teacher probability distribution over the set of candidate output, aggregating, the plurality of portion-specific divergence metrics for the respective output to obtain the first component. As described in MPEP 2106.04(a)(2)(I)(C)(iv). Thus, this claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 4 Step 1: a process, as above. Step 2A Prong 1: The claim recites, inter alia: wherein …, model: this limitation furthers a mathematical calculation in claim 3. Step 2A Prong 2: The judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the teacher probability distributions for each of the portion-specific divergence metrics are generated at least partially in parallel by the machine-learned teacher sequence processing model. These additional elements recite no more than generally linking a judicial exception, the abstract idea of the mathematical concepts of the teacher probability distributions for each of the portion-specific divergence metrics, to a particular technological environment/ Field of Use, are generated at least partially in parallel by the machine-learned teacher sequence processing model. See MPEP 2106.05(h). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 5 Step 1: a process, as above. Step 2A Prong 1: The claim reciting. Inter alia: wherein the multiscale refinement objective comprises: this limitation furthers a mathematical relationship and claim 1. Step 2A Prong 2: The judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the multiscale refinement objective comprises one or more weighting parameters that weight the respective contributions of the first component and the second component. These additional elements recite no more than generally linking a judicial exception, the abstract idea of the mathematical concepts of the teacher probability distributions for each of the portion-specific divergence metrics, to a particular technological environment/ Field of Use, are generated at least partially in parallel by the machine-learned teacher sequence processing model. See MPEP 2106.05(h). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 6 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical relationship as in claim 1. Step 2A Prong 2: The judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the reinforcement learning signal comprises data indicating human feedback on an overall quality of the respective output. These additional elements recite no more than generally linking a judicial exception to a particular technological environment/ Field of Use. Similar to limiting application of the abstract idea to the reinforcement learning signal data is simply an attempt to limit the use of the abstract idea to a particular technological environment, e.g. wherein the reinforcement learning signal comprises data indicating human feedback on an overall quality of the respective output. See MPEP 2106.05(h). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 7 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical relationship as in claim 1. Step 2A Prong 2: The judicial exception is not integrated into a practical application. The additional elements of the claim are as follows: wherein the reinforcement learning signal comprises data indicating a score generated by a machine-learned reward model, wherein the score indicates an overall quality of the respective output. These additional elements recite no more than generally linking a judicial exception, the abstract idea of mathematical relationships to the reinforcement learning signal a particular technological environment/ Field of Use, comprises data indicating a score generated by a machine-learned reward model, wherein the score indicates an overall quality of the respective output. See MPEP 2106.05(h). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 8 Step 1: a process, as above. Step 2A Prong 1: The claim recites, inter alia: wherein evaluating the divergence metric comprises: determining a value of a mixture distribution corresponding to a mixture of a student probability distribution of the machine-learned student sequence processing model and a teacher probability distribution of the machine-learned teacher sequence processing model; computing a first divergence component that characterizes a divergence of the student probability distribution with respect to the mixture distribution; computing a second divergence component that characterizes a divergence of the teacher probability distribution with respect to the mixture distribution; and evaluating the divergence metric based on a combination of the first divergence component and the second divergence component: These limitations recite a high level of mathematical calculations to determine a value corresponding to mathematical formulas, similar to using mathematical formulas to preforming operations in order to obtain a desired value, e.g. determining a value of a mixture distribution corresponding to a mixture of a student probability distribution of the machine-learned student sequence processing model and a teacher probability distribution of the machine-learned teacher sequence processing model, computing a first divergence component… characterizes a divergence of the student probability distribution with respect to the mixture distribution, computing a second divergence component , characterizes a divergence of the teacher probability distribution with respect to the mixture distribution, evaluating the divergence metric based on a combination of the first divergence component and the second divergence component, as described in MPEP 2106.04(a)(2)(I)(C)(ii). Thus, this claim furthers and recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 9 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical calculations as in claim 8. Step 2A Prong 2: wherein evaluating the divergence metric based on the first divergence component and the second divergence component comprises: computing, using a weighting parameter, a weighted combination of the first divergence component and the second divergence component. These additional elements recite no more than generally linking a judicial exception, the abstract idea of mathematical concepts to the evaluation of the divergence metric based on the first divergence and second divergence component to a particular technological environment/ Field of Use, comprises: computing, using a weighting parameter, a weighted combination of the first divergence component and the second divergence component. As described in MPEP 2106.05(h). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 10 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical calculations as in claim 8. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “wherein adjusting the weighting parameter causes the divergence metric to interpolate between a mode-seeking behavior and a mean-seeking behavior.”: These additional elements are merely additional instruction to apply an exception, e.g. “wherein adjusting the weighting parameter causes the divergence metric to interpolate between a mode-seeking behavior and a mean-seeking behavior.” these additional elements do not provide meaningful integration of the judicial exception into a practical application. See MPEP 2106.05(f). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 11 Step 1: a process, as above. Step 2A Prong 1: This claim recites, inter alia: comprising: adjusting the weighting parameter based on a desired output diversity for a type of task. These limitations recite a mentally performable process with the aid of pen and paper by using judgment to adjust observed weight parameters based on evaluating desired output diversity for a type of task. Thus, this claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claim is patent ineligible. Claim 12 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mental process as in claim11. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “comprising: wherein the weight is a learned hyperparameter during training.”: These additional elements recite no more than generally linking a judicial exception, the abstract idea of mental process of the weight to a particular technological environment/ Field of Use, a learned hyperparameter during training. Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 13 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical relationship as in claim1. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “wherein the machine-learned teacher sequence processing model was not trained using reinforcement learning.”: These additional elements recite no more than generally linking a judicial exception, the abstract idea of mathematical relationship to the machine-learned teacher sequence processing model, to a particular technological environment/ Field of Use, where it was not trained using reinforcement learning. Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole does not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 14 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical relationship as in claim 1. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “wherein the machine-learned student sequence processing model was fine-tuned to achieve a baseline threshold of performance before training with the multiscale refinement objective...”: These additional elements are merely additional instruction to apply an exception, because the additional elements recite only the idea of a solution or outcome, e.g. “wherein the machine-learned student sequence processing model was fine-tuned to achieve a baseline threshold of performance before training with the multiscale refinement objective.” these additional elements do not provide meaningful integration of the judicial exception into a practical application. See MPEP 2106.05(f). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 15 Step 1: a process, as above. Step 2A Prong 1: This claim recites, inter alia: wherein: the machine-learned student sequence processing model is characterized by a first number of parameters; the machine-learned teacher sequence processing model is characterized by a second number of parameters; and the second number of parameters is larger than the first number of parameters. These limitations recite a mathematical formula or equations similar a mathematical expression written in word for the calculation of a variable. e.g. a first number of parameters, a second number of parameters, the second number of parameters is larger than the first number of parameters. described in MPEP 2106.04(a)(2)(I)(B). Thus, this claim recites a judicial exception. Step 2A Prong 2 & Step 2B: There are no additional elements recited so this claim does not provide a practical application and is not considered to be significantly more. As such, this claim is patent ineligible. Claim 16 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mathematical formula or equations as in claim 15. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “wherein the second number of parameters is at least 30 times the first number of parameters”: The additional element recites merely the instructions to apply the judicial exception because the only recites the idea of a solution or outcome, e.g. wherein the second number of parameters is at least 30 times the first number of parameters. It is understood that second number of parameters larger than the first number of parameters, therefore. this limitation does not provide meaningful integration to the judicial exception into a practical application. as described in MPEP 2106.05(f). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 do not contain significantly more than the judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 17 Step 1: a process, as above. Step 2A Prong 1: This claim recites, inter alia: “Determining the reinforcement learning signal based on the feedback data.” This limitation recites an abstract idea of mental process using judgement for determining the reinforcement learning signal based on the feedback data, similar to making an observation and evaluating the output data to decide how to fine-tune the application. See MPEP 2106.04(a)(2)(III). Thus, this claim recites a judicial exception. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “generating the respective output using the machine-learned student sequence processing model; The additional element recites merely the instructions to apply the judicial exception because the additional element only recites the idea of a solution or outcome in generality, e.g. generating the respective output using the machine-learned student sequence processing model; Furthermore, this limitation does not provide meaningful integration to the judicial exception into a practical application as described in MPEP 2106.05(f)(I). “receiving, from a client computing system, a request to perform an inference task based on input data; obtaining the respective input from the input data; … returning, to the client computing system and responsive to the request, output data based on the respective output; receiving, from the client computing system, feedback data; and…”: These additional elements are recited at a high level of generality and merely recites insignificant extra-solution activity of data gathering of obtaining a respective input and selecting a particular data source of a respective output from the student machine-learned sequence processing model to be manipulated and transmitted. See MPEP 2106.05(g). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include insignificant extra-solution activities of data gathering by and selecting a particular data source of “Comprising; receiving, from a client computing system, a request to perform an inference task based on input data; obtaining the respective input from the input data, returning, to the client computing system and responsive to the request, output data based on the respective output; receiving, from the client computing system, feedback data” this is well understood routine and conventional activity similar to presenting offers and gathering statistics as described in MPEP 2106.05(d)(II). Additional elements further include invoking computers or other machinery to apply the underlying judicial exception. Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 18 Step 1: a process, as above. Step 2A Prong 1: The claim recites the same abstract idea of mental process as in claim 17. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. “In an offline process and updating the machine-learned student sequence processing model based on the multiscale refinement objective.”: These additional elements are merely instruction to apply an exception, because the additional elements recite only the idea of a solution or outcome, e.g. “updating the machine-learned student sequence processing model based on the multiscale refinement objective.” these additional elements does not provide meaningful integration to the judicial exception into a practical application. See MPEP 2106.05(f). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. “comprising: in an online process, receiving the request and returning the output data; and in an offline process, obtaining the plurality of predictions of the teacher machine-learned sequence processing model …” These additional elements are recited at a high level of generality and merely recites insignificant extra-solution activity of data gathering and transmission of obtaining a respective input and selecting a particular data source of a respective output from the student machine-learned sequence processing model to be manipulated on and offline. See MPEP 2106.05(g). Thus, the way in which the additional elements use or interact with the judicial exception when analyzed with this claim as a whole do not integrate the judicial exception into a practical application. Step 2B: The additional elements from Step 2A Prong 2 include insignificant extra-solution activities of data gathering and output recited by “comprising: in an online process, receiving the request and returning the output data; and in an offline process, obtaining the plurality of predictions of the teacher machine-learned sequence processing model …” These insignificant extra-solution activities are well-understood routine and conventional activities similar to receiving or transmitting data over a network and presenting offers and gathering statistics see MPEP 2106.05(d)(II). Thus, the additional elements, viewed individually or in combination, do not provide an inventive concept or otherwise amount to significantly more than the abstract idea itself. See MPEP 2106.05. Claim 19 Step 1: These claims are directed to “A computing system…comprising:”; therefore, these claims are directed to the statutory category of machine. Step 2A Prong 1: The claim 19 recites a substantially the same abstract ideas as in claim 1, respectfully, as the judicial exception. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. The only substantive difference between claim 1 and claim 19 is directed to “A computing system, comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that are executable by the one or more processors to cause the computing system to perform one or more operations, the operations comprising…” however, mere recitation that a judicial exception is to be performed on a generic computer equipment in their ordinary capacity, i.e. “A computing system, comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that are executable by the one or more processors to cause the computing system to perform one or more operations”, cannot meaningfully integrate the judicial exception into a practical application, See MPEP 2106.05(f). With that, exception, the analysis at this step is substantially the same as that of claim 1. Step 2B: The additional elements from Step 2A Prong 2 of these claims do not contain significantly more than the judicial exception. The only substantive difference between claim 1 and claim 19 is directed to “A computing system, comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that are executable by the one or more processors to cause the computing system to perform one or more operations” mere recitation that a judicial exception is to be performed on a generic computer equipment in their ordinary capacity, i.e. “A computing system, comprising: one or more processors, one or more non-transitory computer-readable media storing instructions…” cannot amount to significantly more than the judicial exception. See MPEP 2106.05(f). With that, exception, the analysis at this step is substantially the same as that of claim 1, respectively. Claim 20 Step 1: These claims are directed to “One or more non-transitory computer-readable media storing…”; therefore, these claims are directed to the statutory category of article of manufacture. Step 2A Prong 1: The claim 20 recites a substantially the same abstract ideas as in claim 1, respectfully, as the judicial exception. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. The judicial exception recited in these claims are not integrated in a practical application. The only substantive difference between claim 1 and claim 20 is directed to “One or more non-transitory computer-readable media storing a machine-learned student sequence processing model that was distilled from a larger teacher machine-learned sequence processing model, wherein the machine-learned model was trained by:” However, mere recitation that a judicial exception is to be performed using generic computer equipment in their ordinary capacity, i.e. “One or more non-transitory computer-readable media storing a machine-learned student sequence processing model that was distilled from a larger teacher machine-learned sequence processing model, wherein the machine-learned model was trained by:” cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f). With that exception, the analysis at this step is substantially the same as that of claims 1, respectively. Step 2A Prong 2: The judicial exception recited in these claims are not integrated into a practical application. The only substantive difference between claim 1 and claim 20 is directed to “One or more non-transitory computer-readable media storing a machine-learned student sequence processing model that was distilled from a larger teacher machine-learned sequence processing model, wherein the machine-learned model was trained by:” However, mere recitation that a judicial exception is to be performed using generic computer equipment in their ordinary capacity, i.e. “One or more non-transitory computer-readable media storing a machine-learned student sequence processing model that was distilled from a larger teacher machine-learned sequence processing model, wherein the machine-learned model was trained by:” cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f). With that exception, the analysis at this step is substantially the same as that of claims 1, respectively. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 – 3, 5 – 13, 15, and 19 - 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lioutas et al (US 2023/ 0222353 A1) (hereinafter Lioutas) in view of Zhang at al. (CN 113222035 A1) (hereinafter Zhang). Regarding claim 1, Lioutas teaches a computer-implemented method for training a machine-learned student sequence processing model, the method comprising: (Lioutas, [0009], and Abstract, “a method of training a student neural network model”. Lioutas, [0067], “The computations of the teacher neural network model 104, 504, student neural network model 106, 506 and adversarial sample generator 102, 502 may be performed by any suitable processing device 1202 of the processing system 1200 or variant thereof. Further, teacher neural network model 104, 504, student neural network model 106, 506 and adversarial sample generator 102, 502 may be use suitable neural network model, including variations such as recurrent neural network models, long short-term memory (LSTM) neural network models.”), obtaining a respective input (Lioutas, [0057], “The original training image data samples x are provided to… the student neural network model 506); obtaining, from the student machine-learned sequence processing model, a respective output corresponding to the respective input; (Lioutas, [0057], “The original training image data samples x are provided to the teacher neural network model 504 and the student neural network model 506 to obtain their respective output predictions… the predicted output of the student neural network model 506 on the original image data samples x.”) generating a multiscale refinement objective configured to jointly distill knowledge from a teacher machine-learned sequence processing model (Lioutas, [0009], “a method of training a student neural network model using adversarial learning and knowledge distillation. The method includes training a generator to generate a respective adversarial data sample for each of a plurality of input data samples by replacing selected parts of the input data samples, the training of the generator being based on an objective of maximizing divergence between output predictions generated by the student neural network and a teacher neural network model for the adversarial data samples. The method further includes training the student neural network model based on objectives of: (i) minimizing divergence between output predictions generated by the student neural network model and the teacher neural network model for the adversarial data samples generated by the generator, and (ii) minimizing divergence between output predictions generated by the student neural network model and the teacher neural network model for the input data samples”), and reinforce preferred behavior of the student machine-learned sequence processing model, (Lioutas, [0036], “ the student neural network model 106 is trained using knowledge distilled from the teacher neural network model ” Wherein the preferred behavior of the student machine-learning processing model is the behavior distilled from then teacher machine-learned processing model). wherein the multiscale refinement objective comprises: a first component based on a divergence metric characterizing, for the respective input, a comparison of a plurality of predictions of the student machine-learned sequence processing model to a plurality of predictions of the teacher machine-learned sequence processing model; and ( Lioutas, [0017], “the divergence between the output predictions generated by the student neural network and a teacher neural network model for the input data samples corresponds to a Kullback-Leibler (KL) divergence.” Wherein the comparison of a plurality of predictions of the student and teacher machine-learned sequence processing model are conducted in the KL metric.) updating the machine-learned student sequence processing model based on the multiscale refinement objective. (Lioutas, [0047] “the student neural network model 106 are then updated using gradient decent with the objective”) Lioutas does not expressly teach a second component based on a reinforcement learning signal associated with the respective output; and However, Zhang teaches a second component based on a reinforcement learning signal associated with the respective output; and (Zhang, [0119], “Use the output of reinforcement learning combined with knowledge distillation to learn sample weights, and construct a loss function by combining the learned sample weights, the output of the teacher network and each student network.”) Because Lioutas and Zhang are analogous art and within the same field of endeavor, specifically computer-implemented methods for machine-learned models based on a knowledge distillation framework it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the techniques of reinforcement learning with knowledge distillation of Zhang, with the multiple objectives of Lioutas for data augmentation with a task to jointly distill to create an efficient method for knowledge distillation. These modifications would have been motivated by the desire to obtain better fault classification results for multi-class imbalanced classification problems based on knowledge distillation and reinforcement learning (Zhang, [003]). Regarding claim 2, the combination of Lioutas and Zhang teaches wherein the divergence metric is evaluated using: a student value generated by the machine-learned student sequence processing model for one or more portions of the respective output based on the respective input (Lioutas, [0043], “the original input text data samples x are provided to the student neural network model 106 to obtain their respective output predictions… In example embodiments logits” Wherein the student value (logit) is generated by the student neural network based on the corresponding input to output. This will then be converted to a probability for the evaluation of the divergence metric (loss function).) the value corresponding to a student probability of the one or more portions of the respective output conditioned on the respective input (Lioutas, [0042], “In example embodiments logits (i.e., the output of the final layer of the neural network before the softmax layer of the neural network) generated by the student neural network model 106 are used as the respective output predictions for determining the KD loss.” Under BRI, examiners interpretation of “the value” is being referred to as the student value; wherein the student value is generated by the student neural model based on the respective input data, and the student value is then SoftMax converted into a probability (corresponding to a student probability). This is then evaluated by KL Divergence for the expected loss. ); and a teacher value generated by the machine-learned teacher sequence processing model for the one or more portions of the respective output based on the respective input (Lioutas, [0043], “the original input text data samples x are provided to the teacher neural network model 106 to obtain their respective output predictions… In example embodiments logits” Wherein the teacher value (logit is generated by the teacher neural network based on the corresponding input to output. This will then be converted to a probability for the evaluation of the divergence metric (loss function).), the teacher value corresponding to a teacher probability of the one or more portions of the respective output conditioned on the respective input (Lioutas, [0042], “In example embodiments logits (i.e., the output of the final layer of the neural network before the softmax layer of the neural network) generated by the teacher neural network model 106 are used as the respective output predictions for determining the KD loss.” Wherein the teacher value(logit) is generated by the teacher neural model based on the respective input data, and the teacher value is then SoftMax converted into a probability (corresponding to a student probability). This is then evaluated by the KL Divergence for the expected loss.) Regarding claim 3, the combination of Lioutas and Zhang teaches, comprising: for each portion of a plurality of portions of the respective output: (Lioutas, [0005],“ zt and zs are the logits (i.e. the output of the neural network)”) determining a portion-specific divergence metric that characterizes a similarity ( Lioutas, [0005],“ An example of a KD loss function used for training a student neural network model is as follows: L.sub.KD =α*H(y,σ(zs;T=1))+(1−α)*H(σ(zt;T=τ),σ(zs;T=τ))  (1) “) between a student probability distribution over a set of candidate output portions and a teacher probability distribution over the set of candidate output portions ( Lioutas, [0005],“ zt and zs are the logits (i.e. the output of the neural network before the last softmax layer) of the teacher neural network model (T) and student neural network model (S).” Wherein z is the candidate output portions) wherein each of the student probability distribution and the teacher probability distribution are conditioned on the respective input ( Lioutas, [0017], “the output predictions generated by the student neural network and a teacher neural network model for the input data samples.”) and one or more portions of the respective output that precede the portion ( Lioutas, [0058], “computed based on a comparison of the true labels and output predictions from the student network model 506”) Furthermore, Lioutas, [0059], “the output predictions output by teacher neural network 504 may be used as the true labels”) aggregating the plurality of portion-specific divergence metrics for the respective output (Lioutas, [0060] “The adversarial loss L.sub.ADV, KD loss L.sub.KD, and CE loss L.sub.CE are combined”) to obtain the first component. (Lioutas, [0060] “to generate a total loss L.sub.total.”) Regarding claim 5, the combination of Lioutas and Zhang teaches the method of claim 1, wherein the multiscale refinement objective comprises one or more weighting parameters that weight the respective contributions of the first component and the second component (Lioutas, [0039], “A generator loss is computed using a generator loss function L.sub.G=−D.sub.KL(T(x′)∥S(x′)) that is designed to minimize a negative Kullback-Leibler (KL) objective, or in other words, to maximize the KL divergence between the teacher and student logits (outputs) on the adversarial text data samples x′. The KL divergence between two discrete distributions P and Q over a probability space” Furthermore, Lioutas, [0030], “As indicated at block 210, the parameters of the text generator 102 (including neural network model 110) are then updated using gradient decent with the objective of maximizing the KL divergence in future training iterations.” Wherein updating the parameters of the generator in turn contributes to the respective parameters of the first component (the plurality of predictions from the respective teacher and student neural network models) distributions and second component (the reinforcement signal associated with the respective out).) Regarding claim 6, the combination of Lioutas and Zhang teaches the method of claim 1, wherein the reinforcement learning signal comprises data indicating human feedback on an overall quality of the respective output. (Lioutas, [0033], “Aspects of the present disclosure can be applied to a variety a different NLP tasks, including for example NLP tasks that support chatbot applications, search engine autocomplete applications, voice assistant applications, language translator applications, sentiment analysis applications, grammar check applications, and email classification and filtering applications, among other things. In all such cases, the teacher neural network model 104 has been trained to map an input text data sample (x) to an output prediction (ŷ) that is appropriate for the application. For example, in the case of sentiment analysis, an input text data sample (x) could be mapped to an output prediction (ŷ) that is selected from the candidate prediction set of “Good”, “Bad”, and “Neutral”.” Wherein the plurality of NLP tasks requires human input and feedback from respective output.) Regarding claim 7, the combination of Lioutas and Zhang, teaches the method of claim 1, wherein the reinforcement learning signal comprises data indicating a score generated by a machine-learned reward model, wherein the score indicates an overall quality of the respective output. (Zhang, [0054], [005], [0056], “Calculate the reward rt. The designed reward is as follows: rt (st, at) = F1 (t); wherein F1 (t) represents the F1 score of the student network in the tth iteration. The state, action (sample weights for each iteration)”.). Regarding claim 8, the combination of Lioutas and Zhang teaches the method of claim 1, wherein evaluating the divergence metric comprises: determining a value of a mixture distribution corresponding to a mixture of a student probability distribution of the machine-learned student sequence processing model and a teacher probability distribution of the machine-learned teacher sequence processing model; Lioutas, [0039]: Line 10 - 16, “A generator loss is computed using a generator loss function L.sub.G=−D.sub.KL(T(x′)∥S(x′)) that is designed to minimize a negative Kullback-Leibler (KL) objective, or in other words, to maximize the KL divergence between the teacher and student logits (outputs) on the adversarial text data samples x′. The KL divergence between two discrete distributions P and Q over a probability space is defined PNG media_image1.png 66 306 media_image1.png Greyscale ”) computing a first divergence component that characterizes a divergence of the student probability distribution with respect to the mixture distribution; (Lioutas, [0042], “as indicated at block 302, the adversarial text data samples x′ are provided to the teacher neural network model 104 and the student neural network model 106 to obtain their respective output predictions, and an adversarial loss L.sub.ADV=D.sub.KL(T(x′)∥S(x′)) is computed.”) computing a second divergence component that characterizes a divergence of the teacher probability distribution with respect to the mixture distribution; and (Lioutas, [0043], “the original input text data samples x are provided to the teacher neural network model 104 and the student neural network model 106 to obtain their respective output predictions, and an KD loss L.sub.KD=D.sub.KL(T(x)|/S(x)) is computed that is based on the KL divergence between the output predictions of the teacher neural network model 104 and the output predictions of the student neural network model 106 on the input text data samples x. In example embodiments logits (i.e., the output of the final layer of the neural network before the softmax layer of the neural network) generated by the teacher neural network model 104 and the student neural network model 106 are used as the respective output predictions for determining the KD loss.”) evaluating the divergence metric based on a combination of the first divergence component and the second divergence component. (Lioutas, [0046], Image 4, “the adversarial loss L.sub.ADV, KD loss L.sub.KD, and CE loss L.sub.CE are combined to generate a total loss L.sub.total. In some examples, the total loss can be an average of the three losses as indicated below, however in other examples the relative weighting of the losses can be adjusted as hyper-parameters: PNG media_image2.png 66 300 media_image2.png Greyscale “) Regarding claim 9, the combination of Lioutas and Zhang teaches the method of claim 8, wherein evaluating the divergence metric based on the first divergence component and the second divergence component comprises: computing, using a weighting parameter, a weighted combination of the first divergence component and the second divergence component. (Lioutas, [0046], Image 4, “ the adversarial loss L.sub.ADV, KD loss L.sub.KD, and CE loss L.sub.CE are combined to generate a total loss L.sub.total. In some examples, the total loss can be an average of the three losses as indicated below, however in other examples the relative weighting of the losses can be adjusted as hyper-parameters: PNG media_image2.png 66 300 media_image2.png Greyscale ”) Regarding claim 10, the combination of Lioutas and Zhang teaches the method of claim 9, wherein adjusting the weighting parameter causes the divergence metric to interpolate between a mode-seeking behavior and a mean-seeking behavior. (Lioutas, [0039], “A generator loss is computed using a generator loss function L.sub.G=−D.sub.KL(T(x′)∥S(x′)) that is designed to minimize a negative Kullback-Leibler (KL) objective, or in other words, to maximize the KL divergence between the teacher and student logits (outputs) on the adversarial text data samples x′.”) Regarding claim 11, the combination of Lioutas and Zhang teaches the method of claim 10, comprising: adjusting the weighting parameter based on a desired output diversity for a type of task. (Lioutas, [0011], “updating parameters of the generator neural network model using gradient decent based on the computed loss.”) Regarding claim 12, the combination of Lioutas and Zhang teaches the method of claim 11, comprising: wherein the weight is a learned hyperparameter during training. (Lioutas, [0046], “In some examples, the total loss can be an average of the three losses as indicated below, however in other examples the relative weighting of the losses can be adjusted as hyper-parameters”) Regarding claim 13, the combination of Lioutas and Zhang teaches, the method of claim 1, wherein the machine-learned teacher sequence processing model was not trained using reinforcement learning. (Zhang, [009], “wherein the Gaussian Bernoulli limited Boltzmann machine parameter obtained by training all samples is the pre-training parameter of the teacher network”) Regarding claim 15, the combination of Lioutas and Zhang teaches the method of claim 1, wherein: the machine-learned student sequence processing model is characterized by a first number of parameters; (Lioutas, [FIG 3] As FIG. 3 details in step 310, updating parameters of the student neural network model using gradient decent) the machine-learned teacher sequence processing model is characterized by a second number of parameters; (Lioutas, [FIG 2] As FIG. 2 details, in step 210 updating the parameters of generator using gradient decent. Under BRI examiner’s interpretation the generator is interpreted to being part of the teacher neural network since its outputs are used to train student neural network) and the second number of parameters is larger than the first number of parameters. (Lioutas, [0034], “The student neural network model 106 is smaller than the teacher neural network model 104 (i.e., has fewer and/or compressed parameters, and/or fewer hidden layers, and/or requires fewer computations to generate a prediction).” Wherein the student neural network model is referred to being the first number of parameters and the teacher neural network model is the second number of parameters.) Regarding claim 19, it is a system claim that corresponds to the method of claim 1. Therefore, it is rejected for the same reason as claim 1 above. Regarding claim 20, it is a non-transitory computer-readable media claim that corresponding to method of claim 1. Therefore, it is rejected for the same reason as claim 1 above. Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Lioutas et al (US 2023/ 0222353 A1) (hereinafter Lioutas) in view of Zhang at al. (CN 113222035 A1) (hereinafter Zhang), as applied in claim 3 and further in view of Shamir at al. (US 2022032041 W) (hereinafter Shamir). Regarding claim 4, the combination of Lioutas and Zhang teaches the method of claim 3. The combination of Lioutas and Zhang does not teach wherein the teacher probability distributions for each of the portion-specific divergence metrics are generated at least partially in parallel by the machine-learned teacher sequence processing model. However, Shamir teaches wherein the teacher probability distributions for each of the portion-specific divergence metrics are generated at least partially in parallel by the machine-learned teacher sequence processing model. (Shamir, [0025], “In a first stage, a training system can apply a first distillation loss (e.g., square loss) in logit space to allow for fast convergence, but not necessarily to the correct minimum (e.g., converging to the logit mean, which for many skewed teacher distributions is farther from the origin than the probability mean).” Shamir, [0027], “The first and second stages can be performed sequentially or simultaneously (e.g., in parallel).”) Because the combination of Lioutas, Zhang, and Shamir are analogous art and within the same field of endeavor, specifically computer-implemented methods for machine-learned models based on a knowledge distillation framework it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Shamir, the fast convergence of gradient based methods into the combinations Lioutas and Zhang (reinforcement learning with knowledge distillation, with multiple objectives) for data augmentation, to enable for joint distillation to create an efficient method for knowledge distillation. These modifications would have been motivated by the desire to improve the efficiency of training models (using fewer training cycles or processing iterations), reduced consumption of computational resources such as processor usage, memory usage, and/or network bandwidth usage and improve the performance of the model and its implementing computing system (Shamir, [0028], [0029]). Claims 14, 17 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lioutas et al (US 2023/ 0222353 A1) (hereinafter Lioutas) in view of Zhang et al. (CN 113222035 A) (hereinafter Zhang), as applied to claims 13 and 1. and further in view of Buduguppa et al. (US 2023/0196030 A1) (hereinafter Buduguppa) Regarding claim 14, the combination of Lioutas and Zhang teaches the method of claim 13. The combination of Lioutas and Zhang does not teach wherein the machine-learned student sequence processing model was fine-tuned to achieve a baseline threshold of performance before training with the multiscale refinement objective. However, Buduguppa teaches the method of claim 13, wherein the machine-learned student sequence processing model was fine-tuned to achieve a baseline threshold of performance before training with the multiscale refinement objective. (Buduguppa, [0046], “In a conventional approach, the starting point of the student model is a smaller, less complex model that has been pretrained on an associated generic task.”) Because the combination of Lioutas and Zhang, with Buduguppa are analogous art and within the same field of endeavor, specifically computer-implemented methods for machine-learned models based on a knowledge distillation framework. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Buduguppa, fine-tuning a machine - learning model relating to knowledge distillation, into the combination of Liuotas and Zhang, (reinforcement learning with knowledge distillation, with multiple objectives) for data augmentation. These modifications would have been motivated by the desire to improve the system by establishing an efficient way to perform knowledge distillation in a contact center setting, which would further improve runtime (Buduguppa [0001], [0046]). Regarding claim 17, the combination of Lioutas and Zhang teaches the method of claim 1. The combination of Lioutas and Zhang teaches determining the reinforcement learning signal based on the feedback data. (Zhang, [0054], [0055], [0056], “) Calculate the reward rt. The designed reward is as follows: rt (st, at) = F1 (t), where F1 (t) represents C is the number of clusters, and K is the number of categories. the F1 score of the student network in the t-th iteration. The state, action (sample weights for each iteration) and the corresponding reward are stored in the experience replay.”) However, the combination of Lioutas and Zhang teaches does not teach comprising: receiving, from a client computing system, a request to perform an inference task based on input data; obtaining the respective input from the input data; generating the respective output using the machine-learned student sequence processing model;, returning, to the client computing system and responsive to the request, output data based on the respective output; receiving, from the client computing system, feedback data; and. Buduguppa teaches, comprising: receiving, from a client computing system, a request to perform an inference task based on input data; obtaining the respective input from the input data; (Budduguppa, [0037], “The behavior models may be used to predict behaviors of, for example, customers or agents, in a variety of situations, thereby allowing embodiments of the present invention to tailor interactions”. Furthermore, Budduguppa, [0037], “such behavior models also may be implemented on customer systems (or, as also used herein, on the “customer-side” of the interaction”.) Wherein the model performs inferences based on customer (user) input data or interaction data. Thereby enabling the predictor model to provide a customized experience similar to obtaining the respective input from the input data.) generating the respective output using the machine-learned student sequence processing model; (Buduguppa, [0063], “recording outputs generated by the candidate student model from the second training data set;” Furthermore, shown in Fig 6. 470) Returning, to the client computing system and responsive to the request, output data based on the respective output receiving, from the client computing system, feedback data; and (Budduguppa, [0031], “the chat server 240 in such a way that a customer communicates with automated chatbots, human agents, or both. In exemplary embodiments, the chat server 240 may perform as a chat orchestration server that dispatches chat conversations among the chatbots and available human agents. The chat server 240 may also be coupled to the knowledge management server 234 and the knowledge systems 238 for receiving suggestions and answers to queries posed by customers during a chat so that, for example, links to relevant articles can be provided.” Wherein the automated chatbot/ human agent returns and response to the request sent by the customer thereby outputting data based respective output and receiving feedback.) Because the combination of Lioutas and Zhang, with Buduguppa are analogous art and within the same field of endeavor, specifically computer-implemented methods for machine-learned models based on a knowledge distillation framework. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Buduguppa, training of Natural Language Processing machine - learning models relating to knowledge distillation, into the combination of Liuotas and Zhang, (reinforcement learning with knowledge distillation, with multiple objectives) for data augmentation. These modifications would have been motivated by the desire to improve the system by establishing an efficient way to perform knowledge distillation in a contact center setting, which would further improve runtime (Buduguppa [0001], [0046]). Regarding claim 18, the combination of Lioutas, Zhang and Buduguppa teaches the method of claim 17, comprising: in an online process, receiving the request and returning the output data; and (Zhang, [0015], [0016], “Obtain online samples; Classify the online samples into one of the C clusters obtained by hierarchical clustering.” Zhang, [0093], [0094], “During the initialization process, each sample is independently assigned to a cluster. Calculate the distance (also called similarity) between the centers of every two clusters; Find the two nearest clusters and group them into one cluster”) in an offline process, obtaining the plurality of predictions of the teacher machine-learned sequence processing model (Zhang, [0087], “Offline modeling”, Zhang, [0100], “Use a Gauss-Bernoulli Restricted Boltzmann Machine (GBM) to train the network based on all samples and samples from each cluster. The GBM parameters obtained from training on all samples are the pre-trained parameters of the teacher network”) updating the machine-learned student sequence processing model based on the multiscale refinement objective. (Lioutas, [0047], “the student neural network model 106 are then updated using gradient decent with the objective”) Claims 16 is rejected under 35 U.S.C. 103 as being unpatentable over Lioutas et al (US 2023/ 0222353 A1) (hereinafter Lioutas), in view of Zhang et al. (CN 113222035 A1) (hereinafter Zhang), as applied in claim 15 and further in view Zhang Tao. et al. (CN 115271064 A) (hereinafter Zhang T.) Regarding claim 16, the combination of Lioutas and Zhang teaches the method of claim 15. The combination of Lioutas and Zhang does not teach wherein the second number of parameters is at least 30 times the first number of parameters. However, Zhang T. teaches wherein the second number of parameters is at least 30 times the first number of parameters. (Zhang T., [0063], “The student model achieves a similar performance to the teacher model with only one-third the number of parameters as the teacher model” ) Because the combination of Lioutas and Zhang with Zhang T. are analogous art and within the same field of endeavor, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang T., method of text distillation into the combination of Lioutas and Zhang, reinforcement learning with a multi-objective knowledge distillation framework to constitute efficient knowledge distillation for the training of a machine – learning model. These modifications would have been motivated by the desire to improve the traditional knowledge distillation algorithm, so that the student model can improve its performance with the smallest possible number of parameters, making it as good as the teacher model in terms of performance (Zhang T., [0005]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Song, (CN 110852426 A), discloses an integrated acceleration method and device pre-training model based on knowledge distilling. Passban, (US 20220076136 A1), discloses an agnostic combinatorial knowledge distillation (CKD) method for transferring trained knowledge of neural models from a complex model (teacher) to a less complex model (student) is described. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAWN BLAIN whose telephone number is (571)270-1815. The examiner can normally be reached Mon - Fri: 8AM - 5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHAWN BLAIN/Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Oct 13, 2023
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month