Prosecution Insights
Last updated: October 04, 2026
Application No. 17/945,978

METHOD AND DEVICE FOR COMPRESSING GENERATIVE PRE-TRAINED LANGUAGE MODELS VIA QUANTIZATION

Final Rejection §101§103
Filed
Sep 15, 2022
Examiner
ALABI, OLUWATOSIN O
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
Huawei Technologies Co., Ltd.
OA Round
2 (Final)
61%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
82%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
138 granted / 226 resolved
+6.1% vs TC avg
Strong +21% interview lift
Without
With
+21.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
22 currently pending
Career history
254
Total Applications
across all art units

Statute-Specific Performance

§101
20.4%
-19.6% vs TC avg
§103
41.2%
+1.2% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
23.9%
-16.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 226 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings were received on 09/15/2022. These drawings are acceptable. Information Disclosure Statement The information disclosure statement (IDS) submitted on 08/07/2024 has been considered by the examiner. Response to Arguments Applicant's arguments filed 7/08/2026 have been fully considered. Regarding the rejection of claims under USC 35 101, 102 and 103 the remarks are directed to amended limitations that have not been previously examined, see the rejection below that addresses the amended claim limitation. Examiner notes the concern for he 101 analysis is that the claims as whole appear to be an improvement on the mathematical concept/relationship. MPEP notes in 2106.05(a) discloses that “[a]n important consideration in determining whether a claim improves technology is the extent to which the claim covers a particular solution to a problem or a particular way to achieve a desired outcome, as opposed to merely claiming the idea of a solution or outcome. McRO, 837 F.3d at 1314-15, 120 USPQ2d at 1102-03; DDR Holdings, 773 F.3d at 1259, 113 USPQ2d at 1107. In this respect, the improvement consideration overlaps with other considerations, specifically the particular machine consideration (see MPEP § 2106.05(b)), and the mere instructions to apply an exception consideration (see MPEP § 2106.05(f)). Thus, evaluation of those other considerations may assist examiners in making a determination of whether a claim satisfies the improvement consideration. It is important to note, the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements. See the discussion of Diamond v. Diehr, 450 U.S. 175, 187 and 191-92, 209 USPQ 1, 10 (1981)) in subsection II, below. In addition, the improvement can be provided by the additional element(s) in combination with the recited judicial exception. See MPEP § 2106.04(d) (discussing Finjan, Inc. v. Blue Coat Sys., Inc., 879 F.3d 1299, 1303-04, 125 USPQ2d 1282, 1285-87 (Fed. Cir. 2018)). Thus, it is important for examiners to analyze the claim as a whole when determining whether the claim provides an improvement to the functioning of computers or an improvement to other technology or technical field.” Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e. an abstract idea) without significantly more. Claim 1: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. a) determining a scaling factor based on a distribution of weightsdetermining, based on a gradient of the training loss, an updated scaling factor for the neural network model. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III) determining, based on a gradient of the training loss, an updated scaling factor for the neural network model, and quantizing the neural network model by quantizing the weights associated with the distribution using the updated scaling factor to obtain a quantized neural network model (Considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). … determining, based on the quantized weights during training of the neural network model, a training loss of the neural network model; and d) determining, based on a gradient of the training loss, an updated scaling factor for the neural network model. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation merely include instructions to implement an abstract idea on a computer, or merely use a computer as a tool to perform an abstract idea; Thus claim limitations amount to mere instructions to apply the judicial exception using a computer/computing environment as a tool, as discussed in MPEP § 2106.05(f).) … a distribution of weights associated with the neural network model; …, an updated scaling factor for the neural network model. Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to generally linking the use of a judicial exception to a particular technological environment or field of use. See 2106.05(h).) a quantized neural network model for processing downstream tasks (Deemed insufficient to transform the judicial exception to a patentable invention because the claimed element generically recites an effect of the judicial exception or claims every mode of accomplishing that effect, and thus amounts to a claim that is merely adding the words "apply it" to the judicial exception, as discussed in MPEP § 2106.05(f).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. First, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use and elements invoking computers or other machinery merely as a tool to perform the claimed process/judicial exception. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 2: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. further comprising: determining, based on the distribution of weights, an average weight magnitude; and determining a clipping factor as a product of the scaling factor and the average weight magnitude, wherein determining, based on the weights in the distribution and the scaling factor, quantized weights is based on the clipping factor. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 3: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. further comprising determining the average weight magnitude by an L1 norm function to the weights associated with the distribution. (Considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 4: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. wherein the weights associated with the distribution are divided into a plurality of value ranges based on the scaling factor, the method further comprising: computing a gradient contribution by each weight of the weights associated with the distribution based on the value range that the respective weight falls in; and computing the gradient of the training loss by aggregating the gradient contributions from the weights associated with the distribution. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 5: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. further comprising: setting an initial value of the scaling factor to one. Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). Alternatively: further comprising: setting an initial value of the scaling factor to one. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to generally linking the use of a judicial exception to a particular technological environment or field of use. See 2106.05(h).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 6: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. further comprising: determining, based on initial values of the weights in the neural network model, an initial value for the scaling factor. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use and elements invoking computers or other machinery merely as a tool to perform the claimed process/judicial exception. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 7: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. Abstract idea noted in claim 1. Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). further comprising determining an optimized scaling factor for a task by carrying out multiple iterations of a) through d). (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to insignificant extra-solution activity (e.g. Performing repetitive calculations), see MPEP § 2106.05(g).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. Secondly, the limitations directed to insufficient to transform the judicial exception to a patentable invention because the recitation is directed to insignificant solution activity for as noted above. The courts have deemed these types of activity as well-known routine and convectional, see evidences noted below: Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values); Bancorp Services v. Sun Life, 687 F.3d 1266, 1278, 103 USPQ2d 1425, 1433 (Fed. Cir. 2012) ("The computer required by some of Bancorp’s claims is employed only for its most basic function, the performance of repetitive calculations, and as such does not impose meaningful limits on the scope of those claims.") These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 8: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. wherein the updated scaling factor is associated with one weight matrix among a plurality of weight matrices in the neural network model, the method further comprising: determining an updated scaling factor for each of the other weight matrices in the plurality of weight matrices in the neural network model by carrying out a) through d) for the respective weight matrix. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 9 Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. applying the updated scaling factors (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). applying the updated scaling factors to the neural network model; …and updating the neural network model by updating the learnable weights. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation merely include instructions to implement an abstract idea on a computer, or merely use a computer as a tool to perform an abstract idea; Thus claim limitations amount to mere instructions to apply the judicial exception using a computer/computing environment as a tool, as discussed in MPEP § 2106.05(f).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 10: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. ….determining a first loss based on pair-wise comparison between first tokens in the set of first token representations and second tokens in the set of second token representations; determining, based on the first loss, a third loss ... (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). wherein the neural network model with the quantized weights associated with the updated scaling factors is included in a student network and the student network is trained with a teacher network, … a set of second token representations by the teacher network; (Claimed limitations are generally linking the use of a judicial exception to a particular technological environment or field of use, as discussed in MPEP § 2106.05(h)) … and the student network is trained with a teacher network, the method further comprising: … determining, based on the first loss, a third loss during training of the student network; updating, based on the third loss, the student network. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation merely include instructions to implement an abstract idea on a computer, or merely use a computer as a tool to perform an abstract idea; Thus claim limitations amount to mere instructions to apply the judicial exception using a computer/computing environment as a tool, as discussed in MPEP § 2106.05(f).) the method further comprising: obtaining, based on an input sequence, a set of first token representations by the student network and a set of second token representations by the teacher network; (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to insignificant extra-solution activity (e.g. Receiving or transmitting data over a network,), see MPEP § 2106.05(g).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. First, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. Second, the limitations directed to insufficient to transform the judicial exception to a patentable invention because the recitation is directed to insignificant solution activity for as noted above. The courts have deemed these types of activity as well-known routine and convectional, see evidences noted below: Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); but see DDR Holdings, LLC v. Hotels.com, L.P., 773 F.3d 1245, 1258, 113 USPQ2d 1097, 1106 (Fed. Cir. 2014) ("Unlike the claims in Ultramercial, the claims at issue here specify how interactions with the Internet are manipulated to yield a desired result‐‐a result that overrides the routine and conventional sequence of events ordinarily triggered by the click of a hyperlink." (emphasis added)); These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 11: Does claim fall within a statutory category? Yes Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. Abstract idea in claim 10. Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). wherein the first loss comprises a student- to-teacher loss and a teacher-to-student loss. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to generally linking the use of a judicial exception to a particular technological environment or field of use. See 2106.05(h).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. First, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 12: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. further comprising: … determining a second loss based on pair-wise comparison between each first logit in the set of first logits and respective second logit in the set of second logits; and determining the third loss based on the first loss and the second loss. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). obtaining, based on the input sequence, a set of first logits from the student network and a set of second logits from the teacher network;. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to insignificant extra-solution activity (e.g. Receiving or transmitting data over a network,), see MPEP § 2106.05(g).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. First, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. Second, the limitations directed to insufficient to transform the judicial exception to a patentable invention because the recitation is directed to insignificant solution activity for as noted above. The courts have deemed these types of activity as well-known routine and convectional, see evidences noted below: Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); but see DDR Holdings, LLC v. Hotels.com, L.P., 773 F.3d 1245, 1258, 113 USPQ2d 1097, 1106 (Fed. Cir. 2014) ("Unlike the claims in Ultramercial, the claims at issue here specify how interactions with the Internet are manipulated to yield a desired result‐‐a result that overrides the routine and conventional sequence of events ordinarily triggered by the click of a hyperlink." (emphasis added)); These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 13: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. wherein the determining of the third loss based on the first loss and the second loss further comprises determining the third loss by aggregating the first loss and the second loss with a tunable factor. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. First, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 14: Does claim fall within a statutory category? Yes. Step 2A Prong 1: Evaluate whether the claim recites a judicial exception. a) determining a scaling factor based on a distribution of weights associated with the neural network model; b) determining quantized weights based on the scaling factor and the weights associated with the distribution; c) determining, …, a training loss of the neural network model; and d) determining, based on a gradient of the training loss, an updated scaling factor for the neural network model. (Considered directed to a Mental Process: Making observations for formulating observations, evaluations and judgements as claimed; see MPEP § 2106.04(a)(2), subsection III; And considered directed to a Mathematical concepts – mathematical relationships, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I)) Step 2A Prong 2: Evaluate whether the claim as a whole integrates the recited judicial exception into a practical application of the exception The preamble is deemed insufficient to transform the judicial exception to a patentable invention because the preamble generally links the use of a judicial exception to a particular technological environment or field of use, see MPEP 2106.05(h). one or more processors; and a non-transitory computer-readable medium, having computer-executable instructions stored thereon, the computer-executable instructions, when executed by one or more processors, causing the one or more processors to facilitate: … determining, based on the quantized weights during training of the neural network model … (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation merely include instructions to implement an abstract idea on a computer, or merely use a computer as a tool to perform an abstract idea; Thus claim limitations amount to mere instructions to apply the judicial exception using a computer/computing environment as a tool, as discussed in MPEP § 2106.05(f).) …. weights associated with the neural network model. (Deemed insufficient to transform the judicial exception to a patentable invention because the recitation is directed to generally linking the use of a judicial exception to a particular technological environment or field of use. See 2106.05(h).) The additional elements do not appear to be sufficient to transform the judicial exception into a practical application at Step 2A as analyzed above. Step 2B: Evaluates whether the claim as a whole/in combination integrates the recited judicial exception into a practical application of the exception The claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception and fail to integrate the abstract into practical application. Specifically, the additional limitations are directed to elements that generally link the use of a judicial exception to a particular technological environment or field of use and elements invoking computers or other machinery merely as a tool to perform the claimed process/judicial exception. These types of claimed elements cannot transform the judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Regarding claim 15, the limitations are similar to claim 2, and rejected under the same rationale. Regarding claim 16, the limitations are similar to claim 4, and rejected under the same rationale. Regarding claim 17, the limitations are similar to claim 8, and rejected under the same rationale. Regarding claim 18, the limitations are similar to claim 9, and rejected under the same rationale. Regarding claim 19, the limitations are similar to claims 10 and 12, and rejected under the same rationale. Regarding claim 20, the limitations are similar to claims 1 and 14, and rejected under the same rationale. Therefore, claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed a judicial exception and does not recite, when claim elements are examined or as an ordered combination, that are directed to what have the courts have identified as "significantly more”, than the identified abstract idea, see MPEP 2106.05. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 7-9, 14-15, 17-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over of Da Costa et al. (US 20230186095, hereinafter ‘Costa’) in view of Lee et al. (NPL: Network Quantization with Element-wise Gradient Scaling, hereinafter ‘Lee’). Regarding independent claim 1, Costa teaches a computer-implemented method for quantizing a neural network model, performed by a processing system, comprising: (in [0092] In order to improve the accuracy of the quantization of weight and activations in the forward pass, a similar principle to the scaling factor described above may be applied to determine a representation for a set of values based on their statistics.…) a) determining a scaling factor based on a distribution of weights associated with the neural network model; b) determining quantized weights based on the scaling factor and the weights associated with the distribution; (in [0092] In order to improve the accuracy of the quantization of weight and activations in the forward pass, a similar principle to the scaling factor described above may be applied to determine a representation for a set of values based on their statistics. In the forward pass, statistics of weights and activations are collected in order to maintain a separate histogram of the weights and activations [a) determining a scaling factor based on a distribution of weights associated with the neural network model; b) determining quantized weights based on the scaling factor and the weights associated with the distribution;], measuring the fraction of the total number of samples of the histogram that are above a given threshold; and adjusting the exponent offset (or exponent bias) accordingly to maintain a predefined fraction of samples above the given threshold. In general, the histograms will comprise a plurality of bins [a) determining a scaling factor based on a distribution of weights associated with the neural network model; b) determining quantized weights based on the scaling factor and the weights associated with the distribution;]. c) determining, based on the quantized weights during training of the neural network model, a training loss of the neural network model; and d) determining, based on a gradient of the training loss, an updated scaling factor for the neural network model. ([0080] The following mathematical description is generally applicable to different configurations of neural network models, for example different optimisers, hyperparameter values, etc. In the present example, a single histogram H.sub.GW is used for collecting statistics of weight gradients (i.e. gradients of the loss function with respect to weights of the network) [determining, based on the quantized weights during training of the neural network model, a training loss of the neural network model; and d) determining, based on a gradient of the training loss, an updated scaling factor for the neural network model], and a second single histogram custom-character.sub.GX is used for collecting statistics of activation gradients (i.e. gradients of the loss function with respect to activations of the network)… [0081] Two alternative methods can be used in the present method to combine the bin count from both histograms. The first method combines the histogram bin values of the two histograms and determines whether the total proportion of bin values exceeding a cut-off bin C defining a threshold T, increasing the loss scaling factor if the following condition is satisfied [determining, based on a gradient of the training loss, an updated scaling factor for the neural network model]:… [0082] The second method compares the proportion of the bin count exceeding a respective cutoff C to a respective critical fraction f separately for each histogram and a joint decision to increase the loss scaling factor is only made if both tests pass (i.e. unanimous vote). Critical fractions f.sub.CX and f.sub.GW and cutoff bins C.sub.GX and C.sub.GW are assumed for activation and weight gradients, respectively. The loss scaling factor is increased if the following condition is met: where f is the critical threshold. Otherwise the scaling factor is reduced [determining, based on a gradient of the training loss, an updated scaling factor for the neural network model]) quantizing the neural network model by quantizing the weights associated with the distribution using the updated scaling factor to obtain a quantized neural network model for processing downstream task (in[0092] In order to improve the accuracy of the quantization of weight and activations in the forward pass, a similar principle to the scaling factor described above may be applied to determine a representation for a set of values based on their statistics…) Additionally, Lee teaches quantizing the neural network model by quantizing the weights associated with the distribution using the updated scaling factor to obtain a quantized neural network model for processing downstream tasks. (in Sec 1. Last Para: …n this paper, we present an element-wise gradient scaling (EWGS) [quantizing the neural network model by quantizing the weights associated with the distribution using the updated scaling factor to obtain a quantized neural network model for processing downstream tasks] that enables better training of a quantized network, compared with the STE, in terms of stability and ac curacy. Given a gradient of discrete values, EWGS adaptively scales up or down each element of the gradient considering its sign and discretization errors between latent and discrete values. The scaled gradient is then used to update the latent value (Fig. 1b). Since optimal scaling factors, which control the extent of EWGS, may vary across weight [quantizing the neural network model by quantizing the weights associated with the distribution using the updated scaling factor to obtain a quantized neural network model for processing downstream tasks] or activation quantizers in different layers, we propose an approach to adjusting the factors adaptively during training… Sec 3: … We design a uniform quantizer Q that converts a full precision input x to a quantized output Q(x), where we denote by x a scalar element of either a weight or an input activation tensor x in a layer. We learn a quantization interval [14, 21] using lower and upper bounds, denoted by l and u, respectively. Specifically, the quantizer first generates a full-precision latent value xn by normalizing and clipping the input value x as follows: … where clip(·,0,1) is a clipping function with lower and up per bounds of 0 and 1, respectively. Note that weight and/or activation quantizers in every quantized layer use separate parameters for the quantization intervals (i.e., l and u). For b-bit quantization, the latent value xn is converted to a discrete value xq using a round function with pre-/post scaling as follows: …Finally, the quantizer outputs a quantized weight QW(x) [quantizing the neural network model by quantizing the weights associated with the distribution using the updated scaling factor to obtain a quantized neural network model for processing downstream tasks] or activation QA(x) as follows: ) Lee and Costa are analogous art because both involve developing information processing techniques using machine learning systems and algorithms. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the prior art for developing information processing techniques of neural networks using element-wise gradient scaling as disclosed by Lee with the method for processing of neural network models using quantization techniques, as disclosed by Costa. One of ordinary skill in the arts would have been motivated to combine the disclosed methods disclosed by Lee and Costa as noted above; Doing so allows for dynamically scaling model parameters for training a quantized network to have improved performance in terms of stability and accuracy, (Lee, abstract). Regarding claim 2, the rejection of claim 1 is incorporated. and Lee further teaches the method according to claim 1, further comprising: determining, based on the distribution of weights, an average weight magnitude; (in Sec. 3.2 …Based on Eq. (7), we could take an average over the absolute values of gradient elements i.e., E[|gxq |] for setting G, but we empirically found that most gradients are concentrated near zero, such that the average value tends to be biased to small gradient elements..) and determining a clipping factor as a product of the scaling factor and the average weight magnitude, wherein determining, based on the weights in the distribution and the scaling factor, quantized weights is based on the clipping factor. (in Sec. 3.1 …full-precision latent value xn by normalizing and clipping the input value x as follows …where clip(·,0,1) is a clipping function with lower and up per bounds of 0 and 1 [and determining a clipping factor as a product of the scaling factor and the average weight magnitude, wherein determining, based on the weights in the distribution and the scaling factor, quantized weights is based on the clipping facto], respectively. Note that weight and/or activation quantizers in every quantized layer use separate parameters for the quantization intervals (i.e., l and u). For b-bit quantization, the latent value xn is converted to a discrete value xq using a round function with pre-/post scaling as follows:…:) Regarding claim 4, the rejection of claim 1 is incorporated and Costa in combination with Lee further teaches the method according to claim 1, wherein the weights associated with the distribution are divided into a plurality of value ranges based on the scaling factor, the method further comprising computing a gradient contribution by each weight of the weights associated with the distribution based on the value range that the respective weight falls I; and computing the gradient of the training loss by aggregating the gradient contributions from the weights associated with the distribution (in [0063] FIG. 4 shows how a neural network may be trained while applying a loss scaling factor L to the loss function 100 based on gradient statistics. As described for FIG. 2, the network processes training data in a forward pass through a series of layers 402, at the end of which a loss function 100 is defined. In this case the loss function is multiplied by a loss scaling factor 420 to obtain a scaled loss function 416. The loss scaling factor may be initialised as any value. Gradients are then computed for scaled loss with respect to the weights and activations. This means that the gradients 406 computed with respect to weights [computing a gradient contribution by each weight of the weights associated with the distribution based on the value range that the respective weight falls in] and activations at each layer are scaled up or down by the same loss scaling factor 420 [wherein the weights associated with the distribution are divided into a plurality of value ranges based on the scaling factor] The gradients are propagated back through the network in a backward pass... [0094] In the forward pass, histograms are collected for activations, gradients with respect to weights [computing a gradient contribution by each weight of the weights associated with the distribution based on the value range that the respective weight falls in] and gradients with respect to outputs, where the goal is to determine an appropriate format for representing these values. As described above for gradients, histograms have at least two bins with the histogram providing an aggregation of all values falling within the ranges [computing the gradient of the training loss by aggregating the gradient contributions from the weights associated with the distribution] indicated by each bin…) Regarding claim 5, the rejection of claim 1 is incorporated and Costa in combination with Lee further teaches the method according to claim 1, further comprising: setting an initial value of the scaling factor to one. (in [0073] FIG. 6 shows a flow chart of how a loss scaling factor L may be updated automatically based on gradient statistics computed periodically during training of a deep learning model... The scaling factor itself is also initialised. For example, the scaling factor may initially be set to 1 [setting an initial value of the scaling factor to one], such that the gradients are not scaled up or down for the first training iterations, and once gradient statistics are known, the scaling factor is adjusted, as will be described below. Regarding claim 6, the rejection of claim 1 is incorporated and Costa in combination with Lee further teaches the method according to claim 1, further comprising: determining, based on initial values of the weights in the neural network model, an initial value for the scaling factor in [0073] FIG. 6 shows a flow chart of how a loss scaling factor L may be updated automatically based on gradient statistics computed periodically during training of a deep learning model... The scaling factor itself is also initialised. For example, the scaling factor may initially be set to 1 [further comprising: determining, based on initial values of the weights in the neural network model, an initial value for the scaling factor], such that the gradients are not scaled up or down for the first training iterations, and once gradient statistics are known, the scaling factor is adjusted, as will be described below. And in [0025] A first aspect disclosed herein provides a computer-implemented method of training, based on a set of training data, a multi-layer neural network comprising a set of network weights, the method comprising: processing the training data in respective forward and backward passes through a sequence of layers of the network, the forward pass comprising computing a set of activations by applying an activation function in dependence on the network weights and training data, and the backward pass comprising: computing gradients of a pre-determined loss function with respect to the network weights and/or computing gradients of the pre-determined loss function with respect to the computed activations of the network, …the gradients with respect to activations computed in the backward pass, and the gradients with respect to weights computed in the backward pass; updating the network weights in dependence on the computed gradients with respect to the weights [further comprising: determining, based on initial values of the weights in the neural network model, an initial value for the scaling factor]; computing a proportion of the subset of values falling above a predefined threshold; and updating the adjustment parameter applied to the subset of machine learning parameters in dependence on the computed proportion. Regarding claim 16, the limitations are similar to those in claim 4, and are thus rejected under the same rationale. Regarding claim 7, the rejection of claim 1 is incorporated and Costa in combination with Lee further teaches the method according to claim 1, further comprising determining an optimized scaling factor for a task by carrying out multiple iterations of a) through d). (in [[0011] Learning is generally based on the iterative update of the parameters of each of the layers, typically through backpropagation. In practice, backpropagation based on gradient descent computes the gradient of the loss with respect to the output of the last layer, and then this gradient is backpropagated using the chain rule of calculus. With backpropagation, each layer receives the gradient of the loss with respect to its output, and uses this quantity to derive the gradient of the loss with respect to the parameters, the weights of that particular layer.. And in [0052] Depending on the form of gradient descent used, the weights may be updated based on one training example at a time, or more commonly based on an aggregated gradient computed for a subset of the training examples, which may be referred to as a minibatch. In this case, an accumulation operation is applied to get an aggregated (e.g. average) gradient to be applied in the respective weight update. Each layer updates their respective weights based on the respective gradients with respect to the weights at that layer, as shown by the multiple weight updates 408 in FIG. 2. The updated weights are then used in the forward pass for a next iteration of training..) Regarding claim 8, the rejection of claim 1 is incorporated and Costa in combination with Lee further teaches the method according to claim 1, wherein the updated scaling factor is associated with one weight matrix among a plurality of weight matrices in the neural network model, the method further comprising: determining an updated scaling factor for each of the other weight matrices in the plurality of weight matrices in the neural network model by carrying out a) through d) for the respective weight matrix. ([0114] In embodiments, a subset of network weights, activations and gradients which are inputs to compute operations in at least one of the forward and backward passes are stored in eight-bit floating-point format, the compute operations comprising at least one of a matrix operation and a convolution operation) Regarding claim 8, the rejection of claim 1 is incorporated and Costa in combination with Lee further teaches the method according to claim 1, applying the updated scaling factors to the neural network model; determining quantized weights associated with the updated scaling factors as learnable weights in the neural network model; and updating the neural network model by updating the learnable weights.. ([0011] Learning is generally based on the iterative update of the parameters of each of the layers, typically through backpropagation. In practice, backpropagation based on gradient descent computes the gradient of the loss with respect to the output of the last layer, and then this gradient is backpropagated using the chain rule of calculus. With backpropagation, each layer receives the gradient of the loss with respect to its output, and uses this quantity to derive the gradient of the loss with respect to the parameters, the weights of that particular layer. These quantities are then used to update the corresponding weights… [0061] One method of loss scaling identifies when a loss scaling factor should be increased based on when clipping events are observed. This may be referred to as ‘Backoff scaling’, as described, for example in Nvidia OpenSeq2Seq documentation, in a section titled ‘Mixed Precision Training’ (https://nvidia.github.io/OpenSeq2Seq/html/mixed-precision.html). Under this method, the loss scaling factor may be increased until a gradient becomes large enough to be clipped, at which point the loss scaling factor is ‘backed off’ to a lower level, from which it progressively increases until the next clipping event. This is based on the premise that clip events need to be avoided..) Regarding claims 14 and 20, the limitations are similar with claim 1 limitations and are rejected under the same rationale. Additionally Char teaches one or more processors; and a non-transitory computer-readable medium, having computer-executable instructions stored thereon, the computer-executable instructions, when executed by one or more processors, causing the one or more processors to facilitate:… And the computer-executable instructions, when executed by one or more processors, causing the one or more processors to facilitate…, in [0020] According to a further example aspect, a processing unit is disclosed. The processing unit includes one or more processing devices and one or more storages operatively connected to the one or more processing devices and storing executable instructions that when executed by the one or more processing devices configure the processing unit to perform on or more of the methods of the preceding aspects. [0021] According to a further example aspect, a computer readable medium is disclosed that stores executable instructions that when executed by one or more processing devices configures the processing device(s) to perform on or more of the methods of the preceding aspects. Regarding claim 15, the limitations are similar to those in claim 2, and are thus rejected under the same rationale. Regarding claims 17 and 18, the limitations are similar to those in claims 8 and 9 respectively, and are thus rejected under the same rationale. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Costa in view of Lee in further view of Moshovos et al (US 20230334285, hereinafter ‘Mos’). Regarding claim 3, the rejection of claim 2 is incorporated. Costa and Lee teach the data processing method and do not expressly disclose claim 3 limitation. Mos teaches the method according to claim 2, further comprising determining the average weight magnitude by an L1 norm function to the weights associated with the distribution. (in [0074] ERT is a dictionary based fine-tuning approach that can use a second order Hessian information to quantize the model's weights to a few (4 to 16) representative values. It can store weights as indexes to those values. Q-BERT can separate the weights of each layer into multiple groups and can quantize each group separately using per group dictionaries each of 4, 8 or 16 entries... [0087] In some embodiments, the representative value can be stored in the memory as a reference (e.g., index or bin number) to the representative value. In some embodiments, GOBO is configured to generate each representative value by quantizing a value (e.g., parameter such as a weight or activation) of the neural network… [0095] Deep Compression uses dictionary compression for CNNs utilizing K-Means with linear initialization for cluster centroids and requiring fine-tuning to regain any accuracy loss…. In some embodiments, importantly, GOBO is configured to minimize L1-Norm within each cluster [further comprising determining the average weight magnitude by an L1 norm function to the weights associated with the distribution] rather than L2. Algorithm 1 summarizes how GOBO compacts each layer according to some embodiments.) Mos, Lee and Costa are analogous art because both involve developing information processing techniques using machine learning systems and algorithms. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the prior art for developing information processing techniques for or quantization of parameters of neural networks, as disclosed by Mos with the method for processing of neural network models using quantization techniques, as collectively disclosed by Lee and Costa. One of ordinary skill in the arts would have been motivated to combine the disclosed methods disclosed by Mos, Lee and Costa as noted above; Doing so allow for improving the energy efficiency of data centers and operating costs and environmental impact by allowing them to perform more computations per unit of time (Mos, 0051-0052). Claims 10 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Costa in view of Lee in further view of Liu et al. (NPL: “BiT: Robustly Binarized Multi-distilled Transformer”, hereinafter ‘Liu’) in further view of Yang et al. (US 11487944, hereinafter ‘Yang’). Regarding claim 10, the rejection of claim 9 is incorporated and Costa further teaches the method according to claim 9, wherein the neural network model with the quantized weights associated with the updated scaling factors … (in [0092] In order to improve the accuracy of the quantization of weight and activations in the forward pass, a similar principle to the scaling factor described above may be applied to determine a representation for a set of values based on their statistics...) Costa does not expressly teach the use of student-teacher networks as claimed in the limitations …the updated scaling factors is included in a student network, and the student network is trained with a teacher network, the method further comprising: obtaining, based on an input sequence, a set of first token representations by the student network and a set of second token representations by the teacher network; determining a first loss based on pair-wise comparison between first tokens in the set of first token representations and second tokens in the set of second token representations; determining, based on the first loss, a third loss during training of the student network; updating, based on the third loss, the student network. Liu does expressly teach the student-teacher network as claimed in the limitations …the updated scaling factors is included in a student network, and the student network is trained with a teacher network, in Sec. 4: …The multi-step distillation follows a quantization schedule, Q = {(b 1 w , b 1 a ),(b 2 w , b 2 a ), . . . ,(b k w , b k a )} with (b 1 w , b 1 a ) > (b 2 w , b 2 a ) > . . . > (b k w , b k a ) 3 . (b k w , b k a ) is the target quantization level, which is in our case binary for both weights and activations. In practice, we find that down to a quantization level of W1A2, we can distill models of reasonable accuracy in single shot, following the best practices outlined in Section 3.2 (See our 1-1-2 baseline results in Table 1). As a result, we follow a fixed quantization schedule, W32A32 → W1A2 → W1A1. This is not necessarily optimal, and how to efficiently find the best quantization schedule is an interesting open problem. We present our initial explorations towards this direction in Section 5.5. Combining the elastic binary activations with multi-distillation we obtain BiT, the robustly binarized multi-distilled transformer…; And in Sec. 3.3: …Learning the scaling [the updated scaling factors is included in a student network,] and threshold parameters, and how to approximate the gradients precisely in the process becomes crucial for the final accuracy. To handle this, we propose the elastic binarization function to rescale and shift the real-valued activations [the updated scaling factors is included in a student network,], where α ∈ R+, β ∈ R: Xi B = αXˆ i B = αbClip(Xi R − β α , 0, 1)e (9) In the function, we initialize α with α ∗ in Sec. 3.1 and β to be 0, and train it with gradients from the final loss. To back-propagate the gradients to α through the discretized binarization function, we follow the practice in Choi et al. (2018); Zhou et al. (2016) to use straight-through estimator (STE) (Bengio et al., 2013) to bypass the incoming gradients to the round function to be the outgoing gradients:… PNG media_image1.png 356 1094 media_image1.png Greyscale And Sec. 4: …This suggests a multi-step approach, where instead of directly distilling from a full-precision teacher to the desired quantization level, we first distill into a model with sufficient precision in order to preserve quality. This model can then be used as a teacher to distill into a further quantized student [and the student network is trained with a teacher network]. This process can be repeated multiple times, while at each step ensuring that the teacher and student models are sufficiently similar, and the performance loss is limited. This multi-distillation approach is sketched in Algorithm 1. Liu, Lee and Costa are analogous art because both involve developing information processing techniques using machine learning systems and algorithms. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the prior art for developing information processing techniques for binarized Multi-distilled Transformer as disclosed by Liu with the method for processing of neural network models using quantization techniques, as collectively disclosed by Lee and Costa. One of ordinary skill in the arts would have been motivated to combine the disclosed methods disclosed by Liu, Lee and Costa as noted above; Doing so allow for developing and implementing fully binarized transformer models that are at a practical level of accuracy, approaching a full-precision BERT baseline on the GLUE language understanding benchmark within as little as 5.9% (Liu, Abstract). While Liu teaches the use of knowledge distillation algorithms for processing a quantization techniques of model information. Liu does not expressly teach the claimed software architecture as claimed in the limitations the method further comprising: obtaining, based on an input sequence, a set of first token representations by the student network and a set of second token representations by the teacher network; determining a first loss based on pair-wise comparison between first tokens in the set of first token representations and second tokens in the set of second token representations; determining, based on the first loss, a third loss during training of the student network; updating, based on the third loss, the student network. Yang teaches the claimed software architecture as claimed in the limitations the method further comprising: obtaining, based on an input sequence, a set of first token representations by the student network and a set of second token representations by the teacher network; determining a first loss based on pair-wise comparison between first tokens in the set of first token representations and second tokens in the set of second token representations; determining, based on the first loss, a third loss during training of the student network; updating, based on the third loss, the student network. (As depicted in Fig.5 PNG media_image2.png 600 586 media_image2.png Greyscale And 7:59-8:2: In one embodiment, the contrastive representation distillation loss is computed as follows: (37) Let the vector representations of an input data sequence x produced by the k-th teacher be f.sup.T.sup.k(x) and by student be f.sup.S(x) [the method further comprising: obtaining, based on an input sequence, a set of first token representations by the student network and a set of second token representations by the teacher network]. A data sequence from the set of input data sequences is treated as a positive example x, and M other randomly sampled data sequences {x′.sub.m}.sub.m=1.sup.M are treated as negative examples. Let the vector representations of the data sequence x′.sub.m be f.sup.S(x′.sub.m). A contrastive loss [determining a first loss based on pair-wise comparison between first tokens in the set of first token representations and second tokens in the set of second token representations] is then utilized to distinguish between the positive and negative examples:… And in 8:41-56: FIG. 5 illustrates an example system for performing the methods described herein. The methods described herein may be implemented in other systems and are not limited to system 500. The system 500 includes a Tag Predictions Module 520, a Loss Calculation Module 535, and a Student Model Optimizer 560. The Tag Prediction Modules 520 applies the teacher models 525 and the student model 530 to input data sequences 510 to obtain the tag predictions. The Loss Calculation Module 535 [determining, based on the first loss, a third loss during training of the student network; updating, based on the third loss, the student network] includes a Distillation Loss Submodule 540 which calculates the distillation losses. In certain embodiments, the Loss Calculation Module 535 also includes a Student Loss Submodule 545 and CRD Loss Submodule 550 for calculating a student loss and a CRD loss, respectively, as described above. The Student Model Optimizer 560 adjusts the parameters of the student model with each iteration to reduce the overall loss [a third loss during training of the student network; updating, based on the third loss, the student network]… And in 4:36-45: …The system obtains a set of input data sequences for use in transferring knowledge from the teacher models to the student model (step 220). An example of input data sequences are text strings. Each input data sequence includes one or more tokens [the method further comprising: obtaining, based on an input sequence, a set of first token representations by the student network and a set of second token representations by the teacher network]. For text strings, the individual words in the string each may be treated as a token. Knowledge can be distilled from various teacher models using only the one set of input data sequences;...) Yang, Liu, Lee and Costa are analogous art because both involve developing information processing techniques using machine learning systems and algorithms. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the prior art for developing distillation approach to implementing information processing techniques in machine learning software that perform natural language processing as disclosed by Yang with the method for processing of neural network models using quantization techniques, as collectively disclosed by Liu, Lee and Costa. One of ordinary skill in the arts would have been motivated to combine the disclosed methods disclosed by Yang, Liu, Lee and Costa as noted above; Doing so allow for including a contrastive representation distillation (CRD) loss in the overall loss function enables the student to distill domain-invariant knowledge from the teacher models and enables the student model to produce vector representations of input data sequences that are domain insensitive or less domain sensitive than they would otherwise be. (Yang, 3:8-21). Regarding claim 11, the rejection of claim 10 is incorporated and Yang further teaches the method according to claim 10, wherein the first loss comprises a student- to-teacher loss and a teacher-to-student loss. (As depicted in Fig. 5 and in 8:41-56: FIG. 5 illustrates an example system for performing the methods described herein. The methods described herein may be implemented in other systems and are not limited to system 500. The system 500 includes a Tag Predictions Module 520, a Loss Calculation Module 535, and a Student Model Optimizer 560. The Tag Prediction Modules 520 applies the teacher models 525 and the student model 530 to input data sequences 510 to obtain the tag predictions. The Loss Calculation Module 535 [wherein the first loss comprises a student- to-teacher loss and a teacher-to-student loss] includes a Distillation Loss Submodule 540 which calculates the distillation losses [wherein the first loss comprises a student- to-teacher loss and a teacher-to-student loss]. In certain embodiments, the Loss Calculation Module 535 also includes a Student Loss Submodule 545 and CRD Loss Submodule 550 for calculating a student loss and a CRD loss, respectively, as described above. The Student Model Optimizer 560 adjusts the parameters of the student model with each iteration to reduce the overall loss …; And in 3:8-21: …In certain embodiments, the overall loss is a function of the aggregate distillation loss, the student loss, and a contrastive representation distillation (CRD) loss. The CRD loss [wherein the first loss comprises a student- to-teacher loss and a teacher-to-student loss] is based on a comparison of the vector representations generated by the teacher models for each of the input data sequences, the vector representations generated by the student model for each of the input data sequences, and the vector representations generated by the student model for negative example data sequences. Including the CRD loss in the overall loss function enables the student to distill domain-invariant knowledge from the teacher models and enables the student model to produce vector representations of input data sequences that are domain insensitive or less domain sensitive than they would otherwise be...) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Yang, Liu, Lee and Coasta for the same reasons disclosed above. Claims 12-13 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Costa in view of Lee in further view of Liu et al. (NPL:”BiT: Robustly Binarized Multi-distilled Transformer”, hereinafter ‘Liu’) in further view of Yang et al. (US 11487944, hereinafter ‘Yang’) in further view of Haidar et al. (US 20220335303, hereinafter ‘Hai’). Regarding claim 12, the rejection of claim 10 is incorporated and Yang further teaches the method according to claim 10, further comprising: obtaining, based on the input sequence, a set of first (in As depicted in Fig.5; And 7:59-8:2: In one embodiment, the contrastive representation distillation loss is computed as follows: (37) Let the vector representations of an input data sequence x produced by the k-th teacher be f.sup.T.sup.k(x) and by student be f.sup.S(x). A data sequence from the set of input data sequences is treated as a positive example x, and M other randomly sampled data sequences {x′.sub.m}.sub.m=1.sup.M are treated as negative examples. Let the vector representations of the data sequence x′.sub.m be f.sup.S(x′.sub.m). A contrastive loss [determining a second loss based on pair-wise comparison between each first ] is then utilized to distinguish between the positive and negative examples:… And in 8:41-56: FIG. 5 illustrates an example system for performing the methods described herein. The methods described herein may be implemented in other systems and are not limited to system 500. The system 500 includes a Tag Predictions Module 520, a Loss Calculation Module 535, and a Student Model Optimizer 560. The Tag Prediction Modules 520 applies the teacher models 525 and the student model 530 to input data sequences 510 to obtain the tag predictions [further comprising: obtaining, based on the input sequence, a set of first ]. The Loss Calculation Module 535 includes a Distillation Loss Submodule 540 which calculates the distillation losses [determining a second loss based on pair-wise comparison between each first ]. In certain embodiments, the Loss Calculation Module 535 [determining the third loss based on the first loss and the second loss] also includes a Student Loss Submodule 545 and CRD Loss Submodule 550 for calculating a student loss and a CRD loss, respectively, as described above. The Student Model Optimizer 560 adjusts the parameters of the student model with each iteration to reduce the overall loss [determining the third loss based on the first loss and the second loss]…) Yang does not expressly teach the data process associated with the distillation models as claimed a set of first logits from the student network and a set of second logits from the teacher network; …each first logit in the set of first logits and respective second logit in the set of second logits; Hai does expressly teach the data process associated with the distillation models as claimed a set of first logits from the student network and a set of second logits from the teacher network; …each first logit in the set of first logits and respective second logit in the set of second logits;, in [0009] As described above, the [t]eacher and the student each typically generate (i.e. predict) an output in the form of logits [set of first logits from the student network and a set of second logits from the teacher network; …each first logit in the set of first logits and respective second logit in the set of second logits], i.e. a non-normalized probability distribution. The predicted logits of the teacher model and the student model are then typically normalized by a softmax function to generate a normalized probability distribution, which may be used as the final prediction of the model (e.g., a model trained to perform a classification task using images as input may generate a normalized probability distribution of (“dog”=0.9, “cat”=0.1))… Hai, Yang, Liu, Lee and Costa are analogous art because both involve developing information processing techniques using machine learning systems and algorithms. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the prior art for developing distillation approach to implementing information processing techniques in machine learning as disclosed by Hai with the method for processing of neural network models using quantization techniques, as collectively disclosed by Yang, Liu, Lee and Costa. One of ordinary skill in the arts would have been motivated to combine the disclosed methods disclosed by Hai, Yang, Liu, Lee and Costa as noted above; Doing so allows for improving knowledge distillation using intermediate representations, (Hai, 0001). Regarding claim 13, the rejection of claim 12 is incorporated and Yang further teaches the method according to claim 12, wherein the determining of the third loss based on the first loss and the second loss further comprises determining the third loss by aggregating the first loss and the second loss with a tunable factor. (in 5:42-60: … The system aggregates the distillation losses of each of the student-teacher model pairs to compute an aggregate distillation loss (step 250)[ wherein the determining of the third loss based on the first loss and the second loss further comprises determining the third loss by aggregating the first loss and the second loss with a tunable factor]. The system computes an overall loss as function of the aggregate distillation loss (step 260). In certain embodiments, the overall loss may be equal to the aggregate distillation loss. In other embodiments, it may also include other losses [wherein the determining of the third loss based on the first loss and the second loss further comprises determining the third loss by aggregating the first loss and the second loss with a tunable factor], such as a student loss or a contrastive representation distillation (CRD) loss, as described below with respect to FIGS. 3A-3B and 4A-4B. The system repeats steps 230-260 for a number of iterations, adjusting the parameters of the student model with each iteration to reduce the overall loss (step 270). The steps may be repeated for a fixed number of iterations or until convergence is achieved. The result is a unified named-entity recognition model (i.e., the student) with the collective predictive capabilities of the teacher models without the need for the annotated training data used to train the teacher models…; And in 7:61-8:30: Let the vector representations of an input data sequence x produced by the k-th teacher be f.sup.T.sup.k(x) and by student be f.sup.S(x). A data sequence from the set of input data sequences is treated as a positive example x, and M other randomly sampled data sequences {x′.sub.m}.sub.m=1.sup.M are treated as negative examples. Let the vector representations of the data sequence x′.sub.m be f.sup.S(x′.sub.m). A contrastive loss is then utilized to distinguish between the positive and negative examples: … Where h(v, v′)=sigmoid(v.sup.Tv′/τ) and τ is a temperature [the second loss with a tunable factor] that adjusts the concentration level. To learn domain-invariant representations on data drawn from D.sub.k, the system maximizes the mutual information [the second loss with a tunable factor] between the student representation and each of the teacher representations by calculating the final CRD loss that as follows:… In contrast to Equation 3 above, which distills knowledge from the k-th teacher with only in-domain data, the CRD loss encourages the model to distill domain invariant knowledge of a teacher using both in-domain and out-domain data. The system calculates the overall loss as a function of the distillation loss, the student loss, and the CRD loss as set forth below: …) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Yang, Liu, Lee and Costa for the same reasons disclosed above. Regarding claim 19, the limitations are similar to those in claims 10 and 12, and are thus rejected under the same rationale. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Moshovos et al. (US 20220092382): teaches a model quantization technique that compresses the vast majority (e.g., 99.9%) of the 32-bit floating-point parameters of state-of-the-art BERT (Bidirectional Encoder Representations from Transformers) models. Shi et al. (US 20170270408) teaches processing the gradient of the cost is considered computing a gradient contribution by each weight of the weights associated with the distribution based on the value range that the respective weight falls in, in [0058] The hardware costs from hardware complexity cost generator 52 are input to hardware complexity cost gradient generator 62. These hardware costs include the bit-depth costs. The hardware cost gradient for each weight [computing a gradient contribution by each weight of the weights associated with the distribution based on the value range that the respective weight falls in] is calculated by hardware complexity cost gradient generator 62, and these gradients collected by weights selector 70. Some gradients, such as a gradient of error, can be back propagated [computing the gradient of the training loss by aggregating the gradient contributions from the weights associated with the distribution] to find the gradient over each parameter using a chain rule. Regularization costs from other regularization generator 56 are input to other regularization gradient generator 66, which generates gradients for regularization costs. Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUWATOSIN ALABI whose telephone number is (571)272-0516. The examiner can normally be reached Monday-Friday, 8:00am-5:00pm EST.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /OLUWATOSIN ALABI/ Primary Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Sep 15, 2022
Application Filed
Oct 16, 2025
Non-Final Rejection mailed — §101, §103
Jan 15, 2026
Response Filed
Jan 15, 2026
Response after Non-Final Action
Jun 04, 2026
Applicant Interview (Telephonic)
Jun 05, 2026
Examiner Interview Summary
Jul 08, 2026
Response Filed
Sep 22, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743621
MODEL UNDERSTANDABILITY
3y 5m to grant Granted Sep 22, 2026
Patent 12743622
NEURONAL ACTIVITY MODULATION OF ARTIFICIAL NEURAL NETWORKS
3y 2m to grant Granted Sep 22, 2026
Patent 12737677
DETERMINATION DEVICE, DETERMINATION METHOD, AND DETERMINATION PROGRAM
3y 6m to grant Granted Sep 15, 2026
Patent 12737593
Systems and Methods Providing a ConjointNet Architecture for Enhanced Conjoint Analysis for Preference Prediction with Representation Learning
3y 6m to grant Granted Sep 15, 2026
Patent 12718154
DATA PROCESSING DEVICE, DATA PROCESSING SYSTEM, AND DATA PROCESSING METHOD
3y 6m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
61%
Grant Probability
82%
With Interview (+21.3%)
3y 11m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 226 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month