Prosecution Insights
Last updated: October 02, 2026
Application No. 18/664,531

QUANTIZATION-AWARE TRAINING FOR MACHINE LEARNING MODEL ADAPTERS

Non-Final OA §101§103
Filed
May 15, 2024
Examiner
ZENG, WENWEI
Art Unit
Tech Center
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
25 currently pending
Career history
18
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on August 2, 2024, and June 17, 2025, were considered by the examiner. The submissions are compliant with the provisions of 37 CFR 1.97. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mental process or math concept) without significantly more. Claim 1: Regarding claim 1, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “1. A processing system comprising: one or more memories comprising processor-executable instructions; and one or more processors configured to execute the processor-executable instructions and cause the processing system to: access a first plurality of weights for a base model; access a second plurality of weights for an adapter model associated with the base model; generate a quantized plurality of weights based on the first plurality of weights, a first quantization scale for the first plurality of weights, and the second plurality of weights; generate a loss based on processing training data using the quantized plurality of weights; generate an updated second plurality of weights based on updating the second plurality of weights based on the loss; and deploy a machine learning model comprising quantized versions of the first plurality of weights and the updated second plurality of weights,” and a system or machine is one of the four statutory categories of invention. In step 2A prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a math concept but for recitation of generic computer components: generate a quantized plurality of weights based on the first plurality of weights, a first quantization scale for the first plurality of weights, and the second plurality of weights; (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0031] from the specification stating “ PNG media_image1.png 1 461 media_image1.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), generate a loss based on processing training data using the quantized plurality of weights; (This recites a mathematical relationship, formula or equation, or mathematical calculation, see specification paragraph [0067] note “At block 420, the quantization training system trains the weights of the adapter based on processing training data using the quantized weights. For example, as discussed above, the quantization training system may process a sample of training data (e.g., the training data 120 of FIG. 1 ) using the quantized model to generate an output, and the output may be compared against the label of the sample to generate a loss. In some aspects, this process is referred to as the forward pass. The quantization training system may then use the loss during a backward pass to compute gradients for the adapter weights,” this recites math since a loss value is a numeric quantity, which then is used compute gradients for weights, see MPEP 2106.04(a)(2), subsection I), generate an updated second plurality of weights based on updating the second plurality of weights based on the loss; (This recites a mathematical concept, see specification paragraphs [0031, 0067], similar to above limitations, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. In step 2A prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has determined that the following additional elements do not integrate this judicial exception into a practical application: A processing system comprising: one or more memories comprising processor-executable instructions; and one or more processors configured to execute the processor-executable instructions and cause the processing system … (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), … to: access a first plurality of weights for a base model; (In step 2A, prong 2, accessing recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), access a second plurality of weights for an adapter model associated with the base model; (In step 2A, prong 2, accessing recites mere data gathering, which is considered insignificant extra-solution activity – see MPEP 2106.05(g)), .. deploy a machine learning model comprising quantized versions of the first plurality of weights and the updated second plurality of weights, (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea. In step 2B of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, additional elements iv and vii recite mere instructions to apply the judicial exception using generic computer components, which are not indicative of significantly more. The additional elements v and vi recite mere data gathering, and are considered insignificant extra-solution activities. In step 2B, these insignificant extra-solution activities are well understood routine and conventional activities, which include receiving or transmitting data over a network from court case Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016), – see MPEP 2106.05(d) (II)(i)), Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim 2: Regarding claim 2, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 2 recites the following additional element: The processing system of claim 1, wherein the first plurality of weights and the first quantization scale are static when the updated second plurality of weights is generated, (In step 2A, prong 2, this is considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), (In step 2B, this is also considered mere instructions to apply an exception using generic computer – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 3: Regarding claim 3, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 3 recites the following abstract ideas: to: generate a scaled plurality of weights based on applying the first quantization scale to the first plurality of weights; (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0031] from the specification stating “ PNG media_image1.png 1 461 media_image1.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), generate an aggregated plurality of weights based on the scaled plurality of weights and the second plurality of weights; (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0031] from the specification stating “In some aspects, the quantization training system 105 can use a b -bit symmetric uniform affine weight quantization, where b is the desired bitwidth of the parameters of the (quantized) aggregated model 125…In some aspects, during training of the adapter model 145, the quantization training system 105 can represent the parameters of the aggregated model 125 using Equation 1 below… s is a quantization scale…”, see MPEP 2106.04(a)(2), subsection I), and generate the quantized plurality of weights based on rounding and clipping the aggregated plurality of weights, (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see equation (1) from [0031] note: PNG media_image1.png 1 461 media_image1.png Greyscale in paragraph [0032] from the specification stating “using Equation 1, the quantization training system 105 may scale the (frozen) parameters W of the base model 110 using the initial (frozen) quantization scale s0, downcast the scaled base model 110 using φ, aggregate (e.g., concatenate) the downcast base model 110 with the parameters A and B of the adapter model 145, round the aggregated parameters to the nearest integer using round (⋅), clip the rounded parameters to values between −2b-1 and 2b-1−1 using clip (⋅), and finally scale the clipped parameters using s (which may be learned during training, or may be fixed)”, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 3 recites the following additional element: The processing system of claim 1, wherein, to generate the quantized plurality of weights, the one or more processors are configured to further execute the processor-executable instructions and cause the processing system… (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 4: Regarding claim 4, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 4 recites the following abstract ideas: …to: generate a downcast plurality of weights based on the first plurality of weights, wherein the downcast plurality of weights is re-used while training the adapter model; (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0033-34] from the specification stating “The quantized version of these integer weights may therefore be represented as WZ*s. However, because model operations (e.g., matrix multiplication, convolution, and the like) allow for this scale s to be pulled outside of the matrix multiplication (or other operation), the quantized version of WZ may not be explicitly computed during inference. Instead, the scale s (along with a scale of the activation data, if applicable) may be multiplied with the output of the matrix multiplication (or other operation) during inference. In some aspects, as discussed above, the downcasting operation φ may be implemented in a variety of ways. For example, in some aspects, φ is an identity operation (e.g., the weights are not downcast). In some aspects, φ(x)=BF16 (x) (e.g., the weights are converted to BF16). In some aspects, the downcasting operation is defined using Equation 2 below PNG media_image2.png 50 655 media_image2.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), and generate the quantized plurality of weights based on the downcast plurality of weights. this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraphs [0033-34] from the specification stating “The quantized version of these integer weights may therefore be represented as WZ*s ... Instead, the scale s (along with a scale of the activation data, if applicable) may be multiplied with the output of the matrix multiplication (or other operation) during inference. In some aspects, as discussed above, the downcasting operation φ may be implemented in a variety of ways. For example, in some aspects, φ is an identity operation (e.g., the weights are not downcast). In some aspects, φ(x)=BF16 (x) (e.g., the weights are converted to BF16). In some aspects, the downcasting operation is defined using Equation 2 below PNG media_image2.png 50 655 media_image2.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 4 recites the following additional element: 4. The processing system of claim 1, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system…(In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 5: Regarding claim 5, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Claim 5 recites the following abstract idea: 5. The processing system of claim 4, wherein, to generate the downcast plurality of weights, … reduce a bitwidth used to store the downcast plurality of weights, as compared to a bitwidth used to store the first plurality of weights, (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, since b is a variable that represents a desired bitwidth of the model, and reducing a bitwidth would mean reducing a variable b within equation 1, see in paragraph [0031] from the specification stating “ PNG media_image1.png 1 461 media_image1.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 5 recites the following additional element: … the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to … (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 6: Regarding claim 6, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Further, claim 6 recites the following abstract idea: 6. The processing system of claim 4, wherein, to generate the downcast plurality of weights, … convert each of the first plurality of weights to an integer format having a target bitwidth for the quantized plurality of weights; (This recites a mental process, since a person can mentally evaluate using pen and paper and convert weights (which are numbers or values) to an integer format, see MPEP 2106.04(a)(2)(III)), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process, but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 6 recites the following additional elements: … the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to … (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), and store the converted first plurality of weights using one or more data structures having at least double the target bitwidth, (In step 2A, prong 2, this recites mere data storing, which is considered an insignificant extra-solution activity – see MPEP 2106.05(g),). In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity which includes storing and retrieving information in memory data from court case Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; – see MPEP 2106.05(d) (II)(iv)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 7: Regarding claim 7, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Claim 7 recites the following abstract idea: …to: convert each of the first plurality of weights to an integer format having a target bitwidth for the quantized plurality of weights; (This recites a mental process, since a person can mentally evaluate using pen and paper and convert weights (which are numbers or values) to an integer format, see MPEP 2106.04(a)(2)(III)), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a mental process, but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 7 recites the following additional elements: The processing system of claim 4, wherein, to generate the downcast plurality of weights, the one or more processors are configured to further execute the processor-executable instructions and cause the processing system … (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), and for each respective weight of the converted first plurality of weights: store the respective weight using a first portion of a data structure having a greater bitwidth than the target bitwidth; (In step 2A, prong 2, storing recites mere data gathering, which is considered an insignificant extra-solution activity – see MPEP 2106.05(g),). In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity which includes storing and retrieving information in memory data from court case Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; – see MPEP 2106.05(d) (II)(iv)), and store a respective fractional portion of the respective weight using a second portion of the data structure, (In step 2A, prong 2, storing recites mere data gathering, which is considered an insignificant extra-solution activity – see MPEP 2106.05(g),). In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity which includes storing and retrieving information in memory data from court case Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; – see MPEP 2106.05(d) (II)(iv)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 8: Regarding claim 8, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Claim 8 recites the following abstract idea: The processing system of claim 1, … generate the quantized plurality of weights based further on a second quantization scale, (This recites a mathematical relationship, formula or equation, see specification in paragraph [0079] note “At block 535, the quantization training system optionally scales the clipped weights based on a second quantization scale (e.g., s in Equation 1). As discussed above, this second scale may be fixed (e.g., static) during training of the adapter, or may be learnable during training,” showing using a second quantization scale involves using math operations on a variable s in equation 1, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 8 recites the following additional element: … wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to … (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 9: Regarding claim 9, it is dependent upon claim 8, and thereby incorporates the limitations of, and corresponding analysis applied to claim 8. Further, claim 9 recites the following abstract idea: … to generate an updated value for the second quantization scale based on the loss, (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0031] from the specification stating “In some aspects, the quantization training system 105 can use a b-bit symmetric uniform affine weight quantization, where b is the desired bitwidth of the parameters of the (quantized) aggregated model 125. In some aspects, b is a hyperparameter. In some aspects, during training of the adapter model 145, the quantization training system 105 can represent the parameters of the aggregated model 125 using Equation 1 below. In Equation 1, Ŵ represents the parameters of the aggregated model 125, s is a quantization scale (which may be a trainable parameter, or may be frozen), φ is a downcasting operation, W is the parameters of the base model 110 (e.g., in original full precision, such as sixteen-bit or thirty-two-bit floating point), s0 is the quantization scale of the base model 110 (e.g., indicated in the quantization parameters 115), and A and B are the trainable parameters of the adapter model 145. PNG media_image1.png 1 461 media_image1.png Greyscale ”, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 9 recites the following additional element: The processing system of claim 8, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system … (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 10: Regarding claim 10, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Claim 10 recites the following abstract idea: … and re-generate the quantized plurality of weights during a corresponding backward pass based on the checkpointed at least one intermediate value, (this recites a mathematical relationship, mathematical formula or equation, or mathematical calculation, see in paragraph [0042] from the specification note “During the backward pass, the quantization training system 105 may re-execute a portion of the forward pass (e.g., using Equation 1 PNG media_image1.png 1 461 media_image1.png Greyscale ) to re-generate these activations or other data.” This shows that backward pass also uses equation 1 to perform math operations when generating weights, see MPEP 2106.04(a)(2), subsection I), If claim limitations, under their broadest reasonable interpretation, cover performance of the limitations as a math concept, but for the recitation of generic computer components, then it falls within the math concept grouping of abstract ideas. Accordingly, the claim “recites” an abstract idea. Further, claim 10 recites the following additional elements: The processing system of claim 1, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to, (In step 2A, prong 2, this recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), (In step 2B, this also recites mere instructions to apply an exception using generic computer or its components – see MPEP 2106.05(f)), …during training of the second plurality of weights: checkpoint at least one intermediate value used to generate the quantized plurality of weights during a forward pass of the training; (see specification in paragraph [0053] note “ the quantization training system 105 may checkpoint (e.g., store or cache) the input (e.g., the feature tensor 205). During the backward pass, this cached version can be retrieved and used to compute updates to the model,” where checkpoint can mean storing or caching data. (In step 2A, prong 2, this recites mere data gathering, which is considered an insignificant extra-solution activity – see MPEP 2106.05(g)), In step 2B, this insignificant extra-solution activity is well understood routine and conventional activity which includes storing and retrieving information in memory data from court case Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; – see MPEP 2106.05(d) (II)(iv)), Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible. Claim 11: Regarding claim 11, in step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A processor-implemented method for training machine learning models, comprising: accessing a first plurality of weights for a base model …”, and a method is one of the four statutory categories of invention. Since claim 11 recites similar limitations as corresponding independent claim 1 listed above, it is rejected for similar reasons under 35 U.S.C. 101. Claims 12 -20: Since claims 12-20 recite similar limitations as corresponding claims 2-10 listed above, they are rejected for similar reasons under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers, T., et al., in "QLORA: Efficient Finetuning of Quantized LLMs," published on May 23, 2023, available at: https://arxiv.org/pdf/2305.14314 , also cited in the August 2, 2024 IDS, (hereafter, Dettmers)., in view of Yuan, J., et al., in "Mobile foundation model as firmware," published on March 12, 2024, version 3 available at: https://arxiv.org/pdf/2308.14363 , (hereafter, Yuan). Claim 1: Regarding claim 1, Dettmers teaches “1. A processing system comprising: one or more memories comprising processor-executable instructions;” and “and one or more processors configured to execute the processor-executable instructions and cause the processing system,” See Dettmers in page 5, section on Paged Optimizers, which note the “ use the NVIDIA unified memory 3 feature [which] does automatic page-to-page transfers between the CPU and GPU for error-free GPU processing in the scenario where the GPU occasionally runs out-of-memory. The feature works like regular memory paging between CPU RAM and the disk. We use this feature to allocate paged memory for the optimizer states which are then automatically evicted to CPU RAM when the GPU runs out-of-memory and paged back into GPU memory when the memory is needed in the optimizer update step.” Here, Dettmers show using memory, including a CPU RAM, GPU, and a disk along with its processors to run the system. Further, see Dettmers in page 6, section on Experimental setup, note “We do, however, perform an analysis of the runtime of paged optimizers for 65B models on 48GB GPUs and find that with a batch size of 16, paged optimizers provide the same training speed as regular optimizers”. Here, Dettmers show using GPU processors for runtime of models, which relate to processors running processor-executable instructions. Further, Dettmers teaches “to: access a first plurality of weights for a base model;” See Dettmers in page 3, section 2. Background, subsection Low-rank Adapters mention "Low-rank Adapter (LoRA) finetuning [28] is a method that reduces memory requirements by using a small set of trainable parameters, often termed adapters, while not updating the full model parameters which remain fixed. Gradients during stochastic gradient descent are passed through the fixed pretrained model weights to the adapter, which is updated to optimize the loss function." Here, Dettmers show that adapter weights are updated but the base model weights remain the same. Further, Dettmers teaches “ access a second plurality of weights for an adapter model associated with the base model;” See Dettmers in page 4, first paragraph describe "For a 7B LLaMA model trained on FLAN v2 with a batch size of 1, with LoRA weights equivalent to commonly used 0.2% of the original model weights [28, 37], the LoRA input gradients have a memory footprint of 567 MB while the LoRA parameters take up only 26 MB. With gradient checkpointing [9], the input gradients reduce to an average of 18 MB per sequence making them more memory intensive than all LoRA weights combined." Here, Dettmers mentions LoRA weights, which relate to the weights from the adapter model or (i.e. second plurality of weights for an adapter model associated with the base model). The adapter model's weights are equivalent to commonly used 0.2% of the original model (i.e. base model's) weights show that the adapter model is associated with the base model. Further, see Dettmers in page 26, in figure 1 caption, note "Figure 6: Breakdown of the memory Breakdown of the memory footprint of different LLaMA models. The input gradient size is for batch size 1 and sequence length 512 and is estimated only for adapters and the base model weights" Here, Dettmers show first set of weights for the base model, and second set of weights for the adapter model. Further, Dettmers teaches “generate a quantized plurality of weights based on the first plurality of weights, a first quantization scale for the first plurality of weights, and the second plurality of weights;” See Dettmers in page 3, section 2. Background, subsection Block-wise k-bit quantization, note “to ensure that the entire range of the low-bit data type is used, the input data type is commonly rescaled into the target data type range through normalization by the absolute maximum of the input elements, which are usually structured as a tensor. For example, quantizing a 32-bit Floating Point (FP32) tensor into a Int8 tensor with range [−127,127]: PNG media_image3.png 207 1085 media_image3.png Greyscale .” Here, Dettmers show that first quantization scale is illustrated as an absolute maximum quantization, and this scale is run on a first set of weights. The tensor contain the set of weights, which are later quantized using equation 1. Also, see Dettmers in page 4, first paragraph describe “For a 7B LLaMA model trained on FLAN v2 with a batch size of 1, with LoRA weights equivalent to commonly used 0.2% of the original model weights [28, 37], the LoRA input gradients have a memory footprint of 567 MB while the LoRA parameters take up only 26 MB. With gradient checkpointing [9], the input gradients reduce to an average of 18 MB per sequence making them more memory intensive than all LoRA weights combined.” Here, Dettmers shows that the adapter model, called LoRA, has its own set of model weights or relate to a second plurality of weights, in addition to the original model (or original model) weights or the first set of weights. Further, Dettmers teaches “generate a loss based on processing training data using the quantized plurality of weights;” See Dettmers in page 8, section 5.1 Experimental Setup, in Training setup, note "To avoid confounding effects from different training objectives, we perform QLoRA finetuning with cross-entropy loss (supervised learning) without reinforcement learning," Here, Dettmers show processing training data or fine tuning with cross-entropy loss. Further, see Dettmers in page 3, section Low-Rank Adapters note for details. Further, see Dettmers in page 3, section 2. Background, subsection Low-rank Adapters mention "Low-rank Adapter (LoRA) finetuning [28] is a method that reduces memory requirements by using a small set of trainable parameters, often termed adapters, while not updating the full model parameters which remain fixed. Gradients during stochastic gradient descent are passed through the fixed pretrained model weights to the adapter, which is updated to optimize the loss function." Here, Dettmers show that adapter weights are updated but the base model weights remain the same, to later optimize a loss function. However, Dettmers did not teach "generate an updated second plurality of weights based on updating the second plurality of weights based on the loss;" or "and deploy a machine learning model comprising quantized versions of the first plurality of weights and the updated second plurality of weights". In an analogous art, Yuan teaches “generate an updated second plurality of weights based on updating the second plurality of weights based on the loss;” See Yuan in page 7, in section 3.3 Multi-path Task Execution, Training Details note for "(3) model weights updating: During each training iteration, we compute the CrossEntropy loss from actual labels and predicted tokens, utilizing it to update the PEFT/MLP parameters." Here, Yuan explicitly describes updating the adapter PEFT model weights using a cross entropy loss, to later produce updated model weights for each training iteration. Further, see Yuan in page 2, Introduction, describe "This vision becomes feasible thanks to recent advancements in the ML community, specifically: (1) The establishment of pre-trained foundation models [92, 105, 117] that capture extensive knowledge from vast Internet data; (2) The development of algorithms to accurately align multimodal data input [57, 114]; (3) The demonstration of parameter-efficient fine-tuning (PEFT) methods like LoRA [65,89] that efficiently adapt pre-trained models to diverse downstream tasks." Here, Yuan shows using a pre-trained foundation model or a base model, along with a parameter-efficient fine-tuning (PEFT) method using LoRA or adapter model, where the latter model contains a second set of weights used for updating based on a loss value described in page 7. Further, Yuan teaches “and deploy a machine learning model comprising quantized versions of the first plurality of weights and the updated second plurality of weights” See Yuan in page 9, from first paragraph through the second to last paragraph describe " However, even in these cases, M4 remains viable for deployment with usable performance... M4 can efficiently preserve the performance with low bit quantization. Figure 7 illustrates the performance comparison of M4 using quantized backbone with respect to the TS-model. As observed, M4 using 8-bit (INT8) and 4-bit (INT4) quantization both achieve nearly lossless accuracy, compared to M4 using 16-bit float representation (FP16)." Here, Yuan shows running M4, which contain both base model and an adapter model with quantized base model weights and updated adapter weights, then deployed to evaluate model performance, shown in figures 5-7 with metrics such as accuracy. Further, see Yuan in figures 5-7 in page 9 by running the quantized models that contain quantized versions of base model weights and updated adapter weights, then evaluating model performance metrics such as accuracy, It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Dettmers, and incorporate with the teachings of Yuan by using the teachings of Dettmers, with Yuan’s teaching of deploying a machine learning model using a first set of quantized weights and second set of updated weights based on the loss. One of ordinary skill in the art would be motivated to do so because by integrating Yuan’s framework into the methods of Dettmers, one with ordinary skill in the art would achieve the goal of providing “instead of building independent foundation models for different possible modalities, M4 provides a unified architecture that maximizes the capability sharing across different modalities, thus being more resource efficient and extensible;” (see Yuan in page 283, section 3. M4 Design and Prototyping, part 3.1 Overview). Claim 11: Regarding claim 11, it comprises of similar additional limitations as independent claim 1, and is rejected under the same rationale under 35 U.S.C. 103. Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers, in view of Yuan, further in view of Lialin, V. et al., in “Scaling down to scale up: A guide to parameter-efficient fine-tuning,” published on March 28, 2023, using version 1, available at: https://arxiv.org/pdf/2303.15647v1, (hereafter, Lialin). Claim 2: Regarding claim 2, Dettmers in view of Yuan, teach the limitations of claim 1. Further, Dettmers teaches “2. The processing system of claim 1, wherein the first plurality of weights and the first quantization scale are static when the updated second plurality of weights is generated,” See Dettmers in page 3, Low-rank Adapters Low-rank section describe the model "Adapter (LoRA) finetuning [28] is a method that reduces memory requirements by using a small set of trainable parameters, often termed adapters, while not updating the full model parameters which remain fixed. Gradients during stochastic gradient descent are passed through the fixed pretrained model weights to the adapter, which is updated to optimize the loss function." Here, Dettmers shows the first set of weights of the base model remain fixed or not changed (i.e. static) while the adapter model and its parameters are updated. Also, see Dettmers in page 4 mention "Since pretrained neural network weights usually have a zero-centered normal distribution with standard deviation σ (see Appendix F), we can transform all weights to a single fixed distribution by scaling σ such that the distribution fits exactly into the range of our data type. For our data type, we set the arbitrary range [−1,1]. As such, both the quantiles for the data type and the neural network weights need to be normalized into this range. The information theoretically optimal data type for zero-mean normal distributions with arbitrary standard deviations σ in the range [−1,1] is computed as follows: (1) estimate the 2k + 1 quantiles of a theoretical N(0,1) distribution to obtain a k-bit quantile quantization data type for normal distributions, (2) take this data type and normalize its values into the [−1,1] range, (3) quantize an input weight tensor by normalizing it into the [−1,1] range through absolute maximum rescaling." Here, Dettmers shows the quantization scaling is performed by the model on the weights. However, Dettmers in view of Yuan, did not teach “... quantization scale are static ...” In an analogous art, Lialin teaches “... quantization scale are static ...” See Lialin in page 10, in section 10.2 LoRA describe “All pre-trained model parameters are kept frozen, and only WA and WB matrices are trainable. The scaling factor is constant and typically equals 1/r. After training, they can be integrated into the original W by just adding the matrix WAWB to the original matrix W.” Here, Lialin shows the scaling factor for the pre-trained model stays constant or is static, and is 1 r . The term 'constant' is construed to have the same meaning as static here. Further, see Lialin in page 4, from section 3. Taxonomy of PEFT: a birds-eye view, part Why add parameters, note “ By saving memory on optimizer states, gradients, and allowing frozen model parameters to be quantized (Dettmers et al., 2022), additive PEFT methods enable the fine-tuning of much larger networks or the use of larger micro batch sizes. Which improves training throughput on GPUs.” Here, Lialin shows that the methods described from 10 can be applied to quantization methods. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reference of Dettmers and Yuan, and incorporate with the teachings of Lialin by using the teachings of Dettmers, with Lialin’s teaching of deploying a machine learning model using a static scale for quantization. One of ordinary skill in the art would be motivated to do so because by integrating Lialin’s framework into the methods of Dettmers and Yuan, one with ordinary skill in the art would achieve “although these methods introduce additional parameters to the network, they achieve significant training time and memory efficiency improvements by reducing the size of the gradients and the optimizer states,” (see Lialin in page 3, section 3. Taxonomy of PEFT: a birds-eye view, part Why add parameters). Claim 12: Regarding claim 12, it comprises of similar additional limitations as corresponding claim 2, and is rejected under the same rationale under 35 U.S.C. 103. Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, further in view of Xia W. et al., in “Chain of LoRa: Efficient fine-tuning of language models via residual learning,” published on January 8, 2024, available at: https://arxiv.org/pdf/2401.04151, (hereafter, Xia), and further in view of Guo, H., et al., in “LQ-LoRA: Low-rank plus quantized matrix decomposition for efficient language model finetuning,” published on Nov 20, 2023, available at: https://arxiv.org/pdf/2311.12023v1, also cited in the August 2, 2024 IDS, (hereafter, Guo). Claim 3: Regarding claim 3, Dettmers in view of Yuan, teach the limitations of claim 1. Further, Dettmers teaches “3. The processing system of claim 1, wherein, to generate the quantized plurality of weights, the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to: generate a scaled plurality of weights based on applying the first quantization scale to the first plurality of weights;” See Dettmers in page 4, section 3. QLoRA Finetuning, in sub-section 4-bit NormalFloat Quantization, describe "Since pretrained neural network weights usually have a zero-centered normal distribution with standard deviation σ (see Appendix F), we can transform all weights to a single fixed distribution by scaling σ such that the distribution fits exactly into the range of our data type. For our data type, we set the arbitrary range [−1,1]. As such, both the quantiles for the data type and the neural network weights need to be normalized into this range." Here, Dettmers mentions transforming weights of a model using a fixed scaled distribution within a range of a specific data type, which relates to producing scaled weights by applying a single fixed distribution by scaling σ (i.e. applying a first quantization scale to the weights). However, Dettmers in view of Yuan, did not teach “generate an aggregated plurality of weights based on the scaled plurality of weights and the second plurality of weights;" or "and generate the quantized plurality of weights based on rounding and clipping the aggregated plurality of weights". In an analogous system, Xia teaches “generate an aggregated plurality of weights based on the scaled plurality of weights and the second plurality of weights,” See Xia in page 3, section 3.2 Chain of LoRA, mention “for a pre-trained LLM weight matrix Wpretrained ∈ Rd× k, we denote the weights update occurred during fine-tuning as ∆W. Ideal adaptation yields the optimal weights W⋆ tailored for the given task and the corresponding optimal weight update ∆W⋆, as shown below. W⋆ = Wpretrained +∆W⋆”. Here, Xia explicitly shows how the weight update is combined from the weights of the pre-trained base model and any weight updates from the adapter model LoRA. Also, see Xia in page 3, from section 3.1 Preliminaries, note “During training, Wfrozen is frozen and only B, A are optimized. At deployment, the learned low-rank matrices can be merged with the frozen weights of the pre-trained model.” Here, Xia mentions the weights can be merged or aggregated. Merge is a term that is construed to be synonymous with aggregated. Also, see Xia in page 3, figure 1, where Xia illustrates the adapter model's weights merging with the model, including the weights, of the frozen pre-trained base model. See Xia in page 2, section LoRA and its variants, for details. Further, see Xia in page 6, in section 5.2 implementation details, describe “in all experiments, we set the rank of LoRA (denoted as ”r”) to 8 and α to 16, where the ratio α/r is employed to scale the weight updates.” Here, Xia shows a scaled set of weights in the adapter model called LoRA model. Later, see Xia in pages 2-3, section 3.1 Preliminaries, Low Rank Adaptation (LoRA) note "Consider a weight matrix Wfrozen from the pre-trained model, the weight update ∆W for task adaptation is represented with a low-rank decomposition BA. The forward pass with LoRA is as follows: PNG media_image4.png 163 640 media_image4.png Greyscale ∆W = 0 at the start of training. During training, Wfrozen is frozen and only B, A are optimized. At deployment, the learned low-rank matrices can be merged with the frozen weights of the pre-trained model.” Here, Xia shows that the variables B and A are the weight matrices of the adapter model LoRA, and these are optimized, and are considered the second set of weights (i.e. second plurality of weights) from the low-rank LoRA adapter model. PNG media_image5.png 903 1432 media_image5.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers and Yuan, and incorporate with the teachings of Xia by using the teachings of Dettmers and Yuan, with Xia’s teaching of generating the aggregated weights from the scaled weights. One of ordinary skill in the art would be motivated to do so because by integrating Xia’s framework into the methods of Dettmers and Yuan, one with ordinary skill in the art would achieve a method that “throughout the fine-tuning process, only the newly introduced lightweight adapters are trained, while the pre-trained model remains frozen and shared across tasks, thus significantly enhancing the practicality and efficiency of adapting large models to diverse tasks,” (see Xia in page 2, second paragraph in section, from section 2. Related work). However, Dettmers in view of Yuan, and further in view of Xia did not teach “and generate the quantized plurality of weights based on rounding and clipping the aggregated plurality of weights.” In an analogous art, Guo teaches “and generate the quantized plurality of weights based on rounding and clipping the aggregated plurality of weights,” See Guo mention in page 2, section 2.2 Weight Quantization of Large Language Models note in "Standard round-to-nearest (RTN) quantization, which quantizes/dequantizes a block of weights as u ≈ s × clamp ( [ 1 s   u ] ; −2b−1,2b−1 −1) with scaling factor s = m a x ( | u | ) 2 b - 1 - 1     max(|u|) / 2b−1−1 and bit size b, has been shown to be effective for quantizing a pretrained LLM’s weights to 8-bits". Here, Guo mentions a round to nearest operation (i.e. rounding) and a clamp step (i.e. clipping) on the block of weights (i.e. aggregated weights). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers, Yuan, and Xia, and incorporate with the teachings of Guo by using the teachings of Dettmers, Yuan, and Xia, with Guo’s teaching of generating the quantized plurality of weights involves rounding and clipping the aggregated plurality of weights. One of ordinary skill in the art would be motivated to do so because by integrating Guo’s framework into the methods of Dettmers, Yuan, and Xia, one with ordinary skill in the art would achieve a method that “apply LQ-LoRA to adapt RoBERTa (Liu et al., 2019) and LLaMA-2 (Touvron et al., 2023b) models and find that it can meaningfully improve upon strong QLoRA (Dettmers et al., 2023a) and GPTQ-LoRA (Frantar et al., 2022; Chai et al., 2023) baselines while enabling users to flexibly set a target memory budget,” (see Guo in page 2, third paragraph, from Introduction). Claim 13: Regarding claim 13, it comprises of similar additional limitations as corresponding claim 3, and is rejected under the same rationale under 35 U.S.C. 103. Claims 4, 5, 14, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, further in view of Sheng Y. et al., in “SLoRA: Scalable Serving of Thousands of LoRA Adapters,” published on May 12, 2024, available at https://proceedings.mlsys.org/paper_files/paper/2024/file/906419cd502575b617cc489a1a696a67-Paper-Conference.pdf , (hereafter, Sheng), further in view of Han Z., et al., in “Parameter-efficient fine-tuning for large models: A comprehensive survey,” published on March 21, 2024 as version 1, available at: https://arxiv.org/abs/2403.14608v1 , (hereafter, Han). Claim 4: Regarding claim 4, Dettmers in view of Yuan, teach the limitations of claim 1. However, Dettmers in view of Yuan, did not teach “The processing system of claim 1, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to: generate a downcast plurality of weights based on the first plurality of weights, wherein the downcast plurality of weights is re-used while training the adapter model; and generate the quantized plurality of weights based on the downcast plurality of weights.” In an analogous art, Sheng teaches “generate a downcast plurality of weights based on the first plurality of weights, wherein the downcast plurality of weights is re-used while training the adapter model;” See Sheng in page 2, sec. 2 Background note "In the training phase, LoRA freezes the weights of a pre-trained base model and adds trainable low-rank matrices to each layer. This approach significantly reduces the number of trainable parameters and memory consumption." Here, Sheng mentions that the method is applied during model training of LoRA models or adapter models. Further, see Sheng in page 2, first full paragraph, from Introduction, note "First, serving many LoRA adapters simultaneously requires efficient memory management. Since GPU memory is limited, we must store adapter weights outside the GPU and dynamically fetch them when needed." Here, Sheng mentions storing the adapter weights, then reuse these weights for training in page 2, section 2. Background. "Re-using" is construed to mean storing the downcasted weights in SRAM or cache memory during the training step so that both the forward and backward passes can immediately access them without triggering redundant type-conversion computations. Note that the examiner construes weights are re-used while training adapter model to mean an inherent property of a parameter efficient fine tuning model, where this type of model relies on keeping weights fixed on the base model, while updating only the weights of the added adapter models. Also, see Sheng in page 2, last paragraph of Introduction, note " When compared to the state-of-the-art parameter-efficient fine-tuning library, Huggingface PEFT, S-LoRA can enhance throughput by up to 30x." Here, Sheng describes that the weights from a PEFT model (called Huggingface PEFT) are re-used because the base model already uses the same fixed weights, which are also then re-used to train the adapter model S-LoRA when the adapter models are added. See figure 2 in Sheng for more details. Further, see Sheng in page 4, section 4.1 Batching, note "While the number of LoRA adapters can be large if we store them in main memory, the number of LoRA adapters needed for the currently running batch is manageable, because the batch size is bounded by the GPU memory. To take advantage of this, we store all LoRA adapters in the main memory and fetch only the LoRA adapters needed for the currently running batch to the GPU RAM when running the inference for that batch." Here, Sheng shows parameters including weights (mentioned from page 2) are re-used or fetched when needed when training. Storing adapters, including their weights, in memory and fetching these adapters when needed is viewed to be similar to re-using weights. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers and Yuan, and incorporate with the teachings of Sheng by using the teachings of Dettmers and Yuan, with Sheng’s teaching of a downcast plurality of weights is re-used while training the adapter model. One of ordinary skill in the art would be motivated to do so because by integrating Sheng’s framework into the methods of Dettmers and Yuan, one with ordinary skill in the art would achieve “In addition to system-level improvements, inference efficiency can be enhanced using algorithm techniques like quantization (Yao et al., 2022; Dettmers et al., 2022; Frantar et al., 2022; Xiao et al., 2023; Lin et al., 2023), sparsification (Frantar & Alistarh, 2023; Zhang et al., 2023b) and model architecture improvements (Shazeer, 2019). These approaches can reduce memory consumption and accelerate the computation, with a minor compromise in model quality,” (see Sheng in page 10, in section 8. Related work, Optimize LLM serving with algorithm techniques). However, Dettmers in view of Yuan, and further in view of Sheng, did not teach “and generate the quantized plurality of weights based on the downcast plurality of weights.” In an analogous art, Han teaches “and generate the quantized plurality of weights based on the downcast plurality of weights.” See Han in page 4 mention "While fine-tuning for a specific downstream task, only the weights of these additional modules or parameters are updated, which results in a substantial reduction in storage, memory, and computational resource requirements". Here, Han shows generating weights from the fine-tuned or quantization steps of a method. See Han in page 11, last paragraph, mention " For example, both Side-Tuning [213] and LST (Ladder-Side Tuning) [167] introduces a learnable network branch parallel to the backbone model. By channeling the backpropagation exclusively through this parallel branch, it circumvents the need to store gradient information for the main model’s weights, thus markedly reducing memory requirements during training." Here, Han shows using the downcast weights, where downcasting operation from specification paragraph [0028] note can be a reducing memory overhead operation, which relates to where Han mentions reducing memory requirements during model training. Downcasting from specification written description paragraph [0028] notes “ In some aspects, the downcasting component 130 is used to downcast the parameters of the aggregated model 125 (e.g., the base model 110 and/or the adapter model 145) during training to enable more efficient storage (e.g., reduced memory overhead) during training of the adapter model 145, as discussed in more detail below. As used herein, downcasting the parameters may generally include reducing the bitwidth used to store the parameters ... This downcasting can reduce memory overhead during training...” which the downcasting operation can be viewed as a form of memory saving reducing memory step or reducing bits of data, or compress data. Also, see Han in page 11, part C. Quantization strategies for PEFT, note " PEQA (Parameter Efficient and Quantization-aware Adaptation) [94] uses a two stage pipeline to achieve parameter-efficient and quantization aware fine-tuning. In the first stage, the pre-trained FFN weight matrix W P Rnˆm is quantized to W “ s ¨ W, where s P Rnˆ1 represents per-channel scales and W denotes the quantized weight.” PNG media_image6.png 193 884 media_image6.png Greyscale Here, Han shows the weights are quantized, which shows producing quantized weights using the downcast weights from the operation from page 11, last paragraph. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers, Yuan, and Sheng, and incorporate with the teachings of Han by using the teachings of Dettmers, Yuan, and Sheng, with Han’s teaching of generating the quantized plurality of weights based on the downcast plurality of weights. One of ordinary skill in the art would be motivated to do so because by integrating Han’s framework into the methods of Dettmers, Yuan, and Sheng, one with ordinary skill in the art would achieve “QA-LoRA uses INT4 quantization and introduces group-wise operators to enable quantization during inference stage, therefore improve the efficiency and accuracy compared with QLoRA,” (see Han in page 11, section C. Quantization Strategies for PEFT). Claim 5: Regarding claim 5, Dettmers in view of Yuan, teach the limitations of claim 4. Further, Dettmers teaches “5. The processing system of claim 4, wherein, to generate the downcast plurality of weights, the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to reduce a bitwidth used to store the downcast plurality of weights, as compared to a bitwidth used to store the first plurality of weights” See Dettmers in page 3, section 2. Background, in Block-wise k-bit Quantization, note "Quantization is the process of discretizing an input from a representation that holds more information to a representation with less information. It often means taking a data type with more bits and converting it to fewer bits, for example from 32-bit floats to 8-bit Integers... For example, quantizing a 32-bit Floating Point (FP32) tensor into a Int8 tensor with range [−127,127]:" Here, Dettmers shows that the system reduced a bitwidth from 32 FP to 8 int values for the quantized weights. Further, see Dettmers in page 2, first half paragraph mention "For a 7B LLaMA model trained on FLAN v2 with a batch size of 1, with LoRA weights equivalent to commonly used 0.2% of the original model weights...In comparison, the 4-bit base model consumes 5,048 MB of memory. This highlights that gradient checkpointing is important but also that aggressively reducing the amount of LoRA parameter yields only minor memory benefits. This means we can use more adapters without significantly increasing the overall training memory footprint" Here, Dettmers show the quantized weights are another set of weights, where the base model weights or the first set of weights, use a different bitwidth or 4-bit. Claim 14: Regarding claim 14, it comprises of similar additional limitations as corresponding claim 4, and is rejected under the same rationale under 35 U.S.C. 103. Claim 15: Regarding claim 15, it comprises of similar additional limitations as corresponding claim 5, and is rejected under the same rationale under 35 U.S.C. 103. Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, further in view of Sheng, further in view of Han, and further in view of Jin, Q., et al., in "Adabits: Neural network quantization with adaptive bit-widths," published from June 13-19, 2020 for a conference, added on web on August 5, 2020, available at: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9157698 , (hereafter, Jin). Claim 6: Regarding claim 6, Dettmers in view of Yuan, further in view of Sheng, and further in view of Han, teach the limitations of claim 4. Further, Dettmers teaches “6. The processing system of claim 4, wherein, to generate the downcast plurality of weights, the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to: convert each of the first plurality of weights to an integer format having a target bitwidth for the quantized plurality of weights;” See Dettmers in page 3, section 2. Background, in Block-wise k-bit Quantization, note “Quantization is the process of discretizing an input from a representation that holds more information to a representation with less information. It often means taking a data type with more bits and converting it to fewer bits, for example from 32-bit floats to 8-bit Integers. To ensure that the entire range of the low-bit data type is used, the input data type is commonly rescaled into the target data type range through normalization by the absolute maximum of the input elements, which are usually structured as a tensor. For example, quantizing a 32-bit Floating Point (FP32) tensor into a Int8 tensor with range [−127,127]:... PNG media_image3.png 207 1085 media_image3.png Greyscale ” Here, Dettmers show the target data type is an integer, where a floating point tensor (like 32 FP) is converted into an integer bitwidth (i.e. target bitwidth is in integer format) for the quantized weights. Also, See Dettmers in pages 1-2, in Introduction, describe "Our method, QLORA, uses a novel high-precision technique to quantize a pretrained model to 4-bit, then adds a small set of learnable Low-rank Adapter weights...that are tuned by backpropagating gradients through the quantized weights." Here, Dettmers show creating quantized weights. However, Dettmers in view of Yuan, further in view of Sheng, and further in view of Han, did not teach “and store the converted first plurality of weights using one or more data structures having at least double the target bitwidth.” In an analogous art, Jin teaches “and store the converted first plurality of weights using one or more data structures having at least double the target bitwidth,” See Jin in page 2149 section 5. Experiments, last paragraph of part 5.1 ImageNet Classification, note "The AdaBits models with the original scheme still need to store full precision weights in order to produce quantized weights in each bit-width. Our results prove that adaptive bit-width is an additional option for adaptive models, which is able to further improve trade-offs between efficiency and accuracy for deep neural networks." Here, Jin shows this model incorporates quantized weights of each bit width type, and applies to storing weights. Further, see Jin in page 2145 section 4. Quantization with Adaptive bit-withs in part 4.1. Benefits and Challenges, mention "Actually, for MobileNet V1/V2, changing the bit-width from 4 bit to 8 bit can enlarge the model size by 1.7× and the BitOPs by 3.2×, while the predictive accuracy can change by 1.5% on the ImageNet dataset with SAT [18]. From this we can see that there is a noticeable trade-off between accuracy and efficiency on quantized models." Here, Jin mentions using ada-bits can store the weights in a data structure that uses twice the target number of bits in this case. From 4-bit to 8-bit relates to storing into data structures having at least double the target bitwidth. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers, Yuan, Sheng, and Han, and incorporate with the teachings of Jin by using the teachings of Dettmers, Yuan, Sheng, and Han, with Jin’s teaching of storing the converted first plurality of weights using one or more data structures having at least double the target bitwidth. One of ordinary skill in the art would be motivated to do so because by integrating Jin’s framework into the methods of Dettmers, Yuan, Sheng, and Han, one with ordinary skill in the art would achieve “Our results prove that adaptive bit-width is an additional option for adaptive models, which is able to further improve trade-offs between efficiency and accuracy for deep neural networks,” (see Jin in page 2149, last paragraph of 5.1 ImageNet Classification). Claim 16: Regarding claim 16, it comprises of similar additional limitations as corresponding claim 6, and is rejected under the same rationale under 35 U.S.C. 103. Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, further in view of Sheng, further in view of Han, and further in view of Demaj, P., et al., in US PG Pub. No. US20200302266-A1, published on September 24, 2020, (hereafter, Demaj). Claim 7: Regarding claim 7, Dettmers in view of Yuan, further in view of Sheng, and further in view of Han, teach the limitations of claim 4. Further, Dettmers teaches “7. The processing system of claim 4, wherein, to generate the downcast plurality of weights, the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to: convert each of the first plurality of weights to an integer format having a target bitwidth for the quantized plurality of weights;” See Dettmers in page 3, section 2. Background, subsection Block-wise k-bit quantization, note “to ensure that the entire range of the low-bit data type is used, the input data type is commonly rescaled into the target data type range through normalization by the absolute maximum of the input elements, which are usually structured as a tensor. For example, quantizing a 32-bit Floating Point (FP32) tensor into a Int8 tensor with range [−127,127]: ... (1) where c is the quantization constant or quantization scale.” Here, Dettmers shows converting the weights into an integer format for the quantized weights. Here, Dettmers show that the tensor contain weights for the model, and the process converts these values from a 32-bit float into an 8-bit integer format. However, Dettmers in view of Sheng, and further in view of Han, did not teach “and for each respective weight of the converted first plurality of weights: store the respective weight using a first portion of a data structure having a greater bitwidth than the target bitwidth; and store a respective fractional portion of the respective weight using a second portion of the data structure.” In an analogous art, Demaj teaches “and for each respective weight of the converted first plurality of weights: See Demaj in paragraph [0024] note "In various embodiments, the term “Initial parameter” may be understood to mean a parameter relating to the configuration of the neural network, e.g. the weights of each layer and the size of the memory area to be allocated for the output data of each layer." Here, Demaj shows the weights of each layer relate to respective weight that uses different portion of a data structure. Also, see Demaj in [0015] note " reduce the size of a memory area of the memory array allocated to the fractional portion and increase the size of the memory area allocated to the integer portion of each new parameter associated with each layer." Where Demaj mentions parameter (which include weights) associated with each layer shows each "...for each respective weight of the converted first plurality of weights..." Further, Demaj teaches” ... store the respective weight using a first portion of a data structure having a greater bitwidth than the target bitwidth;” When Demaj mentions in abstract “adjusting a size of a memory area allocated to the fractional portion and... the integer portion...” and see Demaj in paragraph [0043] note “According to one implementation, if the difference is less than or equal to the second threshold value, the size of the memory area allocated to the integer portion is increased.” Here, Demaj shows allocating a memory for an integer portion (i.e. first portion of a data structure) and can increase the memory area size if the difference is less than or equal to a second threshold, where adjusting a size of memory area shows using a data structure having a greater bitwidth. See Demaj in [0015] for details. Further, Demaj teaches “and store a respective fractional portion of the respective weight using a second portion of the data structure” See Demaj, in paragraph [0008] note "Generally, the output data and the weights of each layer are represented in floating point e.g., over 32 bits, which makes it possible to have a neural network with better performance with regard to predictions. The output data and the weights of each layer may also be represented in fixed point, e.g. over 16 or 8 bits. ... “Fixed point” is understood to mean a representation of a number with a fixed number of decimal places. A fixed-point representation comprises an integer portion, i.e. the bits to the left of the decimal point, and a fractional portion corresponding to the number of bits to the right of the decimal point." Further, see Demaj in paragraph [0024] note "In various embodiments, the term “Initial parameter” may be understood to mean a parameter relating to the configuration of the neural network, e.g. the weights of each layer and the size of the memory area to be allocated for the output data of each layer." Here, Demaj shows the weights of each layer relate to respective weight that uses different portion of a data structure. Further, see Demaj in abstract mention “herein each new parameter of the set of new parameters has its data represented in two portions comprising an integer portion and a fractional portion; implementing the new neural network using a test input data set applied only once to each layer; determining a distribution function or a density function resulting from the set of new parameters for each layer; and based on the determined distribution function or density function, adjusting a size of a memory area allocated to the fractional portion and a size of the memory area allocated to the integer portion of each new parameter associated with each layer.” When Demaj describes data is represented in two portions, which contains an integer portion and a fractional portion, Demaj teaches store the respective weight using a first portion of a data structure and store a respective fractional portion ... using a second portion. Demaj refers to the first portion to be the integer portion, and the second portion to be a fractional portion. Also, see Demaj in [0005] describe "The architecture of a neural network generally comprises a succession of layers each of which takes its inputs from the outputs of the preceding layer. The output data (“features”) are stored in memory areas having a predefined size. The input data are multiplied by at least one weight of a given value for each layer." Here, Demaj shows weights are given for each layer of a neural network and applies these integer and fractional portions to respective weights per layer. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers, Yuan, Sheng, and Han, and incorporate with the teachings of Demaj by using the teachings of Dettmers, Yuan, Sheng, and Han, with Demaj’s teaching of storing both first portion and second portion of a data structure. One of ordinary skill in the art would be motivated to do so because by integrating Demaj’s framework into the methods of Dettmers, Yuan, Sheng, and Han, one with ordinary skill in the art would achieve “Convolutional neural networks (CNN) represent a type of neural network in which the connection pattern between the neurons is inspired by the visual cortex of animals. They allow the effective recognition of objects or persons in images or videos,” (see Demaj in paragraph [0004]). Claim 17: Regarding claim 17, it comprises of similar additional limitations as corresponding claim 7, and is rejected under the same rationale under 35 U.S.C. 103. Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, and further in view of Coelho, Aimee, in “Quantization in LLMs, why does it matter?”, published on Jan 11, 2024; available at: https://medium.com/data-from-the-trenches/quantization-in-llms-why-does-it-matter-7c32d2513c9e , (hereafter, Coelho). Claim 8: Regarding claim 8, Dettmers in view of Yuan, teach the limitations of claim 1. However, Dettmers in view of Yuan, did not teach “8. The processing system of claim 1, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to generate the quantized plurality of weights based further on a second quantization scale.” In an analogous art, Coelho teaches “8. The processing system of claim 1, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to generate the quantized plurality of weights based further on a second quantization scale” See Coelho describe in pages 4-5, introducing Blocks section and Double Quantization section note “To address this, a block-wise approach can be used. To create a block, the tensor is flattened and divided into segments of sizes such as 64 or 128, and each segment can then be scaled individually. This limits the impact of these outliers to only certain blocks. Smaller blocks yield higher accuracy, however the trade-off is that this increases the number of parameters that need to be stored as now there is one scaling factor for every block. Double Quantization : The QLoRA paper [4] offers a solution by proposing Double Quantization. This involves performing a second round of quantization, this time to quantize the scaling factors from the initial quantization of the weights. The 32-bit scale factors are grouped into blocks of 256 and scaled down to 8-bit precision with the introduction of a second round quantization factor. As illustrated below, storing one scaling factor in 32-bit for every block of 64 parameters adds 0.5 bits per parameter (32/64). Instead, using this double quantization to compress the per block scaling factors to 8-bit results in a reduction to only 0.127 bits per parameter (8/64 + 32/(256*64)).” Here, Coelho shows using a second scale for quantization in “a second round quantization factor”, to quantize the weights of the adapter model QLoRA. Since Coelho mentions there is one scaling factor for every block, and many blocks are introduced, this shows there are more than one scaling factor for quantization, which relate to a first quantization factor, second quantization factor, and subsequent ones. Further, see Coelho mention in Efficient Fine-Tuning with Quantization: QLoRA mention “ So far, these techniques focus on taking a pretrained model and compressing the parameters to make inference more efficient. But what about when we want to fine-tune a model for a specific task? … Techniques such as LoRA maintain the original weights frozen and add a small set of trainable weights, [often] called adapters, to train during the fine-tuning. These techniques can also benefit from quantization by loading a quantized version of the base model. QLoRA develops quantization of the parameters down to 4-bit with Double Quantization of the scaling factors down to 8-bit. While this can be used without fine-tuning, they also demonstrate that this 4-bit quantized version can replace the 16-bit base model and LoRa can be used to fine-tune the new parameters at 16-bit.” Here, Coelho shows this method can be applied to a base model and an adapter model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers and Yuan, and incorporate with the teachings of Coelho by using the teachings of Dettmers and Yuan, with Coelho’s teaching of generating the quantized plurality of weights based further on a second quantization scale. One of ordinary skill in the art would be motivated to do so because by integrating Coelho’s framework into the methods of Dettmers and Yuan, one with ordinary skill in the art would achieve a model like in “Figure 3: QLoRA uses a 4-bit quantization of the base transformer model to further improve the efficiency of LoRA for fine-tuning. They also introduce a paging mechanism to transfer the optimizer states to CPU during GPU memory spikes to avoid out-of-memory errors and ease the training of large models on a single machine,” (see Coelho in page 8, figure 3 caption). Claim 18: Regarding claim 18, it comprises of similar additional limitations as corresponding claim 8, and is rejected under the same rationale under 35 U.S.C. 103. Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, further in view of Coelho, and further in view of Frumkin, N. et al., "Jumping through Local Minima: Quantization in the Loss Landscape of Vision Transformers," published for a conference from October 1-6, 2023; available at https://ieeexplore.ieee.org/document/10378150 or pdf at: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10378150&tag=1 ; (hereafter, Frumkin). Claim 9: Regarding claim 9, Dettmers in view of Yuan, further in view of Coelho, teach the limitations of claim 8. However, Dettmers in view of Yuan, further in view of Coelho, did not teach “The processing system of claim 8, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to generate an updated value for the second quantization scale based on the loss.” In an analogous art, Frumkin teaches “9. The processing system of claim 8, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to generate an updated value for the second quantization scale based on the loss,” See Frumkin in 16938 , in section 5.3. Loss Function Choice mention "we compare the infoNCE (contrastive) loss with other common loss functions in Fig. 6. We find mean-squared error (MSE) to be equally (if not more) effective in the initial iterations of Evol-Q. However, as the number of passes grows, MSE does not perform as well as the infoNCE loss. Both cosine similarity and the Kullback–Leibler divergence (KL) fail to improve performance as the number of iterations increases. We postulate that the poor performance of these traditional loss functions is due to overfitting to the calibration dataset. On the other hand, the infoNCE loss is naturally regularized by the negative samples in the batch, allowing for it to preserve the quantization parameters that help discriminate between classes." Here, Frumkin mentions using the infoNCE loss to generate updated value for optimizing a quantization scale over many iterations, where iteration here is construed to mean the number of iterations that optimizes a quantization scaling factor. Here, Frumkin mentions from running a number of iterations, each number like a first iteration, second iteration correlates with a first quantization scale, second quantization scale, and subsequent ones. Also, see Frumkin in 16938, note "Figure 4: A zoomed in section of the landscape in Fig. 3b, where we perform gradient descent and evolutionary search for three initial points. We show the solutions of evolutionary search (X^evol) and gradient descent (X^GD) after 10 iterations." Here, Frumkin shows that this is performed for many iterations. Later, see Frumkin in page 16934, section 3.2 Where to perturb, describe "Our method applies end-to-end quantization meaning that for each attention block we quantize all 3N +1 weight tensors and 6N + 1 intermediary activations where N is number of heads. The quantization scales of all weights and activations can be concatenated and viewed as the vector ∆. This stacked vector is very important for understanding our algorithm – we can perturb the scales for all weights and activations simultaneously by perturbing ∆." Here, Frumkin mentions that this method applies to quantization scales of weights of models. See Frumkin in pages 16932-16933 , Introduction, note in step 3 “In comparison to non-contrastive loss functions such as mean squared error, cosine similarity, and the KL divergence, contrastive losses tend to smooth the loss landscape, as observed in our experiments and supported by recent work [10]. This finding inspires the use of contrastive loss to further facilitate the quantization scale search process. Contrastive losses, specifically the infoNCE loss in this work, also helps in combating overfitting on the small calibration dataset by incorporating negative examples into the loss.” Here, Frumkin shows using a contrastive loss to find a value for a quantization scale in optimizing quantization. See Frumkin in page 16933, section 2. Related work, ViT Quantization for details. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers, Yuan, and Coelho, and incorporate with the teachings of Frumkin by using the teachings of Dettmers, and Yuan, and Coelho, with Frumkin’s teaching of generating an updated value for the second quantization scale based on the loss. One of ordinary skill in the art would be motivated to do so because by integrating Frumkin’s framework into the methods of Dettmers, Yuan, and Coelho one with ordinary skill in the art would achieve “Evol-Q improves the top-1 accuracy of a fully quantized ViT-Base by 10.30%, 0.78%, and 0.15% for 3-bit, 4-bit, and 8-bit weight quantization levels. Extensive experiments on a variety of CNN and ViT architectures further demonstrate its robustness in extreme quantization scenarios,” (see Frumkin in abstract, on page 16932). Claim 19: Regarding claim 19, it comprises of similar additional limitations as corresponding claim 9, and is rejected under the same rationale under 35 U.S.C. 103. Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Dettmers in view of Yuan, further in view of Lee, J. et al., in PG Pub. No. KR20240048196-A, published on April 15, 2024, (hereafter, Lee). Claim 10: Regarding claim 10, Dettmers in view of Yuan, teach the limitations of claim 1. Further, Dettmers teaches “10. The processing system of claim 1, wherein the one or more processors are configured to further execute the processor-executable instructions and cause the processing system to, during training of the second plurality of weights: checkpoint at least one intermediate value used to generate the quantized plurality of weights during a forward pass of the training;” See Dettmers in page 4, first half of paragraph describe "For a 7B LLaMA model trained on FLAN v2 with a batch size of 1, with LoRA weights equivalent to commonly used 0.2% of the original model weights[28, 37], the LoRA input gradients have a memory footprint of 567 MB while the LoRA parameters take up only 26 MB. With gradient checkpointing [9], the input gradients reduce to an average of 18 MB per sequence making them more memory intensive than all LoRA weights combined." Here, Dettmers show using the method of gradient checkpointing the input gradients, which relates to checkpointing the intermediate value. For the weights from the adapter model LoRA, this relates to second set of weights. Further, see Dettmers in page 5, in section QLoRA mention " QLORA has one storage data type (usually 4-bit NormalFloat) and a computation data type (16-bit BrainFloat). We dequantize the storage data type to the computation data type to perform the forward and backward pass, but we only compute weight gradients for the LoRA parameters which use 16-bit BrainFloat." Here, Dettmers use both forward and backward pass for the checkpointing step. However, Dettmers in view of Yuan, did not teach “re-generate the quantized plurality of weights during a corresponding backward pass based on the checkpointed at least one intermediate value.” In an analogous art, Lee teaches “re-generate the quantized plurality of weights during a corresponding backward pass based on the checkpointed at least one intermediate value,” See Lee describe here in [0047] note "Forward propagation can mean calculating and storing variables in order from the input layer to the output layer of a neural network model. Backpropagation may refer to a method of calculating gradients for the parameters of an artificial neural network model. Backpropagation can calculate and store the gradients of the intermediate variables and parameters of the objective function related to each layer of the artificial neural network model from the output layer to the input layer. Weight update may mean replacing the existing weight with a weight determined through backpropagation. The process of learning through the forward propagation step, backpropagation step, and weight update step can be defined as iteration. For example, if an artificial neural network model is learned by repeating 10 times, the iteration of the artificial neural network model is It could be 10." Further, see Lee note in [0048] “Learning of the artificial neural network model according to one embodiment may further include a checkpointing step. Because training an artificial neural network model requires significant computational resources, if work is interrupted due to an unexpected problem in the processor, checkpointing and restart functions may be required to resolve the problem.” Here, Lee shows that weight update is part of the backpropagation step, which is construed to mean a similar concept as backward pass. Further, see Lee note in [0048] “Learning of the artificial neural network model according to one embodiment may further include a checkpointing step. Because training an artificial neural network model requires significant computational resources, if work is interrupted due to an unexpected problem in the processor, checkpointing and restart functions may be required to resolve the problem.” Here, Lee shows that weights are updated with backpropagation (which relates to a backward pass), and the system store intermediate variables which can be used in a checkpointing step. Since this step is iterative, this method repeat in generating weights during a checkpoint step (i.e. re-generating the quantized plurality of weights). Also, see Lee note in [0089] “Checkpointing and weight updates can be performed simultaneously with the backpropagation process of the previous layer, and can be pipelined as much as the dimension of model parallelism.” Here, Lee shows that checkpointing and weight update are a part of the method of backpropagation. Lee here primarily describes backpropagation is run with checkpointing at the same time. Further, see Lee mention in paragraph [0046] “referring to FIG. 2, learning an artificial neural network model may be a process of calculating and determining weights and biases to minimize the error between the final output value and the actual value. Learning of an artificial neural network model according to one embodiment may consist of a forward propagation step, a backward propagation step, and a weight update step.” Here, Lee shows using a backward propagation step, which is synonymous with running a backward pass from checkpointing on intermediate variables. Since both backward propagation and backward pass involve moving from the output layer back to the input layer to compute gradients, these methods show similarity in that backward pass includes backward propagation and weight update (which is described by Lee in paragraph [0047]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Dettmers and Yuan, and incorporate with the teachings of Lee by using the teachings of Dettmers and Yuan, with Lee’s teaching of re-generating the quantized plurality of weights during a corresponding backward pass based on the checkpointed at least one intermediate value. One of ordinary skill in the art would be motivated to do so because by integrating Lee’s framework into the methods of Dettmers and Yuan, one with ordinary skill in the art would achieve “The accuracy of the model is improved by updating state information such as parameters, embedding table, and optimizer state,” (see Lee in paragraph [0055]). Claim 20: Regarding claim 20, it comprises of similar additional limitations as corresponding claim 10, and is rejected under the same rationale under 35 U.S.C. 103. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WenWei Zeng/Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

May 15, 2024
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month