Prosecution Insights
Last updated: October 01, 2026
Application No. 18/422,646

SYSTEM FOR POST-TRAINING QUANTIZATION OF LARGE LANGUAGE MODELS

Non-Final OA §101§103§112
Filed
Jan 25, 2024
Examiner
RAWLINGS, ZANE ALEXANDER
Art Unit
Tech Center
Assignee
Amazon Technologies Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
9 currently pending
Career history
3
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §112
DETAILED ACTION This action is responsive to the Application filed on 1/25/2024. Claims 1-20 are pending in the case. Claims 1, 9 and 15 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 1-20 objected to because of the following informalities: In claims 1, 9, and 15, “the second LLM comprises a first set of integer values and the one or more integer approximations replace corresponding non-linear elements…” should read “the second LLM comprises a first set of integer values and the one or more integer approximations, wherein the one or more integer approximations replace corresponding non-linear elements…” for clarity and proper grammatical structure In claims 2-8 and 16-20, “the hardware processor to execute the first computer-executable instructions to:” is grammatically incorrect and should probably read “wherein the first computer-executable instructions further include:” Additionally, if the claims are amended using the suggested correction, the steps should be put into present participle form, i.e. “determining…” Any alternate corrections that make the introduction to the claims grammatically correct will also suffice. In claims 4, 12, and 18, “determine X data…” where X is first, second, third or fourth, should read “determine a X data set…” so as to prevent misconstruction of the limitation as always determining only a single piece of data In claims 4, 12, and 18, “determine second data using the first data as input to an activation layer” should likely read “determine second data using the first data as input to an activation layer in the first LLM.” The examiner comes to assume that this is a typo and the exclusion of “in the first LLM” was overlooked. However, if this is not a typo the listed claims will be rejected under 35 U.S.C. 112(b) for being indefinite as to which LLM the activation layer belongs In claims 6, 7, and 19, recitation of the methods for determining the linearly quantized model are clear but should be rephrased to prevent misconstruction of the limitations. For example, in claim 6, an example alternative could read: “The system of claim 1, wherein the determination of the linearly quantized model is based on…” In claims 8, 14, and 20 “the integer approximation for the SiLU utilizes the equation” should read “the integer approximation for the SiLU utilizes the equations” as two equations are presented, which is a plurality. In claim 15, the final line of the claim, “replace corresponding non-linear elements” was likely meant to be in accordance with that of claims 1 and 9 and should append “of the first LLM” to the end of the limitation Claims 2-8, 10-14, and 16-20 further inherit the objections of the claims upon which they depend Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-5, 8, 10-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 2, 10, and 16, recite the limitation “determine a first quantization parameter of the first set of quantization parameters based on minimizing a difference between non-quantized output and quantized output from a first layer of the first LLM.” First, it is unclear when this step is to occur in the instructions of the claims upon which these claims depend. Second, this limitation leaves it unclear how the quantization parameters other than the first one are determined. Under the most reasonable interpretation, given the specification, the examiner comes to understand that this limitation occurs during the determination of a first set of quantization parameters in claims 1, 9, and 15. Additionally, the examiner comes to understand that at least one quantization parameter is determined based on minimizing the difference between non-quantized output and quantized output from a first layer of the first LLM, not, potentially, just the first one. The examiner recommends rephrasing of the claims in accordance with these interpretations; one example alternative follows: “The system of claim 1, wherein determining a first set of quantization parameters comprises determining at least one quantization parameter by minimizing the difference between non-quantized output and quantized output from a first layer of the first LLM.” If the interpretations listed above are incorrect, corrections to claims 2, 10, and 15 are required to more clearly define the scope of their limitations in accordance with the issues listed above. Claims 3, 11, and 17 recite limitations relating to determining two sets of values within a first and second normalization layer of the second LLM, expressed using a first and second bit-depth, wherein the second bit-depth is different from the first bit-depth. It is unclear what this “determining” is used for and when. Under the most reasonable interpretation, given the specification, the examiner believes that these limitations should not be phrased as a step of determining. Given paragraph [0094] of the specification, these limitations should merely be stated to represent the end state of the second LLM. The examiner recommends rephrasing of the claims in accordance with this interpretation; one example alternative follows: “The system of claim 1, wherein the second LLM further comprises a first set of values within a first normalization layer and a second set of values within a second normalization layer, the first set of values expressed using a first bit depth and the second set of values expressed using a second bit depth, different from the first bit depth.” If the interpretations listed above are incorrect, corrections to claims 3, 11, and 17 are required to more clearly define the scope of their limitations in accordance with the issues listed above. Claims 4, 12, and 18 recite limitations relating the block process of the first LLM. Given these limitations rely on claims that recite the post-training quantization method of the present invention it is very unclear how they fit in as dependent claims. Under the most reasonable interpretation, given the specification, the examiner comes to understand that these claims are meant to limit the structure of the first LLM. The examiner recommends rephrasing of the claims in accordance with this interpretation; one example alternative follows: “The system of claim 1, wherein the first LLM comprises a block for: determining, based on input to a normalization layer of the first LLM…” If the interpretation listed above is incorrect, corrections to claims 4, 12, and 18 are required to more clearly define the scope of their limitations in accordance with the issues listed above. Claims 5, 13 and 15 recite the limitation “determine a first set of values using a first activation layer of the first [or a first] LLM.” It is unclear how this determination is made “using” a first activation layer as no input is provided. Given the specification and context of other claims, this first set of values is not based on input to the activation layer, but is found within the first activation layer, as in claims 3, 11, and 17. The examiner recommends rephrasing the claims in accordance with this interpretation; one example follows: “determine a first set of values within a first activation layer.” If the interpretation listed above is incorrect, for example, input is meant to be provided to the first activation layer to determine the first set of values, then corrections to claims 5, 13, and 15 are required to more clearly define the scope of their limitations in accordance with the issues listed above. Claims 16-20 further inherit the rejection of the claim upon which they depend, claim 15. Claims 5, 13, and 15 further recite limitations relating the block process of the first LLM. Given these limitations rely on claims that recite the post-training quantization method of the present invention it is very unclear how they fit in as dependent claims. Under the most reasonable interpretation, given the specification, the examiner comes to understand that these claims are meant to limit the structure of the first LLM. The examiner recommends rephrasing of the claims in accordance with this interpretation; one example alternative follows: “The system of claim 1, wherein the first LLM comprises a block for: determining a first set of values using a first activation layer of the first LLM…” If the interpretation listed above is incorrect, corrections to claims 5, 13, and 15 are required to more clearly define the scope of their limitations in accordance with the issues listed above. Claims 16-20 further inherit the rejection of the claim upon which they depend, claim 15. Claims 8, 14 and 20 recite the limitation “determine an integer approximation for a sigmoid-weighted linear unit (SiLU) in the first LLM…” It is unclear if this integer approximation is the same integer approximation performed in the claims upon which these claims depend. Given the specification, the examiner comes to understand that this integer approximation is the same integer approximation performed in the claims upon which these claims depend. The examiner recommends rephrasing of these claims in accordance with this understanding; one example alternative follows: “The system of claim 1, wherein the determining of one or more integer approximations for non-linear elements of the first LLM comprises using the equation…” The examiner notes that it is unclear if the inclusion of the SiLU as part of the claims listed is pertinent to the limitation as the approximation is simply used in claims 1, 9, and 15. If the interpretation listed above is incorrect, corrections to claims 8, 14, and 20 are required to more clearly define the scope of their limitations in accordance with the issues listed above. Claims 8, 14, and 20 do not recite descriptions for all of the elements in the equation presented in the claims. It is unclear what qs and Ss are and what they pertain to. Given the specification, the examiner can come to understand that they are the values determined by the SiLU but there should be adequate description to make their meanings definite. Further, m is referred to as an “output digit” and it is unclear what that term is meant to represent. Given the specification, there is no clear definition or explanation behind the term and therefore appropriate corrections are required to make the scope of the claim clear and definite. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed towards an abstract idea without significantly more. Step 1: Claims 1-8 are directed towards a machine, Claims 9-14 are directed towards a method, and Claims 15-20 are directed towards a machine. Therefore, Claims 1-20 are directed towards one of the 4 statutory categories; process, machine, article of manufacture, or composition of matter. With respect to claim 1: Step 2A Prong 1: The claim is directed to a judicial exception. determine a first large language model (LLM), wherein the first LLM comprises a first set of floating point values (Mental Process and Mathematical Concept: One could determine a first LLM that comprises a set of floating point values using mathematical concepts, mentally or using pen and paper) determine, based on the first LLM, a first set of scaling factors (Mental Process: One could determine based on the first LLM a first set of scaling factors, mentally or using pen and paper) determine, based on the first LLM, a first set of quantization parameters (Mental Process and Mathematical Concept: One could determine based on the first LLM a first set of quantization parameters using mathematical concepts, mentally or using pen and paper) determine, based on the first LLM and the first set of scaling factors, a preprocessed model, wherein a second set of floating point values of the preprocessed model are scaled relative to corresponding values of the first set of floating point values (Mental Process and Mathematical Concept: One could determine, based on the first LLM and first set of scaling factors, a preprocessed model, wherein a second set of floating point values of the preprocessed model are scaled relative to corresponding values of the first set of floating point values using mathematical concepts, mentally or using pen and paper) determine, based on the preprocessed model and the first set of quantization parameters, a linearly quantized model, wherein the linearly quantized model comprises integer quantized values based on one or more of the second set of floating point values (Mental Process and Mathematical Concept: One could determine, based on the preprocessed model and the first set of quantization parameters, a linearly quantized model, wherein the linearly quantized model comprises integer quantized values based on one or more of the second set of floating point values using mathematical concepts, mentally or using pen and paper) determine one or more integer approximations for non-linear elements of the first LLM (Mental Process and Mathematical Concept: One could determine one or more integer approximations for non-linear elements of the first LLM using mathematical concepts, mentally or using pen and paper) determine, based on the linearly quantized model and the one or more integer approximations, a second LLM, wherein the second LLM comprises a first set of integer values and the one or more integer approximations replace corresponding non-linear elements of the first LLM (Mental Process and Mathematical Concept: One could determine, based on the linearly quantized model and the one or more integer approximations, a second LLM, wherein the second LLM comprises a first set of integer values and the one or more integer approximations replace corresponding non-linear elements of the first LLM using mathematical concepts, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: A system comprising: a memory, storing first computer-executable instructions; and a hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: A system comprising: a memory, storing first computer-executable instructions; and a hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 2: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine a first quantization parameter of the first set of quantization parameters based on minimizing a difference between non-quantized output and quantized output from a first layer of the first LLM (Mental Process and Mathematical Concept: One could determine a first quantization parameter of the first set of quantization parameters based on the mathematical concept of minimizing a difference between non-quantized output and quantized output from a first layer of the first LLM, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 3: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine a first set of values within a first normalization layer of the second LLM, wherein the first set of values are expressed using a first bit depth (Mental Process: One could determine a first set of values within a first normalization layer of the second LLM which are expressed using a first bit depth, mentally or using pen and paper) determine a second set of values within a second normalization layer of the second LLM, wherein the second set of values are expressed using a second bit depth that is different from the first bit depth (Mental Process: One could determine a second set of values within a second normalization layer of the second LLM which are expressed using a second bit depth different from the first bit depth, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 4: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine, based on input to a normalization layer of the first LLM: a first set of vector values, a second set of vector values, and a third set of vector values (Mental Process: One could determine based on input to a normalization layer of the first LLM a first, second, and third set of vector values, mentally or using pen and paper) determine a fourth set of vector values and a fifth set of vector values using the first set of vector values and the second set of vector values as input to a rotary position encoding layer of the first LLM (Mental Process: One could determine a fourth and fifth set of vector values using the first and second set of vector values as input to a rotary position encoding layer of the first LLM, mentally or using pen and paper) determine first data using the fourth set of vector values and the fifth set of vector values as inputs to a first batched matrix multiplication function of the first LLM (Mental Process and Mathematical Concept: One could determine first data using the fourth and fifth set of vector value as input to a first batched matrix multiplication function of the first LLM using mathematical concepts, mentally or using pen and paper) determine second data using the first data as input to an activation layer (Mental Process: One could determine a second data using the first data as input to an activation layer, mentally or using pen and paper) determine third data using the second data and the third set of vector values as input to a second batched matrix multiplication function of the first LLM (Mental Process and Mathematical Concept: One could determine third data using the second and third set of vector values as input to a second batched matrix multiplication function of the first LLM using mathematical concepts, mentally or using pen and paper) determine fourth data using the third data as input to a linear down projection layer of the first LLM (Mental Process: One could determine fourth data using the third data as input to a linear down projection layer of the first LLM, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 5: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine a first set of values using a first activation layer of the first LLM (Mental Process: One could determine a first set of values using a first activation layer of the first LLM, mentally or using pen and paper) determine a second set of values using the first set of values as input to a linear down projection layer of the first LLM (Mental Process: One could determine a second set of values using the first set of values as input to a linear down projection layer of the first LLM, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 6: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine the linearly quantized model based on a per-tensor activation quantization using static scaling of the preprocessed model (Mental Process and Mathematical Concept: One could determine the linearly quantized model based on a per-tensor activation quantization using static scale of the preprocessed model by applying mathematical concepts, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 7: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine the linearly quantized model based on a per-channel weight quantization of the preprocessed model (Mental Process and Mathematical Concept: One could determine the linearly quantized model based on a per-channel weight quantization of the preprocessed model) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 8: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 1 via dependency. determine an integer approximation for a sigmoid-weighted linear unit (SiLU) in the first LLM (Mental Process and Mathematical Concept: One could determine an integer approximation for a SiLU in the first LLM by applying mathematical concepts, mentally or using pen and paper) PNG media_image1.png 89 417 media_image1.png Greyscale wherein the integer approximation for the SiLU utilizes the equation: where qx is a quantized input, Sx is a scaling factor associated with the quantized input, and (qe, Se) are determined based on processing inputs with an i_exp algorithm, m is an output digit, and sgn(x) returns a sign of a real number (Mathematical Concept: Merely recited a mathematical formula for performing some calculation, see MPEP 2106.04(a)(2)(I)(B)) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 1… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claims 9-14: See the rejections for claims 1-5 and 8 above. Note that the only difference between claims 9-14 and claims 1-5 and 8 is that claims 9-14 are directed towards the method whereas claims 1-5 and 8 are directed towards a machine that performs the method on a computer. The recited limitation of a “computer-implemented method comprising…” introduced in claim 9 merely falls under: Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). With respect to claim 15: See the rejections for claims 1 and 5 above. Note that the only difference between claim 15 and claims 1 and 5 is that claim 15 incorporates all of the limitations of claims 1 and 5 into a single claim. With respect to claim 16: See the rejection for claim 2 above. With respect to claim 17: Step 2A Prong 1: The claim is directed to a judicial exception, including those inherited from claim 15 via dependency. determine a third set of values within a first normalization layer of the second LLM, wherein the third set of values are expressed using a first bit depth (Mental Process: One could determine a third set of values within a first normalization layer of the second LLM which are expressed using a first bit depth, mentally or using pen and paper) determine a fourth set of values within a second normalization layer of the second LLM, wherein the fourth set of values are expressed using a second bit depth that is different from the first bit depth (Mental Process: One could determine a fourth set of values within a first normalization layer of the second LLM which are expressed using a second bit depth that is different from the first bit depth, mentally or using pen and paper) Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application. Additional Elements: The system of claim 15… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exception. Additional Elements: The system of claim 15… (See above) the hardware processor to execute the first computer-executable instructions to… (Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f)) With respect to claim 18: See the rejection for claim 4 above. With respect to claim 19: See the rejections for claims 6 and 7 above. Note that the only difference between claim 19 and claims 6 and 7 is that claim 19 incorporates all of the limitations of claims 6 and 7 into a single claim. With respect to claim 20: See the rejection for claim 8 above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-2, 6-7, 9-10, 16 and 19 are rejected under 35 U.S.C. as being unpatentable over Kim et. al. (Kim, Sehoon, et al. "I-bert: Integer-only bert quantization." International conference on machine learning. PMLR, 2021.) in view of Xiao et. al. (Xiao, Guangxuan, et al. "Smoothquant: Accurate and efficient post-training quantization for large language models." International conference on machine learning. PMLR, 2023.) Regarding claim 1, Kim et. al. teaches a system… (Abstract: “In this work, we propose I-BERT, a novel quantization scheme for Transformer based models that quantizes the entire inference with integer-only arithmetic.” The examiner notes that the system being proposed is I-BERT which uses a transformer based model called RoBERTa) …comprising: a memory, storying first computer-executable instructions; and a hardware processor to execute the first computer-executable instructions to… (4.2 Latency Evaluation, “We evaluate the latency speedup of INT8 inference of I-BERT, by direct deployment on a Tesla T4 GPU with Turing Tensor Cores that supports accelerated INT8 execution.” The examiner notes that a Tesla T4 GPU has both a memory and a processing device. Additionally, the reference covers any versions of hardware deployments, “work can be extensively deployed on other hardware as well.”) determine a first large language model (LLM)… (4.1 Accuracy Evaluation on GLUE, “While we only test RoBERTa-Base/Large, our method is not restricted to RoBERTa.” The examiner notes that RoBERTa is language model which is denoted as base/large, i.e. the version of RoBERTa used in the I-BERT method is a large language model; RoBERTa-Base/Large is understood to be the first LLM), wherein the first LLM comprises a first set of floating point values (4.1 Accuracy Evaluation on GLUE, “We implement I-BERT on the RoBERTa (Liu et al., 2019) model using (Ott et al., 2019). For the integer-only implementation, we replace all the floating point operations in the original model with the corresponding integer-only operation.” The examiner notes that as above, RoBERTa is the first LLM or “original model.” Additionally, the replacement of all the floating point operations in the original model indicates that RoBERTa comprises a first set of floating point values.) determine, based on the first LLM, a first set of scaling factors (3.1 Basic Quantization Method, “Under uniform symmetric quantization scheme, a real number x is uniformly mapped to an integer value…S is the scaling factor defined as α/(2b−1 −1).” The examiner notes that a scaling factor, S, is determined each time that a value is quantized when quantizing the entire first LLM, i.e. a first set of scaling factors are determined); determine, based on the first LLM, a first set of quantization parameters (3.1 Basic Quantization Method, “where b specifies the quantization bit precision…Q is the quantization operator, Int is the integer map (e.g., round to the nearest integer), clip is the truncation function, α is the clipping parameter used to control the outliers.” The examiner notes that other than the quantization operator, Q, the other variables comprise quantization parameters; There is no explicit limiting definition for quantization parameters, therefore, any parameter used in quantization is understood to be a quantization parameter. Again, note that the quantization parameters are determined each time a value is quantized when quantizing the entire first LLM, i.e. a first set of quantization parameters is determined); PNG media_image2.png 82 366 media_image2.png Greyscale determine, based on the first LLM and the first set of scaling factors (S)… a second set of floating point values… scaled relative to corresponding values of the first set of floating point values (3.1 Basic Quantization Method, “a real number x is uniformly mapped to an integer value q…The formal definition is:” The examiner notes that in the quantization method equation, the real numbers (first set of floating point values) are scaled using the first set of scaling factors S, thus producing a second set of floating point values during quantization of the entire first LLM) PNG media_image2.png 82 366 media_image2.png Greyscale determine, based on… the first set of quantization parameters, a linearly quantized model, wherein the linearly quantized model comprises integer quantized values based on one or more of the second set of floating point values (3.1 Basic Quantization Method, “a real number x is uniformly mapped to an integer value q…The formal definition is: The examiner notes that in the quantization method equation, the scaled set of real numbers (second set of floating point values) is quantized using the integer map Int( ) based on the first set of quantization parameters (see above); Thus, RoBERTa is now a linearly quantized model with respect to its original state (which had floating point values) and now comprises integer quantized values based on one or more of the second set of floating point values) determine one or more integer approximations for non-linear elements of the first LLM (3.2 Non-linear Functions with Integer-Only Arithmetic, Algorithm 1, “we approximate non-linear activation functions, GELU and SoftMax, with polynomials that can be computed with integer-only arithmetic.” The examiner notes that the non-linear activation functions comprise non-linear elements of the first LLM (RoBERTa); thereby, integer approximations are determined for the non-linear elements (GELU and SoftMax) of the first LLM (RoBERTa) so that their operations can be computed using Algorithm 1) and determine, based on the linearly quantized model and the one or more integer approximations, a second LLM, wherein the second LLM comprises a first set of integer values and the one or more integer approximations replace corresponding non-linear elements of the first LLM (4.2 Accuracy Evaluation on GLUE, “For the integer-only implementation, we replace all the floating point operations in the original model with the corresponding integer-only operations.” Further, see 3.2 Non-linear Functions with Integer-only Arithmetic, “we approximate non-linear activation functions, GELU and SoftMax, with polynomials that can be computed with integer-only arithmetic.” The examiner notes that, as above, RoBERTa was reconfigured into a linearly quantized model which comprises integer quantized values. Therefore, the non-linear functions in the linearly quantized version of RoBERTa are integer-approximated and replaced with their integer-only approximations to form a second LLM that comprises the integer values of the linearly quantized model and the integer approximations of the non-linear functions; thus, this version of RoBERTa is now understood to qualify as a second LLM) Kim et. al. does not distinctly disclose: Separating the steps of scaling and quantizing to first produce a pre-processed model with scaled values that is then used with the quantization parameters to produce a linearly quantized model. However, Xiao et. al. teaches: determine… a preprocessed model, wherein a… set of... values of the preprocessed model are scaled (1 Introduction, “SmoothQuant proposes a mathematically equivalent per-channel scaling transformation that significantly smooths the magnitude across the channels, making the model quantization-friendly.” The examiner comes to understand this as creating a pre-processed model that just has scaling done to the values before the model is quantized) and determine, based on the preprocessed model…a linearly quantized model (1 Introduction, “we implement three efficiency levels of quantization settings for SmoothQuant.” The examiner comes to understand that based on the pre-processed or value-scaled model, a linearly quantized model is created via multiple versions of quantization) Before the effective filing date of the claimed invention it would have been obvious to one or ordinary skill in the art to combine the method for post-training quantization of an LLM in Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) with the technique for producing a pre-processed model with scaled values (before quantization) of Xiao in order to allow for implementations with various quantization schemes (Xiao, 1 Introduction, “SmoothQuant is compatible with various quantization schemes.”) and smooth the values of the LLM before quantization so that they are easier to quantize (Xiao, Figure 2(b), “The smoothed activation X’ and the adjusted weight W’ are both easy to quantize”). Regarding claim 2, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above and Xiao et. al. further teaches the limitation: determine a first quantization parameter… (5.5 Ablation Study, Migration Strength, “We need to find a suitable migration strength α”) …of the first set of quantization parameters based on minimizing a difference between non-quantized output and quantized output from a first layer of the first LLM (5.5 Ablation Study, Migration Strength, “When α is too small (<0.4), the activations are hard to quantize; when α is too large (>0.6), the weights will be hard to quantize. Only when we choose α from the sweet spot (0.4-0.6) can we get small quantization errors for both weights and activations, and maintain the model performance after quantization.” The examiner comes to interpret that getting small quantization errors and maintaining model performance after quantization are analogous to minimizing the difference between non-quantized output and quantized output from a first layer of the first LLM. Again, without an explicit definition of quantization parameters, the migration strength is considered a quantization parameter as it affects how the quantization is performed) Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao with the technique for determining a quantization parameter based on minimizing the difference between quantized and non-quantized outputs of an LLM in Xiao in order to make the quantization of both the activations and weights easy (Xiao, Figure 10, “A suitable migration strength α (sweet spot) makes both activations and weights easy to quantize.”) and maintain performance of the LLM after quantization (Xiao, 5.5 Ablation Study, Migration Strength, “…and maintain the model performance after quantization.”) Regarding claim 6, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above and Xiao et. al further teaches the limitation: determine the linearly quantized model based on a per-tensor activation quantization… (2 Preliminaries, Quantization, “The per-tensor quantization uses a single step size for the entire matrix.” The examiner notes that as described in the reference this per-tensor quantization is used for both activation and weight quantization, i.e. covering per-tensor activation quantization) …using static scaling of the preprocessed model (2 Preliminaries, Quantization, Equation (1) and “We can calculate ∆ offline with the activations of some calibration samples, what we call static quantization.” The examiner notes that as in equation 1, ∆ is used to scale the floating point values before they are quantized, i.e. using a ∆ that is calculated offline and not changed at run-time; note that with no explicit definition of static scaling, the examiner interprets this as static scaling of the pre-processed model) Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao with the technique for performing per-tensor activation quantization using static scaling in Xiao in order to make the method more efficient by not requiring re-calculation of scaling factors for each quantization (2 Preliminaries, Quantization, “can calculate ∆ offline with the activations of some calibration samples… per-tensor quantization uses a single step size for the entire matrix” and Figure 3, “Per-tensor quantization is the most efficient to implement.”) Regarding claim 7, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above and Xiao et. al further teaches the limitation: determine the linearly quantized model based on a per-channel weight quantization of the preprocessed model (2 Preliminaries, Quantization, “We can further enable finer-grained quantization by using different quantization step sizes for… each output channel of weights (per-channel quantization)” The examiner comes to interpret quantization using different step sizes for each output channel of weights as per-channel weight quantization. See figure 3 (b)) Before the effective filing date of the claimed invention it would have been obvious to one or ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao with the technique for performing per-channel weight quantization in Xiao in order to preserve the accuracy of the LLM (Xiao, 3 Review of Quantization Difficulty, 1. Activations are hard to quantize than weights, “shown that quantizing weights of LLMs with INT8 or even INT4 does not degrade the accuracy, which echoes our observation.”) Regarding claim 9, see the rejection for claim 1 above. Note that the only difference between claim 9 and claim 1 is that claim 9 is directed towards a computer-implemented method whereas claim 1 is directed towards a system that performs that method, i.e. covering the method. Regarding claim 10, see the rejection for claim 2 above. Note that the only difference between claim 10 and claim 2 is that claim 10 is directed towards a computer-implemented method, introduced in claim 9, whereas claim 2 is directed towards a system that performs that method, introduced in claim 1, i.e. covering the method. Regarding claim 16, see the rejection for claim 2 above. Note that the only difference between claim 16 and claim 2 is that claim 16 depends on a system comprising the limitations of claims 1 and 5 whereas claim 2 depends on a system comprising the limitations of just claim 1. See the rejection for claim 5 below which incorporates the additional limitations into its rejection. Regarding claim 19, see the rejections for claims 6 and 7 above. Note that the only difference between claim 19 and claims 6 and 7 is that claim 19 depends on a system comprising the limitations of claims 1 and 5 whereas claims 6 and 7 depend on a system comprising the limitations of just claim 1. See the rejection for claim 5 below which incorporates the additional limitations into its rejection. Further, note that claim 19 merely incorporates the limitations of claims 6 and 7 into a single claim using an ‘and,’ without changing the scope of the claims. Claims 3, 11, and 17 are rejected under 35 U.S.C 103 as being unpatentable over Kim et. al. (Kim, Sehoon, et al. "I-bert: Integer-only bert quantization." International conference on machine learning. PMLR, 2021.) in view of Xiao et. al. (Xiao, Guangxuan, et al. "Smoothquant: Accurate and efficient post-training quantization for large language models." International conference on machine learning. PMLR, 2023.) further in view of Xu et. al. (Xu, Junhao, et al. "Mixed precision quantization of transformer language models for speech recognition." ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021). Regarding claim 3, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above and Xiao et. al. further teaches the limitation: …a first normalization layer (Figure 6, 1st LayerNorm in the figure) of the second LLM (Kim et. al., See the rejection for claim 1 above); and …a second normalization layer (Figure 6, 2nd LayerNorm in the figure) of the second LLM (Kim et. al., See the rejection for claim 1 above)… See the rationale for combining Kim et. al. with Xiao et. al. in the rejection for claim 1 above. Further, the combination of Kim et. al. and Xiao et. al. does not appear to distinctly disclose the limitation: determine a first set of values within a first normalization layer… wherein the first set of values are expressed using a first bit depth; and determine a second set of values within a second normalization layer… wherein the second set of values are expressed using a second bit depth that is different from the first bit depth However, Xu et. al. teaches this limitation: determine a first set of values (Section 2. Transformer LMS, Equation (3), ylt) within a first normalization layer (Section 2. Transformer LMS, Equation (4), LayerNorm. Note that y is put into the normalization, i.e. the first set of values y, is within a first normalization layer (LayerNorm of Eq. 5))… …wherein the first set of values (y) are expressed using a first bit depth (Fig. 1, “For the first transformer module positioned right after the embedding and position encoding layers, its multi-head attention layer (green) uses binary quantization.” The examiner comes to understand this to mean that y, the output of the attention layer, has binary quantization (1-bit depth), which is then input into the first LayerNorm. The first bit depth is thereby, 1-bit); and determine a second set of values (Section 2. Transformer LMS, Equation (5), slt) within a second normalization layer (Section 2. Transformer LMS, Equation (6), LayerNorm. Note that s is put into the normalization, i.e. the second set of values s, is within a second normalization layer (LayerNorm of Eq. 6))… …wherein the second set of values (s) are expressed using a second bit depth that is different from the first bit depth (Fig. 1, “its feed forward layer (orange) uses 4-bit quantization precision.” The examiner comes to understand this to mean that s, the output of the feed forward layer, has 4-bit quantization (4-bit depth), which is then input into the second LayerNorm. The second bit depth is thereby, 4-bits, different from the first bit depth which is 1-bit.) Before the effective filing date of the claimed invention it would have been obvious to one of ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao with the technique for using mixed-precision in different normalization layers of Xu in order to achieve model size compression without any degradation in performance. (Xu, Abstract, “Experiments conducted on Penn Treebank (PTB) and a Switchboard corpus trained LF-MMI TDNN system suggest the proposed mixed precision Transformer quantization techniques achieved model size compression ratios of up to 16 times over the full precision baseline with no recognition performance degradation.”) Regarding claim 11, see the rejection for claim 3 above. Note that the only difference between claim 11 and claim 3 is that claim 11 is directed towards a computer-implemented method, introduced in claim 9, whereas claim 3 is directed towards a system that performs that method, introduced in claim 1, i.e. covering the method. Regarding claim 17, see the rejection for claim 3 above. Note that the only difference between claim 17 and claim 3 is that claim 17 depends on a system comprising the limitations of claims 1 and 5 whereas claim 3 depends on a system comprising the limitations of just claim 1. See the rejection for claim 5 below which incorporates the additional limitations into its rejection. Claims 4, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et. al. (Kim, Sehoon, et al. "I-bert: Integer-only bert quantization." International conference on machine learning. PMLR, 2021.) in view of Xiao et. al. (Xiao, Guangxuan, et al. "Smoothquant: Accurate and efficient post-training quantization for large language models." International conference on machine learning. PMLR, 2023.) further in view of Su et. al. (Su, Jianlin, et al. "Roformer: Enhanced transformer with rotary position embedding." arXiv preprint arXiv:2104.09864 (2021)) Regarding claim 4, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above and Xiao et. al. further teaches the limitations to: determine, based on input to a normalization layer (Figure 6, ‘LayerNorm,’ The examiner notes that LayerNorm is understood be to a normalization layer, as it is represented in almost all transformer architectures. Further, see Kim 3.6 Integer-only LayerNorm as an example) of the first LLM… …a first set of vector values (Figure 6, ‘Q’, The examiner comes to understand this as being the query vector of the transformer model, i.e. the first set of vector values), …a second set of vector values (Figure 6, ‘K’, The examiner comes to understand this as being the key vector of the transformer model, i.e. the second set of vector values), and a …third set of vector values (Figure 6, ‘V’, The examiner comes to understand this as being the value vector of the transformer model, i.e. the third set of vector values) determine first data using the… set of vector values (Q) and the… set of vector values (K) as inputs to a first batched matrix multiplication function (Figure 6, ‘BMM,’ See the first BMM below the block for Q, K, and V. The examiner notes that BMM is understood to stand for “Batched Matrix Multiplication.” The two sets of vector values Q and K are input into the BBM block and the BMM block outputs the first data) of the first LLM; determine second data using the first data (Output of the first BMM) as input to an activation layer (Figure 6, ‘SoftMax,’ The examiner notes that SoftMax is an activation layer. The first data (output of the first BMM block) is input into the SoftMax block and the second data (Output of the SoftMax block) is determined) determine third data using the second data (Output of the SoftMax block) and the third set of vector values (V) as input to a second batched matrix multiplication function (Figure 6, ‘BMM,’ See the second BMM block below the SoftMax block. The second data (output from the SoftMax block) and third set of vector values (V) are input into the second BMM block and the second BMM block outputs the third data) of the first LLM: and determine fourth data using the third data (Output of the second BMM) as input to a linear down projection layer (Figure 6, ‘Projection,’ The examiner notes that the projection layer here is a linear down projection layer to turn the high-dimensional output from the SoftMax activation function into the dimensional size of the original model. See the rejection for claim 5) of the first LLM. See the rationale for combining Kim et. al with Xiao et. al. in the rejection for claim 1 above. Further, the combination of Kim et. al. and Xiao et. al. does not appear to distinctly disclose the limitation: determine a fourth set of vector values and a fifth set of vector values using the first set of vector values and the second set of vector values as input to a rotary position encoding layer of the first LLM However, Su et. al. teaches this limitation: determine a fourth set of vector values and a fifth set of vector values using the first set of vector values (Q) and the second set of vector values (K) as input to a rotary position encoding layer (1 Introduction, “We introduce a novel method, namely Rotary Position Embedding (RoPE).” 2.1 Preliminary, “The self-attention first incorporates position information to the word embeddings and transforms them into queries, keys, and value representations,” this is understood to be the same as determining the 3 sets of vector values, which are understood to be Query, Key, and Value vectors, as above. See the whole of section 3, including Equation 14 (General RoPE equation) and figure 1, which exemplifies using the first set of vector values (Q) and second set of vector values (K) to determine a fourth and fifth set of vector values (Position Encoded Query/Key in figure 1). The examiner notes that the fourth and fifth set of vector values could be put into the first BMM block of Xiao instead of Q and K with no unexpected outcomes) of the first LLM Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao and further modified with the transformer block mapping in Xiao with the technique for Rotary Position Embedding in Su in order to increase the accuracy of the system. (Su, 4.5.4 Results, “RoFormer outperforms WoBERT by an absolute improvement of 1.5%,” Table 5) Regarding claim 12, see the rejection for claim 4 above. Note that the only difference between claim 12 and claim 4 is that claim 12 is directed towards a computer-implemented method, introduced in claim 9, whereas claim 4 is directed towards a system that performs that method, introduced in claim 1, i.e. covering the method. Regarding claim 18, see the rejection for claim 4 above. Note that the only difference between claim 18 and claim 4 is that claim 18 depends on a system comprising the limitations of claims 1 and 5 whereas claim 4 depends on a system comprising the limitations of just claim 1. See the rejection for claim 5 below which incorporates the additional limitations into its rejection. Claims 5, 13, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et. al. (Kim, Sehoon, et al. "I-bert: Integer-only bert quantization." International conference on machine learning. PMLR, 2021.) in view of Xiao et. al. (Xiao, Guangxuan, et al. "Smoothquant: Accurate and efficient post-training quantization for large language models." International conference on machine learning. PMLR, 2023.) further in view of Vaswani (Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017)) Regarding claim 5, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above, but does not distinctly disclose the limitations to: determine a first set of values using a first activation layer of the first LLM; and determine a second set of values using the first set of values as input to a linear down projection layer of the first LLM However, Vaswani et. al. teaches these limitations: determine a first set of values using a first activation layer of the first LLM (3.3 Position Wise Feed-Forward Networks, Equation 2, “This consists of two linear transformations with a ReLU activation in between.” The examiner notes that the ReLU activation is understood to be the first activation layer and this method is being applied to FFNs which transformers are a type of. The ReLU activation layer determines the first set of values); and determine a second set of values using the first set of values as input to a linear down projection layer of the first LLM (3.3 Position Wise Feed-Forward Networks, Equation 2, “The dimensionality of input and output is dmodel = 512, and the inner-layer has dimensionality dff = 2048.” The examiner comes to understand that the values from the ReLU activation (first set of values) in the citation above are processed using the second linear transformation layer. This second liner transformation layer takes the high-dimensional output from the inner-layer (ReLU) and down-projects it to the dimensional size of the original model, i.e. a second set of values is determined using the first set of values from the activation layer by processing them as input to a linear transformation that reduces dimensionality of the input. Given the specification, this linear dimension reducing transformation is considered a linear down projection layer. The examiner notes that the first linear transformation layer has no relevance to the interpretation above and does not change the use/meaning.) Before the effective filing date of the claimed invention it would have been obvious to one of ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao with the technique for performing linear down projection after activation in Vaswani in order to allow the transformer to apply an activation function while maintaining the dimensionality of the original model (Vaswani, 3.3 Position-wise Feed-Forward Networks, “The dimensionality of input and output is dmodel = 512.”) Regarding claim 13, see the rejections for claims 5 and 6 above. Note that the only difference between claim 13 and claims 5 and 6 is that claim 13 is directed towards a computer-implemented method, introduced in claim 9, whereas claims 5 and 6 are directed towards a system that performs that method, introduced in claim 1, i.e. covering the method. Further, note that claim 13 merely incorporates the limitations of claims 5 and 6 into a single claim using an ‘and,’ without changing the scope of the claims. Regarding claim 15, see the rejections for claims 1 and 5 above. Note that the only difference between claim 15 and claims 1 and 5 is that claim 15 incorporates the limitations of claims 1 and 5 into a single claim, without changing the scope of claims. Claims 8, 14, and 20 are rejected under 35 U.S.C 103 as being unpatentable over Kim et. al. (Kim, Sehoon, et al. "I-bert: Integer-only bert quantization." International conference on machine learning. PMLR, 2021.) in view of Xiao et. al. (Xiao, Guangxuan, et al. "Smoothquant: Accurate and efficient post-training quantization for large language models." International conference on machine learning. PMLR, 2023.) further in view Elfwing et. al. (Elfwing, Stefan, Eiji Uchibe, and Kenji Doya. "Sigmoid-weighted linear units for neural network function approximation in reinforcement learning." Neural networks 107 (2018): 3-11) Regarding claim 8, Kim et. al. as modified by Xiao et. al. teaches all of the limitations of the system of claim 1 as cited above and Kim et. al. further teaches the limitation to: determine an integer approximation… (Algorithm 3, 3.5 Integer-only SoftMax, “Algorithm 3 describes the integer-only computation of the SoftMax function using i_exp.” The examiner comes to understand that i_exp is an integer approximation function for exponential functions, which in this case is being done on SoftMax) in the first LLM… …wherein the integer approximation… where qx is a quantized input (Algorithm 3, q), Sx (Algorithm 3, S) is a scaling factor associated with the quantized input, and (qe, Se) (Algorithm 3, Input, q and S) are determined based on processing inputs with an i_exp algorithm (Algorithm 3, I-Exp(q, s)) Further, the combination of Kim et. al. and Xiao et. al. does not appear to distinctly disclose the limitation: PNG media_image1.png 89 417 media_image1.png Greyscale determine an… approximation for a sigmoid-weighted linear unit (SiLU)… wherein the… approximation for the SiLU utilizes the equation: where… m is an output digit, and sgn(x) returns a sign of a real number. However, Elfwing et. al. teaches this limitation: PNG media_image1.png 89 417 media_image1.png Greyscale determine an… approximation for a sigmoid-weighted linear unit (SiLU)… (2. Method, “we propose the SiLU as an activation function for neural network function approximation…”) wherein the… approximation for the SiLU utilizes the equation: where… m is an output digit, and sgn(x) returns a sign of a real number. (2. Method, “The activation ak of the kth SiLU for input zk is computed by the sigmoid function multiplied by its input,” see equation 9. The examiner notes that anyone of ordinary skill in the art could re-write equation 9 of Elfwing, using equation 8 of Elfwing, to produce equation 23 in the specification of the present application. Then, one could apply the i_exp function from Kim et. al. in order to create equation set 1 from the present application, suited towards determining a linear approximation using a sigmoid-weighted linear unit. Finally, one would realize that the integer ratios for scale can be precomputed and are bounded by 2, allowing for them to be rescaled by 2/2m, where m is the output digit. The examiner notes that this combination of operations could be done by anyone of ordinary skill in the art to compute the equation claimed above) Before the effective filing date of the claimed invention it would have been obvious to one of ordinary skill in the art to combine the post-training LLM quantization method of Kim (including determining a first LLM, a set of scaling factors, and a set of quantization parameters to scale floating point values using the scaling factors and quantize the scaled values based on quantization parameters to produce an LLM with the integer quantized values and integer approximations of non-linear elements) as modified with the technique for first determining a pre-processed (scaled model) in Xiao with the method for performing a function approximation using a SiLU of Elfwing in order to improve the performance of the combined system (Elfwing, 1. Introduction, “After we first proposed the SiLU (Elfwing, Uchibe, & Doya, 2017), Ramachandran, Zoph, and Le (2017) recently performed a comprehensive comparison between the SiLU, the rectifier linear unit (ReLU; Hahnloser, Sarpeshka, Mahowald, Douglas, & Seung, 2000), and 6 other activation functions in the supervised learning domain. They found that the SiLU consistently outperformed the other activation functions when tested in 3 deep architectures on CIFAR-10/100 (Krizhevsky, 2009), in 5 deep architectures on ImageNet(Dengetal.,2009),andon4testsetsfor English-to-German machine translation) Regarding claim 14, see the rejection for claim 8 above. Note that the only difference between claim 14 and claim 8 is that claim 14 is directed towards a computer-implemented method, introduced in claim 9, whereas claim 8 is directed towards a system that performs that method, introduced in claim 1, i.e. covering the method. Regarding claim 20, see the rejection for claim 8 above. Note that the only difference between claim 20 and claim 8 is that claim 20 depends on a system comprising the limitations of claims 1 and 5 whereas claim 8 depends on a system comprising the limitations of just claim 1. See the rejection for claim 15 above which incorporates the combination of limitations into its rejection. Citation of Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to the applicant’s disclosure. Touvron et. al discloses a LLaMA language model architecture that incorporates the rotary position encoding layer of the RoFormer reference into its disclosure. Lin et. al. discloses a method for mixed-precision quantization of the weights of an LLM. Frantar et. al. proposes a GPTQ post-training quantization method for generative pre-trained transformer models that can dramatically compress some of the largest LLMs to 3 and 4 bits. Conclusion The prior art made of record and not relied upon is considered pertinent to Applicant's disclosure. Applicant is required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Zane A Rawlings whose telephone number is (571)270-3372. The examiner can normally be reached M-F, 8am to 5pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Z.A.R./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Jan 25, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month