DETAILED ACTION
This non-final office action is responsive to application 18/332,683 as submitted on June 9th 2023.
Claim status is currently pending and under examination for claims 1-20 of which independent claims are 1, 17 and 20.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent Claims 1, 17 and 20
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, independent claim 1, under the broadest reasonable interpretation, recites the following limitations that are abstract ideas:
for a network weight value of the matrix of network weight values, determining a representation quantity (R) of bits to be used for representing the network weight value during multiplication with a corresponding vector value of the input vector, based at least in part on a magnitude of the corresponding vector value;
The “determining” step involves identifying a quantity of bits to represent a network weight value based on a magnitude which amounts to no more than observations, evaluations, and judgments that can be performed in the human mind or with the use of a physical aid (e.g., pen and paper). The claim recites the step of determining a representation quantity of bits at a high degree of generality, thus the step is not required to have any specific level of complexity that would preclude the step from being mental processes. Therefore, the “determining” step is considered to be mental processes, see MPEP § 2106.04(a)(2)(III).
Therefore, the independent claim recites a judicial exception. Independent claims 17 and 20 recite similar limitations corresponding to claim 1, therefore the same subject matter eligibility analysis is applied.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, the judicial exception recited above is not integrated into a practical application. The claims recite the following additional elements, but these additional elements are not sufficient to integrate the judicial exception into a practical application:
during execution of a machine learning model, receiving an input vector for multiplication with a matrix of network weight values, (MPEP § 2106.05(g) necessary data gathering and insignificant extra-solution activity to the judicial exception)
wherein each network weight value of the matrix of network weight values is stored in computer memory using a stored quantity (S) of bits; (MPEP § 2106.05(g) storing and retrieving information in memory and insignificant extra-solution activity to the judicial exception)
and retrieving R bits of the network weight value from the computer memory for multiplication with the corresponding vector value. (MPEP § 2106.05(g) storing and retrieving information in memory and insignificant extra-solution activity to the judicial exception)
a logic subsystem; (claim 17) (MPEP § 2106.05(f) mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea)
and a storage subsystem including computer memory, the storage subsystem holding instructions executable by the logic subsystem to: (claim 17) (MPEP § 2106.05(f) mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea)
wherein R is a smaller number of bits than S; (claim 20) (MPEP § 2106.05(f) mere instructions to implement an abstract idea on a computer, or generally links exception to a technological environment)
retrieving R bits of the network weight value from the computer memory for multiplication with the first corresponding vector value, such that at least some bits of the network weight value are stored in the computer memory and not retrieved for multiplication with the first corresponding vector value; (claim 20) (MPEP § 2106.05(g) storing and retrieving information in memory and insignificant extra-solution activity to the judicial exception)
and retrieving S bits of a second network weight value from the computer memory for multiplication with a second corresponding vector value of the input vector, the second corresponding vector value being larger than the first corresponding vector value. (claim 20) (MPEP § 2106.05(g) storing and retrieving information in memory and insignificant extra-solution activity to the judicial exception)
The “receiving” step amounts to mere data gathering and is recited at a high level of generality, thus adding insignificant extra-solution activity to the judicial exception – see MPEP § 2106.05(g). Under MPEP § 2106.05(d), such additional elements have been found by the courts to not integrate a judicial exception into a practical application.
The “wherein each network weight value of the matrix …” step amounts to merely storing information and is recited at a high level of generality, thus adding insignificant extra-solution activity to the judicial exception – see MPEP § 2106.05(g). Under MPEP § 2106.05(d), such additional elements have been found by the courts to not integrate a judicial exception into a practical application.
The “retrieving” steps amount to merely retrieving information in memory and are recited at a high level of generality, thus adding insignificant extra-solution activity to the judicial exception – see MPEP § 2106.05(g). Under MPEP § 2106.05(d), such additional elements have been found by the courts to not integrate a judicial exception into a practical application.
The “wherein R is a smaller number of bits than S” step is recited at a high-level of generality such that the limitation amounts to no more than mere instructions to “apply” the judicial exception on a computer. It can also be viewed as nothing more than an attempt to generally link the use of the judicial exception to the technological environment of computers, see MPEP § 2106.05(f).
The remaining additional elements are recited at a high-level of generality such that they amount to no more than mere instructions to “apply” an exception using a generic component. Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, see MPEP § 2106.05(f).
Therefore, the above limitations do not integrate the judicial exception into a practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to amount to significantly more than the judicial exception.
In regards to the “receiving” step, this step adds insignificant extra-solution activity. An extra-solution activity is a well-understood, routine and conventional (WURC) activity per MPEP § 2106.05(d)(II), “the courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network, e.g., using the Internet to gather data.” The “receiving” step does not integrate the judicial exception into a practical application and does not amount to significantly more.
In regards to the “wherein each network weight value of the matrix …” step, this step adds insignificant extra-solution activity. An extra-solution activity is a well-understood, routine and conventional (WURC) activity per MPEP § 2106.05(d)(II), “the courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity … iv. Storing and retrieving information in memory.” The step does not integrate the judicial exception into a practical application and does not amount to significantly more.
In regards to the “retrieving” steps, the steps add insignificant extra-solution activity. An extra-solution activity is a well-understood, routine and conventional (WURC) activity per MPEP § 2106.05(d)(II), “the courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity … iv. Storing and retrieving information in memory.” The “retrieving” steps do not integrate the judicial exception into a practical application and do not amount to significantly more.
In regards to the “wherein R is a smaller number of bits than S” step and the remaining additional elements, the limitations are recited so generically such that they amount to no more than mere instructions to “apply” the judicial exception on a computer using generic computer components. Mere instructions to apply a judicial exception cannot provide an inventive concept. See MPEP § 2106.05(f).
Therefore, independent claims 1, 17 and 20 are not patent eligible.
Dependent Claims 2-16 and 18-19
The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than a judicial exception.
Claim limitation
Examiner analysis
2, 18. The method of claim 1, wherein R is a smaller number of bits than S, such that at least some bits of the network weight value are stored in the computer memory and not retrieved for multiplication with the corresponding vector value.
This is merely additional information about one or more previously identified mental processes.
3. The method of claim 1, further comprising, for a second network weight value of the matrix of network weight values, determining a second representation quantity (R’) of bits to be used for representing the second network weight value during multiplication with a second corresponding vector value of the input vector, based at least in part on a magnitude of the second corresponding vector value.
This is a mental process akin to a human evaluation/judgment/observation.
4. The method of claim 3, wherein the corresponding vector value is smaller than the second corresponding vector value, and wherein R is a smaller number of bits than R’.
This is merely additional information about one or more previously identified mental processes.
5. The method of claim 4, wherein R’ is equal to S, such that all stored bits of the second network weight value are retrieved for multiplication with the second corresponding vector value.
This is merely additional information about one or more previously identified mental processes.
6. The method of claim 1, wherein the corresponding vector value of the input vector is equal to zero, and R is equal to zero.
This is merely additional information about one or more previously identified mental processes.
7. The method of claim 1, wherein R is further determined based at least in part on a difference between the corresponding vector value and other vector values of the input vector.
This is merely additional information about one or more previously identified mental processes.
8. The method of claim 7, wherein R is relatively larger based at least in part on determining that the corresponding vector value is relatively higher than an average vector value of the input vector.
This is merely additional information about one or more previously identified mental processes.
9. The method of claim 1, wherein R is further determined based at least in part on a power consumption policy of the computing device.
This is merely additional information about one or more previously identified mental processes.
10. The method of claim 1, wherein the matrix of network weight values are stored in the computer memory such that network weight values in a same row of the matrix are stored in a same memory row or memory page of the computer memory.
This is merely additional information about one or more previously identified mental processes.
11. The method of claim 10, wherein the same memory row or memory page is organized beginning with most significant bits for each network weight value, and ending with least significant bits for each network weight value.
This is merely additional information about one or more previously identified mental processes.
12. The method of claim 10, wherein the same memory row or memory page is organized beginning with sign bits, followed by exponent bits, and ending with mantissa bits for each network value stored in the same memory row or memory page.
This is merely additional information about one or more previously identified mental processes.
13. The method of claim 1, wherein the machine learning model is a transformer-based language model.
This is merely additional information about one or more previously identified mental processes.
14. The method of claim 13, wherein the input vector represents tokens of a natural language input.
This is merely additional information about one or more previously identified mental processes.
15. The method of claim 13, wherein the input vector represents intermediate result vectors inside hidden layers of the transformer-based language model.
This is merely additional information about one or more previously identified mental processes.
16. The method of claim 1, wherein the computer memory includes one or more of dynamic random-access memory (DRAM), flash memory, and static random-access memory (SRAM).
This is merely additional information about one or more previously identified mental processes.
19. The computing system of claim 17, wherein the instructions are further executable to, for a second network weight value of the matrix of network weight values, determine a second representation quantity (R’) of bits to be used for representing the second network weight value during multiplication with a second corresponding vector value of the input vector, based at least in part on a magnitude of the second corresponding vector value,
wherein the corresponding vector value is smaller than the second corresponding vector value, and wherein R is a smaller number of bits than R’.
The “determine” step is a mental process akin to a human evaluation/judgment/observation.
The remaining steps are merely additional information about one or more previously identified mental processes.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Lin, Ji, et al. “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration” arXiv preprint arXiv:2306.00978v1 (2023)
Park (US 20220309314 A1)
Claims 1, 3, 7-10 and 13-17 are rejected under 35 U.S.C. 103 as being unpatentable over Lin / Park.
Regarding Claim 1, Lin teaches:
A method for computer memory access, comprising ((P. 5, Sec. 2.3, ¶1) “for AWQ (left), all memory accesses are contiguous, resulting in a 2× end-to-end speedup for LLaMA models.”):
… receiving an input vector for multiplication with a matrix of network weight values ((P. 1-2, Sec. 1, Last Paragraph) “To find the salient weight channels, the insight is that we should refer to the activation distribution instead of the weight distribution, despite we are doing weight-only quantization: weight channels corresponding to larger activation magnitudes are more salient since they process more important features”
(P. 4, Sec. 2.3, ¶2) “we instead model these layers as matrix-matrix (MM) multiplications.”
(P. 8, Sec. 3.3) “We attribute this efficiency improvement to our formulation of linear layers in LLMs as matrix-matrix instead of matrix-vector product.”
See Figure 2 on P. 4 depicting a matrix of weights. See (P. 4, Sec. 2.2, ¶1) describing activations are retrieved from a calibration set.
Linear layers of an LLM are modeled as matrix-matrix multiplications. Weight channels corresponding to larger activation magnitudes are more salient (important) since the weight channels process more important features. To process important features, weight channels corresponding to larger activation magnitudes must be multiplied together, therefore the matrix-matrix multiplications in the LLM multiply weights (‘matrix of network weight values’) with activations (‘input vector’).),
wherein each network weight value of the matrix of network weight values is stored in computer memory using a stored quantity (S) of bits ((P. 2, Sec. 2, ¶1) “Low-bit weight-only quantization [17, 39], which quantizes weights into low-bit integers (typically ≤4 bits). The decoding stage dominates LLMs’ total runtime [17, 32], which is highly memory-bounded with a single batch size. Given that the memory is dominated by weights, we focus on weight-only quantization”
(P. 3-4, Sec. 2.2, ¶3) “We automatically search for an optimal (per input channel) scaling factor that minimizes the output difference after quantization for a certain layer. … W is the original weights in FP16”
Original weights W (‘matrix of network weight values’) are stored in memory in FP16 (Half-precision floating-point format, 16 bits) before being quantized into low-bit integers. Therefore, each weight (‘network weight value’) of matrix W is stored in computer memory using a stored quantity S of bits (S being 16 bits).);
for a network weight value of the matrix of network weight values, determining a representation quantity (R) of bits to be used for representing the network weight value during multiplication with a corresponding vector value of the input vector, based at least in part on a magnitude of the corresponding vector value (The Examiner interprets “representation quantity (R) of bits” according to its broadest reasonable interpretation (BRI) in view of the Applicant’s specification as encompassing “quantized weights”. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0015], (see excerpt below).
Applicant’s written description at [0015]: “a representation quantity (R) of bits that should be retrieved from computer memory as a representation of the network weight value. This may differ from a stored quantity (S) of bits used to store the network weight value in the computer memory. For instance, R may be relatively higher than the R values for other network weight values in cases where the corresponding vector value is relatively higher, and the corresponding vector values for the other network weight values are relatively lower, which serves to reduce quantization error as described above. However, in some cases R may be significantly less than S, or even reduced to zero, when quantization is relatively less likely to have a significant effect on the vector-matrix product – e.g., because the corresponding vector value is relatively low, or equal to zero.”
(P. 2, Sec. 2, ¶1) “Low-bit weight-only quantization [17, 39], which quantizes weights into low-bit integers (typically ≤4 bits). … Given that the memory is dominated by weights, we focus on weight-only quantization.”
(P. 2, Sec. 1, ¶1) “To avoid mixed precision, we follow the activation-awareness principle and design a per-channel scaling method to search for the optimal scaling factors that minimize the quantization error, and quantize all the weights”
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized. … The quantized performance generally gets better as the scaling factor increases since important weights are better represented”
(P. 4, Sec. 2.2, ¶2) “the optimal scale should be related to: 1. Activation magnitude: as mentioned above, we are relying on input activation magnitude to pick out the salient weight channels. Therefore, we should consider the magnitude of X to protect the salient weight channels. We compute the average activation magnitude
s
x
=
m
e
a
n
c
_
o
u
t
|
X
|
as the importance factor. 2. Weight magnitude: to minimize the quantization loss of non-salient weights, we should flatten their distribution so that it is easier for quantization … We determine the final scaling factors by a function considering both of the factors
s
=
f
(
s
x
,
s
w
)
.”
Activations (‘input vector’) have magnitudes, therefore each activation (‘corresponding vector value of the input vector’) has a magnitude. Activation magnitudes are used to determine salient (important) weight channels (weights). Weight channels with larger activation magnitudes have a larger scaling factor, therefore these weights are more precisely quantized than non-salient weight channels (weight channels with smaller activation magnitudes). Activation magnitudes are therefore used to determine a scaling factor that determines how precise quantization is performed to convert FP16 weights to low-bit integers (to minimize quantization error), and therefore the quantized low-bit integers are a representation quantity (R) of bits that represent network weight values (weights). The quantized low-bit integers are 3 or 4 bits (see Page 1, Abstract) and are a different quantity of bits than original weights stored in FP16 (stored quantity S). See (P. 4, Sec. 2.3, ¶2) describing matrix-matrix (MM) multiplications.);
and retrieving R bits of the network weight value from the computer memory for multiplication with the corresponding vector value ((P. 5, Sec. 2.3, ¶1) “In Figure 2, we assume that the weights are stored in row major (i.e. IC×OC). Each OC has different scales and zero points and every two ICs within each OC share dequantization parameters (i.e. g = 2). For GPTQ with reordering (right), since these g = 2 ICs are not continuous, irregular DRAM accesses are required to fetch scaling factors and zero points when dequantizing each weight. However, for AWQ (left), all memory accesses are contiguous, resulting in a 2× end-to-end speedup for LLaMA models.”
(P. 10, Sec. 4, Last Paragraph) “Our AWQ kernels are executed on tensor cores, suitable for both context and generation phases in LLM inference, and do not require hardware-inefficient reordering.”
(P. 2, Sec. 1, ¶2) “We employ reorder-free online dequantization to efficiently convert low-bit weights to FP16”
Lin discloses Figure 2 (reproduced below) on P. 4 depicting quantized weights (low-bit weights) stored in contiguous memory. During online dequantization, quantized low-bit weights (‘R bits of the network weight value’) are retrieved to be converted into FP16, thereby retrieving R bits of the network weight value from computer memory. Online dequantization is used to perform LLM inference, therefore, quantized weights (after being dequantized) are retrieved to be multiplied with activations (matrix-matrix multiplications).
PNG
media_image1.png
344
458
media_image1.png
Greyscale
).
However, Lin does not teach during execution of a machine learning model, receiving an input vector for multiplication with a matrix of network weight values, which is taught by Park:
during execution of a machine learning model, receiving an input vector for multiplication with a matrix of network weight values ([0035] “a dynamic neural network quantization logic may be configured at runtime to change the quantization, masking, and/or neural network pruning based on operating conditions, such as temperature, power consumption, utilization of processing units, etc. of an AI processor, SoC having an AI processor, memory accessed by an AI processor, and/or other peripherals of an AI processor. Some embodiments may include configuring the dynamic neural network quantization logic for quantization of activation and weight values based on a number of dynamic bits for dynamic quantization.”
[0045] “The AI processor 124 may receive and store activation values at an activation buffer 206 and weight values at a weight buffer 204. Generally, the MAC array 200 may receive the activation values from the activation buffer 206 and the weight values from the weight buffer 204, and process the activation and weight values by multiplying and accumulating the activation and weight values.”
[0007] “dynamically adjusting the AI quantization level for the segment of the neural network may include adjusting the AI quantization level for quantizing weight values and activation values to be processed by the segment of the neural network.”
During runtime of a neural network, operating conditions can change, requiring changing how weights and activations are quantized, therefore affecting the multiplication result of processed (‘received’) weights and activations.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the quantization method of Lin with the technique disclosed by Park to perform dynamic quantization during model runtime. By performing dynamic quantization during model runtime, quantization of weights and activations can be adjusted based on operating conditions, thereby enabling more accurate weight and activation multiplications and reducing quantization error.
Regarding Claim 3, the combined quantization method of Lin / Park teaches:
The method of claim 1, further comprising, for a second network weight value of the matrix of network weight values, determining a second representation quantity (R’) of bits to be used for representing the second network weight value during multiplication with a second corresponding vector value of the input vector, based at least in part on a magnitude of the second corresponding vector value ((P. 2, Sec. 1, ¶1) “To avoid mixed precision, we follow the activation-awareness principle and design a per-channel scaling method to search for the optimal scaling factors that minimize the quantization error, and quantize all the weights”
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized. … The quantized performance generally gets better as the scaling factor increases since important weights are better represented”
See (P. 4, Sec. 2.3, ¶2) describing matrix-matrix (MM) multiplications.
Activations (‘input vector’) have magnitudes, therefore an activation (‘second corresponding vector value of the input vector’) has a magnitude. A weight channel (‘second network weight value’) with a smaller activation magnitude has a smaller scaling factor, therefore the weight channel (weight) is less precisely quantized (is non-salient) than salient weight channels (weight channels with larger activation magnitudes). The activation magnitude of the non-salient weight channel is therefore used to determine a scaling factor that determines how precise quantization is performed to convert a FP16 weight to a low-bit integer (to minimize quantization error), and therefore the quantized low-bit integer is a second representation quantity of bits that represents a second network weight value.).
Regarding Claim 7, the combined quantization method of Lin / Park teaches:
The method of claim 1, wherein R is further determined based at least in part on a difference between the corresponding vector value and other vector values of the input vector ((P. 4, Sec. 2.2, ¶2) “the optimal scale should be related to: 1. Activation magnitude: as mentioned above, we are relying on input activation magnitude to pick out the salient weight channels. Therefore, we should consider the magnitude of X to protect the salient weight channels”
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized. … The quantized performance generally gets better as the scaling factor increases since important weights are better represented”
An important (salient) weight is more precisely quantized if it has a larger activation magnitude (‘corresponding vector value’) and thus a scaling factor larger than 1. Weight channels (weights) that are non-salient have smaller activation magnitudes (‘other vector values of the input vector’), therefore the weight channels have a scaling factor less than 1, (and therefore those weights are less important and less precisely quantized). Salient and non-salient weights therefore have different activation magnitudes, and this difference between activation magnitudes (‘difference between corresponding vector values’) determines the scaling factor that determines how precisely weights are quantized (and therefore quantized weights (‘R representation quantity of bits’) is determined based on a difference between corresponding vector values (activation magnitudes)).).
Regarding Claim 8, the combined quantization method of Lin / Park teaches:
The method of claim 7, wherein R is relatively larger based at least in part on determining that the corresponding vector value is relatively higher than an average vector value of the input vector (The Examiner interprets “R is relatively larger” according to its BRI in view of the Applicant’s specification as encompassing “a larger activation magnitude”. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0015], (see excerpt below).
Applicant’s written description at [0015]: “R may be relatively higher than the R values for other network weight values in cases where the corresponding vector value is relatively higher, and the corresponding vector values for the other network weight values are relatively lower, which serves to reduce quantization error as described above”
(P. 4, Sec. 2.2, ¶2) “the optimal scale should be related to: 1. Activation magnitude: as mentioned above, we are relying on input activation magnitude to pick out the salient weight channels. Therefore, we should consider the magnitude of X to protect the salient weight channels. We compute the average activation magnitude
s
x
=
m
e
a
n
c
_
o
u
t
|
X
|
as the importance factor. … We determine the final scaling factors by a function considering both of the factors
s
=
f
(
s
x
,
s
w
)
.”
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized”
A scaling factor is determined by calculating an average activation magnitude (‘average vector value of the input vector’). A salient weight has a larger activation magnitude and scaling factor (larger than 1), therefore a salient weight has a ‘corresponding vector value’ (activation magnitude) that is higher than an average activation magnitude, and therefore R (a quantized salient weight) is relatively larger (higher) since non-salient weights have a lower activation magnitude (lower corresponding vector value).).
Regarding Claim 9, the combined quantization method of Lin / Park teaches:
The method of claim 1, wherein R is further determined based at least in part on a power consumption policy of the computing device ([0035] “a dynamic neural network quantization logic may be configured at runtime to change the quantization, masking, and/or neural network pruning based on operating conditions, such as temperature, power consumption”
[0057] “the dynamic quantization controller 208 may use the AI quantization level and/or operating conditions as inputs to an algorithm that may output a number of dynamic bits to use for quantization of activation and weight values”
[0027] “The term “dynamic bit(s)” is used herein to refer to bits of an activation value and/or a weight value for configuring the dynamic neural network quantization logics for quantization of activation and weight values”
Power consumption is used to determine a number of dynamic bits (representation quantity (R) of bits) to use for quantizing weights.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin / Park with the technique disclosed by Park to use power consumption to determine a number of bits to use for quantizing weights. By using power consumption to determine a number of bits to use for quantizing weights, weights can have their bit-width adjusted based on how much power is being consumed by a processor, thereby minimizing processor energy usage.
Regarding Claim 10, the combined quantization method of Lin / Park teaches:
The method of claim 1, wherein the matrix of network weight values are stored in the computer memory such that network weight values in a same row of the matrix are stored in a same memory row or memory page of the computer memory ((P. 5, Sec. 2.3, ¶1) “In Figure 2, we assume that the weights are stored in row major (i.e. IC×OC). Each OC has different scales and zero points and every two ICs within each OC share dequantization parameters (i.e. g = 2). … for AWQ (left), all memory accesses are contiguous, resulting in a 2× end-to-end speedup for LLaMA models.”
(P. 10, Sec. 4, Last Paragraph) “Our AWQ kernels are executed on tensor cores, suitable for both context and generation phases in LLM inference, and do not require hardware-inefficient reordering.”
(P. 2, Sec. 1, ¶2) “We employ reorder-free online dequantization to efficiently convert low-bit weights to FP16”
Lin discloses Figure 2 (reproduced above) on P. 4 depicting quantized weights stored in rows in contiguous memory. After performing quantization, quantized weights are not reordered in memory, so weights stored in each row remain stored in the same row before and after quantization. During online dequantization, quantized weights are retrieved to be converted into FP16, however, weights still remain in the same order, therefore network weight values (quantized weights) in a same row of a weight matrix are stored in a same memory row of computer memory.).
Regarding Claim 13, the combined quantization method of Lin / Park teaches:
The method of claim 1, wherein the machine learning model is a transformer-based language model ((P. 1, Sec. 1, ¶1-3) “Large language models (LLMs) based on transformers [48] have shown excellent performance on various benchmarks [5, 58, 47, 43]. However, the large model size leads to the high serving costs. … Low-bit weight quantization for LLMs can save memory but is hard. … In this paper, we propose Activation-aware Weight Quantization (AWQ), a hardware-friendly low-bit weight-only quantization method for LLMs.”).
Regarding Claim 14, the combined quantization method of Lin / Park teaches:
The method of claim 13, wherein the input vector represents tokens of a natural language input ((P. 7, Sec. 3.2, ¶2) “AWQ improves the responses compared to the round-to-nearest (RTN) baseline for INT4-g128 quantization, leading to more reasonable answers. … In the second example, AWQ correctly answers the question (the artist of the painting), while RTN does not provide any information about the artist.”
Lin discloses Figure 5 on P. 8 (reproduced below) depicting a visual reasoning example of a natural language question asked and an answer generated by using activation-aware quantization (AWQ). To generate an answer based on a natural language question (input), an input vector of activations must use tokens of the natural language question for an LLM to generate a response, therefore an input vector of activations represents tokens of a natural language input.
PNG
media_image2.png
188
897
media_image2.png
Greyscale
).
Regarding Claim 15, the combined quantization method of Lin / Park teaches:
The method of claim 13, wherein the input vector represents intermediate result vectors inside hidden layers of the transformer-based language model ((P. 8, Sec. 3.3) “We attribute this efficiency improvement to our formulation of linear layers in LLMs as matrix-matrix instead of matrix-vector product.”
(P. 3, Sec. 2.2, Last Paragraph) “We automatically search for an optimal (per input channel) scaling factor that minimizes the output difference after quantization for a certain layer.”
(P. 2, Sec. 1, ¶1) “weight channels corresponding to larger activation magnitudes are more salient since they process more important features”
Salient weight channels corresponding to larger activation magnitudes process more important features, therefore, matrix-matrix multiplications between weights and activations are performed to process the important features. By performing multiplication between weights and activations (‘input vector’), the activations represent intermediate result vectors inside hidden layers of a transformer-based language model (LLM) since matrix multiplication occurs inside hidden layers of an LLM in order to process important features to generate an output.).
Regarding Claim 16, the combined quantization method of Lin / Park teaches:
The method of claim 1, wherein the computer memory includes one or more of dynamic random-access memory (DRAM), flash memory, and static random-access memory (SRAM) ((P. 5, Sec. 2.3, Last Paragraph) “For GPTQ with reordering (right), since these g = 2 ICs are not continuous, irregular DRAM accesses are required to fetch scaling factors and zero points when dequantizing each weight. However, for AWQ (left), all memory accesses are contiguous, resulting in a 2× end-to-end speedup for LLaMA models.”).
Regarding Claim 17, the rejection of claim 1 is incorporated. The difference in scope being:
A computing system, comprising: a logic subsystem ((P. 10, Sec. 4, Last Paragraph) “Our AWQ kernels are executed on tensor cores”);
and a storage subsystem including computer memory, the storage subsystem holding instructions executable by the logic subsystem to (See Figure 2 on P. 4 depicting weights stored in computer memory (‘storage subsystem’). Stored executable instructions are implied by using tensor cores (‘logic subsystem’) to execute AWQ kernels.).
The following are the references relied upon in the rejections below:
Rajagopal, Aditya, et al. "Multi-precision policy enforced training (MuPPET): A precision-switching strategy for quantised fixed-point training of CNNs." International Conference on Machine Learning. PMLR, 2020.
Claims 2, 4-5 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Lin / Park / Rajagopal.
Regarding Claims 2 and 18, the combined quantization method of Lin / Park teaches: The method of claim 1, however the combination does not teach wherein a representation quantity (R) of bits is a smaller number of bits that S, such that at least some bits of a network weight value are stored in computer memory and not retrieved for multiplication with the corresponding vector value, which is taught by Rajagopal:
wherein R is a smaller number of bits than S, such that at least some bits of the network weight value are stored in the computer memory and not retrieved for multiplication with the corresponding vector value ((P. 3, Sec. 3.2, ¶1) “At epoch j, MuPPET performs mixed precision training where the weights are stored in an FP32 master copy and are quantised to the desired fixed-point precision (
q
j
) on-the-fly.”
(P. 1, Abstract) “This work pushes the boundary of quantised training by employing a multilevel optimisation approach that utilises multiple precisions including low-precision fixed-point representations resulting in a novel training strategy MuPPET”
(P. 4, Sec. 3.2.1, ¶2) “During the forward and backward passes of the training process, the weights and feature maps are both quantised, and the multiplication operations are performed at the same low precision.”
Weights are stored in a 32-bit floating-point representation (FP32) master copy, where FP32 is a stored quantity (S) of bits. Weights are quantized to be low-precision fixed-point representations (therefore the low-precision representations are a representation quantity (R) of bits that are a smaller number of bits than S (FP32)). A quantized weight and quantized feature map (‘corresponding vector value’) are multiplied together at the same low precision, therefore not all the bits of a FP32 weight (original weight before being quantized) is used for multiplication (and therefore at least some bits of a weight are stored in computer memory and not retrieved for multiplication).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin /Park with the quantization technique disclosed by Rajagopal to use quantized weights to perform multiplication operations. By using quantized weights to perform multiplication operations, quantized weights speed up computation by using low-precision numbers and less bits to perform multiplication, thereby reducing memory usage and allowing hardware to perform math calculations faster.
Regarding Claim 4, the combined quantization method of Lin / Park teaches:
The method of claim 3, wherein the corresponding vector value is smaller than the second corresponding vector value ((P. 4, Sec. 2.2, ¶2) “the optimal scale should be related to: 1. Activation magnitude: as mentioned above, we are relying on input activation magnitude to pick out the salient weight channels.
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized. … The quantized performance generally gets better as the scaling factor increases since important weights are better represented”
Weight channels have different activation magnitudes (‘corresponding vector values’). Salient weight channels have larger activation magnitudes and non-salient weight channels have smaller activation magnitudes (‘corresponding vector value is smaller’).),
However, the combination does not teach a representation quantity (R) of bits is a smaller number of bits than a second representation quantity (R’) of bits, which is taught by Rajagopal:
and wherein R is a smaller number of bits than R’ ((P. 1, Abstract) “As networks grow in size and complexity, training time can be reduced through low-precision data representations and computations, however, in doing so the final accuracy suffers due to the problem of vanishing gradients. Existing state-of-the-art methods combat this issue by means of a mixed-precision approach utilising two different precision levels, FP32 (32-bit floating-point) and FP16/FP8 (16-/8-bit floating-point), leveraging the hardware support of recent GPU architectures for FP16 operations to obtain performance gains. This work pushes the boundary of quantised training by employing a multilevel optimisation approach that utilizes multiple precisions including low-precision fixed-point representations resulting in a novel training strategy MuPPET;”
(P. 3, Sec. 3.2, ¶1) “At epoch j, MuPPET performs mixed precision training where the weights are stored in an FP32 master copy and are quantised to the desired fixed-point precision (
q
j
) on-the-fly.”
(P. 4, Sec. 3.1, ¶3) “Starting from the N-th problem, the inputs, weights and activations of the CNN model f are quantised with precision qN, which is the lowest precision in the system and represents the coarsest version of the model. Each of the N levels progressively employs higher precision until the first level is reached”
Weights are quantized from 32-bit floating-point representation (FP32) until reaching a desired fixed-point precision. Weights are quantized to be represented in a low-precision data representation and the weights are mixed-precision, therefore, some weights have a low-precision (representation quantity (R) of bits is smaller) and some weights have a higher fixed-point precision (second representation quantity (R’) of bits), and therefore R is a smaller number of bits than R’.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin / Park with the quantization technique disclosed by Rajagopal to perform mixed-precision quantization. By performing mixed-precision quantization, low-precision weights can be used to save memory and high-precision weights can be used to maintain accuracy when performing calculations, thereby creating a balance between memory efficiency and model accuracy.
Regarding Claim 5, the combined quantization method of Lin / Park / Rajagopal teaches:
The method of claim 4, wherein R’ is equal to S, such that all stored bits of the second network weight value are retrieved for multiplication with the second corresponding vector value ((P. 2, Sec. 1, ¶2) “We employ reorder-free online dequantization to efficiently convert low-bit weights to FP16”
(P. 5, Sec. 2.3, Last Paragraph) “Weight-only quantization methods require online weight dequantization.”
((P. 2, Sec. 1, ¶1) “weight channels corresponding to larger activation magnitudes are more salient since they process more important features”
During online dequantization, a quantized low-bit weight (second representation quantity R’ of bits) is retrieved to be converted into FP16, therefore the quantized low-bit weight (R’) has the same number of bits as FP16 (stored quantity S of bits), and therefore all stored bits of the quantized weight are retrieved for multiplication with an activation (‘second corresponding vector value’) to process important features.).
Regarding Claim 19, the combined quantization method of Lin / Park teaches:
The computing system of claim 17, wherein the instructions are further executable to, for a second network weight value of the matrix of network weight values, determine a second representation quantity (R’) of bits to be used for representing the second network weight value during multiplication with a second corresponding vector value of the input vector, based at least in part on a magnitude of the second corresponding vector value ((P. 2, Sec. 1, ¶1) “To avoid mixed precision, we follow the activation-awareness principle and design a per-channel scaling method to search for the optimal scaling factors that minimize the quantization error, and quantize all the weights”
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized. … The quantized performance generally gets better as the scaling factor increases since important weights are better represented”
See (P. 4, Sec. 2.3, ¶2) describing matrix-matrix (MM) multiplications.
Activations (‘input vector’) have magnitudes, therefore an activation (‘second corresponding vector value of the input vector’) has a magnitude. A weight channel (‘second network weight value’) with a smaller activation magnitude has a smaller scaling factor, therefore the weight channel (weight) is less precisely quantized (is non-salient) than salient weight channels (weight channels with larger activation magnitudes). The activation magnitude of the non-salient weight channel is therefore used to determine a scaling factor that determines how precise quantization is performed to convert a FP16 weight to a low-bit integer (to minimize quantization error), and therefore the quantized low-bit integer is a second representation quantity of bits that represents a second network weight value.),
wherein the corresponding vector value is smaller than the second corresponding vector value ((P. 4, Sec. 2.2, ¶2) “the optimal scale should be related to: 1. Activation magnitude: as mentioned above, we are relying on input activation magnitude to pick out the salient weight channels.
Weight channels have different activation magnitudes (‘corresponding vector values’). Salient weight channels have larger activation magnitudes and non-salient weight channels have smaller activation magnitudes (‘corresponding vector value is smaller’).),
However, the combination does not teach a representation quantity (R) of bits is a smaller number of bits than a second representation quantity (R’) of bits, which is taught by Rajagopal:
and wherein R is a smaller number of bits than R’ ((P. 1, Abstract) “As networks grow in size and complexity, training time can be reduced through low-precision data representations and computations, however, in doing so the final accuracy suffers due to the problem of vanishing gradients. Existing state-of-the-art methods combat this issue by means of a mixed-precision approach utilising two different precision levels, FP32 (32-bit floating-point) and FP16/FP8 (16-/8-bit floating-point), leveraging the hardware support of recent GPU architectures for FP16 operations to obtain performance gains. This work pushes the boundary of quantised training by employing a multilevel optimisation approach that utilizes multiple precisions including low-precision fixed-point representations resulting in a novel training strategy MuPPET;”
(P. 3, Sec. 3.2, ¶1) “At epoch j, MuPPET performs mixed precision training where the weights are stored in an FP32 master copy and are quantised to the desired fixed-point precision (
q
j
) on-the-fly.”
(P. 4, Sec. 3.1, ¶3) “Starting from the N-th problem, the inputs, weights and activations of the CNN model f are quantised with precision qN, which is the lowest precision in the system and represents the coarsest version of the model. Each of the N levels progressively employs higher precision until the first level is reached”
Weights are quantized from 32-bit floating-point representation (FP32) until reaching a desired fixed-point precision. Weights are quantized to be represented in a low-precision data representation and the weights are mixed-precision, therefore, some weights have a low-precision (representation quantity (R) of bits is smaller) and some weights have a higher fixed-point precision (second representation quantity (R’) of bits), and therefore R is a smaller number of bits than R’.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin / Park with the quantization technique disclosed by Rajagopal to perform mixed-precision quantization. By performing mixed-precision quantization, low-precision weights can be used to save memory and high-precision weights can be used to maintain accuracy when performing calculations, thereby creating a balance between memory efficiency and model accuracy.
Regarding Claim 20, Lin teaches:
A method for computer memory access, comprising ((P. 5, Sec. 2.3, ¶1) “for AWQ (left), all memory accesses are contiguous, resulting in a 2× end-to-end speedup for LLaMA models.”)):
… receiving an input vector for multiplication with a matrix of network weight values ((P. 1-2, Sec. 1, Last Paragraph) “To find the salient weight channels, the insight is that we should refer to the activation distribution instead of the weight distribution, despite we are doing weight-only quantization: weight channels corresponding to larger activation magnitudes are more salient since they process more important features”
(P. 4, Sec. 2.3, ¶2) “we instead model these layers as matrix-matrix (MM) multiplications.”
(P. 8, Sec. 3.3) “We attribute this efficiency improvement to our formulation of linear layers in LLMs as matrix-matrix instead of matrix-vector product.”
See Figure 2 on P. 4 depicting a matrix of weights. See (P. 4, Sec. 2.2, ¶1) describing activations are retrieved from a calibration set.
Linear layers of an LLM are modeled as matrix-matrix multiplications. Weight channels corresponding to larger activation magnitudes are more salient (important) since the weight channels process more important features. To process important features, weight channels corresponding to larger activation magnitudes must be multiplied together, therefore the matrix-matrix multiplications in the LLM multiply weights (‘matrix of network weight values’) with activations (‘input vector’).),
wherein each network weight value of the matrix of network weight values is stored in computer memory using a stored quantity (S) of bits ((P. 2, Sec. 2, ¶1) “Low-bit weight-only quantization [17, 39], which quantizes weights into low-bit integers (typically ≤4 bits). The decoding stage dominates LLMs’ total runtime [17, 32], which is highly memory-bounded with a single batch size. Given that the memory is dominated by weights, we focus on weight-only quantization”
(P. 3-4, Sec. 2.2, ¶3) “We automatically search for an optimal (per input channel) scaling factor that minimizes the output difference after quantization for a certain layer. … W is the original weights in FP16”
Original weights W (‘matrix of network weight values’) are stored in memory in FP16 (Half-precision floating-point format, 16 bits) before being quantized into low-bit integers. Therefore, each weight (‘network weight value’) of matrix W is stored in computer memory using a stored quantity S of bits (S being 16 bits).);
for a network weight value of the matrix of network weight values, determining a representation quantity (R) of bits to be used for representing the network weight value during multiplication with a first corresponding vector value of the input vector, based at least in part on a magnitude of the first corresponding vector value (The Examiner interprets “representation quantity (R) of bits” according to its BRI in view of the Applicant’s specification as encompassing “quantized weights”. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0015].
(P. 2, Sec. 2, ¶1) “Low-bit weight-only quantization [17, 39], which quantizes weights into low-bit integers (typically ≤4 bits). … Given that the memory is dominated by weights, we focus on weight-only quantization.”
(P. 2, Sec. 1, ¶1) “To avoid mixed precision, we follow the activation-awareness principle and design a per-channel scaling method to search for the optimal scaling factors that minimize the quantization error, and quantize all the weights”
(P. 3, Sec. 2.2, ¶2) “A straightforward way to protect the outlier weight channels is to multiply the channel by a certain scaling ratio (>1) such that they can be precisely quantized. … The quantized performance generally gets better as the scaling factor increases since important weights are better represented”
(P. 4, Sec. 2.2, ¶2) “the optimal scale should be related to: 1. Activation magnitude: as mentioned above, we are relying on input activation magnitude to pick out the salient weight channels. Therefore, we should consider the magnitude of X to protect the salient weight channels. We compute the average activation magnitude
s
x
=
m
e
a
n
c
_
o
u
t
|
X
|
as the importance factor. 2. Weight magnitude: to minimize the quantization loss of non-salient weights, we should flatten their distribution so that it is easier for quantization … We determine the final scaling factors by a function considering both of the factors
s
=
f
(
s
x
,
s
w
)
.”
Activations (‘input vector’) have magnitudes, therefore an activation (‘first corresponding vector value of the input vector’) has a magnitude. Activation magnitudes are used to determine salient (important) weight channels (weights). Weight channels with larger activation magnitudes have a larger scaling factor, therefore these weights are more precisely quantized than non-salient weight channels (weight channels with smaller activation magnitudes). Activation magnitudes are therefore used to determine a scaling factor that determines how precise quantization is performed to convert FP16 weights to low-bit integers (to minimize quantization error), and therefore the quantized low-bit integers are a representation quantity (R) of bits that represent network weight values (weights). The quantized low-bit integers are 3 or 4 bits (see Page 1, Abstract) and are a different quantity of bits than original weights stored in FP16 (stored quantity S). See (P. 4, Sec. 2.3, ¶2) describing matrix-matrix (MM) multiplications.),
wherein R is a smaller number of bits than S (Quantized low-bit integers (R) are 3 or 4 bits, which is less than FP16 (stored quantity S).);
retrieving R bits of the network weight value from the computer memory for multiplication with the first corresponding vector value ((P. 5, Sec. 2.3, ¶1) “In Figure 2, we assume that the weights are stored in row major (i.e. IC×OC). Each OC has different scales and zero points and every two ICs within each OC share dequantization parameters (i.e. g = 2). For GPTQ with reordering (right), since these g = 2 ICs are not continuous, irregular DRAM accesses are required to fetch scaling factors and zero points when dequantizing each weight. However, for AWQ (left), all memory accesses are contiguous, resulting in a 2× end-to-end speedup for LLaMA models.”
(P. 10, Sec. 4, Last Paragraph) “Our AWQ kernels are executed on tensor cores, suitable for both context and generation phases in LLM inference, and do not require hardware-inefficient reordering.”
(P. 2, Sec. 1, ¶2) “We employ reorder-free online dequantization to efficiently convert low-bit weights to FP16”
Lin discloses Figure 2 (reproduced above) on P. 4 depicting quantized weights (low-bit weights) stored in contiguous memory. During online dequantization, quantized low-bit weights (‘R bits of the network weight value’) are retrieved to be converted into FP16, thereby retrieving R bits of the network weight value from computer memory. Online dequantization is used to perform LLM inference, therefore, quantized weights (after being dequantized) are retrieved to be multiplied with an activation (‘first corresponding vector value’) in matrix-matrix multiplication.),
and retrieving S bits of a second network weight value from the computer memory for multiplication with a second corresponding vector value of the input vector ((P. 1-2, Sec. 1, Last Paragraph) “To find the salient weight channels, the insight is that we should refer to the activation distribution instead of the weight distribution, despite we are doing weight-only quantization: weight channels corresponding to larger activation magnitudes are more salient since they process more important features”
When a quantized low-bit weight is dequantized into FP16, an FP16 weight (stored quantity S bits) is used to perform matrix-matrix multiplication. When performing matrix-matrix multiplication, the FP16 weight is multiplied by an activation magnitude to process important features. These multiplications are performed for each weight and activation magnitude (since there is a distribution of activation magnitudes, the ‘input vector’ being a distribution of activation magnitudes and a ‘second corresponding vector value’ being an activation magnitude from the distribution).),
the second corresponding vector value being larger than the first corresponding vector value (In a distribution of activation magnitudes, weight channels can have larger activation magnitudes (the weight channels are more salient), or weight channels can have smaller activation magnitudes (the weights channels are non-salient). Therefore, a larger activation magnitude (‘second corresponding vector value’) can be larger than a first corresponding vector value (smaller activation magnitude).).
However, Lin does not teach during execution of a machine learning model, receiving an input vector for multiplication with a matrix of network weight values, which is taught by Park:
during execution of a machine learning model, receiving an input vector for multiplication with a matrix of network weight values ([0035] “a dynamic neural network quantization logic may be configured at runtime to change the quantization, masking, and/or neural network pruning based on operating conditions, such as temperature, power consumption, utilization of processing units, etc. of an AI processor, SoC having an AI processor, memory accessed by an AI processor, and/or other peripherals of an AI processor. Some embodiments may include configuring the dynamic neural network quantization logic for quantization of activation and weight values based on a number of dynamic bits for dynamic quantization.”
[0045] “The AI processor 124 may receive and store activation values at an activation buffer 206 and weight values at a weight buffer 204. Generally, the MAC array 200 may receive the activation values from the activation buffer 206 and the weight values from the weight buffer 204, and process the activation and weight values by multiplying and accumulating the activation and weight values.”
[0007] “dynamically adjusting the AI quantization level for the segment of the neural network may include adjusting the AI quantization level for quantizing weight values and activation values to be processed by the segment of the neural network.”
During runtime of a neural network, operating conditions can change, requiring changing how weights and activations are quantized, therefore affecting the multiplication result of processed (‘received’) weights and activations.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the quantization method of Lin with the technique disclosed by Park to perform dynamic quantization during model runtime. By performing dynamic quantization during model runtime, quantization of weights and activations can be adjusted based on operating conditions, thereby enabling more accurate weight and activation multiplications and reducing quantization error.
Furthermore, the combined quantization method of Lin / Park does not teach retrieving R bits of a network weight value such that at least some bits of the network weight value are stored in the computer memory and not retrieved for multiplication with the first corresponding vector value, which is taught by Rajagopal:
retrieving R bits of the network weight value from the computer memory for multiplication with the first corresponding vector value such that at least some bits of the network weight value are stored in the computer memory and not retrieved for multiplication with the first corresponding vector value ((P. 3, Sec. 3.2, ¶1) “At epoch j, MuPPET performs mixed precision training where the weights are stored in an FP32 master copy and are quantised to the desired fixed-point precision (
q
j
) on-the-fly.”
(P. 1, Abstract) “This work pushes the boundary of quantised training by employing a multilevel optimisation approach that utilises multiple precisions including low-precision fixed-point representations resulting in a novel training strategy MuPPET”
(P. 4, Sec. 3.2.1, ¶2) “During the forward and backward passes of the training process, the weights and feature maps are both quantised, and the multiplication operations are performed at the same low precision.”
Weights are stored in a 32-bit floating-point representation (FP32) master copy, where FP32 is a stored quantity of bits. Weights are quantized to be low-precision fixed-point representations (therefore the low-precision representations are a representation quantity (R) of bits that are a smaller number of bits than FP32). A quantized weight and quantized feature map (‘first corresponding vector value’) are multiplied together at the same low precision, therefore not all the bits of a FP32 weight (original weight before being quantized) is used for multiplication (and therefore at least some bits of a weight are stored in computer memory and not retrieved for multiplication).);
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin /Park with the quantization technique disclosed by Rajagopal to use quantized weights to perform multiplication operations. By using quantized weights to perform multiplication operations, quantized weights speed up computation by using low-precision numbers and less bits to perform multiplication, thereby reducing memory usage and allowing hardware to perform math calculations faster.
The following are the references relied upon in the rejections below:
Kundu (US 20230252299 A1)
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Lin / Park / Kundu.
Regarding Claim 6, the combined quantization method of Lin / Park teaches: The method of claim 1, however the combination does not teach wherein the corresponding vector value of the input vector is equal to zero and R is equal to zero, which is taught by Kundu:
wherein the corresponding vector value of the input vector is equal to zero, and R is equal to zero ([0019] “Input tensors or weight tensors can include zero valued elements that do not impact the output of the dot product. DNN accelerators can exploit the sparsity in such input tensors or weigh tensors to accelerate deep learning operations in DNNs, which can lead to higher speedup or throughput as well as less energy consumption. For instance, a DNN accelerator may include a sparsity acceleration unit that can be used to skip the processing of the zero valued activations or zero valued weights by the PEs.”
[0022] “The products from the sequence of multiplications may be accumulated to generate a single data point. zero valued activations or zero valued weight would not contribute to the result of the accumulation as the product of the activation or weight with any number would be zero.”
[0074] “The local memory 410 may store activation operands and weight operands in a compressed format so that nonzero valued activations and nonzero valued weights are stored but zero valued activations and zero valued weights are not stored”
When performing multiplications between weights and activations (‘input vector’), a zero valued activation (‘vector value of the input vector is equal to zero’) multiplied with any weight would result in zero. Therefore, when an activation is zero, multiplying a zero valued activation and weight is skipped, therefore a weight is not represented since it is skipped (and therefore a representation quantity (R) of bits for the weight is 0).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin / Park with the technique disclosed by Kundu to reduce weight representation when an activation value is zero. By reducing weight representation when an activation value is zero, less energy can be consumed by not having to perform multiplication with zero valued activations, thereby accelerating machine learning operations.
The following are the references relied upon in the rejections below:
Imani, Mohsen, et al. "Floatpim: In-memory acceleration of deep neural network training with high precision." Proceedings of the 46th international symposium on computer architecture. 2019.
Claims 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Lin / Park / Imani.
Regarding Claim 11, the combined quantization method of Lin / Park teaches: The method of claim 10, however, the combination does not teach organizing a memory row or memory page beginning with most significant bits, which is taught by Imani:
wherein the same memory row or memory page is organized beginning with most significant bits for each network weight value, and ending with least significant bits for each network weight value ((P. 804, Sec. 2.2, ¶2-3) “Digital processing in-memory achieves maximum performance when the operands are present in the same row because, in this configuration, all the bits of an operand are accessible by all the bits of the other operand. This increases the flexibility in implementing operations in memory. … this PIM architecture can provide significant speedup with large parallelism. PIM can support addition and multiplications in parallel, irrespective of the number of rows. For example, to add values stored in different columns of memory, it takes the same amount of time for PIM to process the addition in a single row or all memory rows.”
(P. 806, Sec. 4.1, ¶2) “It writes all convolution weights in a single row and then copies them in other rows using the row-parallel write operation that happens just in two cycles. This method enables the input values to be multiplied with any convolution weights stored in another column”
(P. 809, Sec. 6) “This work represents the very first implementation of floating point addition and multiplication in the crossbar memory. A floating point number consists of a binary number string with three different parts: a sign bit, an exponent part, and a fractional value. For example, the IEEE 754 32-bit floating point notation consists of a sign bit, eight exponent bits, and 23 fractional bits. The first bit in the floating point notation (A32) represents the sign bit, where ‘0’ represents a positive number. The next eight bits represent the exponent of the binary numbers (A31, . . . ,A24), ranging from -126 to 127. The following 23 bits (A23, . . . ,A1) represent the fractional part, also known as mantissa, which has a value between 1 and 2.”
Convolution weights (‘each network weight value’) are stored in a single row. Weights are floating-point numbers consisting of a sign bit, exponent, and mantissa. The sign bit is the leftmost bit and represents whether a floating-point number is positive or negative, therefore a sign bit is the most significant bit for each convolution weight. The mantissa represents a fractional value and the last bits of the mantissa are the rightmost bits, therefore the last bits of the mantissa are the least significant bits for each convolution weight. Therefore, a memory row is organized from most significant bits (sign bit) to least significant bits (last bits of mantissa) for each convolution weight.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin / Park with the technique disclosed by Imani to store weights in the same memory row. By storing weights in the same memory row, weights can be easily accessible by another operand when performing multiplication, thereby enabling parallel access and speeding up calculations.
Regarding Claim 12, the combined quantization method of Lin / Park teaches: The method of claim 10, however, the combination does not teach organizing a memory row or memory page beginning with sign bits, followed by exponent bits, and ending with mantissa bits for each network value, which is taught by Imani:
wherein the same memory row or memory page is organized beginning with sign bits, followed by exponent bits, and ending with mantissa bits for each network value stored in the same memory row or memory page ((P. 806, Sec. 4.1, ¶2) “It writes all convolution weights in a single row and then copies them in other rows using the row-parallel write operation that happens just in two cycles. This method enables the input values to be multiplied with any convolution weights stored in another column”
(P. 804, Sec. 2.2, ¶2) “Digital processing in-memory achieves maximum performance when the operands are present in the same row because, in this configuration, all the bits of an operand are accessible by all the bits of the other operand”
(P. 809, Sec. 6) “This work represents the very first implementation of floating point addition and multiplication in the crossbar memory. A floating point number consists of a binary number string with three different parts: a sign bit, an exponent part, and a fractional value. For example, the IEEE 754 32-bit floating point notation consists of a sign bit, eight exponent bits, and 23 fractional bits. The first bit in the floating point notation (A32) represents the sign bit, where ‘0’ represents a positive number. The next eight bits represent the exponent of the binary numbers (A31, . . . ,A24), ranging from -126 to 127. The following 23 bits (A23, . . . ,A1) represent the fractional part, also known as mantissa, which has a value between 1 and 2.”
Convolution weights (‘each network value’) are stored in a single row in memory. Weights are floating-point numbers consisting of a sign bit, exponent, and mantissa. The sign bit is followed by exponent bits and ends with mantissa bits for each convolution weight, therefore a memory row is organized beginning with sign bits, following by exponent bits, and ending with mantissa bits for each weight in the row.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined quantization method of Lin / Park with the technique disclosed by Imani to store weights in the same memory row. By storing weights in the same memory row, weights can have the same bit order and be easily accessible by another operand when performing multiplication, thereby enabling parallel access and speeding up calculations.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Pareek et al. (US 20200089472 A1) teaches quantizing weights and activations to perform multiply and accumulate operations by generating shared exponents to reduce the number of bits used to represent each operand.
Lee et al. (US 20210303972 A1) teaches determining if layers of a neural network need to be quantized by using profile information (statistics related to weights and activations) and then quantizing layers by using a scaling factor.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PEDRO J MORALES/Examiner, Art Unit 2124
/MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124