Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
2. This action is in response to the original filing on 02/26/2024. Claims 1-20 are pending and have been considered below.
Information Disclosure Statement
3. The information disclosure statement (IDS(s)) submitted on 02/26/2024, 07/16/2025 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
4. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
5. Claims 1-5, 7, 9, 12, 15, and 17-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Wang et al. (U.S. Patent Application Pub. No. US 20190050710 A1).
Claim 1: Wang teaches a system comprising:
at least one processor (i.e. an electronic device includes one or more processors; para. [0006]); and
at least one memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising (i.e. memory storing one or more programs; the one or more programs are configured to be executed by the one or more processors and the one or more programs include instructions for performing or causing performance of the operations of any of the methods described herein; para. [0006]):
identifying a set of candidate quantization configurations (i.e. different configurations of bit-width and layer combinations are prepared as candidates for evaluation. In some embodiments, the quantization candidates are generated by using different combinations of bit-widths for all the layers; instead, the weights from different layers are clustered based on their values, and the quantization candidates are generated by using different combinations of bit-widths for all the clusters; para. [0062]) for a component of a trained machine learning model (i.e. the base model is either the full-precision trained model 106′ or the pruned slender full precision model 106″, with their respective sets of weights Wi (or Wri) and bias bi (or bri) for each layer i of the full-precision model … the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers, wherein each layer of the plurality of layers (e.g., one or more convolution layers, a pooling layer, an activation layer, etc.) has a respective set of parameters (e.g., a set of weights for coupling the layer and its next layer, a set of network bias parameters for the layer, etc.); para. [0067, 0071]);
for each candidate quantization configuration in the set of candidate quantization configurations: applying the candidate quantization configuration to the component (i.e. For qb1 = 1 to 8: // quantization bit-width candidates for weights For qb2 = 1 to 8: // quantization bit-width for layer response W.sub.q,=Q.sub.Non(W.sub.i, qb1), b.sub.q,i=Q.sub.Non(b.sub.i, qb1) Stat.sub.q,=Q.sub.Non(W.sub.q,i×x.sub.i+b.sub.i, qb2) … expressing the respective set of parameters of the first layer with the first reduced bit-width includes performing non-uniform quantization (e.g., logarithmic quantization Qnon( . . . )) on the respective set of parameters of the first layer to generate a first set of quantized parameters for the first layer; para. [0067, 0076]),
after applying the candidate quantization configuration, executing the trained machine learning model on an input dataset (i.e. For each candidate model, the validation data set is used as input in a forward pass through the candidate model; para. [0062, 0065]) to obtain candidate output values for the candidate quantization configuration (i.e. For each candidate model, the validation data set is used as input in a forward pass through the candidate model, and the statistical distribution of the response values in each layer is collected; para. [0062, 0067, 0075]), and
determining a loss associated with the candidate quantization configuration based on a comparison between the candidate output values and reference output values for the component (i.e. the computing device collects a respective baseline statistical distribution of activation values for the first layer (e.g., statistics_i) as the validation data set is forward propagated as input through the first neural network model, while the respective sets of parameters of the plurality of layers are expressed with the original bit-width (e.g., 32-bit) of the first neural network model; the computing device collects a respective modified statistical distribution of activation values for the first layer (e.g., statq,i) as the validation data set is forward propagated as input through the first neural network model, while the respective set of parameters of the first layer are expressed with a first reduced bit-width (e.g., Wq,i and bq,i) that are smaller than the original bit-width of the first neural network model; determining a predefined divergence (e.g., Inf_tmp=InformationLoss (statq,I , statistics_i)) between the respective modified statistical distribution of activation values for the first layer and the respective baseline statistical distribution of activation values for the first layer; para. [0075]);
selecting, for the component, a candidate quantization configuration from the set of candidate quantization configurations based on the determined losses associated with the set of candidate quantization configurations (i.e. the computing device identifies a minimum value of the first reduced bit-width for which a reduction in the predefined divergence due to a further reduction of bit-width for the first layer is below a predefined threshold; para. [0067, 0075]); and
quantizing at least part of the trained machine learning model using the selected candidate quantization configuration for the component (i.e. The result of the above process is the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i, and the set of quantized bias bopt,i with the optimal quantization bit-width(s) for each layer i. The adaptive bit-width model 112 is thus obtained. In addition, the optimal quantization bit-width for the layer response qb2 is also obtained; para. [0067, 0068, 0073]).
Claim 2: Wang teaches the system of claim 1. Wang further teaches wherein the applying of the candidate quantization configuration to the component causes the trained machine learning model to be executed with the component having quantized parameters (i.e. a calibration data set is used as input in a forward propagation pass through the different candidate models (e.g., with different bit-width combinations for the parameters (e.g., weights and bias) of the different layers, and for the intermediate results (e.g., layer responses)); para. [0065, 0067, 0075]), and the reference output values are obtained by executing the trained machine learning model with the component having unquantized parameters (i.e. the computing device collects a respective baseline statistical distribution of activation values for the first layer (e.g., statistics_i) as the validation data set is forward propagated as input through the first neural network model, while the respective sets of parameters of the plurality of layers are expressed with the original bit-width (e.g., 32-bit) of the first neural network model; para. [0075]).
Claim 3: Wang teaches the system of claim 1. Wang further teaches wherein the component is one of a plurality of components of the trained machine learning model, the selected candidate quantization configuration is a first quantization configuration, and the trained machine learning model is quantized using the first quantization configuration for the component and one or more further quantization configurations selected for one or more further components of the plurality of components (i.e. different configurations of bit-width and layer combinations are prepared as candidates for evaluation. The quantization candidates are generated by using different combinations of bit-widths for all the layers; para. [0062, 0067, 0074]).
Claim 4: Wang teaches the system of claim 3. Wang further teaches wherein the trained machine learning model comprises a neural network, the plurality of components comprises a plurality of layers of the neural network (i.e. the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers; para. [0071]), and, for each candidate quantization configuration, the candidate output values are candidate feature map values (i.e. Each layer i of the L-layer neural network is described as y.sub.p×q=g(W.sub.p*q×x.sub.q×1+b.sub.p×1), where y is a layer response vector, W is weight matrix, x is an input vector, g(x)=max(0,x) is an activation function and b is a bias vector. (Note that p and q are the dimensions of the layer and will be different in different layers) … the statistical distribution of the response values in each layer; para. [0049, 0062, 0067, 0075]), the reference output values being reference feature map values for one of the plurality of layers (i.e. the computing device collects a respective baseline statistical distribution of activation values for the first layer (e.g., statistics_i) as the validation data set is forward propagated as input through the first neural network model, while the respective sets of parameters of the plurality of layers are expressed with the original bit-width (e.g., 32-bit) of the first neural network model; para. [0064, 0075]).
Claim 5: Wang teaches the system of claim 4. Wang further teaches wherein the first quantization configuration comprises first bit settings for quantizing parameters of a first layer of the plurality of layers (i.e. The result of the above process is the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i, and the set of quantized bias bopt,i with the optimal quantization bit-width(s) for each layer i. The adaptive bit-width model 112 is thus obtained. In addition, the optimal quantization bit-width for the layer response qb2 is also obtained; para. [0067, 0073]), and the one or more further quantization configurations comprise second bit settings for quantizing parameters of one or more further layers of the neural network, the first bit settings being different than the second bit settings (i.e. different configurations of bit-width and layer combinations are prepared as candidates for evaluation … a first layer (e.g., i=2) of the plurality of layers in the reduced neural network model (e.g., model 112) has a first reduced bit-width (e.g., 4-bit) that is smaller than the original bit-width (e.g., 32-bit) of the first neural network model, a second layer (e.g., i=3) of the plurality of layers in the reduced neural network model (e.g., model 112) has a second reduced bit-width (e.g., 6-bit) that is smaller than the original bit-width of the first neural network model, and the first reduced bit-width is distinct from the second reduced bit-width in the reduced neural network model; para. [0062, 0067, 0074]).
Claim 7: Wang teaches the system of claim 1. Wang further teaches wherein the set of candidate quantization configurations is a first set (i.e. different configurations of bit-width and layer combinations are prepared as candidates for evaluation. In some embodiments, the quantization candidates are generated by using different combinations of bit-widths for all the layers; instead, the weights from different layers are clustered based on their values, and the quantization candidates are generated by using different combinations of bit-widths for all the clusters; para. [0062]), the component of the trained machine learning model is a first component (i.e. the base model is either the full-precision trained model 106′ or the pruned slender full precision model 106″, with their respective sets of weights Wi (or Wri) and bias bi (or bri) for each layer i of the full-precision model … the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers, wherein each layer of the plurality of layers (e.g., one or more convolution layers, a pooling layer, an activation layer, etc.) has a respective set of parameters (e.g., a set of weights for coupling the layer and its next layer, a set of network bias parameters for the layer, etc.); para. [0067, 0071]), the candidate output values are first candidate output values (i.e. For each candidate model, the validation data set is used as input in a forward pass through the candidate model, and the statistical distribution of the response values in each layer is collected; para. [0062, 0067, 0075]), the reference output values are first reference output values (i.e. the computing device collects a respective baseline statistical distribution of activation values for the first layer (e.g., statistics_i) as the validation data set is forward propagated as input through the first neural network model, while the respective sets of parameters of the plurality of layers are expressed with the original bit-width (e.g., 32-bit) of the first neural network model; the computing device collects a respective modified statistical distribution of activation values for the first layer (e.g., statq,i) as the validation data set is forward propagated as input through the first neural network model, while the respective set of parameters of the first layer are expressed with a first reduced bit-width (e.g., Wq,i and bq,i) that are smaller than the original bit-width of the first neural network model; determining a predefined divergence (e.g., Inf_tmp=InformationLoss (statq,I , statistics_i)) between the respective modified statistical distribution of activation values for the first layer and the respective baseline statistical distribution of activation values for the first layer; para. [0075]), and the selected candidate quantization configuration is a first candidate quantization configuration (i.e. the computing device identifies a minimum value of the first reduced bit-width for which a reduction in the predefined divergence due to a further reduction of bit-width for the first layer is below a predefined threshold; para. [0067, 0075]), the operations further comprising:
identifying a second set of candidate quantization configurations for a second component (i.e. different configurations of bit-width and layer combinations are prepared as candidates for evaluation. In some embodiments, the quantization candidates are generated by using different combinations of bit-widths for all the layers; instead, the weights from different layers are clustered based on their values, and the quantization candidates are generated by using different combinations of bit-widths for all the clusters; para. [0062, 0067]) of the trained machine learning model (i.e. the base model is either the full-precision trained model 106′ or the pruned slender full precision model 106″, with their respective sets of weights Wi (or Wri) and bias bi (or bri) for each layer i of the full-precision model … the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers, wherein each layer of the plurality of layers (e.g., one or more convolution layers, a pooling layer, an activation layer, etc.) has a respective set of parameters (e.g., a set of weights for coupling the layer and its next layer, a set of network bias parameters for the layer, etc.); para. [0067, 0071]);
for each candidate quantization configuration in the second set of candidate quantization configurations (i.e. For qb1 = 1 to 8: // quantization bit-width candidates for weights For qb2 = 1 to 8: // quantization bit-width for layer response W.sub.q,=Q.sub.Non(W.sub.i, qb1), b.sub.q,i=Q.sub.Non(b.sub.i, qb1) Stat.sub.q,=Q.sub.Non(W.sub.q,i×x.sub.i+b.sub.i, qb2) … expressing the respective set of parameters of the first layer with the first reduced bit-width includes performing non-uniform quantization (e.g., logarithmic quantization Qnon( . . . )) on the respective set of parameters of the first layer to generate a first set of quantized parameters for the first layer; para. [0067, 0076]), determining the loss associated with the candidate quantization configuration based on a comparison between second candidate output values and second reference output values for the second component (i.e. the computing device collects a respective baseline statistical distribution of activation values for the first layer (e.g., statistics_i) as the validation data set is forward propagated as input through the first neural network model, while the respective sets of parameters of the plurality of layers are expressed with the original bit-width (e.g., 32-bit) of the first neural network model; the computing device collects a respective modified statistical distribution of activation values for the first layer (e.g., statq,i) as the validation data set is forward propagated as input through the first neural network model, while the respective set of parameters of the first layer are expressed with a first reduced bit-width (e.g., Wq,i and bq,i) that are smaller than the original bit-width of the first neural network model; determining a predefined divergence (e.g., Inf_tmp=InformationLoss (statq,I , statistics_i)) between the respective modified statistical distribution of activation values for the first layer and the respective baseline statistical distribution of activation values for the first layer; para. [0075]), the second candidate output values obtained by applying the candidate quantization configuration to the second component prior to execution of the trained machine learning model (i.e. For each candidate model, the validation data set is used as input in a forward pass through the candidate model, and the statistical distribution of the response values in each layer is collected; para. [0062, 0067, 0075]); and
selecting, for the second component, a second candidate quantization configuration from the second set of candidate quantization configurations based on the determined losses associated with the second set of candidate quantization configurations (i.e. the computing device identifies a minimum value of the first reduced bit-width for which a reduction in the predefined divergence due to a further reduction of bit-width for the first layer is below a predefined threshold; para. [0067, 0075]),
wherein the trained machine learning model is quantized using the first candidate quantization configuration for the first component and the second candidate quantization configuration for the second component (i.e. The result of the above process is the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i, and the set of quantized bias bopt,i with the optimal quantization bit-width(s) for each layer i. The adaptive bit-width model 112 is thus obtained. In addition, the optimal quantization bit-width for the layer response qb2 is also obtained; para. [0067, 0068, 0073]).
Claim 9: Wang teaches the system of claim 1. Wang further teaches wherein each candidate quantization configuration in the set of candidate quantization configurations comprises bit settings for quantizing parameters of the component of the trained machine learning model (i.e. The adaptive bit-width of the model refers to the characteristic that the respective bit-width for storing the set of parameters (e.g., weights and bias) for each layer of the model is specifically selected for that set of parameters (e.g., in accordance with the distribution and range of the parameters). Specifically, the validation data set is used as input in a forward pass through the pruned slender full-precision network (e.g., model 106″), and the statistical distribution of the response values in each layer is collected. Then different configurations of bit-width and layer combinations are prepared as candidates for evaluation; para. [0062, 0065, 0066, 0067]).
Claim 12: Wang teaches the system of claim 1. Wang further teaches wherein the loss is determined based on a loss function (i.e. the predefined measure of information loss is the Jensen-Shannon Divergence that measures the difference between two statistical distributions; para. [0064-0067]), and the selecting of the candidate quantization configuration from the set of candidate quantization configurations comprises: detecting that the selected candidate quantization configuration results in a lowest value for the loss function with respect to the component of the trained machine learning model (i.e. If Inf_tmp < Inf_min: W.sub.opt,i=W.sub.q,i, b.sub.opt,i=b.sub.q,i qb.sub.opt,=qb2 Inf_min = Inf_tmp The result of the above process is the set of quantized weights W.sub.opt,i with the optimal quantization bit-width(s) for each layer i, and the set of quantized bias b.sub.opt,i with the optimal quantization bit-width(s) for each layer i. The adaptive bit-width model 112 is thus obtained. In addition, the optimal quantization bit-width for the layer response qb2 is also obtained; para. [0067]).
Claim 15: Wang teaches the system of claim 1. Wang further teaches comprising: generating output comprising the selected candidate quantization configuration; and causing the output to be transmitted to a user device (i.e. the model deployment system 104 receives a reduced model as generated in accordance with the techniques described herein from the model generation system 102 over the network or through other file or data transmission means; para. [0017]).
Claim 17: Wang teaches the system of claim 1. Wang further teaches wherein the input dataset comprises unlabeled sample data (i.e. the training includes supervised training, unsupervised training, or semi-supervised training … the candidate is evaluated based on the amount of information loss that has resulted from the quantization applied to the candidate model. In some embodiments, Jensen-Shannon divergence between the two statistical distributions for each layer (or for the model as a whole) is used to identify the optimal bit-widths with the least information loss for that layer; para. [0018, 0062, 0067]).
Claim 18: Wang teaches the system of claim 1. Wang further teaches wherein the input dataset comprises unlabeled (i.e. the training includes supervised training, unsupervised training, or semi-supervised training … the candidate is evaluated based on the amount of information loss that has resulted from the quantization applied to the candidate model. In some embodiments, Jensen-Shannon divergence between the two statistical distributions for each layer (or for the model as a whole) is used to identify the optimal bit-widths with the least information loss for that layer; para. [0018, 0062, 0067]) sample images (i.e. if the model 112 is trained for a computer vision task, the real-world input data may be an image or a set of image features; para. [0024]).
Claims 19 and 20 are similar in scope to Claim 1 and are rejected under a similar rationale.
Claim Rejections – 35 USC § 103
6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
7. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Lee et al. (U.S. Patent Application Pub. No. US 20190042948 A1).
Claim 6: Wang teaches the system of claim 3. Wang further teaches wherein the trained machine learning model comprises a neural network (i.e. the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers; para. [0071]).
Wang does not explicitly teach the plurality of components comprises a plurality of channels within a layer of the neural network.
However, Lee teaches wherein the trained machine learning model comprises a neural network (i.e. pre-trained floating-point neural network; para. [0005]), and the plurality of components comprises a plurality of channels within a layer of the neural network (i.e. fig. 8, determine fractional lengths of a bias and a weight for each channel among the parameters for each channel based on a result of performing a convolution operation, and generate a fixed-point quantized neural network in which the bias and the weight for each channel have the determined fractional lengths; para. [0016, 0102]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Lee. One would have been motivated to make this modification because it improves model accuracy at reduced bit widths.
8. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Donnelly et al. (U.S. Patent Application Pub. No. US 20220044109 A1).
Claim 8: Wang teaches the system of claim 1. Wang further teaches wherein the component comprises a layer of a neural network (i.e. the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers; para. [0071]), the layer is associated with a threshold function (i.e. The activation function can be linear, rectified linear unit, sigmoid, hyperbolic tangent, or other types … Each layer i of the L-layer neural network is described as y.sub.p×q=g(W.sub.p*q×x.sub.q×1+b.sub.p×1), where y is a layer response vector, W is weight matrix, x is an input vector, g(x)=max(0,x) is an activation function and b is a bias vector. (Note that p and q are the dimensions of the layer and will be different in different layers); para. [0045, 0049]), and the threshold function is applied to obtain at least and at least a subset of the reference output values (i.e. Each layer i of the L-layer neural network is described as y.sub.p×q=g(W.sub.p*q×x.sub.q×1+b.sub.p×1), where y is a layer response vector, W is weight matrix, x is an input vector, g(x)=max(0,x) is an activation function and b is a bias vector. (Note that p and q are the dimensions of the layer and will be different in different layers) … y.sub.i=g(W.sub.i×x.sub.i+b.sub.i) x.sub.i+1=y.sub.i Statistics_i=Statistics_i ∪ y.sub.i; para. [0049, 0067]).
Wang does not explicitly teach the threshold function is applied to obtain at least a subset of the candidate output values.
However, Donnelly teaches wherein the component comprises a layer of a neural network (i.e. the quantized neural network layer 100 is configured to receive a layer input 112 from a previous neural network layer 110 in the quantized neural network; para. [0037-0041]), the layer is associated with a threshold function (i.e. using an activation function, e.g., the TanH or ReLU function, to generate the unquantized activation tensor 142; para. [0047]), and the threshold function is applied to obtain at least a subset of the candidate output values (i.e. The activation quantization engine 150 is configured to receive the unquantized activation tensor 142 and to process the unquantized activation tensor 142 to generate the quantized layer output 152; para. [0049]) and at least a subset of the reference output values (i.e. using an activation function, e.g., the TanH or ReLU function, to generate the unquantized activation tensor 142; para. [0047]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Donnelly. One would have been motivated to make this modification because it improves the accuracy of the candidate configuration evaluation.
9. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Sun et al. (U.S. Patent Application Pub. No. US 20210064976 A1).
Claim 10: Wang teaches the system of claim 9. Wang further teaches wherein the parameters comprise weights (i.e. the set of parameters (e.g., weights and bias) for each layer of the model; para. [0062]), and each weight is quantized (i.e. a respective set of quantized parameters (e.g., quantized weights and bias parameters), and each quantized parameter is expressed with the preferred values of the respective reduced bit-widths for the layer as determined through the multiple iterations; para. [0073]) to be represented by a combination of exponent bits and mantissa bits.
Wang does not explicitly teach each weight is quantized to be represented by a combination of exponent bits and mantissa bits.
However, Sun teaches wherein the parameters comprise weights (i.e. the FP GEMM modules 402 may each be configured to handle floating point numbers, e.g., activations and weights, having the 8-bit floating point format 200 in the (1, 4, 3) configuration during forward propagation neural network operations; para. [0035]), and each weight is quantized (i.e. Precision conversion modules 406 are configured to convert floating point numbers in the FP32 format or other formats such as, e.g., the FP16 format, to the FP8 format; para. [0037]) to be represented by a combination of exponent bits and mantissa bits (i.e. FPU 102 is configured as a hybrid FPU which can interchangeably use an 8-bit floating point format 200, e.g., the first floating point format, comprising a sign bit 202, four exponent bits 204 and three mantissa bits 206 in a (1, 4, 3) configuration for forward propagation neural network operations; para. [0029]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Sun. One would have been motivated to make this modification because it reduces the time to inference a trained neural network model without significant loss of accuracy from the expected performance of the model.
10. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Mellempudi et al. (U.S. Patent Application Pub. No. US 20210342692 A1), and further in view of Sarma et al. (U.S. Patent Application Pub. No. US 20200348909 A1).
Claim 11: Wang teaches the system of claim 1. Wang does not explicitly teach receiving, from a user device, a quantization request comprising a selected bit precision for quantization of the component, wherein the set of candidate quantization configurations is identified based on the selected bit precision, the set of candidate quantization configurations comprising different combinations of exponent bits and mantissa bits that satisfy the selected bit precision.
However, Mellempudi teaches receiving, from a user device, a quantization request comprising a selected bit precision for quantization of the component, wherein the set of candidate quantization configurations is identified based on the selected bit precision (i.e. The quantization controller 206 is configured to receive a request to quantize a message and a request to send the message to one or more receiver computing devices 102 b. In the illustrative embodiment, the request indicative of a quantization level is received from the application 202 of the sender computing node 102 a via the quantization library 204. It should be appreciated that, in some embodiments, the quantization controller 206 may receive the quantization request via a middleware library such as MPI; para. [0024, 0051, 0079]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Mellempudi. One would have been motivated to make this modification because it reduces candidate search burden.
However, Sarma teaches the set of candidate quantization configurations comprising different combinations of exponent bits and mantissa bits that satisfy the selected bit precision (i.e. 8-bit floating-point format 400 includes a single bit for sign bit 401, 4-bits for exponent 403, and 3-bits for mantissa 405. Sign bit 401, exponent 403, and mantissa 405 take up a total of 8-bits and can be used to represent a floating-point number. Similarly, 8-bit floating-point format 410 includes a single bit for sign bit 411, 5-bits for exponent 413, and 2-bits for mantissa 415. Sign bit 411, exponent 413, and mantissa 415 take up a total of 8-bits and can be used to represent a floating-point number; para. [0061]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Wang and Mellempudi to include the feature of Sarma. One would have been motivated to make this modification because its lower bit formats increase processing and performance bandwidth and power efficiency.
11. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Seo et al. (U.S. Patent Application Pub. No. US 20230129133 A1).
Claim 13: Wang teaches the system of claim 1. Wang further teaches wherein the quantizing of the trained machine learning model comprises quantizing parameters of the trained machine learning model (i.e. different configurations of bit-width and layer combinations are prepared as candidates for evaluation. In some embodiments, the quantization candidates are generated by using different combinations of bit-widths for all the layers; instead, the weights from different layers are clustered based on their values, and the quantization candidates are generated by using different combinations of bit-widths for all the clusters; para. [0062, 0073]), the operations further comprising: storing the quantized parameters in memory of a processing device (i.e. the model deployment system 300 includes one or more processing units (or “processors”) 302, memory 304; para. [0036-0038]).
Wang does not explicitly teach storing on-chip memory.
However, wherein the quantizing of the trained machine learning model comprises quantizing parameters of the trained machine learning model (i.e. To achieve high accuracy with very low-precision quantization, weights of the DNN are quantized during training; para. [0085]), the operations further comprising: storing the quantized parameters (i.e. Aided by 16×HCGS compression and with 6-bit weight quantization, all parameters of LSTMs for TIMIT/TED-LIUM/LibriSpeech are stored on-chip in <300-kB SRAM; para. [0052]) in on-chip memory (i.e. To enable this efficiently, each row of four matrices Wxi, Wxƒ, Wxo, and Wxc is stored in a staggered manner (same for Wh*) in on-chip SRAM arrays; para. [0104]) of a processing device (i.e. Weights can be stored on-chip (e.g., SRAM cache of mobile processors); para. [0005]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Seo. One would have been motivated to make this modification because to improve the energy efficiency of neural network hardware, off-chip memory access and communication need to be minimized. To that end, it becomes crucial to store most or all weights on-chip through sparsity/compression, weight quantization, and network size reduction.
12. Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Nagel et al. (U.S. Patent Application Pub. No. US 20200302299 A1).
Claim 14: Wang teaches the system of claim 1. Wang further teaches wherein the trained machine learning model comprises a neural network (i.e. the device obtains (502) a first neural network model (e.g., a trained full-precision model 160′, or a pruned full-precision model 160″) that includes a plurality of layers; para. [0071]), and the operations further comprise: performing prior to obtaining the reference output values and prior to the quantization of the trained machine learning model (i.e. the computing device collects a respective baseline statistical distribution of activation values for the first layer (e.g., statistics_i) as the validation data set is forward propagated as input through the first neural network model, while the respective sets of parameters of the plurality of layers are expressed with the original bit-width (e.g., 32-bit) of the first neural network model; the computing device collects a respective modified statistical distribution of activation values for the first layer (e.g., statq,i) as the validation data set is forward propagated as input through the first neural network model, while the respective set of parameters of the first layer are expressed with a first reduced bit-width (e.g., Wq,i and bq,i) that are smaller than the original bit-width of the first neural network model; determining a predefined divergence (e.g., Inf_tmp=InformationLoss (statq,I , statistics_i)) between the respective modified statistical distribution of activation values for the first layer and the respective baseline statistical distribution of activation values for the first layer; para. [0062-0064, 0075]).
Wang does not explicitly teach performing batch normalization folding.
However, Nagel teaches wherein the trained machine learning model comprises a neural network (i.e. FIG. 5 illustrates a method 500 for training a neural network on a large bit-width computing device and using cross layer rescaling and quantization to transform the neural network into a form suitable for execution on a small bit-width computing device; para. [0070]), and the operations further comprise: performing batch normalization folding (i.e. embodiments may use batch normalization to further reduce errors from quantizing a (e.g., a 32-bit model) neural network into a quantized (e.g., an INT8 model) neural network. Batch normalization may improve inference times in addition to reducing quantization errors. Folding the batch normalization parameters into a preceding layer improves inference time because per channel scale and shift operations do not have to be performed; para. [0050, 0053, 0072]) prior to obtaining the output values and prior to the quantization of the trained machine learning model (i.e. In block 508, the processor may quantize the weights or weight tensors form a large bit-width (e.g., FP32, etc.) form into small bit-width (e.g., INT8) representations to yield a neural network suitable for implementation on a small-bit width processor. The processor may use any of a variety of know methods of quantizing neural network weights to generate a quantized version of the neural network in block 508; para. [0072-0074]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Nagel. One would have been motivated to make this modification because it reduces quantization error and improves inference speed.
13. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Wang in view of Mellempudi et al. (U.S. Patent Application Pub. No. US 20210342692 A1).
Claim 16: Wang teaches the system of claim 1. Wang further teaches wherein the quantization of the trained machine learning model comprises generating a new instance of the trained machine learning model that comprises the selected candidate quantization configuration for the component, generating the new instance of the trained machine learning model (i.e. The device generates (506) a reduced neural network model (e.g., model 112) that includes the plurality of layers, wherein each layer of two or more the plurality of layers includes a respective set of quantized parameters (e.g., quantized weights and bias parameters), and each quantized parameter is expressed with the preferred values of the respective reduced bit-widths for the layer as determined through the multiple iterations; para. [0073]).
Wang does not explicitly teach receiving, from a user device, a quantization request; and
generating, in response to receiving the quantization request.
However, Mellempudi teaches receiving, from a user device, a quantization request (i.e. The quantization controller 206 is configured to receive a request to quantize a message and a request to send the message to one or more receiver computing devices 102 b. In the illustrative embodiment, the request indicative of a quantization level is received from the application 202 of the sender computing node 102 a via the quantization library 204. It should be appreciated that, in some embodiments, the quantization controller 206 may receive the quantization request via a middleware library such as MPI; para. [0024, 0051, 0079]); and generating, in response to receiving the quantization request, the new instance (i.e. The quantizer 208 is configured to determine a quantization level for a message in response to receipt of the request to send the message and receipt of the request to quantize the message from the application 202. The quantizer 208 is further configured to quantize the message based on the quantization level to generate a quantized message including one or more quantized values; para. [0025]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Wang to include the feature of Mellempudi. One would have been motivated to make this modification because it reduces candidate search burden.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Ovtcharov et al. (Pub. No. US 20200302271 A1), a computing architecture diagram that shows aspects of the configuration of a computing system disclosed herein that is capable of quantizing activations and weights during ANN training and inference.
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TAN H TRAN/Primary Examiner, Art Unit 2141