Prosecution Insights
Last updated: October 01, 2026
Application No. 18/748,071

METHOD FOR FINDING AT LEAST ONE OPTIMAL POST-TRAINING QUANTIZATION MODEL AND A NON-TRANSITORY MACHINE-READABLE MEDIUM

Non-Final OA §103
Filed
Jun 19, 2024
Priority
Jul 12, 2023 — provisional 63/513,127
Examiner
GURMU, MULUEMEBET
Art Unit
Tech Center
Assignee
MediaTek Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
398 granted / 496 resolved
+20.2% vs TC avg
Strong +18% interview lift
Without
With
+17.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
27 currently pending
Career history
520
Total Applications
across all art units

Statute-Specific Performance

§101
18.2%
-21.8% vs TC avg
§103
68.1%
+28.1% vs TC avg
§102
3.3%
-36.7% vs TC avg
§112
1.4%
-38.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 496 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-20 are present in this application. Claims 1-20 are pending in this office action. This office action is NON-FINAL. Drawings The Drawings filed on 06/19/24 are acceptable for examination purposes. Specification The Specification filed on 06/19/24 is acceptable for examination purposes. Information Disclosure Statement The information disclosure statements (IDS) filed on 03/21/25 has been considered by the Examiner and made of record in the application file. Claim Rejections 35 U.S.C. §103 6. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: Claims 1-11 and 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over Sriram et al. (US 2022/0044114 A1) in view of JANG et al. (US 2023/0161558 A1) . Regarding claim 1, Sriram teaches a method for finding at least one optimal post-training quantization (PTQ) model comprising, (See Sriram paragraph [0062], post-training quantization PTQ 112 to generate a trained machine-learning model (also referred to herein as a trained model or a trained neural network model) 114): converting and optimizing a floating-point machine learning model into a converted machine learning model, (See Sriram Abstract one or more weights of a trained model are represented by low bit integer numbers instead of using full floating point precision. Changing precision of the one or more weights is performed by first quantizing all weights); applying a plurality of PTQ settings to generate a plurality of PTQ models, (See Sriram paragraph [0158], by applying post-training quantization (PTQ) on a quantization-aware training (QAT) model to generate a lower-bit quantized model), and Sriram does not explicitly disclose evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model. However, JANG teaches evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify to evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 2, Sriram taught the method according to claim 1 as described above. Sriram further teaches wherein each of the plurality of PTQ settings, comprises a precision setting, (See Sriram paragraph [0004], performing low-precision quantization while training a neural network using both quantization-aware training (QAT) and post-training quantization (PTQ) to generate a trained model), a quantization error minimization algorithm, (See Sriram paragraph [0063], using one or more processors, first applies QAT 108 by considering quantization errors when training a model) and a calibration scheme, and the method further comprises, (See Sriram paragraph [0058], range/scale factors for activations are calibrated again using the PTQ process): configuring a plurality of precision settings, a plurality of quantization error minimization algorithms, (See Sriram paragraph [0063], A training graph may be modified to simulate the lower precision behavior in the forward pass of the training process, and thus introduces the quantization errors as part of the training loss, which the optimizer tries to minimize during the training) and a plurality of calibration schemes, (See Sriram paragraph [0109], The calibrator calibrates a model when building an INT8 engine. While PTQ provides an easy way to model quantization errors): performing Cartesian product on the plurality of precision settings, (See Sriram paragraph [0063], A training graph may be modified to simulate the lower precision behavior in the forward pass of the training process), the plurality of quantization error minimization algorithms, (See Sriram paragraph [0063], the quantization errors as part of the training loss, which the optimizer tries to minimize during the training), and the plurality of calibration schemes to form the plurality of PTQ settings, (See Sriram paragraph [0109], The calibrator calibrates a model when building an INT8 engine. While PTQ provides an easy way to model quantization errors): Sriram does not explicitly disclose sorting the plurality of PTQ settings based on lexicographical order; and storing the sorted plurality of PTQ settings in the storage. However, JANG teaches sorting the plurality of PTQ settings based on lexicographical order; (See JANG paragraph [0080], a quantization scheme of an artificial neural network…The operations in FIG. 3 may be performed in the sequence…Many of the operations shown in FIG. 3 may be performed in parallel or concurrently) and storing the sorted plurality of PTQ settings in the storage, (See JANG paragraph [0122],The registers 510 may store first bit streams into which input data corresponding to a first M-dimensional vector is encoded using a predetermined quantization scheme, and second bit streams into which a weight parameter corresponding to a second M-dimensional vector is encoded using the predetermined quantization scheme). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify sorting the plurality of PTQ settings based on lexicographical order; and storing the sorted plurality of PTQ settings in the storage of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 3, Sriram taught the method according to claim 1 as described above. Sriram further teaches wherein each of the plurality of PTQ settings, comprises a precision setting, (See Sriram paragraph [0004], performing low-precision quantization while training a neural network using both quantization-aware training (QAT) and post-training quantization (PTQ) to generate a trained model), a quantization error minimization algorithm, and or calibration scheme (See Sriram paragraph [0063], using one or more processors, first applies QAT 108 by considering quantization errors when training a model): and when applying the plurality of PTQ settings to generate a plurality of PTQ models, (See Sriram paragraph [0158], by applying post-training quantization (PTQ) on a quantization-aware training (QAT) model to generate a lower-bit quantized model). Sriram does not explicitly disclose skipping at least one redundant operation if at least two PTQ settings have the same precision setting, quantization error minimization algorithm, or calibration scheme. However, JANG teaches skipping at least one redundant operation if at least two PTQ settings have the same precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), quantization error minimization algorithm, or calibration scheme, (See JANG paragraph [0039], a general quantization scheme, positive and negative quantization levels may be unequally assigned (e.g., −1, 0, 1, 2, etc.), which may lead to an occurrence of an error and a reduction in performance at a low-precision quantization level). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify skipping at least one redundant operation if at least two PTQ settings have the same precision setting, quantization error minimization algorithm, or calibration scheme of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 4, Sriram taught the method according to claim 3 as described above. Sriram further teaches comprises skipping quantizing the constant weight tensors of the converted machine learning model, See Sriram paragraph [0075], pruning the pre-trained model 206 (e.g., eliminating unnecessary values in the weight tensor) would get the most compute savings possible without compromising accuracy, using a function such as tlt-prune), based on the precision setting of the second PTQ setting, See Sriram paragraph [0075], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision); Sriram does not explicitly disclose comprises: wherein a first PTQ and a second PTQ setting have the same precision setting, skipping at least one redundant operation. However, JANG teaches wherein a first PTQ and a second PTQ setting have the same precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), skipping at least one redundant operation (See JANG paragraph [0038], the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein a first PTQ and a second PTQ setting have the same precision setting, skipping at least one redundant operation of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 5, Sriram taught the method according to claim 3 as described above. Sriram further teaches skipping quantizing the constant weight tensors, (See Sriram paragraph [0075], pruning the pre-trained model 206 (e.g., eliminating unnecessary values in the weight tensor) would get the most compute savings possible without compromising accuracy, using a function such as tlt-prune), of the converted machine learning model based on the precision setting of the second PTQ setting, (See Sriram paragraph [0075], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision); skipping running the quantization error minimization algorithm of the second PTQ setting to compensate weight quantization error, (See Sriram paragraph [0109], PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range); and skipping collecting tensor statistics from PTQ calibration dataset, (See Sriram paragraph [0624], PTQ performs quantization on the one or more weights and one or more activation values of the intermediate trained model by ignoring the one or more activation values of the intermediate trained model…on statistics against a calibration dataset). Sriram does not explicitly disclose wherein a first PTQ and a second PTQ setting have the same precision setting, and the same quantization error minimization algorithm, skipping at least one redundant operation comprises. However, JANG teaches wherein a first PTQ and a second PTQ setting have the same precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), and the same quantization error minimization algorithm, See JANG paragraph [0039], a general quantization scheme, positive and negative quantization levels may be unequally assigned (e.g., −1, 0, 1, 2, etc.), which may lead to an occurrence of an error and a reduction in performance at a low-precision quantization level), skipping at least one redundant operation (See JANG paragraph [0038], the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein a first PTQ and a second PTQ setting have the same precision setting, skipping at least one redundant operation of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 6, Sriram taught the method according to claim 3 as described above. Sriram further teaches skipping quantizing the constant weight tensors, (See Sriram paragraph [0075], pruning the pre-trained model 206 (e.g., eliminating unnecessary values in the weight tensor) would get the most compute savings possible without compromising accuracy, using a function such as tlt-prune), of the converted machine learning model based on the precision setting of the second PTQ setting, (See Sriram paragraph [0075], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision); skipping running the quantization error minimization algorithm of the second PTQ setting to compensate weight quantization error, (See Sriram paragraph [0109], PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range); skipping collecting tensor statistics from PTQ calibration dataset, (See Sriram paragraph [0109], there are cases where this scaling cannot preserve the statistics of the model weights. One such example is illustrated in FIG. 3. The model was trained with tensors represented in FP32 mode…applying both PTQ and QAT to train a neural network and output a trained model may improve the models' performance on object detection); and skipping running the calibration scheme of the second PTQ setting based on the collected tensor statistics to calibrate the activation tensors of the converted machine learning model, (See Sriram paragraph [0109], While PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range. However, there are cases where this scaling cannot preserve the statistics of the model weights…The model was trained with tensors represented in FP32 mode and calibrated using an INT8 entropy calibrator). Sriram does not explicitly disclose wherein a first PTQ and a second PTQ setting have the same precision setting, the same quantization error minimization algorithm, and the same calibration scheme, skipping at least one redundant operation comprises. However, JANG teaches wherein a first PTQ and a second PTQ setting have the same precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), the same quantization error minimization algorithm, See JANG paragraph [0039], a general quantization scheme, positive and negative quantization levels may be unequally assigned (e.g., −1, 0, 1, 2, etc.), which may lead to an occurrence of an error and a reduction in performance at a low-precision quantization level), and the same calibration scheme, skipping at least one redundant operation comprises, (See JANG paragraph [0038], the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein a first PTQ and a second PTQ setting have the same precision setting, the same quantization error minimization algorithm, and the same calibration scheme, skipping at least one redundant operation comprises of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 7, Sriram taught the method according to claim 1 as described above. Sriram further teaches a calibration scheme, and applying the plurality of PTQ settings to generate a plurality of PTQ models further comprises, (See Sriram paragraph [0158], by applying post-training quantization (PTQ) on a quantization-aware training (QAT) model to generate a lower-bit quantized model): (a) obtaining a new PTQ setting from a storage, (See Sriram paragraph [0617], generate the second trained model is performed by applying post-training quantization (PTQ): (c)quantizing the constant weight tensors of the converted machine learning model based on the precision setting of the new PTQ setting, go to (e), (See Sriram paragraph [0060], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision); (e)running the quantization error minimization algorithm of the new PTQ setting to compensate weight quantization error, (See Sriram paragraph [0109], The calibrator calibrates a model when building an INT8 engine. While PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range. However, there are cases where this scaling cannot preserve the statistics of the model weights); (f) collecting tensor statistics from PTQ calibration dataset, (See Sriram paragraph [0109], there are cases where this scaling cannot preserve the statistics of the model weights. One such example is illustrated in FIG. 3. The model was trained with tensors represented in FP32 mode…applying both PTQ and QAT to train a neural network and output a trained model may improve the models' performance on object detection); (g) running the calibration scheme of the new PTQ setting based on the collected tensor statistics to calibrate the activation tensors of the converted machine learning model, (See Sriram paragraph [0109], While PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range. However, there are cases where this scaling cannot preserve the statistics of the model weights…The model was trained with tensors represented in FP32 mode and calibrated using an INT8 entropy calibrator); and (h) quantizing the converted machine learning model to generate a PTQ model, (See Sriram paragraph [0060], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision). Sriram does not explicitly disclose wherein each of the plurality of post-training quantization settings comprises a precision setting, a quantization error minimization algorithm, (b)checking whether the precision setting of the new PTQ setting is the same as the precision setting of the last obtained PTQ setting, if so, go to (d), else, go to (c), (d)checking whether the quantization error minimization algorithm of the new PTQ setting is the same as the quantization error minimization algorithm of the last obtained PTQ setting, if so, go to (f); else, go to (e). However, JANG teaches wherein each of the plurality of post-training quantization settings comprises a precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), a quantization error minimization algorithm, (See JANG paragraph [0039], a general quantization scheme, positive and negative quantization levels may be unequally assigned (e.g., −1, 0, 1, 2, etc.), which may lead to an occurrence of an error and a reduction in performance at a low-precision quantization level), (b)checking whether the precision setting of the new PTQ setting is the same as the precision setting of the last obtained PTQ setting, if so, go to (d), else, go to (c), (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)); (d) checking whether the quantization error minimization algorithm of the new PTQ setting is the same as the quantization error minimization algorithm of the last obtained PTQ setting, if so, go to (f); else, go to (e), (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein each of the plurality of post-training quantization settings comprises a precision setting, (b)checking whether the precision setting of the new PTQ setting is the same as the precision setting of the last obtained PTQ setting, if so, go to (d), else, go to (c), (d)checking whether the quantization error minimization algorithm of the new PTQ setting is the same as the quantization error minimization algorithm of the last obtained PTQ setting, if so, go to (f); else, go to (e). of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 8, Sriram taught the method according to claim 7 as described above. Sriram further teaches wherein after quantizing the constant weight tensors of the converted machine learning model based on the precision setting of the new PTQ setting, go to (d) instead of going to (e), (See Sriram paragraph [0060], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision. Typically, DNN training and inference have relied on the Institute of Electrical and Electronics Engineers (IEEE) single-precision floating-point format, using 32 bits to represent the floating-point model weights and activation tensors). Regarding claim 9, Sriram taught the method according to claim 7 as described above. Sriram does not explicitly disclose wherein further comprises: (f-1) checking whether the calibration scheme of the new PTQ setting is the same as the calibration scheme of the last obtained PTQ setting, if so, go to (h); else, go to (g). However, JANG teaches wherein further comprises: (f-1) checking whether the calibration scheme of the new PTQ setting is the same as the calibration scheme of the last obtained PTQ setting, if so, go to (h); else, go to (g), (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein further comprises: (f-1) checking whether the calibration scheme of the new PTQ setting is the same as the calibration scheme of the last obtained PTQ setting, if so, go to (h); else, go to (g) of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 10, Sriram taught the method according to claim 7 as described above. Sriram does not explicitly disclose wherein further comprises: (i) checking whether all the PTQ settings are obtained, if not, go to(a). However, JANG teaches wherein further comprises: (i) checking whether all the PTQ settings are obtained, if not, go to(a), (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein further comprises: (i) checking whether all the PTQ settings are obtained, if not, go to(a) of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 11, Sriram taught the method according to claim 7 as described above. Sriram further teaches (k) comparing all the evaluation results to find at least one optimal PTQ model, (See Sriram paragraph [0052], the pre-trained model 206 includes parameters that are represented by full precision floating point numbers (e.g., 32-bit representation) and the resulting trained model 208 includes parameters that are represented by lower-bit integers compared to the full precision floating point numbers. In an embodiment, lower-bit integers are 8-bit integers). Sriram does not explicitly disclose wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, comprises: (i) evaluating the new PTQ model based on at least one predetermined indirect metric, (j) checking if all the PTQ setting are obtained, if so, go to (k), else, go to(a). However, JANG teaches wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization), comprises: (i) evaluating the new PTQ model based on at least one predetermined indirect metric, , (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization); (j) checking if all the PTQ setting are obtained, if so, go to (k), else, go to(a); (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, comprises: (i) evaluating the new PTQ model based on at least one predetermined indirect metric, (j) checking if all the PTQ setting are obtained, if so, go to (k), else, go to(a).of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 15, Sriram teaches a non-transitory machine-readable medium for storing a program code, wherein when loaded and executed by a processor, (See Sriram paragraph [0111], a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium), the program code instructs the processor to execute, (See Sriram paragraph [0111], a plurality of computer-readable instructions executable by one or more processors): converting and optimizing a floating-point machine learning model into a converted machine learning model, (See Sriram Abstract one or more weights of a trained model are represented by low bit integer numbers instead of using full floating point precision. Changing precision of the one or more weights is performed by first quantizing all weights); applying a plurality of post-training quantization (PTQ) settings to generate a plurality of PTQ models, (See Sriram paragraph [0158], by applying post-training quantization (PTQ) on a quantization-aware training (QAT) model to generate a lower-bit quantized model). Sriram does not explicitly disclose evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model. However, JANG teaches evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify to evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 16, Sriram taught the non-transitory machine-readable medium according to claim 15 as described above. Sriram further teaches wherein each of the plurality of post-training quantization settings comprises a precision setting, (See Sriram paragraph [0004], performing low-precision quantization while training a neural network using both quantization-aware training (QAT) and post-training quantization (PTQ) to generate a trained model), a quantization error minimization algorithm, (See Sriram paragraph [0063], using one or more processors, first applies QAT 108 by considering quantization errors when training a model), and a calibration scheme, and the processor executes, (See Sriram paragraph [0058], range/scale factors for activations are calibrated again using the PTQ process): configuring a plurality of precision settings, a plurality of quantization error minimization algorithms, (See Sriram paragraph [0063], A training graph may be modified to simulate the lower precision behavior in the forward pass of the training process, and thus introduces the quantization errors as part of the training loss, which the optimizer tries to minimize during the training) and a plurality of calibration schemes, (See Sriram paragraph [0109], The calibrator calibrates a model when building an INT8 engine. While PTQ provides an easy way to model quantization errors): performing Cartesian product on the plurality of precision settings, See Sriram paragraph [0063], A training graph may be modified to simulate the lower precision behavior in the forward pass of the training process), the plurality of quantization error minimization algorithms, (See Sriram paragraph [0063], the quantization errors as part of the training loss, which the optimizer tries to minimize during the training), and the plurality of calibration schemes to form the plurality of PTQ settings, (See Sriram paragraph [0109], The calibrator calibrates a model when building an INT8 engine. While PTQ provides an easy way to model quantization errors). Sriram does not explicitly disclose sorting the plurality of PTQ settings based on lexicographical order; and storing the sorted plurality of PTQ settings in the storage. However, JANG teaches sorting the plurality of PTQ settings based on lexicographical order; (See JANG paragraph [0080], a quantization scheme of an artificial neural network…The operations in FIG. 3 may be performed in the sequence…Many of the operations shown in FIG. 3 may be performed in parallel or concurrently) and storing the sorted plurality of PTQ settings in the storage, (See JANG paragraph [0122],The registers 510 may store first bit streams into which input data corresponding to a first M-dimensional vector is encoded using a predetermined quantization scheme, and second bit streams into which a weight parameter corresponding to a second M-dimensional vector is encoded using the predetermined quantization scheme). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify to sorting the plurality of PTQ settings based on lexicographical order; and storing the sorted plurality of PTQ settings in the storage of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 17, Sriram taught the non-transitory machine-readable medium according to claim 15 as described above. Sriram further teaches wherein each of the plurality of post-training quantization settings comprises a precision setting, (See Sriram paragraph [0004], performing low-precision quantization while training a neural network using both quantization-aware training (QAT) and post-training quantization (PTQ) to generate a trained model), a quantization error minimization algorithm, (See Sriram paragraph [0063], using one or more processors, first applies QAT 108 by considering quantization errors when training a model): and a calibration scheme, and when applying the plurality of PTQ settings to generate a plurality of PTQ models, (See Sriram paragraph [0158], by applying post-training quantization (PTQ) on a quantization-aware training (QAT) model to generate a lower-bit quantized model). Sriram does not explicitly disclose the processor skips at least one redundant operation if at least two PTQ settings have the same precision setting, quantization error minimization algorithm, or calibration scheme. However, JANG teaches the processor skips at least one redundant operation if at least two PTQ settings have the same precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), quantization error minimization algorithm, or calibration scheme, (See JANG paragraph [0039], a general quantization scheme, positive and negative quantization levels may be unequally assigned (e.g., −1, 0, 1, 2, etc.), which may lead to an occurrence of an error and a reduction in performance at a low-precision quantization level). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify the processor skips at least one redundant operation if at least two PTQ settings have the same precision setting, quantization error minimization algorithm, or calibration scheme of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 18, Sriram taught the non-transitory machine-readable medium according to claim 15 as described above. Sriram further teaches a calibration scheme, and when applying the plurality of PTQ settings to generate a plurality of PTQ models, the processor further executes, (See Sriram paragraph [0158], by applying post-training quantization (PTQ) on a quantization-aware training (QAT) model to generate a lower-bit quantized model): (a)obtaining a new PTQ setting from a storage, (See Sriram paragraph [0617], generate the second trained model is performed by applying post-training quantization (PTQ): (c)quantizing the constant weight tensors of the converted machine learning model based on the precision setting of the new PTQ setting, and go to (e), (See Sriram paragraph [0060], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision); (e) running the quantization error minimization algorithm of the new PTQ setting to compensate weight quantization error, (See Sriram paragraph [0109], The calibrator calibrates a model when building an INT8 engine. While PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range. However, there are cases where this scaling cannot preserve the statistics of the model weights); (f) collecting tensor statistics from PTQ calibration dataset, (See Sriram paragraph [0109], there are cases where this scaling cannot preserve the statistics of the model weights. One such example is illustrated in FIG. 3. The model was trained with tensors represented in FP32 mode…applying both PTQ and QAT to train a neural network and output a trained model may improve the models' performance on object detection); (g) running the calibration scheme of the new PTQ setting based on the collected tensor statistics to calibrate the activation tensors of the converted machine learning model, (See Sriram paragraph [0109], While PTQ provides an easy way to model quantization errors, there is an inherent assumption that the weights of the trained model can be effectively scaled to a smaller range. However, there are cases where this scaling cannot preserve the statistics of the model weights…The model was trained with tensors represented in FP32 mode and calibrated using an INT8 entropy calibrator); and (h) quantizing the converted machine learning model to generate a PTQ model, (See Sriram paragraph [0060], Quantization is the process of transforming deep learning models to use parameters and computations at a lower precision). Sriram does not explicitly disclose wherein each of the plurality of post-training quantization settings comprises a precision setting, a quantization error minimization algorithm, and (b)checking whether the precision setting of the new PTQ setting is the same as the precision setting of the last obtained PTQ setting, if so, go to (d), else, go to (c); (d) checking whether the quantization error minimization algorithm of the new PTQ setting is the same as the quantization error minimization algorithm of the last obtained PTQ setting, if so, go to (f); else, go to (e). However, JANG teaches wherein each of the plurality of post-training quantization settings comprises a precision setting, (See JANG paragraph [0038], The same name may be used to describe an element included in the examples described above and an element having a common function. Unless otherwise mentioned, the descriptions on the examples may be applicable to the following examples and thus, duplicated descriptions will be omitted for conciseness), a quantization error minimization algorithm, (See JANG paragraph [0039], a general quantization scheme, positive and negative quantization levels may be unequally assigned (e.g., −1, 0, 1, 2, etc.), which may lead to an occurrence of an error and a reduction in performance at a low-precision quantization level), and (b)checking whether the precision setting of the new PTQ setting is the same as the precision setting of the last obtained PTQ setting, if so, go to (d), else, go to (c), (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)); (d) checking whether the quantization error minimization algorithm of the new PTQ setting is the same as the quantization error minimization algorithm of the last obtained PTQ setting, if so, go to (f); else, go to (e); (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein each of the plurality of post-training quantization settings comprises a precision setting, a quantization error minimization algorithm, and (b)checking whether the precision setting of the new PTQ setting is the same as the precision setting of the last obtained PTQ setting, if so, go to (d), else, go to (c); (d) checking whether the quantization error minimization algorithm of the new PTQ setting is the same as the quantization error minimization algorithm of the last obtained PTQ setting, if so, go to (f); else, go to (e) of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 19, Sriram taught the non-transitory machine-readable medium according to claim 18 as described above. Sriram does not explicitly disclose wherein the processor further executes: (i) checking whether all the PTQ settings are obtained, if not, go to(a). However, JANG teaches wherein the processor further executes: (i) checking whether all the PTQ settings are obtained, if not, go to(a), (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein the processor further executes: (i) checking whether all the PTQ settings are obtained, if not, go to(a). of JANG in order to reduce the number of bits necessary to represent information. Regarding claim 20, Sriram taught the non-transitory machine-readable medium according to claim 18 as described above and (k) comparing all the evaluation results to find at least one optimal PTQ model, (See J Sriram paragraph [0052], the pre-trained model 206 includes parameters that are represented by full precision floating point numbers (e.g., 32-bit representation) and the resulting trained model 208 includes parameters that are represented by lower-bit integers compared to the full precision floating point numbers. In an embodiment, lower-bit integers are 8-bit integers). Sriram does not explicitly disclose when wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, the processor further executes: (i) evaluating the new PTQ model based on at least one predetermined indirect metric; (j) checking if all the PTQ setting are obtained, if so, go to (k), else, go to(a). However, JANG teaches when wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization), the processor further executes: (i) evaluating the new PTQ model based on at least one predetermined indirect metric, (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization); (j) checking if all the PTQ setting are obtained, if so, go to (k), else, go to(a); (See JANG paragraph [0052], a quantization range may be trained in the same manner as learned step-size quantization (LSQ). Centered Symmetric Quantization (CSQ) may not be limited to QAT because it has mostly to do with resulting quantization levels, whether obtained from QAT or post-training quantization (PTQ or even data-free quantization (DFQ)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify when wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, the processor further executes: (i) evaluating the new PTQ model based on at least one predetermined indirect metric; (j) checking if all the PTQ setting are obtained, if so, go to (k), else, go to(a) of JANG in order to reduce the number of bits necessary to represent information. Claims 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Sriram et al. (US 2022/0044114 A1) in view of JANG et al. (US 2023/0161558 A1) and further in view of BUCS et al. (US 2023/0325662 A1). Regarding claim 12, Sriram taught the method according to claim 1 as described above. Sriram together with JANG does not explicitly disclose wherein the at least one predetermined indirect metric comprises: signal-to-quantization-noise ratio (SQNR), mean absolute error (MAE), mean squared error (MSE), or cosine similarity. However, BUCS teaches wherein the at least one predetermined indirect metric comprises: signal-to-quantization-noise ratio (SQNR), (See BUCS paragraph [0120], the quantization error may be determined based on a mean squared error and/or a mean average error and/or a peak signal to noise ratio), mean absolute error (MAE), mean squared error (MSE), or cosine similarity, (See BUCS paragraph [0096], the Mean Squared Error (MSE), Mean Average Error (MAE)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify to wherein the at least one predetermined indirect metric comprises: signal-to-quantization-noise ratio (SQNR) of BUCS for determining a representative input data set for post-training quantization of artificial neural networks. Regarding claim 13, Sriram taught the method according to claim 12 as described above. Sriram does not explicitly disclose wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, further comprises: However, JANG teaches wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, further comprises, (See JANG paragraph [0065], a quantization range determined by a step size optimized…a quantization result may be obtained within a predetermined gradient with respect to an input value included in a quantization range by the quantization). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein evaluating the plurality of PTQ models based on at least one predetermined indirect metric to find at least one optimal PTQ model, further comprises of JANG in order to reduce the number of bits necessary to represent information. Sriram together with JANG does not explicitly disclose executing at least one of the following operations to obtain a plurality of evaluation results, calculating SQNR difference between the floating-point machine learning model and each of the plurality of PTQ models, calculating the MAE between the floating-point machine learning model and each of the plurality of PTQ models; calculating the MSE between the floating-point machine learning model and each of the plurality of PTQ models; calculating cosine similarity between the floating-point machine learning model and each of the plurality of PTQ models; and the evaluating step further comprises, comparing all the evaluation results to find at least one optimal PTQ model. However, BUCS teaches executing at least one of the following operations to obtain a plurality of evaluation results, (See BUCS paragraph [0063], he quantized ANN may be deployed, executed and analyzed on the hardware. If, however, quantization does not provide satisfactory results): calculating SQNR difference between the floating-point machine learning model and each of the plurality of PTQ models, (See BUCS paragraph [0096], the quantization error may be calculated (in error calculation block 112) between the original floating-point input data array and its de-quantized also floating-point counterpart. Herein, the exact mathematical formulation for error calculation may be selected by the user from a set of possible metrics, including the Mean Squared Error (MSE), Mean Average Error (MAE), Peak Signal to Noise Ratio); calculating the MAE between the floating-point machine learning model and each of the plurality of PTQ models; (See BUCS paragraph [0096], the exact mathematical formulation for error calculation may be selected by the user from a set of possible metrics, including the Mean Squared Error (MSE), Mean Average Error (MAE), calculating the MSE between the floating-point machine learning model and each of the plurality of PTQ models, (See BUCS paragraph [0096], the quantization error may be calculated (in error calculation block 112) between the original floating-point input data array and its de-quantized also floating-point counterpart. Herein, the exact mathematical formulation for error calculation may be selected by the user from a set of possible metrics, including the Mean Squared Error (MSE)); calculating cosine similarity between the floating-point machine learning model and each of the plurality of PTQ models; and the evaluating step further comprises, (See BUCS paragraph [0096], the quantization error may be calculated (in error calculation block 112) between the original floating-point input data array and its de-quantized also floating-point counterpart): comparing all the evaluation results to find at least one optimal PTQ model, (See BUCS paragraph [0063], Finding the right settings for PTQ may involve numerous trial-and-error cycles. Moreover, in certain deployment frameworks, quantization calibration needs to use a target hardware emulation environment to gather statistics, which may slow down the calibration process by several orders of magnitudes compared to the time one might expect for calibration on a workstation). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify to executing at least one of the following operations to obtain a plurality of evaluation results, calculating SQNR difference between the floating-point machine learning model and each of the plurality of PTQ models, calculating the MAE between the floating-point machine learning model and each of the plurality of PTQ models; calculating the MSE between the floating-point machine learning model and each of the plurality of PTQ models; calculating cosine similarity between the floating-point machine learning model and each of the plurality of PTQ models; and the evaluating step further comprises, comparing all the evaluation results to find at least one optimal PTQ model of BUCS for determining a representative input data set for post-training quantization of artificial neural networks. Regarding claim 14, Sriram taught the method according to claim 12 as described above. Sriram together with JANG does not explicitly disclose wherein the at least one optimal PTQ model is any one or a combination of: a PTQ model has the closest SQNR to the floating-point machine learning model, a PTQ model which can obtain the smallest MAE, a PTQ model which can obtain the smallest MSE, a PTQ model which can obtain the optimal cosine similarity. However, BUCS teaches wherein the at least one optimal PTQ model is any one or a combination of: a PTQ model has the closest SQNR to the floating-point machine learning model, (See BUCS paragraph [0096], the quantization error may be calculated (in error calculation block 112) between the original floating-point input data array and its de-quantized also floating-point counterpart. Herein, the exact mathematical formulation for error calculation may be selected by the user from a set of possible metrics, including the Mean Squared Error (MSE), Mean Average Error (MAE), Peak Signal to Noise Ratio (PSNR)), a PTQ model which can obtain the smallest MAE, (See BUCS paragraph [0096], the Mean Squared Error (MSE), Mean Average Error (MAE)), a PTQ model which can obtain the smallest MSE, a PTQ model which can obtain the optimal cosine similarity, (See paragraph [0096], the quantization error may be calculated (in error calculation block 112) between the original floating-point input data array and its de-quantized also floating-point counterpart. Herein, the exact mathematical formulation for error calculation may be selected by the user from a set of possible metrics, including the Mean Squared Error (MSE)). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify to wherein the at least one optimal PTQ model is any one or a combination of: a PTQ model has the closest SQNR to the floating-point machine learning model, a PTQ model which can obtain the smallest MAE, a PTQ model which can obtain the smallest MSE, a PTQ model which can obtain the optimal cosine similarity of BUCS for determining a representative input data set for post-training quantization of artificial neural networks. Conclusions/Points of Contacts The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See form PTO-892. GU et al. (US 2023/0252294 A1) The practice has found that among the three schemes, in the post-quantization, the full-precision model is directly subjected to post-quantization, and a good recognition effect of the quantized model cannot be guaranteed. This is because errors caused by quantization are not taken into account during training of the full-precision model. However, a model often requires extremely high precision, and the errors caused by model quantization lead to wrong recognition results and bring immeasurable losses. GOLLANAPALLI et al. (US 2023/0068381 A1) the disclosure provide a novel post-training quantization method for a Deep Neural Network (DNN) model compression and fast inference by generating data (self-generated data) for quantizing weights and activations at lower bit precision without access to training/validation dataset. While existing methods of post-training quantization require access to training dataset to quantize the weights and activations or require to retrain the entire DNN model for a random number of epochs to adjust the weights and the activations. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUEMEBET GURMU whose telephone number is (571)270-7095. The examiner can normally be reached M-F 9am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tony Mahmoudi can be reached at 5712724078. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MULUEMEBET GURMU/Primary Examiner, Art Unit 2163
Read full office action

Prosecution Timeline

Jun 19, 2024
Application Filed
Sep 01, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705549
SYSTEM AND METHOD OF SEGMENTING DATA AND FORECASTING BY A COMBINATION OF MODELS TRAINED ON SEGMENTED DATA
3y 8m to grant Granted Aug 11, 2026
Patent 12705213
Use of Disaggregated Storage by a Distributed Storage System to Facilitate Performance of Data Management Features that Operate at Distributed Scale
2y 5m to grant Granted Aug 11, 2026
Patent 12688169
METHOD, DEVICE, AND COMPUTER PROGRAM PRODUCT FOR DATA MIGRATION
2y 1m to grant Granted Jul 21, 2026
Patent 12664471
MACHINE LEARNING MODEL ANALYSIS
3y 6m to grant Granted Jun 23, 2026
Patent 12650955
SYSTEM AND METHOD FOR EDITING A FILE-BACKED DATABASE TABLE
2y 0m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
98%
With Interview (+17.5%)
3y 1m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 496 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month