DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to a patent application filed on April 18th, 2024. Claims 1-20 are pending in the current application.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 6-9, 13-15, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Hosseini et al. (Herein referred to as Hosseini) (Cyclic Sparsely Connected Architectures for Compact Deep Convolutional Neural Networks) in view of Chung et al. (Herein referred to as Chung) (U.S. Patent Application No. US 20180341851 A1) and in further view of Oh et al. (Herein referred to as Oh) (Weight Equalizing Shift Scaler-Coupled Post-training Quantization)
Regarding claim 1, Hosseini teaches a method comprising: converting a plurality of functions or function call instructions of a first neural network (NN) model into a plurality of graph modules; (“Contrary to pruning methods that follow a train-prune routine, compact models can be considered as a prune-train routine in which the model is forced to be trained on an already-sparsified basis.”, pg. 3, right column, under “3) Compact Models”; See also Fig. 3 on pg. 3) (The compact models correspond to a first neural network which converts functions into a graph module.) analyzing a relationship between one or more inputs and one or more outputs of the plurality of graph modules; (“Every input node from a graph represents a 2-D channel, and each edge is weighted with a 2-D kernel, the connection of which corresponds to a 2-D convolution operation. Every output node is a 2-Dchannel that is resulted by summing over all of its connecting kernels operated on their connected inputs.” pg. 3, right column, under “3) Compact Models”; See also Fig. 3 on pg. 3) generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the plurality of graph modules corresponding to the first NN model, by mapping the one or more inputs and the one or more outputs of the plurality of graph modules to each other based on the relationship (“The objective of this Section is hence to formulate a directed acyclic graph composed of a few layers where all Input nodes are connected via an equal number of paths to all Output nodes of the graph.”, pg. 4, left column, under “A. Problem Statement and Formulation”) (Inspired by the first neural network, a DCNN in the form of a DAG is formulated based on connecting the input and output nodes of the graph, corresponding to a second neural network.)
However, Hosseini does not explicitly teach adding a plurality of markers to the plurality of graph modules in the second NN model; generating calibration data by collecting input values and output values of each of the plurality of graph modules using the plurality of markers nor determining, based on the calibration data, a scale value and an offset value applicable to the second NN model nor determining, for each graph module of the second NN model, an optimal value for the scale value or the offset value by performing a quantization simulation for one or more candidates among optimization candidates of the scale value or the offset value.
Chung teaches adding a plurality of markers to the plurality of graph modules in the second NN model; (“In some embodiments, performance data can be collected from probes that provide data on system performance with respect to a specified performance goal”, Paragraph 33) (The probes of Chung correspond to a plurality of markers, which when combined with the data from the graph modules and model of Hosseini, fully teaches the limitation.) and generating calibration data by collecting input values and output values of each of the plurality of graph modules using the plurality of markers (“During the monitoring phase, the instrumented training code can be profiled for communication and computation characteristics. This can be done by using performance analyzing tools relying on known data analytics functions such as probes, and/or software patches (an example of which is discussed with reference to FIG. 4), and changes to run-time control parameters. In some embodiments, data analytic probes are inserted into the program code, providing workload performance profiling statistics/data on the running system.”, Paragraph 49) (The control parameters correspond to input and output values collected as data by probes in the system; the probes corresponding to a plurality of markers. In combination with the plurality of graph modules of Hosseini, the limitation is fully taught.)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the neural networks of Hosseini, with the probes of Chung. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for live tuning, as detailed in Chung. (“During system monitoring, readings from the performance data probes 455 are provided to and received e.g.,
by tuning server 150. These readings can reflect performance statistics such as bandwidth utilization, memory usage, and power/wattage consumed. Using known performance monitoring tools, data collection can also include performance data 450 cataloging system software events 435 and hardware counter events 445”, Paragraph 62)
However, the combination does not teach determining, based on the calibration data, a scale value and an offset value applicable to the second NN model nor determining, for each graph module of the second NN model, an optimal value for the scale value or the offset value by performing a quantization simulation for one or more candidates among optimization candidates of the scale value or the offset value.
Oh teaches determining, based on the calibration data, a scale value and an offset value applicable to the second NN model (“applying the element-wise quantization function qw defined as follows… where the boundaries [min, max] are nudged by small amounts so that value 0.0 is exactly representable as an integer zero point, z, after quantization: min ← s ∗ min/s , max←max+s∗min/s−min. The function · rounds to the nearest integer. The scale, s, indicates the step size of quantization.”, pg. 2, second-to-last paragraph; See also Equations 1, 2, and 3 on pg. 2) (The variable z corresponds to an offset value, with the scale variable corresponding to a scale value. The data of the Hosseini-Chung combination, corresponding to calibration data, scaled according to Oh’s disclosure fully teaches this limitation.) and determining, for each graph module of the second NN model, an optimal value for the scale value or the offset value by performing a quantization simulation for one or more candidates among optimization candidates of the scale value or the offset value. (“In detail, QAT [quantization-aware training] is to simulate quantization during the time of training from scratch or fine-tuning, enabling to apply the effect of quantization simultaneously… In detail, weights are first channel-wise shifted by si (Line 7 in Algorithm 1), followed by uniform affine layer-wise quantization (FP32→UINT8) and dequantization (UINT8→FP32) process (Line8-10) as emulating inference-time quantization, and the n channel-wise shifted in the reversed direction (Line 11). Lastly, the quantization error can be measured in l2-norm between original weights and WES-coupled fake quantized weights so that bigger quantization error can be more penalized (Line 12).”, pg. 3, left column, under “Iterative search for an optimal total range”; See also Algorithm 1 on pg. 3) (These processes are done implicitly to find an optimal shift scale and offset.)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the neural networks of Hosseini, as modified by Chung, with the scale and offset values of Oh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for a quantized model with a low quantization error. (“distribution of the channel with a relatively narrow range is stretched out so it comes to having low quantization error after WES-coupled quantization. The process so far is performed in compile time, and a resultant quantized model is saved to a binary file.”, pg. 3, right column, first paragraph)
Regarding claim 13, Hosseini teaches a neural network (NN) model in a form of a directed acyclic graph (DAG); (“The objective of this Section is hence to formulate a directed acyclic graph composed of a few layers where all Input nodes are connected via an equal number of paths to all Output nodes of the graph.”, pg. 4, left column, under “A. Problem Statement and Formulation”) (A DCNN in the form of a DAG is formulated based on connecting the input and output nodes of the graph, the nodes corresponding to graph modules.)
However, Hosseini does not explicitly teach adding a plurality of markers to a plurality of graph modules included in a neural network (NN) model, nor collecting input values and output values of each of the plurality of graph modules using the plurality of markers so as to generate calibration data; nor determining, based on the calibration data, a scale value and an offset value applicable to the NN model; nor determining an optimal value for the scale value or the offset value by performing a quantization simulation of one or more candidates among optimization candidates for the scale value or the offset value for each graph module of the NN model.
Chung teaches adding a plurality of markers to a plurality of graph modules included in a neural network (NN) model, (“In some embodiments, performance data can be collected from probes that provide data on system performance with respect to a specified performance goal”, Paragraph 33) (The probes of Chung correspond to a plurality of markers, which when combined with the data from the graph modules and model of Hosseini, fully teaches the limitation.) and collecting input values and output values of each of the plurality of graph modules using the plurality of markers so as to generate calibration data; (“During the monitoring phase, the instrumented training code can be profiled for communication and computation characteristics. This can be done by using performance analyzing tools relying on known data analytics functions such as probes, and/or software patches (an example of which is discussed with reference to FIG. 4), and changes to run-time control parameters. In some embodiments, data analytic probes are inserted into the program code, providing workload performance profiling statistics/data on the running system.”, Paragraph 49) (The control parameters correspond to input and output values collected as data by probes in the system; the probes corresponding to a plurality of markers. In combination with the plurality of graph modules of Hosseini, the limitation is fully taught.)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the neural networks of Hosseini, with the probes of Chung. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for live tuning, as detailed in Chung. (“During system monitoring, readings from the performance data probes 455 are provided to and received e.g.,
by tuning server 150. These readings can reflect performance statistics such as bandwidth utilization, memory usage, and power/wattage consumed. Using known performance monitoring tools, data collection can also include performance data 450 cataloging system software events 435 and hardware counter events 445”, Paragraph 62)
However, the combination does not explicitly teach determining, based on the calibration data, a scale value and an offset value applicable to the NN model nor determining an optimal value for the scale value or the offset value by performing a quantization simulation of one or more candidates among optimization candidates for the scale value or the offset value for each graph module of the NN model.
Oh teaches determining, based on the calibration data, a scale value and an offset value applicable to the NN model (“applying the element-wise quantization function qw defined as follows… where the boundaries [min,max] are nudged by small amounts so that value 0.0 is exactly representable as an integer zero point, z, after quantization: min ← s ∗ min/s , max←max+s∗min/s−min. The function · rounds to the nearest integer. The scale, s, indicates the step size of quantization.”, pg. 2, second-to-last paragraph; See also Equations 1, 2, and 3 on pg. 2) (The variable z corresponds to an offset value, with the scale variable corresponding to a scale value.) and determining an optimal value for the scale value or the offset value by performing a quantization simulation of one or more candidates among optimization candidates for the scale value or the offset value for each graph module of the NN model. (“In detail, QAT [quantization-aware training] is to simulate quantization during the time of training from scratch or fine-tuning, enabling to apply the effect of quantization simultaneously… In detail, weights are first channel-wise shifted by si (Line 7 in Algorithm 1), followed by uniform affine layer-wise quantization (FP32→UINT8) and dequantization (UINT8→FP32) process (Line8-10) as emulating inference-time quantization, and the n channel-wise shifted in the reversed direction (Line 11). Lastly, the quantization error can be measured in l2-norm between original weights and WES-coupled fake quantized weights so that bigger quantization error can be more penalized (Line 12).”, pg. 3, left column, under “Iterative search for an optimal total range”; See also Algorithm 1 on pg. 3) (These processes are done implicitly to find an optimal shift scale and offset.)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the neural networks of Hosseini, as modified by Chung, with the scale and offset values of Oh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for a quantized model with a low quantization error. (“distribution of the channel with a relatively narrow range is stretched out so it comes to having low quantization error after WES-coupled quantization. The process so far is performed in compile time, and a resultant quantized model is saved to a binary file.”, pg. 3, right column, first paragraph)
Regarding claim 20 Hosseini teaches A non-volatile computer-readable storage medium storing instructions (While Hosseini’s disclose never explicitly cite a non-transitory computer readable storage medium, one would implicitly need one to distribute the method of Hosseini.) the instructions, when executed by one or more processors, causing the one or more processors to perform steps comprising: a plurality of graph modules included in a neural network (NN) model in a form of a directed acyclic graph (DAG); (“The objective of this Section is hence to formulate a directed acyclic graph composed of a few layers where all Input nodes are connected via an equal number of paths to all Output nodes of the graph.”, pg. 4, left column, under “A. Problem Statement and Formulation”) (Q DCNN in the form of a DAG is formulated based on connecting the input and output nodes of the graph, the nodes corresponding to graph modules.)
However, Hosseini does not explicitly teach adding a plurality of markers to a plurality of graph modules included in a neural network (NN) model in a form of a directed acyclic graph (DAG), nor collecting input values and output values of each of the plurality of graph modules using the plurality of markers so as to generate calibration data; nor determining, based on the calibration data, a scale value and an offset value applicable to the NN model; nor determining an optimal value for the scale value or the offset value by performing a quantization simulation of one or more candidates among optimization candidates for the scale value or the offset value for each graph module of the NN model.
Chung teaches adding a plurality of markers to a plurality of graph modules included in a neural network (NN) model, (“In some embodiments, performance data can be collected from probes that provide data on system performance with respect to a specified performance goal”, Paragraph 33) (The probes of Chung correspond to a plurality of markers, which when combined with the data from the graph modules and model of Hosseini, fully teaches the limitation.) and collecting input values and output values of each of the plurality of graph modules using the plurality of markers so as to generate calibration data; (“During the monitoring phase, the instrumented training code can be profiled for communication and computation characteristics. This can be done by using performance analyzing tools relying on known data analytics functions such as probes, and/or software patches (an example of which is discussed with reference to FIG. 4), and changes to run-time control parameters. In some embodiments, data analytic probes are inserted into the program code, providing workload performance profiling statistics/data on the running system.”, Paragraph 49) (The control parameters correspond to input and output values collected as data by probes in the system; the probes corresponding to a plurality of markers. In combination with the plurality of graph modules of Hosseini, the limitation is fully taught.)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the neural networks of Hosseini, with the probes of Chung. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for live tuning, as detailed in Chung. (“During system monitoring, readings from the performance data probes 455 are provided to and received e.g.,
by tuning server 150. These readings can reflect performance statistics such as bandwidth utilization, memory usage, and power/wattage consumed. Using known performance monitoring tools, data collection can also include performance data 450 cataloging system software events 435 and hardware counter events 445”, Paragraph 62)
However, the combination does not explicitly teach determining, based on the calibration data, a scale value and an offset value applicable to the NN model; nor determining an optimal value for the scale value or the offset value by performing a quantization simulation of one or more candidates among optimization candidates for the scale value or the offset value for each graph module of the NN model.
Oh teaches determining, based on the calibration data, a scale value and an offset value applicable to the NN model (“applying the element-wise quantization function qw defined as follows… where the boundaries [min,max] are nudged by small amounts so that value 0.0 is exactly representable as an integer zero point, z, after quantization: min ← s ∗ min/s , max←max+s∗min/s−min. The function · rounds to the nearest integer. The scale, s, indicates the step size of quantization.”, pg. 2, second-to-last paragraph; See also Equations 1, 2, and 3 on pg. 2) (The variable z corresponds to an offset value, with the scale variable corresponding to a scale value.) and determining an optimal value for the scale value or the offset value by performing a quantization simulation of one or more candidates among optimization candidates for the scale value or the offset value for each graph module of the NN model. (“In detail, QAT [quantization-aware training] is to simulate quantization during the time of training from scratch or fine-tuning, enabling to apply the effect of quantization simultaneously… In detail, weights are first channel-wise shifted by si (Line 7 in Algorithm 1), followed by uniform affine layer-wise quantization (FP32→UINT8) and dequantization (UINT8→FP32) process (Line8-10) as emulating inference-time quantization, and the n channel-wise shifted in the reversed direction (Line 11). Lastly, the quantization error can be measured in l2-norm between original weights and WES-coupled fake quantized weights so that bigger quantization error can be more penalized (Line 12).”, pg. 3, left column, under “Iterative search for an optimal total range”; See also Algorithm 1 on pg. 3) (These processes are done implicitly to find an optimal shift scale and offset.)
Therefore, it would have been considered obvious to one of ordinary skill in the art,
prior to the current application’s filing date, to combine the neural networks of Hosseini, as modified by Chung, with the scale and offset values of Oh. One would be motivated to combine the teachings, prior to the filing date of the current application, as this allows for a quantized model with a low quantization error. (“distribution of the channel with a relatively narrow range is stretched out so it comes to having low quantization error after WES-coupled quantization. The process so far is performed in compile time, and a resultant quantized model is saved to a binary file.”, pg. 3, right column, first paragraph)
Regarding claim 2, Hosseini, as modified by Chung and Oh, teaches the method of claim 1, further comprising: determining the optimal value for the scale value or the offset value from a first graph module to a last graph module of the plurality of graph modules, based on a relationship between graph modules included in the second NN model. (“in 8-bit quantized TFlite models, both the weight tensor F and the input fmap X are in 8-bit precision and the convolution operation between the two operand tensors followed by a quantizing rectified linear unit (ReLU) should result in another tensor Y that has 8-bit values. In order to do so, for a given layer with 8-bit quantized weights F and all possible 8-bit quantized values of X that are resulted from the dataset, an 8-bit offset scalar α and a 32-bit scaling factor scalar β are calculated and fine-tuned based on the min/max values of Y during the training.”, pg. 10, left column, under “1) Quantization” (Hosseini))
Regarding claim 3, Hosseini, as modified by Chung and Oh, teaches the method of claim 1, further comprising: determining the optimal value for the offset value for the plurality of graph modules included in the second NN model, and then determining the optimal value for the scale value for the second NN model reflecting the optimal value for the offset value for each of the plurality of graph modules. (“Next, a scale (sw) and a zero point (zw) are derived by uniform affine layer-wise quantization (Fig.2(c)). The scale specifies the step size of the quantizer and the zero point is an integer mapped to the floating point zero value.”, pg. 2, right column, fourth paragraph (Oh))
Regarding claim 7, Hosseini, as modified by Chung and Oh, teaches the method of claim 1, wherein the optimization candidates for the scale value include the scale value and the optimization candidates for the offset value include the offset value. (“the boundaries [min, max] are nudged by small amounts so that value 0.0 is exactly representable as an integer zero point, z, after quantization: min ← s ∗ min/s, max←max+s∗min/s−min. The function · rounds to the nearest integer. The scale, s, indicates the step size of quantization… Based on our experiments, 4 bits are enough to represent the candidates of channel-wise shift scales,”, pg. 2, right column, fourth paragraph; (Oh))
Regarding claim 8, Hosseini, as modified by Chung and Oh, teaches the methods of claim 1 wherein the scale value is generated for an input parameter, an output parameter, and a weight parameter of the plurality of graph modules, respectively, and wherein the offset value is generated for the input parameter and the output parameter of the plurality of graph modules, respectively (“In addition, the channel-wise scales, (maxi − mini)/(2bits − 1), in depth-wise convolution are sometimes quite small so that as quantizing the bias by the multiplication of the scale of input and scale of weight… An integer format of a channel-wise shift scale factor Si to rescale the original range for i-th output channel tensor into the total range is”, pg. 2, right column, under “2. Bias overflow”; pg. 3, left column, under “Initialization of a total range”; See also Algorithm 2 on pg. 4 (Oh)) (Algorithm 2 indicates a scale and offset for the weights input, and output on lines 5 and 6.)
Regarding claim 9, Hosseini, as modified by Chung and Oh, teaches the methods of claim 1 wherein the scale value and the offset value are obtained by an equation below,
PNG
media_image1.png
47
286
media_image1.png
Greyscale
where max means a maximum value among the input values and output values collected for the calibration data, min means a minimum value among the input values and output values collected for calibration data, and bitwidth means a target quantization bitwidth. (See Equation 2 on pg. 2 of Oh’s disclosure, or down below.)
PNG
media_image2.png
100
319
media_image2.png
Greyscale
Oh’s Equations 1-3
Regarding claim 14, Hosseini, as modified by Chung and Oh, teaches the method of claim 13, further comprising: determining the optimal value for the scale value or the offset value from a first graph module to a last graph module of the plurality of graph modules, based on a connection relationship between graph modules included in the NN model. (“in 8-bit quantized TFlite models, both the weight tensor F and the input fmap X are in 8-bit precision and the convolution operation between the two operand tensors followed by a quantizing rectified linear unit (ReLU) should result in another tensor Y that has 8-bit values. In order to do so, for a given layer with 8-bit quantized weights F and all possible 8-bit quantized values of X that are resulted from the dataset, an 8-bit offset scalar α and a 32-bit scaling factor scalar β are calculated and fine-tuned based on the min/max values of Y during the training.”, pg. 10, left column, under “1) Quantization” (Hosseini))
Regarding claim 15, Hosseini, as modified by Chung and Oh, teaches the method of claim 13, further comprising: determining the optimal value for the offset value for the plurality of graph modules included in the NN model, and then determining the optimal value for the scale value for the NN model reflecting the optimal value for the offset value for each of the plurality of graph modules. (“Next, a scale (sw) and a zero point (zw) are derived by uniform affine layer-wise quantization (Fig.2(c)). The scale specifies the step size of the quantizer and the zero point is an integer mapped to the floating point zero value.”, pg. 2, right column, fourth paragraph (Oh))
Claim(s) 4, 5, 16, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Hosseini, in view of Chung, in further view of Oh, and in further view of Wu et al. (Herein referred to as Wu) (EasyQuant: Post-training Quantization via Scale Optimization)
Regarding claims 4 and 16, Hosseini, as modified by Chung and Oh, teaches the methods of claims 1 and 13 respectively, but does not explicitly teach calculating a cosine similarity of a first computation result value of each graph module of the second NN model and a second computation result value of performing the quantization simulation using each candidate included in the optimization candidates, and selecting the candidate with a highest cosine similarity value included in the optimization candidates as the optimal value.
Wu teaches calculating a cosine similarity of a first computation result value of each graph module of the second NN model (the NN model in the case for claim 16) (“From Equation (2), it can be seen that the scale factors actually control the thresholds clipped in quantization process, which affects the cosine similarity of convolutional results between original output feature map (Ol) and quantization inference feature map ( ˆOl) to a great extent.”, pg. 5, above “3.2 Scale Optimization”; See Equation 2 on pg. 5) and a second computation result value of performing the quantization simulation using each candidate included in the optimization candidates, (See Equation 3 on pg. 5) and selecting the candidate with a highest cosine similarity value included in the optimization candidates as the optimal value. (“we first formulate the quantized convolutional process as an optimization problem target at maximizing the cosine similarity between FP32 and INT8 outputs.”, pg. 2, second paragraph)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine the neural networks of Hosseini, as modified by Chung and Oh, with the cosine similarity calculation of Wu. One would have been motivated to combine the teachings prior to the application’s filing date, as this allows for the optimization of the scales of weights and activations. (“In this paper, we introduce an efficient and simple post-training quantization method via effectively optimizing the scales of weights and activations. The proposed scale optimization method is named EasyQuant (EQ).”, pg. 2, second paragraph)
Regarding claims 5 and 17, Hosseini, as modified by Chung, Oh, and Wu teaches the methods of claims 4 and 16 respectively, wherein the cosine similarity is calculated after performing dequantization on a result of the quantization simulation using each of the optimization candidates. (“Therefore, the whole linear quantization forward convolution and dequant operation in the l-th layer can be described as [See Equation 2 on pg. 5] where ∗ denotes convolution operation.”, pg. 5 (Wu))
Claim(s) 6, 10 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Hosseini, in view of Chung, in further view of Oh, and in further view of Subramanian et al. (Herein referred to as Subramanian) (Practical Quantization in PyTorch)
Regarding claims 6 and 18, Hosseini, as modified by Chung and Oh, teaches the methods of claims 1 and 13 respectively, but does not explicitly teach wherein the optimization candidates for the scale value are selected according to a predetermined number within a certain range comprising the scale value, and wherein the optimization candidates for the offset value are selected from a predetermined number within a certain range comprising the offset value.
Subramanian teaches the optimization candidates for the scale value are selected according to a predetermined number within a certain range comprising the scale value, and wherein the optimization candidates for the offset value are selected from a predetermined number within a certain range comprising the offset value. (“The mapping function is parameterized by the scaling factor S and zero-point Z.
PNG
media_image3.png
46
104
media_image3.png
Greyscale
S is simply the ratio of the input range to the output range where [a, cB] is the clipping range of the input, i.e. the boundaries of permissible inputs. [aq, Bq] is the range in quantized output space that it is mapped to...
PNG
media_image4.png
34
120
media_image4.png
Greyscale
Z acts as a bias to ensure that a 0 in the input space maps perfectly to a 0 in the quantized space.”, pg. 2) (In the Subramanian disclosure, the scaling factor and zero-point, corresponding to a scale and offset value, implicitly have certain range of preselected number which can be the optimization candidates.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine the neural networks of Hosseini, as modified by Chung and Oh, with the specific optimization candidates and the ranges of values of Subramanian. One would have been motivated to combine the teachings prior to the application’s filing date, as this allows for calibration. (“The process of choosing the input clipping range is known as calibration. The simplest technique (also the default in PyTorch) is to record the running minimum and maximum values and assign them to
PNG
media_image5.png
8
11
media_image5.png
Greyscale
and B. TensorRT also uses entropy minimization (KL divergence), mean-square-error minimization, or percentiles of the input range.”, pg. 2)
Regarding claim 10, Hosseini, as modified by Chung and Oh, teaches the methos of claim 1, but does not teach a convolution operation in the second NN model is expressed as:
PNG
media_image6.png
53
498
media_image6.png
Greyscale
where feature_infp represents an input feature map parameter in a form of floating-point, weightfp represents a weight parameter in a form of floating-point, of represents the offset value for an input feature map, sf represents the scale value for the input feature map, sw represents the scale value for a weight, and ⌊ ⌋ represents round and clip operations.
Subramanian teaches a convolution operation similar in purpose to the second NN model as expressed above. Subramanian’s equation is shown below:
PNG
media_image7.png
28
198
media_image7.png
Greyscale
where the quantization operation is clipped and rounded, r corresponds to an input feature map, and S and Z are quantization parameters, such as offset and scaling factor. While not the same equation, the idea of the equation is the same as the claimed equation, as both equations attempt to use asymmetric/affine quantization for the activations/inputs and symmetric quantization for the weights, thereby teaching the claim’s limitations.
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine the neural networks of Hosseini, as modified by Chung and Oh, with the affine quantization of Subramanian. One would have been motivated to combine the teachings, prior to the application filing date, as “Affine quantization leads to more computationally expensive inference when used for weight tensors [3].” (pg. 2)
Claim(s) 11 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Hosseini, in view of Chung, in further view of Oh, and in further view of Chen et al. (Herein referred to as Chen) (Binarized Neural Architecture Search for Efficient Object Recognition)
Regarding claim 11, Hosseini, as modified by Chung and Oh, teaches the method of claim 1, but does not explicitly teach generating, based on optimal values of the scale value and the offset value, a third neural network (NN) model comprising a quantized weight parameter in a form of integer, based on the second NN model.
Chen teaches generating, based on optimal values of the scale value and the offset value, a third neural network (NN) model comprising a quantized weight parameter in a form of integer, based on the second NN model. (“Denote by Ltrain and Lval the training loss and the validation loss, respectively. Both losses are determined by not only the architecture α but also the binarized weights ˆX in the network. The goal for the warm-up step is to find ˆX∗ and α∗ that minimize the validation loss Lval(ˆX∗,α∗), where the weights ˆX∗ associated with the architecture are obtained by minimizing the training loss…”, pg. 7, left column, bottom paragraph) (In BNAS, the training data is used to determine a quantized weight parameter for the network architecture. In combination with the training data of the combination, the limitation is taught.)
Therefore, it would have been considered obvious to one of ordinary skill in the art, prior to the current application’s filing date, to combine the neural networks and data of Hosseini, as modified by Chung and Oh, with the training of a network architecture of Chen. One would have been motivated to combine the teachings, prior to the filing date of the current application, as this allows for the finding of a weight that minimizes the loss value. (“The goal for the warm-up step is to find ˆX∗ and α∗ that minimize the validation loss Lval(ˆX∗,α∗), where the weights ˆX∗ associated with the architecture are obtained by minimizing the training loss.”, pg. 7, left column, bottom paragraph)
Regarding claim 12, Hosseini, as modified by Chung, Oh, and Chen teaches the method of claim 11, wherein a convolution operation in the third NN model is expressed as:
PNG
media_image8.png
35
314
media_image8.png
Greyscale
where feature_outint represents an output feature map parameter in a form of integer, feature_inint represents an input feature map parameter in a form of integer, and weightint represents a weight parameter in a form of integer. (See Equation 13 on pg. 9 of Chen’s disclosure, or See the Equation below)
PNG
media_image9.png
40
119
media_image9.png
Greyscale
(Fhl+1 corresponds to an output feature map parameter, Fgl corresponds to an input feature map parameter, and Xil corresponds to a weight parameter, all in the form of integers.)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Tyler E Iles whose telephone number is (571)272-5442. The examiner can normally be reached 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.E.I./Patent Examiner, Art Unit 2122
/KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122