DETAILED ACTION
Status of Claims
Claim(s) 1-4, 7-13, and 16-24 are pending and are examined herein.
Claim(s) 8-10, 17-18, and 24 have been Amended. Claim(s) 5-6 and 14-15 were previously Cancelled.
Claim(s) 1-4, 7-13, and 16-24 are rejected under 35 U.S.C. § 103.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/12/2026 has been entered.
Response to Arguments
Applicant's arguments, with respect to the rejection under 35 U.S.C. § 103 filed on 06/12/2026 (see remarks Pp. 11-23) have been fully considered but are moot in view of the new ground of rejection.
The examiner refers to the updated rejection under 35 U.S.C. § 103 for more details.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION. —The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 10-13 and 16-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, for pre-AIA the applicant regards as the invention.
Regarding Currently Amended Claim 10, Claim 10 recites the limitation “the iterative training process comprising, for at least the first feature map, extracting scalar samples from the first feature map, generating a flattened and detached array of output feature map values from the scalar samples, and performing a re-estimation operation using the array to adjust quantization boundaries of the first LUT” lines 30-34, which renders the claim indefinite. The terms “flattened” and “detached array” are not defined in the claim, and paragraph [0064] of the specification merely repeats the claim language without providing any definition or example to clarify these terms. It is unclear what is meant by “flattened” in the context of feature map values, whether it refers to a specific data structure transformation or a separate processing step. Similarly, it is unclear whether “detached” describes the way of extracting scalar samples from the feature map or refers to a separate operation that is not clearly defined in the applicant disclosure. As a result, the metes and bounds of the claimed invention are unclear, and one of ordinary skill in the art would not be reasonably apprised of the scope of the claimed invention.
For examination purposes, the Examiner broadly interprets the “flattened and detached array” as the results of extracting scalar sample from the feature map and performing data transformation to the extracted representation.
It is noted that a substantially similar limitation raising similar indefiniteness issues was previously recited in dependent claim 21 and addressed in the Final Office Action mailed on 02/21/2025.
Regarding Dependent Claims 11-13 and 16-18, these claims depend from a rejected claim 10 and therefore inherit the deficiencies of the respective parent claim.
In view of the above, Examiner respectfully requests that Applicant thoroughly review the claims for compliance with the requirements set forth under 35 U.S.C. § 112.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-2, 4, 7-9, 19-20, and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Leibovich et al., (Pub. No.: US 20190102673 A1) in view of Fu et al., (NPL: "Don’t waste your bits! squeeze activations and gradients for deep neural networks via tinyscript." (2020)), and further in view of Zhang et al., (NPL: "Lq-nets: Learned quantization for highly accurate and compact deep neural networks." (2018)). Hereinafter, the combination of Leibovich, Fu, and Zhang.
Regarding Previously Presented Claim 1,
Leibovich discloses the following:
A method comprising: (Leibovich, [0193] “a method 2300 to provide online activation compression with K-means, in accordance with an embodiment. One or more operations of method 2300 may be performed by logic (e.g., incorporated in a processor, GPU, GPGPU, etc.) including those discussed with reference to the other figures herein.” [0196] “Example 15 includes a method comprising: compressing one or more activation functions for a convolutional network based on non-uniform quantization, wherein the non-uniform quantization for each layer of the convolutional network is performed offline, wherein an activation function for a specific layer of the convolutional network is quantized during runtime.”)
processing, using at least one processor of an electronic device, input data using a first layer of a neural network to generate a first feature map and a second layer of the neural network to generate a second feature map, wherein the first layer and the second layer are consecutive layers of the neural network; (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0182] “As discussed above, some embodiments relate to online activation compression with K-means. In one embodiment, the memory footprint of activations is reduced by quantizing the activation values using non-uniform quantization.”) [Examiner’s Note: Leibovich teaches sequential processing where each layer generates activations (feature maps) that serves as input to the next consecutive layer, which corresponds to the first/second consecutive layers.]
representing, using the at least one processor, first feature data of the first feature map generated by the first layer of the neural network using first index values, the first index values corresponding to multiple records of a first look up table (LUT), each of the multiple records of the first LUT comprising a representation of a first quantization level of a nonuniform distribution of quantization levels of the first feature map; (Leibovich, [0184] “Some embodiments compute a non-linear quantization scheme offline (e.g., for each layer activations). Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I.” [0186] “These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0188]-[0189] “The algorithm extracts K centers of a given data, ... the distribution of each layer is determined offline, as well as determining a non-uniform quantization scheme by K-Means algorithm. Thereafter, the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online).”) [Examiner’s Note: Leibovich teaches replacing feature values (activation values) with indexes pointing to LUT records where each record holds non-uniform quantization level (k-mean center) based on the layer’s activation values distribution.]
representing, using the at least one processor, second feature data of the second feature map generated by the second layer of the neural network using second index values, the second index values corresponding to multiple records of a second LUT, each of the multiple records of the second LUT comprising a representation of a second quantization level of the non-uniform distribution of quantization levels of the second feature map, wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values; (Leibovich, [0184] “Some embodiments compute a non-linear quantization scheme offline (e.g., for each layer activations). Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I.” [0186] “These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0188]-[0189] “The algorithm extracts K centers of a given data, ... the distribution of each layer is determined offline, as well as determining a non-uniform quantization scheme by K-Means algorithm. Thereafter, the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online).” [0191] “FIGS. 22A, 22B, and 22C present three examples for activations of different layers, according to some embodiments (e.g., which may be derived from Googlenet). One can notice the different scaling and appearance. In addition, the high density around 0 and especially many zero values (due to preceding ReLU layer which zeroed smaller than zero values). The higher density around small values suggests that there might be smarter quantization than a simple uniform one as utilized by one or more embodiments.” [0194] “An operation 2304 determines whether the convolutional network is offline and if it is operation 2306 performs non-uniform quantization as discussed above. After operation 2306, operation 2308 determines whether the convolutional network is online, and if it is online operation 2310 quantizes the activation function for a specific layer of convolutional network, and continues with the next layer during runtime.” Further see [0196].) [Examiner’s Note: Leibovich teaches that different layers have different activation distribution (FIGS. 22A-22C show different distributions), and each layer gets its own separately determined non-uniform quantization scheme. It is noted that the broadest reasonable interpretation in light of the specification, the limitation reciting “wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values” is broadly interpreted as suggesting that different feature data produced by different layers leads to different quantization levels.]
storing, using the at least one processor, the first index values and the second index values in at least one memory of the electronic device; (Leibovich, [0195] “Example 2 includes the apparatus of example 1, further comprising memory to store an index corresponding to the quantized activation function during runtime. Example 3 includes the apparatus of example 2, wherein the index is to be stored in in a lookup table.” [0184]-[0185] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V.”) [Examiner’s Note: Leibovich teaches storing index values in memory of the computing device. Further see [0031]-[0032].]
regenerating, using the at least one processor, the second feature data of the second feature map by cross-referencing the second index values with the second LUT; (Leibovich, [0184]-[0186] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V. These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0189] “the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online). Such compression techniques reduce the memory bandwidth (thus reduces power consumption) and improve performance. Notably, and in contrast to the GEMMLOWP technique, one or more embodiments discussed herein do not target lowering bit width for compute, but only reduce the memory bandwidth. Then, the compressed activations are decompressed to the original bit width to perform the computation.” [0192] “decompression (F:I→V). While decompression may be simply implemented by a LUT (Look Up Table), compression may still need a sequence of comparators (e.g., at most log2N in a binary search implementation).” [0195] “Example 7 includes the apparatus of example 1, wherein the quantized activation function is to be decompressed during runtime.”) [Examiner’s Note: Leibovich teaches dequantizing/decompressing the quantized activations using the stored index values. Further see [0031]-[0032].] and
processing, using the at least one processor, the regenerated second feature data using a third layer of the neural network; (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0184]-[0186] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V. These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.”) [Examiner’s Note: Leibovich teaches that decompressed activations serve as input for the subsequent layer in the sequential forward pass.]
wherein the representation of the first quantization level of the non-uniform distribution of quantization levels of the first feature map comprised in the multiple records of the first LUT and the representation of the second quantization level of the non-uniform distribution of quantization levels of the second feature map comprised in the multiple records of the second LUT... (Leibovich, [0195] “Example 1 includes an apparatus comprising: logic, at least a portion of which is in hardware, to compress one or more activation functions for a convolutional network based on non-uniform quantization, wherein the non-uniform quantization for each layer of the convolutional network is to be performed offline, wherein an activation function for a specific layer of the convolutional network is to be quantized during runtime. Example 2 includes the apparatus of example 1, further comprising memory to store an index corresponding to the quantized activation function during runtime. Example 3 includes the apparatus of example 2, wherein the index is to be stored in in a lookup table. Example 4 includes the apparatus of example 1, wherein compression of the one or more activation functions is to reduce memory bandwidth usage for processing information between layers of the convolutional network.”)
As explained above, While Leibovich teaches the iterative training process of the neural network and the repeated quantization loop (see [0168] and fig. 23), Leibovich does not appear to explicitly suggest that the quantization levels comprised in the multiple records of the first and second LUTs “are estimated in an iterative training process based on repeated extraction and re-estimation until stable LUTs are obtained.”
However, Leibovich in view of Fu teaches the following:
representing, ..., first feature data of the first feature map generated by the first layer of the neural network using first index values, ... representing, ..., second feature data of the second feature map generated by the second layer of the neural network using second index values, (Fu, [Abstract] “In this work, we introduce TINYSCRIPT, which applies a non-uniform quantization algorithm to both activations and gradients. TINYSCRIPT models the original values by a family of Weibull distributions and searches for “quantization knobs” that minimize quantization variance.” [P. 1, Section: 1] “For one thing, activations (a.k.a. feature maps) in forward propagation need to be stored for gradient computation in backward propagation, which leads to large memory footprints.” [P. 3, Section: 3.1] “Figure 2 illustrates the quantization for activations and gradients with an example of fully connected layer. In forward pass, the l-th layer takes as input the activations v(l−1) from previous layer and generates v(l) to the next layer. Then v(l−1) is quantized to reduce memory usage.”) the first index values corresponding to multiple records of a first look up table (LUT), each of the multiple records of the first LUT comprising a representation of a first quantization level of a nonuniform distribution of quantization levels of the first feature map; ... the second index values corresponding to multiple records of a second LUT, each of the multiple records of the second LUT comprising a representation of a second quantization level of the non-uniform distribution of quantization levels of the second feature map, wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values; (Fu, [P. 4, Section: 3.2] “As a result, we determine to assume a prior on p(x). In each iteration, we compute statistics of the tensor, such as mean and standard deviation, and then estimate p(x) with the prior.” [P. 5, Section: 3.3] “Algorithm 2 shows how we compute the quantization points on-the-fly. Optimal points are pre-computed (line 1-8). Motivated by three-sigma rule of thumb, we set M = 3σ heuristically (line 5). For tensor x, we first fit distributions by Algorithm 1, then we lookup the optimal points and scale by λ (line 9-11). Finally the quantization points are reorganized in ascending order (line 16).” [P. 5, Section: 3.4] “Algorithm 2: lines 1-11 “9: function QUANCALC(dist
k
,
λ
, table
Q
) 10: Lookup
s
^
*
←
Q
[
k
,
n
/
2
]
return
s
^
*
∗ λ” [P. 5, Section: 3.4] “Second, we apply the bucketing strategy where each tensor is sliced into buckets and each bucket is quantized individually. Such a bucketing technique is widely used to decrease the quantization variance...” [Pp. 8-9, Section: 5.3] “Finally, we compare the compression performance of TINYSCRIPT and Vanilla with the results in Figure 8. Although the existing uniform quantization conveys a possibility to train DNNs with a small level n so that memory consumption and communication cost can be reduced, our non-uniform quantization further decreases n to one-half or even one-fourth.” [P. 9, Section: 6] “This work proposes TINYSCRIPT, a non-uniform quantization framework for DNNs. TINYSCRIPT models the values with a family of Weibull distributions and leverages the distribution property to achieve minimum quantization variance. Empirical results show that TINYSCRIPT obtains lower quantization variance than the uniform-based mechanism, and therefore achieves a higher compression rate with model accuracy on par with full precision training.” Further see Algorithm 1: Lines 14-17: Compute positive and negative statistics for tensor x)
wherein the representation of the first quantization level of the non-uniform distribution of quantization levels of the first feature map comprised in the multiple records of the first LUT and the representation of the second quantization level of the non-uniform distribution of quantization levels of the second feature map comprised in the multiple records of the second LUT are estimated in an iterative training process based on repeated extraction and re-estimation until stable LUTs are obtained. (Fu, [Abstract] “In this work, we introduce TINYSCRIPT, which applies a non-uniform quantization algorithm to both activations and gradients. TINYSCRIPT models the original values by a family of Weibull distributions and searches for “quantization knobs” that minimize quantization variance.” [P. 3, Section: 3.1] “Figure 2 illustrates the quantization for activations and gradients with an example of fully connected layer. In forward pass, the l-th layer takes as input the activations v(l−1) from previous layer and generates v(l) to the next layer. Then v(l−1) is quantized to reduce memory usage.” [P. 4, Section: 3.1] “Given the number of quantization levels n, an ideal quantization mechanism is expected to minimize Equation1. There are two critic issues: (1) a good estimation for the distribution p(x), and (2) a method to search quantization points s for a given p(x). In the rest of the paper, we assume tensors to be quantized are flattened as vectors and ||·|| denotes the L2 norm for simplicity.” [P. 4, Section: 3.2] “In each iteration, we compute statistics of the tensor, such as mean and standard deviation, and then estimate p(x) with the prior.” [P. 5, Section: 3.3] “A general method of searching optimal quantization points can be: (1) randomly initialize ˆs, e.g., uniform quantization points; (2) iteratively update ˆ s via gradient descent by Equation 2 until converged. Algorithm 2 shows how we compute the quantization points on-the-fly. Optimal points are pre-computed (line 1-8). Motivated by three-sigma rule of thumb, we set M = 3σ heuristically (line 5). For tensor x, we first fit distributions by Algorithm 1, then we lookup the optimal points and scale by λ (line 9-11). Finally the quantization points are reorganized in ascending order (line 16).” [P. 5, Section: 3.5] “Second, we apply the bucketing strategy where each tensor is sliced into buckets and each bucket is quantized individually. Such a bucketing technique is widely used to decrease the quantization variance ... We finish this section by summarizing the complexity of TINYSCRIPT. To quantize a tensor, we need to (1) compute statistics of each bucket, (2) fit Weibull distributions, (3) calculate optimal points, and finally (4) perform the quantization.” [P. 8, Section: 5.3] “We compare the convergence of TINYSCRIPT, uniform based quantization (Vanilla) 4, and full precision (FP32) over Cifar10 and ImageNet datasets. Figure 7 presents the convergence curves and Table 1, 2 list the best performances”...etc.) [Examiner’s Note: Fu disclose the non-uniform quantization of activations and gradients of neural network layers using quantization points of lookup table (Q), the paper describes the process of quantization involving iterative extraction of feature statistics and iterative computing and recomputing of non-uniform quantization points until convergence. The quantization points represent the quantization levels from the fitted distribution and the entries in the pre-computed table stores the set of optimal quantization points of a given distribution (i.e., records of the LUT).]
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skilled in the art of machine learning to modify the method of Leibovich to incorporate the TINYSCRIPT quantization methods as taught by Fu. One would have been motivated to make such a combination in order to achieve minimum quantization variance and a higher compression rate with model accuracy on par with full precision training. Doing so would reduce memory consumption and communication cost (Fu [Section: 6]).
As outlined above, while Leibovich implicitly describes and teaches per-layer non-uniform quantization schemes with FIGS. 22A-22C confirming different layers have different distributions, yielding different LUTs with specific levels per layer, Leibovich in view of Fu does not appear to explicitly state: “wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values.”
However, it would have been obvious in view of Zhang. Hereinafter, Zhang, in combination with Leibovich and Fu, teaches the limitation:
wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values; (Zhang, [P. 2, Section: 1] “Moreover, the distributions of weights and activations in different networks and even different network layers may differ a lot. We believe a better quantizer should be made adaptive to the weights and activations to gain more flexibility.” [Pp. 5-7, Section: 3.2] “In Fig. 1 we present the statistical distributions of the weights and activations (after batch normalization (BN) and Rectified Linear Unit (ReLU) layers) in a trained floating-point network. It can be seen that the distributions can be complex and differ across layers, and a uniform quantizer is not optimal for them... To get better network quantizers and improve the accuracy of a quantized network, we propose to jointly train the network and its quantizers. The insight behind is that if the optimizers are learnable and optimized through network training, they can not only minimize the quantization error, but also adapt to the training goal thus improving the final accuracy... In practice, we apply layer-wise quantizers for activations (i.e., one quantizer per layer) and channel-wise quantizers for weights (one quantizer for each conv filter).” [P. 11, Section: 4.2] “Figure 4 presents the weight and activation statistics in two layers of a trained ResNet-20 model before (i.e., the floating-point values) and after quantization using our method... It can be seen that our learned quantizers are not uniform ones and they differ at different layers. Statistical results with more bits can be found in the supplementary material.”) [Examiner’s Note: Zhang clearly describes the adaptive quantization of layers weights and activations. The paper explicitly states that learned quantizers differ across layers, which suggests that feature distribution differ across layers and that each layer’s quantizer is adaptively derived from that layer’s own data distribution. This corresponds to the broadest reasonably interpretation in light of the specification of the limitation suggesting that the difference in quantization levels between layers is based on the difference in the feature data distribution.]
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, and Zhang to incorporate the method for learning the quantizers applies to both network weights and activations with arbitrary-bit precision as taught by Zhang. One would have been motivated to make such a combination in order to reduce DNN model size and improve inference efficiency for practical applications. Doing so would enable fast inference on some resource-constrained devices such as mobile phones (Zhang [Section: 5]).
Regarding Previously Presented Claim 2, the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above, and further teaches:
wherein each of the multiple records of the first LUT includes one index value of the first index values and a single quantization level identified by the one index value. (Leibovich, [0184]-[0188] “the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I ... the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I ... These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table... The algorithm extracts K centers of a given data, and clusters all data points to those K centers, allowing storage of a cluster index per data point, which only requires log2K bits.”) [Examiner’s Notes: Leibovich defines the index values paired with associated k-mean derived quantization values, where each index maps to one quantization value.]
Regarding Original Claim 4, the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above, and further teaches:
wherein the memory of the electronic device comprises on-chip dynamic random access memory (DRAM) or static random access memory (SRAM). (Leibovich, [0036] “The memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, flash memory device, phase-change memory device, or some other memory device having suitable performance to serve as process memory.” Further described in [0064] and [0140].)
Regarding Original Claim 7,
the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above, and further teaches:
wherein the input data is associated with one or more images or videos. (Leibovich, [0152] “the input to a convolution layer can be a multidimensional array of data that defines the various color components of an input image.” [0158] “FIG. 16A-16B illustrate an exemplary convolutional neural network. FIG. 16A illustrates various layers within a CNN. As shown in FIG. 16A, an exemplary CNN used to model image processing can receive input 1602 describing the red, green, and blue (RGB) components of an input image. The input 1602 can be processed by multiple convolutional layers (e.g., first convolutional layer 1604, second convolutional layer 1606). The output from the multiple convolutional layers may optionally be processed by a set of fully connected layers 1608.” Further see [0195].)
Regarding Currently Amended Claim 8,
the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above, and further teaches:
Claim 8 recites substantially similar limitations as claim 1, claim 8 further introduce the concept of processing feature map as input into subsequent layer.
Leibovich in view of Fu teaches: processing, using the at least one processor, the first feature map using the second layer of the neural network to generate the second feature map; regenerating, using the at least one processor, the second feature data of the second feature map by cross-referencing the second index values with the second LUT; and processing, using the at least one processor, the regenerated second feature data. (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0184]-[0186] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V. These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0189] “the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online)... the compressed activations are decompressed to the original bit width to perform the computation.”)
Regarding Currently Amended Claim 9,
the combination of Leibovich, Fu, and Zhang teaches the elements of claim 8 as outlined above, and further teaches:
wherein the regenerated second feature data is processed using the third layer of the neural network or a fourth layer of the neural network. (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0158]-[0160] “FIG. 16A-16B illustrate an exemplary convolutional neural network. FIG. 16A illustrates various layers within a CNN. As shown in FIG. 16A, an exemplary CNN used to model image processing can receive input 1602 describing the red, green, and blue (RGB) components of an input image. The input 1602 can be processed by multiple convolutional layers (e.g., first convolutional layer 1604, second convolutional layer 1606). The output from the multiple convolutional layers may optionally be processed by a set of fully connected layers 1608. Neurons in a fully connected layer have full connections to all activations in the previous layer, as previously described for a feedforward network. The output from the fully connected layers 1608 can be used to generate an output result from the network. The activations within the fully connected layers 1608 can be computed using matrix multiplication instead of convolution. Not all CNN implementations are make use of fully connected layers 1608. For example, in some implementations the second convolutional layer 1606 can generate output for the CNN. The convolutional layers are sparsely connected, which differs from traditional neural network configuration found in the fully connected layers 1608. Traditional neural network layers are fully connected, such that every output unit interacts with every input unit. However, the convolutional layers are sparsely connected because the output of the convolution of a field is input (instead of the respective state value of each of the nodes in the field) to the nodes of the subsequent layer, as illustrated. The kernels associated with the convolutional layers perform convolution operations, the output of which is sent to the next layer. The dimensionality reduction performed within the convolutional layers is one aspect that enables the CNN to scale to process large images.”) [Examiner’s Note: the multiple of neural network layers include the next layer being processed using the output data (feature data) from the previous layer.]
Regarding Previously Presented Claim 19,
The claim recites substantially similar limitation as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 1 is directed to a method, and claim 19 is directed to a non-transitory machine-readable medium containing instructions that when executed cause at least one processor of an electronic device ... .
Leibovich also discloses “In various embodiments, the operations discussed herein, e.g., with reference to FIG. 1 et seq., may be implemented as hardware (e.g., logic circuitry), software, firmware, or combinations thereof, which may be provided as a computer program product, e.g., including one or more tangible (e.g., non-transitory) machine-readable or computer-readable medium having stored thereon instructions (or software procedures) used to program a computer to perform a process discussed herein. The machine-readable medium may include a storage device such as those discussed with respect to FIG. 1 et seq. Additionally, such computer-readable media may be downloaded as a computer program product, wherein the program may be transferred from a remote computer (e.g., a server) to a requesting computer (e.g., a client) by way of data signals provided in a carrier wave or other propagation medium via a communication link (e.g., a bus, a modem, or a network connection).” See [0200]-[0201].
Regarding Previously Presented Claim 20,
The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale.
Regarding Currently Amended Claim 24,
The claim recites substantially similar limitations as corresponding claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale.
Claim(s) 3 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Leibovich, Fu, and Zhang as outlined above and further in view of Folliot et al., (Pub. No.: US 20210012208 A1).
Regarding Previously Presented Claim 3, the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above.
While Leibovich describes the bit precision of the memory, the combination of Leibovich, Fu, and Zhang does not appear to explicitly suggest:
wherein a size of the first LUT and a size of the second LUT each corresponds to a bit precision of the memory of the electronic device.
However, it would have been obvious in view of Folliot. Hereinafter, Folliot, in combination with Leibovich, Fu, and Zhang, teaches the limitation:
wherein a size of the first LUT and a size of the second LUT each corresponds to a bit precision of the memory of the electronic device. (Folliot, [0043] “However, it is preferable, in this case, to limit the number of bits to which the input and output data are quantized so as to limit the size of the lookup table and therefore the memory space required to subsequently store it in the device intended to implement the neural network.” [0105] “Thus, for a quantization to 8 bits, 256 values are obtained for the lookup table TF.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, Zhang, and Folliot, to incorporate the methods for quantize the input and output data of each layer as taught by Folliot. One would have been motivated to make such a combination in order to improve processing speed and to decrease the memory space required to store these data (Folliot [0007]).
Regarding Previously Presented Claim 23,
The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Claim(s) 22 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Leibovich, Fu, and Zhang as described above, and further in view of Hsu et al., (Pub. No.: US 11861452 B1).
Regarding Previously Presented Claim 22,
the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above, and further teaches:
As explained above, while Leibovich teaches the concept of reduce the number of representing bits for activations and implicitly discuss the representation as smaller than the original data, the combination of Leibovich, Fu, and Zhang does not appear to explicitly state:
wherein: the first feature data of the first feature map comprises first data; the first index values and the first LUT comprise second data; and the second data is smaller than the first data.
However, it would have been obvious in view of Hsu. Hereinafter, Hus, in combination with Leibovich, Fu, and Zhang, teaches the limitation:
wherein: the first feature data of the first feature map comprises first data; (Hus, [Col. 5, Lines 45-50] “Method 400 begins with operation 402 receiving, at an input to a softmax layer of a neural network from an intermediate layer of the neural network, a non-normalized output comprising a plurality of intermediate network decision values.”) the first index values and the first LUT comprise second data; (Hus, [Col. 5, Lines 5-10] “In the lookup table according to various embodiments, the index of the lookup table represents the distance between the current input and the maximum possible value of the input.” [Col.5, Lines 63-70] “A corresponding lookup table value is then requested from a lookup table in operation 406 using the difference between the intermediate network decision value and the maximum network decision value for each intermediate network decision value of the plurality of intermediate network decision values.”) and the second data is smaller than the first data. (Hus, [Col. 6, Lines 25-32] “In other embodiments, matching bits values for inputs and outputs to the softmax layer are used (e.g., eight bits, 24 bits, etc.). In other embodiments, with significant reduction in the number of table entry values, the number of output bits can be smaller than the number of input bits.” Further see [Col. 5, Lines 10-20]) Examiner’s Note: the examiner notes the non-normalized output of intermediate network decision values would read on the feature data of the feature map “first data.” The index of the lookup table would correspond to the “second data.” The quantized representation of output values using the index of lookup table is smaller than the input value.]
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, and Zhang before them, to incorporate the method of generating a single compact lookup table for a quantized softmax layer as taught by Hus. One would have been motivated to make such a combination in order to enable improvements to a device by reducing memory resources for softmax operations and further reducing the associated processing resources for softmax operations when compared with similar operations using larger tables or deconstructed index entries (Hus [Col. 2, Lines 10-15]).
Claim(s) 21 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Leibovich, Fu, and Zhang as described above, and further in view of Cai et al., (NPL: "Learning a single tucker decomposition network for lossy image compression with multiple bits-per-pixel rates" (2018)).
Regarding Previously Presented Claim 21, the combination of Leibovich, Fu, and Zhang teaches the elements of claim 1 as outlined above, and further teaches:
As outlined above, the combination of Leibovich, Fu, and Zhang teaches an iterative non-uniform quantization scheme to refine and find the optimal quantization levels for representing feature data using LUT of index values, the combination of Leibovich, Fu, and Zhang does not appear to explicitly suggest:
extracting scalar samples from the first feature map for at least the first layer; estimating a distribution of feature map values for the first layer based on the scalar samples as an array of feature map values performing an estimation operation based on the array of feature map values; and adjusting quantization boundaries of the first LUT based on the estimation operation to obtain a revised LUT.
However, Cai, in combination with Leibovich, Fu, and Zhang, teaches the limitations:
extracting scalar samples from the first feature map for at least the first layer; (Cai, [P. 3, Section: A] “The encoder E(·) first converts an input image xi into a latent feature representation zi = E(xi). Then, the quantizer Q(·) quantizes the features into discrete values ˆzi = Q(zi), which can be losslessly encoded into a bitstream for transmission or storage.” [P. 3, Col. 2, Section: B] “Given the latent image representation
z
i
, we use
{
Y
,
U
(
1
)
,
U
(
2
)
,
U
(
3
)
}
=
T
(
z
i
)
to decompose the features into 3 orthogonal matrices
{
U
(
3
)
}
n
=
1
3
and a core tensor Y, and then quantize the decomposed components to generate bitstream.” [P. 6, Section: B] “Tucker decomposition aims to decompose an N-order tensor
X
∈
R
I
1
×
I
2
×
…
×
I
N
as an affiliation of N orthogonal bases
U
n
∈
R
I
n
×
R
n
n
=
1
N
and the associated core tensor
Y
∈
R
R
1
×
R
2
×
…
×
R
N
,.... With a set of training images, we can easily compute p(|Y|), the probability density function (PDF) of the positive core tensor |Y |... More specifically, we first scan the positive core tensor |Y | in raster order, then utilize the decision boundaries
b
q
0
M
to divide the core tensor into M non-overlapping chunks. For each chunk
C
m
(
m
∈
[
1
,
M
]
)
, instead of using
Y
^
q
m
as the quantized values, we define a new quantizer for symbol
|
Y
i
|
(the i-th element of
|
Y
|
)..) estimating a distribution of feature map values for the first layer based on the scalar samples as an array of feature map values; (Cai, [P. 2, Col.2 2nd Paragraph] “The key component of TDNet is a novel tucker decomposition layer (TDL), which decomposes the latent image representation into a set of projection matrices and a compact core tensor. By changing the rank of core tensor and its quantization levels, we can easily adjust the bpp rate of latent image representation, and thus a single CNN model can be trained to compress and reconstruct images under multiple bpp rates. Besides, we propose an iterative non-uniform quantization strategy to obtain the optimal quantization boundaries based on the distribution of encoding coefficients. A coarse-to-fine training strategy is introduced to train a stable TDNet and reconstruct the decompressed images.” [P. 6, Col. 2, Section: B] “Quantization and de-quantization: Since the core tensor Y has both positive and negative values, we take one bit to represent the sign of the original value. Let |Y| denotes the absolute value of the core tensor. With a set of training images, we can easily compute p(|Y|), the probability density function (PDF) of the positive core tensor |Y|. The optimal quantizer can be solved as follows by minimizing the quantization error: …. See equation (16).” [P. 6, Col. 2, Section: B] “The optimal quantizer can be solved as follows by minimizing the quantization error: ….see equation (16). Given a number M of decision intervals, the optimal quantizer is expected to find the set of decision boundaries
{
b
q
}
0
M
and quantized values
{
Y
^
q
}
1
M
. Solving the partial derivative of Eq.(16), we could have: …. See equation (17), The optimal solutions of Eq.(17) can be easily solved by the Lloyd‘s algorithm [38], outputting the optimal quantizer Q(·) with decision boundaries
{
b
q
}
0
M
and quantized values
{
Y
^
q
}
1
M
.”) [Examiner’s Note: examiner interpreted the scalar samples as an array of feature map values as the obtained and decomposed laten representation. The probability density function of tensor values correspond to the estimating a distribution of feature map values.] performing a estimation operation based on the array of feature map values; (Cai(2), [P. 6, Col. 2, Section: B] “The optimal quantizer can be solved as follows by minimizing the quantization error: ….see equation (16). Given a number M of decision intervals, the optimal quantizer is expected to find the set of decision boundaries
{
b
q
}
0
M
and quantized values
{
Y
^
q
}
1
M
. Solving the partial derivative of Eq.(16), we could have: …. See equation (17), The optimal solutions of Eq.(17) can be easily solved by the Lloyd‘s algorithm [38], outputting the optimal quantizer Q(·) with decision boundaries
{
b
q
}
0
M
and quantized values
{
Y
^
q
}
1
M
.”) and adjusting quantization boundaries of the first LUT based on the estimation operation to obtain a revised LUT. (Cai(2), [P. 6, Col.2, Section: V] “With the initialized TDL, we can use Eq.(4) or Eq. (3) to jointly fine-tune the encoder-TDL-decoder network by minimizing the loss function. The latent image representations at multiple bpp rates will be taken into consideration during the training process. In each training epoch, we first decide which group of ranks and decision boundaries will be used by calculating gˆ = mod(epoch, G), then take this group of desired output ranks and decision boundaries
{
R
1
,
R
2
,
R
3
,
{
b
q
}
0
M
}
g
^
+
1
(ˆg+1) to update the TDL and fine-tune the parameters
{
Ω
,
Θ
,
Π
}
of the whole network. To obtain G groups of optimal decision boundaries
{
{
b
q
}
0
M
}
g
=
1
G
and network parameters
{
Ω
,
Θ
,
Π
}
, an iterative training scheme can be used, i.e., fix the encoder-decoder network to update the TDL decision boundaries
{
{
b
q
}
0
M
}
g
=
1
G
by solving Eq.(17), and fix the TDL to update the network parameters
{
Ω
,
Θ
,
Π
}
. Such an alternative optimization process continues till the loss function in Eq.(4) or Eq. (3) converges. After the TDNet converges, we can use the optimal decision boundaries
{
{
b
q
}
0
M
}
g
=
1
G
and network parameters
{
Ω
,
Θ
,
Π
}
to compress and reconstruct images with different bpp rates. The overall all-in-one training scheme is summarized as Algorithm 2.”) [Examiner’s Note: examiner interpreted adjusting quantization boundaries based on estimation operation as the iterative ranking, updating, fine-tune to obtain the optimal quantization boundaries until loss function converges. The examiner refers to the process described in algorithm 1 and 2 on Page 6.]
Accordingly, it would have been prima facie obvious to one having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, and Zhang, to incorporate the proposed tucker decomposition framework involving an iterative non-uniform quantization scheme to optimize the quantizer, and a coarse-to-fine training strategy as taught by Cai. One would have been motivated to make such a combination in order to achieve the objective of multiple bpp rates with a single network. Doing so would provide highly competitive performance (Cai [VII]).
Claim(s) 10-11, 13, and 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Leibovich et al., (Pub. No.: US 20190102673 A1) in view of Fu et al., (NPL: "Don’t waste your bits! squeeze activations and gradients for deep neural networks via tinyscript." (2020)), further in view of Cai et al., (NPL: "Learning a single tucker decomposition network for lossy image compression with multiple bits-per-pixel rates." (2018)), and further in view of Zhang et al., (NPL: "Lq-nets: Learned quantization for highly accurate and compact deep neural networks." (2018)). Hereinafter, the combination of Leibovich, Fu, Cai, and Zhang.
Regarding Currently Amended Claim 10,
Leibovich discloses the following:
An electronic device comprising: at least one memory configured to store instructions; and at least one processing device configured when executing the instructions to: (Leibovich, [0031]-[0034] “FIG. 1 is a block diagram of a processing system 100, according to an embodiment. In various embodiments the system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single processor desktop system, a multiprocessor workstation system, or a server system having a large number of processors 102 or processor cores 107... the processor 102 includes cache memory 104. Depending on the architecture, the processor 102 can have a single internal cache or multiple levels of internal cache. In some embodiments, the cache memory is shared among various components of the processor 102.”)
process input data using a first layer of a neural network to generate a first feature map and a second layer of the neural network to generate a second feature map, wherein the first layer and the second layer are consecutive layers of the neural network; (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0182] “As discussed above, some embodiments relate to online activation compression with K-means. In one embodiment, the memory footprint of activations is reduced by quantizing the activation values using non-uniform quantization.”) [Examiner’s Note: Leibovich teaches sequential processing where each layer generates activations (feature maps) that serves as input to the next consecutive layer, which corresponds to the first/second consecutive layers.]
represent first feature data of the first feature map generated by the first layer of the neural network using first index values, the first index values corresponding to multiple records of a first look up table (LUT), each of the multiple records of the first LUT comprising a representation of a first quantization level of a non-uniform distribution of quantization levels of the first feature map; (Leibovich, [0184] “Some embodiments compute a non-linear quantization scheme offline (e.g., for each layer activations). Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I.” [0186] “These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0188]-[0189] “The algorithm extracts K centers of a given data, ... the distribution of each layer is determined offline, as well as determining a non-uniform quantization scheme by K-Means algorithm. Thereafter, the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online).”) [Examiner’s Note: Leibovich teaches replacing feature values (activation values) with indexes pointing to LUT records where each record holds non-uniform quantization level (k-mean center) based on the layer’s activation values distribution.]
represent second feature data of the second feature map generated by the second layer of the neural network using second index values, the second index values corresponding to multiple records of a second LUT, each of the multiple records of the second LUT comprising a representation of a second quantization level of the non-uniform distribution of quantization levels of the second feature map, wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values; (Leibovich, [0184] “Some embodiments compute a non-linear quantization scheme offline (e.g., for each layer activations). Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I.” [0186] “These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0188]-[0189] “The algorithm extracts K centers of a given data, ... the distribution of each layer is determined offline, as well as determining a non-uniform quantization scheme by K-Means algorithm. Thereafter, the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online).” [0191] “FIGS. 22A, 22B, and 22C present three examples for activations of different layers, according to some embodiments (e.g., which may be derived from Googlenet). One can notice the different scaling and appearance. In addition, the high density around 0 and especially many zero values (due to preceding ReLU layer which zeroed smaller than zero values). The higher density around small values suggests that there might be smarter quantization than a simple uniform one as utilized by one or more embodiments.” [0194] “An operation 2304 determines whether the convolutional network is offline and if it is operation 2306 performs non-uniform quantization as discussed above. After operation 2306, operation 2308 determines whether the convolutional network is online, and if it is online operation 2310 quantizes the activation function for a specific layer of convolutional network, and continues with the next layer during runtime.” Further see [0196].) [Examiner’s Note: Leibovich teaches that different layers have different activation distribution (FIGS. 22A-22C show different distributions), and each layer gets its own separately determined non-uniform quantization scheme. It is noted that the broadest reasonable interpretation in light of the specification, the limitation reciting “wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values” is broadly interpreted as suggesting that different feature data produced by different layers leads to different quantization levels.]
store the first index values and the second index values in the at least one memory; (Leibovich, [0195] “Example 2 includes the apparatus of example 1, further comprising memory to store an index corresponding to the quantized activation function during runtime. Example 3 includes the apparatus of example 2, wherein the index is to be stored in in a lookup table.” [0184]-[0185] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V.”) [Examiner’s Note: Leibovich teaches storing index values in memory of the computing device. Further see [0031]-[0032].]
regenerate the second feature data of the second feature map by cross-referencing the second index values with the second LUT; (Leibovich, [0184]-[0186] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V. These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0189] “the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online). Such compression techniques reduce the memory bandwidth (thus reduces power consumption) and improve performance. Notably, and in contrast to the GEMMLOWP technique, one or more embodiments discussed herein do not target lowering bit width for compute, but only reduce the memory bandwidth. Then, the compressed activations are decompressed to the original bit width to perform the computation.” [0192] “decompression (F:I→V). While decompression may be simply implemented by a LUT (Look Up Table), compression may still need a sequence of comparators (e.g., at most log2N in a binary search implementation).” [0195] “Example 7 includes the apparatus of example 1, wherein the quantized activation function is to be decompressed during runtime.”) [Examiner’s Note: Leibovich teaches dequantizing/decompressing the quantized activations using the stored index values. Further see [0031]-[0032].] and
process the regenerated second feature data using a third layer of the neural network; (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0184]-[0186] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V. These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.”) [Examiner’s Note: Leibovich teaches that decompressed activations serve as input for the subsequent layer in the sequential forward pass.]
wherein the representation of the first quantization level of the non-uniform distribution of quantization levels of the first feature map comprised in the multiple records of the first LUT and the representation of the second quantization level of the non-uniform distribution of quantization levels of the second feature map comprised in the multiple records of the second LUT... (Leibovich, [0195] “Example 1 includes an apparatus comprising: logic, at least a portion of which is in hardware, to compress one or more activation functions for a convolutional network based on non-uniform quantization, wherein the non-uniform quantization for each layer of the convolutional network is to be performed offline, wherein an activation function for a specific layer of the convolutional network is to be quantized during runtime. Example 2 includes the apparatus of example 1, further comprising memory to store an index corresponding to the quantized activation function during runtime. Example 3 includes the apparatus of example 2, wherein the index is to be stored in in a lookup table. Example 4 includes the apparatus of example 1, wherein compression of the one or more activation functions is to reduce memory bandwidth usage for processing information between layers of the convolutional network.”)
While Leibovich teaches the iterative training process of the neural network and the repeated quantization loop (see [0168] and fig. 23), Leibovich does not appear to explicitly suggest that the quantization levels comprised in the multiple records of the first and second LUTs “are estimated in an iterative training process based on repeated extraction and re-estimation until stable LUTs are obtained.” Further, Leibovich does not appear to explicitly teach:
the iterative training process comprising, for at least the first feature map, extracting scalar samples from the first feature map, generating a flattened and detached array of output feature map values from the scalar samples, and performing a re-estimation operation using the array to adjust quantization boundaries of the first LUT.
However, Leibovich in view of Fu teaches the following:
representing, ..., first feature data of the first feature map generated by the first layer of the neural network using first index values, ... representing, ..., second feature data of the second feature map generated by the second layer of the neural network using second index values, (Fu, [Abstract] “In this work, we introduce TINYSCRIPT, which applies a non-uniform quantization algorithm to both activations and gradients. TINYSCRIPT models the original values by a family of Weibull distributions and searches for “quantization knobs” that minimize quantization variance.” [P. 1, Section: 1] “For one thing, activations (a.k.a. feature maps) in forward propagation need to be stored for gradient computation in backward propagation, which leads to large memory footprints.” [P. 3, Section: 3.1] “Figure 2 illustrates the quantization for activations and gradients with an example of fully connected layer. In forward pass, the l-th layer takes as input the activations v(l−1) from previous layer and generates v(l) to the next layer. Then v(l−1) is quantized to reduce memory usage.”) the first index values corresponding to multiple records of a first look up table (LUT), each of the multiple records of the first LUT comprising a representation of a first quantization level of a nonuniform distribution of quantization levels of the first feature map; ... the second index values corresponding to multiple records of a second LUT, each of the multiple records of the second LUT comprising a representation of a second quantization level of the non-uniform distribution of quantization levels of the second feature map, wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values; (Fu, [P. 4, Section: 3.2] “As a result, we determine to assume a prior on p(x). In each iteration, we compute statistics of the tensor, such as mean and standard deviation, and then estimate p(x) with the prior.” [P. 5, Section: 3.3] “Algorithm 2 shows how we compute the quantization points on-the-fly. Optimal points are pre-computed (line 1-8). Motivated by three-sigma rule of thumb, we set M = 3σ heuristically (line 5). For tensor x, we first fit distributions by Algorithm 1, then we lookup the optimal points and scale by λ (line 9-11). Finally the quantization points are reorganized in ascending order (line 16).” [P. 5, Section: 3.4] “Algorithm 2: lines 1-11 “9: function QUANCALC(dist
k
,
λ
, table
Q
) 10: Lookup
s
^
*
←
Q
[
k
,
n
/
2
]
return
s
^
*
∗ λ” [P. 5, Section: 3.4] “Second, we apply the bucketing strategy where each tensor is sliced into buckets and each bucket is quantized individually. Such a bucketing technique is widely used to decrease the quantization variance...” [Pp. 8-9, Section: 5.3] “Finally, we compare the compression performance of TINYSCRIPT and Vanilla with the results in Figure 8. Although the existing uniform quantization conveys a possibility to train DNNs with a small level n so that memory consumption and communication cost can be reduced, our non-uniform quantization further decreases n to one-half or even one-fourth.” [P. 9, Section: 6] “This work proposes TINYSCRIPT, a non-uniform quantization framework for DNNs. TINYSCRIPT models the values with a family of Weibull distributions and leverages the distribution property to achieve minimum quantization variance. Empirical results show that TINYSCRIPT obtains lower quantization variance than the uniform-based mechanism, and therefore achieves a higher compression rate with model accuracy on par with full precision training.” Further see Algorithm 1: Lines 14-17: Compute positive and negative statistics for tensor x.)
wherein the representation of the first quantization level of the non-uniform distribution of quantization levels of the first feature map comprised in the multiple records of the first LUT and the representation of the second quantization level of the non-uniform distribution of quantization levels of the second feature map comprised in the multiple records of the second LUT are estimated in an iterative training process based on repeated extraction and re-estimation until stable LUTs are obtained. (Fu, [Abstract] “In this work, we introduce TINYSCRIPT, which applies a non-uniform quantization algorithm to both activations and gradients. TINYSCRIPT models the original values by a family of Weibull distributions and searches for “quantization knobs” that minimize quantization variance.” [P. 3, Section: 3.1] “Figure 2 illustrates the quantization for activations and gradients with an example of fully connected layer. In forward pass, the l-th layer takes as input the activations v(l−1) from previous layer and generates v(l) to the next layer. Then v(l−1) is quantized to reduce memory usage.” [P. 4, Section: 3.1] “Given the number of quantization levels n, an ideal quantization mechanism is expected to minimize Equation1. There are two critic issues: (1) a good estimation for the distribution p(x), and (2) a method to search quantization points s for a given p(x). In the rest of the paper, we assume tensors to be quantized are flattened as vectors and ||·|| denotes the L2 norm for simplicity.” [P. 4, Section: 3.2] “In each iteration, we compute statistics of the tensor, such as mean and standard deviation, and then estimate p(x) with the prior.” [P. 5, Section: 3.3] “A general method of searching optimal quantization points can be: (1) randomly initialize ˆs, e.g., uniform quantization points; (2) iteratively update ˆ s via gradient descent by Equation 2 until converged. Algorithm 2 shows how we compute the quantization points on-the-fly. Optimal points are pre-computed (line 1-8). Motivated by three-sigma rule of thumb, we set M = 3σ heuristically (line 5). For tensor x, we first fit distributions by Algorithm 1, then we lookup the optimal points and scale by λ (line 9-11). Finally the quantization points are reorganized in ascending order (line 16).” [P. 5, Section: 3.5] “Second, we apply the bucketing strategy where each tensor is sliced into buckets and each bucket is quantized individually. Such a bucketing technique is widely used to decrease the quantization variance ... We finish this section by summarizing the complexity of TINYSCRIPT. To quantize a tensor, we need to (1) compute statistics of each bucket, (2) fit Weibull distributions, (3) calculate optimal points, and finally (4) perform the quantization.” [P. 8, Section: 5.3] “We compare the convergence of TINYSCRIPT, uniform based quantization (Vanilla) 4, and full precision (FP32) over Cifar10 and ImageNet datasets. Figure 7 presents the convergence curves and Table 1, 2 list the best performances”...etc.) [Examiner’s Note: Fu disclose the non-uniform quantization of activations and gradients of neural network layers using quantization points of lookup table (Q), the paper describes the process of quantization involving iterative extraction of feature statistics and iterative computing and recomputing of non-uniform quantization points until convergence. The quantization points represents the quantization levels from the fitted distribution and the entries in the pre-computed table stores the set of optimal quantization points of a given distribution (i.e., records of the LUT).]
Leibovich and Fu are from the same field of endeavor and their disclosure generally relates to (Neural network quantization).
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skilled in the art of machine learning to modify the method of Leibovich to incorporate the TINYSCRIPT quantization methods as taught by Fu. One would have been motivated to make such a combination in order to achieve minimum quantization variance and a higher compression rate with model accuracy on par with full precision training. Doing so would reduce memory consumption and communication cost (Fu [Section: 6]).
While Fu, in combination with Leibovich, defines the tensor being quantized (i.e., feature map) as being flattened as vectors during iterative training process including search/obtaining feature statistics and estimation process, Leibovich in view of Fu does not appear to explicitly suggest:
the iterative training process comprising, for at least the first feature map, extracting scalar samples from the first feature map, generating a flattened and detached array of output feature map values from the scalar samples, and performing a re-estimation operation using the array to adjust quantization boundaries of the first LUT.
However, Cai, in combination with Leibovich and Fu, teaches the limitations:
for at least the first feature map, extracting scalar samples from the first feature map, generating a flattened and detached array of output feature map values from the scalar samples, (Cai, [P. 3, Section: A] “The encoder E(·) first converts an input image xi into a latent feature representation zi = E(xi). Then, the quantizer Q(·) quantizes the features into discrete values ˆzi = Q(zi), which can be losslessly encoded into a bitstream for transmission or storage.” [P. 3, Col. 2, Section: B] “Given the latent image representation
z
i
, we use
{
Y
,
U
(
1
)
,
U
(
2
)
,
U
(
3
)
}
=
T
(
z
i
)
to decompose the features into 3 orthogonal matrices
{
U
(
3
)
}
n
=
1
3
and a core tensor Y, and then quantize the decomposed components to generate bitstream.” [P. 6, Section: B] “Tucker decomposition aims to decompose an N-order tensor
X
∈
R
I
1
×
I
2
×
…
×
I
N
as an affiliation of N orthogonal bases
U
n
∈
R
I
n
×
R
n
n
=
1
N
and the associated core tensor
Y
∈
R
R
1
×
R
2
×
…
×
R
N
,.... With a set of training images, we can easily compute p(|Y|), the probability density function (PDF) of the positive core tensor |Y |... More specifically, we first scan the positive core tensor |Y | in raster order, then utilize the decision boundaries
b
q
0
M
to divide the core tensor into M non-overlapping chunks. For each chunk
C
m
(
m
∈
[
1
,
M
]
)
, instead of using
Y
^
q
m
as the quantized values, we define a new quantizer for symbol
|
Y
i
|
(the i-th element of
|
Y
|
)..) and performing a re-estimation operation using the array to adjust quantization boundaries of the first LUT based on repeated extraction and re-estimation until stable LUTs are obtained. (Cai, [P. 2, Col.2 2nd Paragraph] “The key component of TDNet is a novel tucker decomposition layer (TDL), which decomposes the latent image representation into a set of projection matrices and a compact core tensor. By changing the rank of core tensor and its quantization levels, we can easily adjust the bpp rate of latent image representation, and thus a single CNN model can be trained to compress and reconstruct images under multiple bpp rates. Besides, we propose an iterative non-uniform quantization strategy to obtain the optimal quantization boundaries based on the distribution of encoding coefficients. A coarse-to-fine training strategy is introduced to train a stable TDNet and reconstruct the decompressed images.” [P. 6, Col. 2, Section: B] “Quantization and de-quantization: Since the core tensor Y has both positive and negative values, we take one bit to represent the sign of the original value. Let |Y| denotes the absolute value of the core tensor. With a set of training images, we can easily compute p(|Y|), the probability density function (PDF) of the positive core tensor |Y|. The optimal quantizer can be solved as follows by minimizing the quantization error: …. See equation (16).” [P. 6, Col. 2, Section: B] “The optimal quantizer can be solved as follows by minimizing the quantization error: ….see equation (16). Given a number M of decision intervals, the optimal quantizer is expected to find the set of decision boundaries
{
b
q
}
0
M
and quantized values
{
Y
^
q
}
1
M
. Solving the partial derivative of Eq.(16), we could have: …. See equation (17), The optimal solutions of Eq.(17) can be easily solved by the Lloyd‘s algorithm [38], outputting the optimal quantizer Q(·) with decision boundaries
{
b
q
}
0
M
and quantized values
{
Y
^
q
}
1
M
.” [P. 6, Col.2, Section: V] “With the initialized TDL, we can use Eq.(4) or Eq. (3) to jointly fine-tune the encoder-TDL-decoder network by minimizing the loss function. The latent image representations at multiple bpp rates will be taken into consideration during the training process. In each training epoch, we first decide which group of ranks and decision boundaries will be used by calculating gˆ = mod(epoch, G), then take this group of desired output ranks and decision boundaries
{
R
1
,
R
2
,
R
3
,
{
b
q
}
0
M
}
g
^
+
1
(ˆg+1) to update the TDL and fine-tune the parameters
{
Ω
,
Θ
,
Π
}
of the whole network. To obtain G groups of optimal decision boundaries
{
{
b
q
}
0
M
}
g
=
1
G
and network parameters
{
Ω
,
Θ
,
Π
}
, an iterative training scheme can be used, i.e., fix the encoder-decoder network to update the TDL decision boundaries
{
{
b
q
}
0
M
}
g
=
1
G
by solving Eq.(17), and fix the TDL to update the network parameters
{
Ω
,
Θ
,
Π
}
. Such an alternative optimization process continues till the loss function in Eq.(4) or Eq. (3) converges. After the TDNet converges, we can use the optimal decision boundaries
{
{
b
q
}
0
M
}
g
=
1
G
and network parameters
{
Ω
,
Θ
,
Π
}
to compress and reconstruct images with different bpp rates. The overall all-in-one training scheme is summarized as Algorithm 2.”) [Examiner’s Note: the claimed “flattened and detached array” is broadly interpreted as the obtained and decomposed latent representation by scanning in raster reordering into n-dimensional vector. The probability density function of tensor values correspond to the estimating a distribution of feature map values. It is further noted that the iterative ranking, updating, fine-tune to obtain the optimal quantization boundaries until loss function converges being interpreted as adjusting quantization boundaries based on re-estimation operation. The examiner refers to the process described in algorithm 1 and 2 on Page 6.]
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, and Cai, to incorporate the proposed tucker decomposition framework involving an iterative non-uniform quantization scheme to optimize the quantizer, and a coarse-to-fine training strategy as taught by Cai. One would have been motivated to make such a combination in order to achieve the objective of multiple bpp rates with a single network. Doing so would provide highly competitive performance (Cai [VII]).
As outlined above, while Leibovich implicitly describes and teaches per-layer non-uniform quantization schemes with FIGS. 22A-22C confirming different layers have different distributions, yielding different LUTs with specific levels per layer, the combination of Leibovich, Fu, and Cai does not appear to explicitly suggest: “wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values.”
However, it would have been obvious in view of Zhang. Hereinafter, Zhang, in combination with Leibovich, Fu, and Cai, teaches the limitation:
wherein the second quantization level differs from the first quantization level based on differences between the first feature data associated with the first index values and the second feature data associated with the second index values; (Zhang, [P. 2, Section: 1] “Moreover, the distributions of weights and activations in different networks and even different network layers may differ a lot. We believe a better quantizer should be made adaptive to the weights and activations to gain more flexibility.” [Pp. 5-7, Section: 3.2] “In Fig. 1 we present the statistical distributions of the weights and activations (after batch normalization (BN) and Rectified Linear Unit (ReLU) layers) in a trained floating-point network. It can be seen that the distributions can be complex and differ across layers, and a uniform quantizer is not optimal for them... To get better network quantizers and improve the accuracy of a quantized network, we propose to jointly train the network and its quantizers. The insight behind is that if the optimizers are learnable and optimized through network training, they can not only minimize the quantization error, but also adapt to the training goal thus improving the final accuracy... In practice, we apply layer-wise quantizers for activations (i.e., one quantizer per layer) and channel-wise quantizers for weights (one quantizer for each conv filter).” [P. 11, Section: 4.2] “Figure 4 presents the weight and activation statistics in two layers of a trained ResNet-20 model before (i.e., the floating-point values) and after quantization using our method... It can be seen that our learned quantizers are not uniform ones and they differ at different layers. Statistical results with more bits can be found in the supplementary material.”) [Examiner’s Note: Zhang clearly describes the adaptive quantization of layers weights and activations. The paper explicitly states that learned quantizers differ across layers, which suggest that feature distribution differ across layers and that each layer’s quantizer is adaptively derived from that layer’s own data distribution. This correspond to the broadest reasonably interpretation in light of the specification of the limitation suggesting that the difference in quantization levels between layers is based on the difference in the feature data distribution.]
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, Cai, and Zhang to incorporate the method for learning the quantizers applies to both network weights and activations with arbitrary-bit precision as taught by Zhang. One would have been motivated to make such a combination in order to reduce DNN model size and improve inference efficiency for practical applications. Doing so would enable fast inference on some resource-constrained devices such as mobile phones (Zhang [Section: 5]).
Regarding Previously Presented Claim 11, the combination of Leibovich, Fu, Cai, and Zhang teaches the elements of claim 10 as outlined above, and further teaches:
wherein each of the multiple records of the first LUT includes one index value of the first index values and a single quantization level identified by the one index value. (Leibovich, [0184]-[0188] “the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I ... the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I ... These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table... The algorithm extracts K centers of a given data, and clusters all data points to those K centers, allowing storage of a cluster index per data point, which only requires log2K bits.”) [Examiner’s Notes: Leibovich defines the index values paired with associated k-mean derived quantization values, where each index maps to one quantization value.]
Regarding Original Claim 13, the combination of Leibovich, Fu, Cai, and Zhang teaches the elements of claim 10 as outlined above, and further teaches:
wherein the memory of the electronic device comprises on-chip dynamic random access memory (DRAM) or static random access memory (SRAM). (Leibovich, [0036] “The memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, flash memory device, phase-change memory device, or some other memory device having suitable performance to serve as process memory.” Further described in [0064] and [0140].)
Regarding Original Claim 16, the combination of Leibovich, Fu, Cai, and Zhang teaches the elements of claim 10 as outlined above, and further teaches:
wherein the input data is associated with one or more images or videos. (Leibovich, [0152] “the input to a convolution layer can be a multidimensional array of data that defines the various color components of an input image.” [0158] “FIG. 16A-16B illustrate an exemplary convolutional neural network. FIG. 16A illustrates various layers within a CNN. As shown in FIG. 16A, an exemplary CNN used to model image processing can receive input 1602 describing the red, green, and blue (RGB) components of an input image. The input 1602 can be processed by multiple convolutional layers (e.g., first convolutional layer 1604, second convolutional layer 1606). The output from the multiple convolutional layers may optionally be processed by a set of fully connected layers 1608.” Further see [0195].)
Regarding Currently Amended Claim 17, the combination of Leibovich, Fu, Cai, and Zhang teaches the elements of claim 10 as outlined above, and further teaches:
Claim 17 recites substantially similar limitations as claim 10, claim 17 further introduce the concept of processing feature map as input into subsequent layer.
Leibovich in view of Fu teaches: process the first feature map using the second layer of the neural network to generate the second feature map; regenerate the second feature data of the second feature map by cross-referencing the second index values with the second LUT; and process the regenerated second feature data. (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0184]-[0186] “Then, during runtime, the activations of a specific layer is quantized and only the indexes are stored in memory, as denoted by a value-to-index function: F:V→I... In this way, the memory bandwidth usage is reduced. When computing the next layer, the indexes are translated by an index-to-value function: F:I→V. These functions can be implemented simply by a lookup table in an a embodiment, where the indexes and associated values are stored in the lookup table.” [0189] “the quantization scheme may be used to compress and decompress the activations of each layer during runtime (or online)... the compressed activations are decompressed to the original bit width to perform the computation.”)
Regarding Currently Amended Claim 18, the combination of Leibovich, Fu, Cai, and Zhang teaches the elements of claim 17 as outlined above, and further teaches:
wherein the regenerated second feature data is processed using the third layer of the neural network or a fourth layer of the neural network. (Leibovich, [0026] “In general, deep network models have consisted of several layers which are computed sequentially. The output of the i-th layer (a.k.a., activations) is the input for the i+1-th layer. Such a process is named: “forward pass”.” [0158]-[0160] “FIG. 16A-16B illustrate an exemplary convolutional neural network. FIG. 16A illustrates various layers within a CNN. As shown in FIG. 16A, an exemplary CNN used to model image processing can receive input 1602 describing the red, green, and blue (RGB) components of an input image. The input 1602 can be processed by multiple convolutional layers (e.g., first convolutional layer 1604, second convolutional layer 1606). The output from the multiple convolutional layers may optionally be processed by a set of fully connected layers 1608. Neurons in a fully connected layer have full connections to all activations in the previous layer, as previously described for a feedforward network. The output from the fully connected layers 1608 can be used to generate an output result from the network. The activations within the fully connected layers 1608 can be computed using matrix multiplication instead of convolution. Not all CNN implementations are make use of fully connected layers 1608. For example, in some implementations the second convolutional layer 1606 can generate output for the CNN. The convolutional layers are sparsely connected, which differs from traditional neural network configuration found in the fully connected layers 1608. Traditional neural network layers are fully connected, such that every output unit interacts with every input unit. However, the convolutional layers are sparsely connected because the output of the convolution of a field is input (instead of the respective state value of each of the nodes in the field) to the nodes of the subsequent layer, as illustrated. The kernels associated with the convolutional layers perform convolution operations, the output of which is sent to the next layer. The dimensionality reduction performed within the convolutional layers is one aspect that enables the CNN to scale to process large images.”)
Claim(s) 12 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Leibovich, Fu, Cai, and Zhang as outlined above and further in view of Folliot et al., (Pub. No.: US 20210012208 A1).
Regarding Previously Presented Claim 12, the combination of Leibovich, Fu, Cai, and Zhang teaches the elements of claim 10 as outlined above.
While Leibovich describes the bit precision of the memory, the combination of Leibovich, Fu, Cai, and Zhang does not appear to explicitly suggest:
wherein a size of the first LUT and a size of the second LUT each corresponds to a bit precision of the at least one memory.
However, it would have been obvious in view of Folliot. Hereinafter, Folliot, in combination with Leibovich, Fu, Cai, and Zhang, teaches the limitation:
wherein a size of the first LUT and a size of the second LUT each corresponds to a bit precision of the at least one memory. (Folliot, [0043] “However, it is preferable, in this case, to limit the number of bits to which the input and output data are quantized so as to limit the size of the lookup table and therefore the memory space required to subsequently store it in the device intended to implement the neural network.” [0105] “Thus, for a quantization to 8 bits, 256 values are obtained for the lookup table TF.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Leibovich, Fu, Cai, Zhang, and Folliot, to incorporate the methods for quantize the input and output data of each layer as taught by Folliot. One would have been motivated to make such a combination in order to improve processing speed and to decrease the memory space required to store these data (Folliot [0007]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
(Foreign Reference: WO 2019080483 A1) – “JIANG, FAN” relates to “Neural network computation acceleration method and system based on non-uniform quantization and look-up table.”
[Abstract] “The method (300) comprises: performing non-uniform quantization on parameters of each layer of a neural network (S310); performing non-uniform quantization on inputs of each layer of the neural network (S320); constructing a look-up table for each layer by multiplying each of the quantization values of the parameters of said layer by each of the quantization values of the inputs of said layer (S330); and when forward computation of the neural network is to be performed, looking up a result of multiplication computation of the parameters and inputs of each layer in the look-up table of said layer, and performing computation layer by layer until all computation is done (S340). The method performs non-uniform quantization on all parameters and inputs of a neural network, and further adopts a look-up table to replace multiplication computation, thus accelerating computation of the neural network.”
NPL: Kim, Doyun, et al. "Convolutional neural network quantization using generalized gamma distribution." (2018).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SADIK ALSHAHARI whose telephone number is (703)756-4749. The examiner can normally be reached Monday - Friday, 9 a.m. 6 p.m. ET.
Examiner interviews are available via telephone, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached on (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.A.A./Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121