Prosecution Insights
Last updated: October 02, 2026
Application No. 17/547,158

METHOD AND SYSTEM FOR BIT QUANTIZATION OF ARTIFICIAL NEURAL NETWORK

Non-Final OA §103
Filed
Dec 09, 2021
Priority
Feb 25, 2019 — RE 10-2019-0022047 +3 more
Examiner
KAPOOR, DEVAN
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
DeepX Co., Ltd.
OA Round
4 (Non-Final)
7%
Grant Probability
At Risk
4-5
OA Rounds
0m
Est. Remaining
18%
With Interview

Examiner Intelligence

Grants only 7% of cases
7%
Career Allowance Rate
1 granted / 14 resolved
-47.9% vs TC avg
Moderate +11% lift
Without
With
+11.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
29 currently pending
Career history
47
Total Applications
across all art units

Statute-Specific Performance

§101
34.0%
-6.0% vs TC avg
§103
57.4%
+17.4% vs TC avg
§102
5.8%
-34.2% vs TC avg
§112
2.2%
-37.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the application filed on 5/16/2026. Claims 1-9, 11-21 are pending and have been examined. This action is Final. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Response to Arguments Argument 1: Applicant argues that amended claims 1-9 and 11-20 are patent eligible because the claims do not merely recite mathematical quantization concepts or generic computer implementation. Instead, the claims tie a per-layer, accuracy-gated iterative bit-width reduction process, producing different final bit sizes for different neural-network layers, to a specific hardware accelerator having memory that stores the resulting non-uniformly quantized weight kernels, a multiplier array connected to an adder tree, and an accumulator circuit connected to the adder-tree output. Applicant contends that this ordered combination improves neural-network inference by permitting more aggressive quantization of tolerant layers while preserving greater precision in sensitive layers, thereby reducing memory use and computational cost while maintaining target accuracy. Applicant further argues that the claims are tied to a particular machine, that the hardware elements must be considered together rather than individually, and that the claimed process cannot practically be performed mentally or with pen and paper for a useful multilayer neural network. Response to Argument 1: The examiner has considered the argument above. In light of the amendments, the claims are no longer rejection under 101 and are eligible. The rejection is withdrawn. Argument 2: The applicant argues that none of the cited references, individually or in combination, teaches or suggests selecting weight kernels from different layers, repeatedly reducing each kernel’s bit size subject to network-accuracy gating, and obtaining different final bit sizes for the respective layers. The applicant asserts that Appuswamy merely teaches a multiplier, adder-tree, and feedback-accumulation datapath without teaching how quantized weights are produced; Jouppi teaches staging weights and performing uniform eight-bit inference rather than per-layer non-uniform quantization; and Guo partitions weights and converts them to a predetermined power-of-two-or-zero representation rather than progressively reducing the bit width of individual layer kernels based on accuracy. The applicant maintains that Kang, Dally, Ha, and Zhou were cited only for dependent limitations and do not cure this deficiency, and that merely combining the references would still produce uniformly quantized weights rather than the claimed layer-specific results. The applicant applies the same argument to analogous independent claims 14 and 20, and additionally argues that claims 9 and 19 require arithmetic bit width corresponding to the network’s quantization-bit count, while claim 15 requires retaining the last acceptable bit size after a further reduction causes accuracy to fall below the target, and thus limitations allegedly absent from the cited art. Response to Argument 2: The applicant’s arguments have been fully considered but are not persuasive because they are directed primarily to alleged deficiencies in the previously applied Guo-based combination, whereas the present rejection relies on Wang for the claimed per-layer, accuracy-controlled quantization process. Wang teaches separately determining reduced bit widths for different neural-network layers through multiple iterations using validation data and a predefined performance threshold, determining an optimal quantization bit width for each layer, and assigning different reduced bit widths to first and second layers. Thus, Wang teaches selecting and repeatedly quantizing first- and second-layer weight kernels based on network performance and obtaining different final bit sizes for the respective layers. Jouppi further teaches the claimed accelerator architecture, including memory that stores and supplies neural-network weights, a multiplier-based processing unit that performs convolution, and accumulators that receive resulting partial sums. For claims 9 and 19, Appuswamy further teaches parallel multipliers feeding an adder tree and accumulation circuitry that combines newly generated partial sums with previously stored partial sums, while Jouppi teaches reduced-bit arithmetic having defined operand, product, and accumulator widths. For claim 15, Lin further teaches progressively reducing layer-specific precision, evaluating performance against a threshold, and selecting the last acceptable bit width when a further reduction causes performance to fall below the threshold. Accordingly, the references are relied upon in combination for their respective teachings, and a person of ordinary skill in the art would have been motivated to apply Wang’s adaptive layer-specific bit-width technique to Jouppi’s accelerator to reduce model size and computational requirements while maintaining acceptable accuracy. Therefore, the rejection under 103 is maintained. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 2, 5, 8 14, and 20 is/are rejected under 35 U.S.C. § 103 as being unpatentable over in view of NPL reference, “In-Datacenter Performance Analysis of a Tensor Processing Unit.” by Jouppi et. al. (referred herein as Jouppi) in view of US 20190050710 A1, by Wang et. al. (referred herein Wang). Regarding claim 1, Jouppi teaches: A hardware accelerator configured to process a quantized artificial neural network, the hardware accelerator comprising: a memory configured to store weight kernels of the quantized artificial neural network, wherein the quantized artificial neural network includes a plurality of layers, ([Jouppi, page 1] “This paper evaluates a custom ASIC - called a Tensor Processing Unit (TPU) - deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN)”, and [Jouppi, page 3] “The weights for the matrix unit are staged through an on-chip Weight FIFO that reads from an off-chip 8 GiB DRAM called Weight Memory … Read_Weights reads weights from Weight Memory into the Weight FIFO as input to the Matrix Unit ”, wherein the examiner interprets Jouppi’s custom TPU ASIC that accelerates neural-network inference, together with the Weight Memory and on-chip Weight FIFO that store and stage neural-network weights for input to the Matrix Unit, to be the same as a hardware accelerator configured to process a quantized artificial neural network and comprising a memory configured to store weight kernels of a quantized artificial neural network having a plurality of layers, because they are both specialized neural-network inference hardware having memory that stores and supplies layer weight values to a processing unit for execution of a multilayer neural network.) a processing unit configured to perform convolution using weight kernel of the first layer and the weight kernel of the second layer, the processing unit including a plurality of multipliers operably connected to at least one adder tree; ([Jouppi, page 3]… “ Matrix Multiply Unit is the heart of the TPU. It contains 256x256 MACs that can perform 8-bit multiplyand-adds on signed or unsigned integers … MatrixMultiply/Convolve causes the Matrix Unit to perform a matrix multiply or a convolution from the Unified Buffer into the Accumulators.”, wherein the examiner interprets the “Matrix Multiply Unit”, which receives stored weights and performs matrix multiplication or convolution during neural-network inference, to be the same as a processing unit configured to perform convolution using the weight kernel of the first layer and the weight kernel of the second layer, because they are both neural-network processing units that receive stored layer-weight values and apply those weight values in convolution operations as the respective layers of the neural network are executed. The examiner further interprets Jouppi’s 256-by-256 array of multiply-accumulate units, which performs a plurality of parallel multiplications and additions to produce partial sums, to correspond to a processing unit including a plurality of multipliers operably connected to addition circuitry, because both structures multiply numerous weight and input values in parallel and combine the resulting products into summed convolution results.) an accumulator circuit operably connected to an output of the adder tree, the accumulator circuit including an adder and accumulator. ([Jouppi, page 3] “The 16-bit products are collected in the 4 MiB of 32-bit Accumulators below the matrix unit … The matrix unit produces one 256-element partial sum per clock cycle,” and [the convolution result is written] “into the Accumulators”, wherein the examiner interprets the “Accumulators” positioned below the Matrix Multiply Unit and configured to receive the products and partial sums generated by the Matrix Unit to be the same as an accumulator circuit operably connected to an output of the adder tree, because they are both accumulation storage circuits connected to the output of neural-network multiplication-and-summation circuitry and configured to receive and retain the resulting partial sums.) Jouppi does not teach and wherein the quantized artificial neural network is quantized by: selecting a weight kernel of a first layer of the plurality of layers, repeatedly executing bit quantization on the weight kernel of the first layer to progressively reduce bit size of the weight kernel of the first layer to a first bit size responsive to the accuracy of the artificial neural network being greater than or equal to a target value, selecting a weight kernel of a second layer of the plurality of layers, and repeatedly executing the bit quantization on the one weight kernel of the second layer to progressively reduce a bit size of the weight kernel of the second layer to a second bit size responsive to the accuracy of the artificial neural network being greater than or equal to the target value, the second bit size different from the first bit size;. Wang teaches: and wherein the quantized artificial neural network is quantized by: selecting a weight kernel of a first layer of the plurality of layers, ([Wang, [0005]] “generating a reduced neural network model that includes the plurality of layers, wherein each layer of two or more the plurality of layers includes a respective set of quantized parameters , and each quantized parameter is expressed with the preferred values of the respective reduced bit - widths for the layer as determined through the multiple iterations”, wherein the examiner interprets the generation of a reduced neural-network model in which each of two or more layers includes a respective set of quantized parameters expressed using a preferred reduced bit width determined for that layer to be the same as selecting a weight kernel of a first layer of the plurality of layers, because they are both directed to identifying and separately processing the weight parameters associated with a particular layer of a multilayer neural network for layer-specific quantization.) repeatedly executing bit quantization on the weight kernel of the first layer to progressively reduce bit size of the weight kernel of the first layer to a first bit size responsive to the accuracy of the artificial neural network being greater than or equal to a target value, ([Wang, [0005]] “preferred values of the respective reduced bit widths are determined through multiple iterations of forward propagation through the first neural network model using a validation data set while each of two or more layers … is expressed with different degrees of quantization … until a predefined information loss threshold is met by respective response statistics”, wherein the examiner interprets the determination of preferred reduced bit widths through multiple iterations of forward propagation using a validation data set, while layers are expressed using different degrees of quantization until a predefined information-loss threshold is met, to be the same as repeatedly executing bit quantization on the weight kernel of the first layer to progressively reduce its bit size to a first bit size responsive to the accuracy of the artificial neural network being greater than or equal to a target value, because they are both iterative quantization procedures that evaluate progressively reduced precision for layer parameters against a predetermined network-performance threshold and retain a reduced bit width that satisfies the required performance level.) selecting a weight kernel of a second layer of the plurality of layers, and ([Wang, 0006] “The result of the above process is the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i”, and [Wang, [0074]] “ a first layer (e.g., i = 2 ) of the plurality of layers in the reduced neural network model (e.g., model 112 ) has a first reduced bit - width (e.g., 4 - bit ) that is smaller than the original bit - width (e.g., 32 - bit ) of the first neural network model , a second layer (e.g. , i = 3 ) of the plurality of layers in the reduced neural network model (e.g., model 112 ) has a second reduced bit - width (e.g., 6 - bit ) that is smaller than the original bit - width of the first neural network model , and the first reduced bit - width is distinct from the second reduced bit - width in the reduced neural network model.”, wherein the examiner interprets the determination of optimal quantized weights for each layer and its express identification of a first layer having a first reduced bit width and a second layer having a second reduced bit width to be the same as selecting a weight kernel of a second layer of the plurality of layers, because they are both directed to separately identifying and processing the weight parameters of another layer, distinct from the first layer, for layer-specific bit-width determination.) repeatedly executing the bit quantization on the one weight kernel of the second layer to progressively reduce a bit size of the weight kernel of the second layer to a second bit size responsive to the accuracy of the artificial neural network being greater than or equal to the target value, ([Wang, [0005]] “using respective reduced bit-widths for storing the respective sets of parameters of different layers … while each of two or more layers … is expressed with different degrees of quantization corresponding to different reduced bit widths until a predefined information loss threshold is met”, wherein the examiner interprets the use of respective reduced bit widths for the parameter sets of different layers, determined through multiple iterations while two or more layers are expressed with different degrees of quantization until the predefined information-loss threshold is met, to be the same as repeatedly executing bit quantization on the weight kernel of the second layer to progressively reduce its bit size to a second bit size responsive to the accuracy of the artificial neural network being greater than or equal to the target value, because they are both iterative, performance-controlled quantization procedures separately applied to the weight parameters of a second layer to determine an acceptable reduced precision for that layer.) the second bit size different from the first bit size; [Wang, [0074]] “ a first layer (e.g., i = 2 ) of the plurality of layers in the reduced neural network model (e.g., model 112 ) has a first reduced bit - width (e.g., 4 - bit ) that is smaller than the original bit - width (e.g., 32 - bit ) of the first neural network model , a second layer (e.g. , i = 3 ) of the plurality of layers in the reduced neural network model (e.g., model 112 ) has a second reduced bit - width (e.g., 6 - bit ) that is smaller than the original bit - width of the first neural network model , and the first reduced bit - width is distinct from the second reduced bit - width in the reduced neural network model.”, wherein the examiner interprets the first layer having a first reduced bit width of, for example, four bits and its second layer having a second reduced bit width of, for example, six bits, with the first reduced bit width expressly described as distinct from the second reduced bit width, to be the same as the second bit size being different from the first bit size, because they both expressly assign different final quantization precisions to the weight parameters of different neural-network layers.) Jouppi, Wang, and the instant application are analogous art because they are all directed to hardware-based processing of quantized multilayer neural networks using reduced-precision weight parameters. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the Tensor Processing Unit accelerator disclosed by Jouppi to include the adaptive bit-width model with quantized parameters disclosed by Wang. One would be motivated to do so to efficiently reduce the memory and computational footprint of Jouppi’s neural-network accelerator while preserving an acceptable level of model accuracy, as suggested by Wang ([Wang, [0004]] “reducing model size and computational footprint of the machine learning models and keeping same accuracy.”). Claim 20 is analogous to claim 1, aside from claim type, and thus the same rejection applies. Regarding claim 2, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Wang further teaches wherein in the quantized artificial neural network, weight kernels of the plurality of layers are sequentially quantized based on an amount of computation or an amount of memory. ([Wang, [0004]] “reducing model size and computational footprint of the machine learning models and keeping same accuracy,” AND [Wang, [0005]] “preferred values of the respective reduced bit widths are determined through multiple iterations of forward propagation through the first neural network model…each layer of two or more the plurality of layers includes a respective set of quantized parameters”, wherein the examiner interprets Wang’s iterative determination of respective reduced bit widths for the parameter sets of multiple layers to reduce model size or computational footprint to be the same as sequentially quantizing the weight kernels of the plurality of layers based on an amount of computation or an amount of memory, because they are both layer-specific quantization processes that select reduced weight precision according to computational or memory-resource considerations.) Regarding claim 5, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi further teaches wherein the memory includes at least one of a buffer memory, a register memory, or a cache memory. ([Jouppi, page 3] “The weights for the matrix unit are staged through an on-chip Weight FIFO that reads from an off-chip 8 GiB DRAM called Weight Memory…The intermediate results are held in the 24 MiB on-chip Unified Buffer, which can serve as inputs to the Matrix Unit,” wherein the examiner interprets the on-chip Weight FIFO and Unified Buffer for staging weights and storing intermediate neural-network results to be the same as a memory including at least one of a buffer memory, a register memory, or a cache memory, because the Unified Buffer expressly constitutes buffer memory and the Weight FIFO provides temporary on-chip storage that supplies weight data to the processing unit.) Regarding claim 8, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi further teaches wherein the memory further includes at least one of a weight kernel cache or an input feature map cache. ([Jouppi, page 3] “The weights for the matrix unit are staged through an on-chip Weight FIFO that reads from an off-chip 8 GiB DRAM called Weight Memory,…The intermediate results are held in the 24 MiB on-chip Unified Buffer, which can serve as inputs to the Matrix Unit,” wherein the examiner interprets Jouppi's on-chip Weight FIFO that stages weight values from Weight Memory to be the same as a weight kernel cache, and the Unified Buffer that stores intermediate results serving as inputs to the Matrix Unit to be the same as an input feature map cache, because they are both on-chip memory structures configured to temporarily store neural-network weight kernels or feature-map data for use by the processing unit.) Regarding claim 14, Jouppi teaches: A method for quantizing bits of a multi-layered artificial neural network having a plurality of layers, the method being executed by a hardware accelerator configured to process a quantized artificial neural network for acceleration, the method comprising: ([Jouppi, page 1] “This paper evaluates a custom ASIC- called a Tensor Processing Unit (TPU) - deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN),” wherein the examiner interprets Jouppi’s operation of a custom TPU ASIC to accelerate neural-network inference to be the same as a method executed by a hardware accelerator configured to process a quantized multilayer artificial neural network for acceleration, because they are both directed to executing neural-network processing using specialized accelerator hardware.) … storing the weight kernel of the first layer and the weight kernel of the second layer in a memory of the hardware accelerator; ([Jouppi, page 3] “The weights for the matrix unit are staged through an on-chip Weight FIFO that reads from an off-chip 8 GiB DRAM called Weight Memory” and “Read_Weights reads weights from Weight Memory into the Weight FIFO as input to the Matrix Unit,” wherein the examiner interprets storing and staging neural-network weights in Weight Memory and the Weight FIFO for use by the Matrix Unit to be the same as storing the weight kernels of the first and second layers in a memory of the hardware accelerator, because they are both directed to retaining and supplying layer-specific neural-network weights to accelerator processing circuitry.) and performing convolution using the weight kernels stored in the memory, the convolution being performed by a processing unit of the hardware accelerator; ([Jouppi, page 3] “Matrix Multiply Unit is the heart of the TPU. It contains 256x256 MACs that can perform 8-bit multiply-and-adds on signed or unsigned integers” and “MatrixMultiply/Convolve causes the Matrix Unit to perform a matrix multiply or a convolution from the Unified Buffer into the Accumulators,” wherein the examiner interprets Jouppi’s Matrix Multiply Unit performing convolution using weights supplied from Weight Memory to be the same as performing convolution using stored weight kernels by a processing unit of the hardware accelerator, because they are both accelerator operations that apply stored layer-weight values in neural-network convolution.) the processing unit including a plurality of multipliers operably connected to at least one adder tree; ([Jouppi, page 3] “It contains 256x256 MACs that can perform 8-bit multiply-and-adds on signed or unsigned integers” and “The matrix unit produces one 256-element partial sum per clock cycle,” wherein the examiner interprets Jouppi’s array of multiply-accumulate units that performs parallel multiplications and combines the products into partial sums to be the same as a plurality of multipliers operably connected to addition circuitry, because both structures multiply numerous weights and input values in parallel and combine the resulting products into summed convolution outputs.) and an accumulator circuit operably connected to an output of the adder tree, the accumulator circuit including an adder and accumulator. ([Jouppi, page 3] “The 16-bit products are collected in the 4 MiB of 32-bit Accumulators below the matrix unit…The matrix unit produces one 256-element partial sum per clock cycle,…MatrixMultiply/Convolve causes the Matrix Unit to perform a matrix multiply or a convolution from the Unified Buffer into the Accumulators,” wherein the examiner interprets Jouppi’s Accumulators receiving products and partial sums from the Matrix Unit to be the same as an accumulator circuit operably connected to the output of multiplication-and-addition circuitry and including accumulation functionality, because they are both circuits that receive, add, and retain successive partial convolution results.) Jouppi does not teach selecting a weight kernel of a first layer of the plurality of layers; repeatedly executing bit quantization on the weight kernel of the first layer to progressively reduce bit size of the weight kernel of the first layer to a first bit size responsive to accuracy of the artificial neural network being greater than or equal to a target value; selecting a weight kernel of a second layer of the plurality of layers; and repeatedly executing the bit quantization on the weight kernel of the second layer to progressively reduce a bit size of the weight kernel of the second layer to a second bit size responsive to the accuracy of the artificial neural network being greater than or equal to the target value, the second bit size different from the first bit size. Wang teaches: selecting a weight kernel of a first layer of the plurality of layers; ([Wang, [0005]] “generating a reduced neural network model that includes the plurality of layers, wherein each layer of two or more the plurality of layers includes a respective set of quantized parameters, and each quantized parameter is expressed with the preferred values of the respective reduced bit-widths for the layer as determined through the multiple iterations,” wherein the examiner interprets the separate treatment of the respective quantized parameters and reduced bit width of each neural-network layer to be the same as selecting a weight kernel of a first layer, because they are both directed to identifying and separately processing the weight parameters associated with a particular layer for layer-specific quantization.) repeatedly executing bit quantization on the weight kernel of the first layer to progressively reduce bit size of the weight kernel of the first layer to a first bit size responsive to accuracy of the artificial neural network being greater than or equal to a target value; ([Wang, [0005]] “preferred values of the respective reduced bit widths are determined through multiple iterations of forward propagation through the first neural network model using a validation data set while each of two or more layers … is expressed with different degrees of quantization … until a predefined information loss threshold is met by respective response statistics,” wherein the examiner interprets the determination of a preferred reduced bit width through multiple iterations using validation data while testing different degrees of quantization until a predefined information-loss threshold is met to be the same as repeatedly executing bit quantization on the first-layer weight kernel to progressively reduce its bit size to a first bit size responsive to network accuracy satisfying a target value, because they are both iterative, performance-controlled quantization processes that retain an acceptable reduced precision.) selecting a weight kernel of a second layer of the plurality of layers; ([Wang, [0006]] “The result of the above process is the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i,” AND [Wang, [0074]] “a first layer … has a first reduced bit-width (e.g., 4-bit) … [and] a second layer … has a second reduced bit-width (e.g., 6-bit),” wherein the examiner interprets the determination of optimal quantized weights for each layer and express identification of a second layer having its own reduced bit width to be the same as selecting a weight kernel of a second layer, because they are both directed to separately identifying and processing the weight parameters of another layer for layer-specific quantization.) repeatedly executing the bit quantization on the weight kernel of the second layer to progressively reduce a bit size of the weight kernel of the second layer to a second bit size responsive to the accuracy of the artificial neural network being greater than or equal to the target value, ([Wang, [0005]] “using respective reduced bit-widths for storing the respective sets of parameters of different layers … while each of two or more layers … is expressed with different degrees of quantization corresponding to different reduced bit widths until a predefined information loss threshold is met,” wherein the examiner interprets the use of multiple iterations and different degrees of quantization to determine respective reduced bit widths for different layers to be the same as repeatedly quantizing the second-layer weight kernel to a second bit size responsive to accuracy satisfying the target value, because they are both iterative, performance-controlled quantization procedures separately applied to a second layer.) the second bit size different from the first bit size; ([Wang, [0074]] “a first layer … has a first reduced bit-width (e.g., 4-bit) … a second layer … has a second reduced bit-width (e.g., 6-bit) … and the first reduced bit-width is distinct from the second reduced bit-width,” wherein the examiner interprets the expressly distinct reduced bit widths for the first and second layers to be the same as the second bit size being different from the first bit size, because they both assign different final quantization precisions to different neural-network layers.) Jouppi, Wang, and the instant application are analogous art because they are all directed to hardware-accelerated processing of quantized multilayer neural networks using layer-specific reduced-precision weight parameters. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the neural-network acceleration method disclosed by Jouppi to include the adaptive layer-specific bit-width quantization disclosed by Wang. One would have been motivated to do so to efficiently reduce the memory and computational footprint of Jouppi’s neural-network accelerator while preserving acceptable model accuracy, as suggested by Wang ([Wang, [0004]] “reducing model size and computational footprint of the machine learning models and keeping same accuracy.”). Claim(s) 3, and 11-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jouppi in view of US-11593625-B2 by Kang et. al. (referred herein as Kang). Regarding claim 3, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi and Wang do not teach wherein the processing unit is further configured to process the quantized artificial neural network by at least one of a computational cost bit quantization method, a forward bit quantization method, or a backward bit quantization method. Kang teaches: wherein the processing unit is further configured to process the quantized artificial neural network by at least one of a computational cost bit quantization method, a forward bit quantization method, or a backward bit quantization method. ([Kang, col. 1, lines 49-60] “In a general aspect, a processor implemented method includes performing training or an inference operation with a neural network, by obtaining a parameter for the neural network in a floating-point format, applying a fractional length of a fixed-point format to the parameter in the floating-point format, performing an operation with an integer arithmetic logic unit (ALU) to determine whether to round off a fixed point based on a most significant bit among bit values to be discarded after a quantization process, and performing an operation of quantizing the parameter in the floating-point format to a parameter in the fixed-point format, based on a result of the operation with the ALU,” AND [Kang, col. 8, lines 5-12] “in order to implement a neural network within an allowable accuracy loss while sufficiently reducing the number of operations in the above devices, the floating-point format parameters processed in the neural network may be quantized. The parameter quantization may signify a conversion of a floating-point format parameter having high precision to a fixed-point format parameter having lower precision,” wherein the examiner interprets Kang’s quantization of neural-network parameters during training or inference to reduce the number of operations and convert higher-precision parameters to lower-precision parameters to be the same as processing the quantized artificial neural network by at least one of a computational-cost, forward, or backward bit-quantization method, because they are both directed to applying reduced-precision quantization during neural-network training or inference to decrease computational requirements.) Jouppi, Wang, Kang, and the instant application are analogous art because they are all directed to processing quantized artificial neural networks using reduced-precision parameters to improve computational efficiency.) It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware accelerator of claim 1 disclosed by Jouppi and Wang to include the processor-implemented training or inference quantization disclosed by Kang. One would have been motivated to do so to efficiently reduce the number of operations required to process the neural network while maintaining an allowable level of accuracy, as suggested by Kang ([Kang, col. 8, lines 5-12] “within an allowable accuracy loss while sufficiently reducing the number of operations”). Regarding claim 11, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi and Wang do not teach further comprising: an output activation map cache configured to store a result value of convolution of the processing unit. Kang teaches further comprising: an output activation map cache configured to store a result value of convolution of the processing unit. ([Kang, col 10, lines 62-67] “The memory 220 is hardware configured to store various pieces of data processed in the neural network inference apparatus 20. For example, the memory 220 may store data that has been processed and data that is to be processed in the neural network inference apparatus 20. Furthermore, the memory 220 may store applications and drivers that are to be executed by the neural network inference apparatus 20.” AND ([Kang, col 6, lines 50-52] “In addition, weighted connections may further include kernels for convolutional layers and/or recurrent connections for recurrent layers.”, wherein the examiner interprets the “memory is hardware configured to store various pieces of data processed in the neural network inference apparatus” and “kernels for convolutional layers” to be the same as “an output activation map cache configured to store a result” because both are regarding storage and parts of a neural network (including an activation map)). Jouppi, Wang, Kang, and the instant application are analogous art because they are all directed to optimally quantizing an artificial neural network while maintaining a target performance metric. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware of claim 1 disclosed by Jouppi and Wang to include the hardware configured to store various pieces of data processed by a NN taught by Kang. One would be motivated to do so to efficiently utilize kernels for convolutional and recurrent layers, as suggested by Kang (Kang, [Kang, col 6, lines 50-52] “In addition, weighted connections may further include kernels for convolutional layers and/or recurrent connections for recurrent layers.”). Regarding claim 12 , Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi and Wang do not teach wherein the processing unit further includes a plurality of convolution processing units. Kang teaches wherein the processing unit further includes a plurality of convolution processing units. ([Kang, col 6, lines 50-52] “In addition, weighted connections may further include kernels for convolutional layers and/or recurrent connections for recurrent layers.”, wherein the examiner interprets kernels for convolutional layers to be the same as a plurality of convolution processing units, as both are directed to performing convolution operations). Jouppi, Wang, Kang, and the instant application are analogous art because they are all directed to optimally quantizing an artificial neural network while maintaining a target performance metric. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware of claim 1 disclosed by Jouppi and Wang to include the weighted connectivity taught by Kang. One would be motivated to do so to efficiently utilize kernels for convolutional and recurrent layers, as suggested by Kang (Kang, [Kang, col 6, lines 50-52] “In addition, weighted connections may further include kernels for convolutional layers and/or recurrent connections for recurrent layers.”). Claim(s) 4, 6, and 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jouppi in view of Wang further in view of NPL reference “DEEP COMPRESSION: COMPRESSING DEEP NEURAL NETWORKS WITH PRUNING, TRAINED QUANTIZATION AND HUFFMAN CODING” by Dally et. al. (referred herein as Dally). Regarding claim 4, Jouppi and Wang teach The hardware accelerator of claim 2 (see rejection for claim 2). Jouppi and Wang do not teach wherein the amount of computation or the amount of memory of the quantized artificial neural network are relatively reduced compared to those before quantization, and a number of bits of weight kernels stored in the memory is reduced. Dally teaches wherein the amount of computation or the amount of memory of the quantized artificial neural network are relatively reduced compared to those before quantization, and a number of bits of weight kernels stored in the memory is reduced. Dally teaches wherein the amount of computation or the amount of memory of the quantized artificial neural network are relatively reduced compared to those before quantization, and a number of bits of weight kernels stored in the memory is reduced. Dally teaches wherein the amount of computation or the amount of memory of the quantized artificial neural network are relatively reduced compared to those before quantization, and a number of bits of weight kernels stored in the memory is reduced. ([Dally, page 1] “Quantization then reduces the number of bits that represent each connection from 32 to 5,” AND [Dally, page 3, section 3] “Network quantization and weight sharing further compresses the pruned network by reducing the number of bits required to represent each weight,”, wherein the examiner interprets Dally’s reduction of the number of bits required to represent and store each neural-network weight, thereby compressing the network, to be the same as relatively reducing the amount of computation or memory compared with before quantization and reducing the number of bits of weight kernels stored in memory, because they are both directed to using lower-precision weight representations to decrease neural-network storage and computational requirements.) Jouppi, Wang, Dally, and the instant application are analogous art because they are all directed to efficiently storing and processing quantized neural-network weights using reduced-precision representations. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware accelerator of claim 2 disclosed by Jouppi and Wang to include the reduced-bit weight representation disclosed by Dally. One would have been motivated to do so to efficiently reduce the memory and computational resources required to store and process neural-network weight kernels, as suggested by Dally ([Dally, page 3, section 3] “reducing the number of bits required to represent each weight.”). Regarding claim 6, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi and Wang do not teach wherein a size of a data bit of a data path through which data of a specific layer among the plurality of layers is transmitted is reduced in a unit of bits. Dally teaches wherein a size of a data bit of a data path through which data of a specific layer among the plurality of layers is transmitted is reduced in a unit of bits. ([Dally, page 1] “Quantization then reduces the number of bits that represent each connection from 32 to 5,” wherein the examiner interprets Dally’s reduction in the number of bits representing each neural-network connection to be the same as reducing, in units of bits, the size of data transmitted along a data path associated with a specific layer, because neural-network connections carry layer-specific weight data and quantization reduces the bit width used to represent and transmit that data.) Jouppi, Wang, Dally, and the instant application are analogous art because they are all directed to reduced-precision storage and transmission of neural-network data for efficient hardware-based inference. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware accelerator of claim 1 disclosed by Jouppi and Wang to use the reduced-bit connection representation disclosed by Dally. One would have been motivated to do so to reduce the amount of data transferred through the accelerator data path and thereby decrease memory bandwidth and computational resource requirements, as suggested by Dally ([Dally, page 1] “reduces the number of bits that represent each connection from 32 to 5”). Regarding claim 7, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi and Wang do not teach wherein in the quantized artificial neural network, the bit quantization is executed to reduce a storage size of the memory configured to store the weight kernels. Dally teaches wherein, in the quantized artificial neural network, the bit quantization is executed to reduce a storage size of the memory configured to store the weight kernels. ([Dally, page 12, section 9] “quantizing the network using weight sharing, and then applying Huffman coding. We highlight our experiments on AlexNet which reduced the weight storage by 35× without loss of accuracy,” wherein the examiner interprets Dally's quantization of neural-network weights that reduces weight storage by 35× to be the same as executing bit quantization to reduce a storage size of the memory configured to store the weight kernels, because they are both directed to reducing the memory required to store neural-network weight kernels through reduced-precision quantization.) Jouppi, Wang, Dally, and the instant application are analogous art because they are all directed to efficient storage and processing of quantized neural-network weight kernels. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware accelerator of claim 1 disclosed by Jouppi and Wang to include the weight-storage reduction technique disclosed by Dally. One would have been motivated to do so to reduce the memory required to store neural-network weight kernels while maintaining model accuracy, as suggested by Dally ([Dally, page 12, section 9] “reduced the weight storage by 35× without loss of accuracy.”). Claim(s) 9, 19, and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jouppi in view of Wang further in view of US10621489B2, by Appuswamy et. al. (referred herein as Appuswamy). Regarding claim 9, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi further teaches the adder tree being configured to sum result values of the elementwise multiplication using arithmetic having a bit width corresponding to a number of quantization bits of the quantized artificial neural network, ([Jouppi, page 3] “It contains 256x256 MACs that can perform 8-bit multiply-and-adds on signed or unsigned integers,…The 16-bit products are collected in the 4 MiB of 32-bit Accumulators,…The matrix unit produces one 256-element partial sum per clock cycle,” wherein the examiner interprets Jouppi’s summation of products generated from defined reduced-bit integer operands using correspondingly defined arithmetic widths to be the same as summing the elementwise multiplication results using arithmetic having a bit width corresponding to the quantization-bit precision of the quantized neural network.) Wang further teaches using quantized operands including weight kernel of the first layer and the weight kernel of the second layer, ([Wang, [0006]] “the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i,” AND [Wang, [0074] “describing a first layer having a first reduced bit width and a second layer having a second reduced bit width,” wherein the examiner interprets Wang’s quantized weights associated with respective first and second layers to be the same as quantized operands including the weight kernel of the first layer and the weight kernel of the second layer.) Jouppi and Wang do not teach wherein the plurality of multipliers forms an array of multipliers and the at least one adder tree has an input operably connected to an output of the array of multipliers, the array of multipliers being configured to perform elementwise multiplication … and wherein the accumulator circuit is configured to accumulate the summed result values output by the adder tree Appuswamy teaches wherein the plurality of multipliers forms an array of multipliers and the at least one adder tree has an input operably connected to an output of the array of multipliers. Appuswamy teaches: wherein the plurality of multipliers forms an array of multipliers and the at least one adder tree has an input operably connected to an output of the array of multipliers, the array of multipliers being configured to perform elementwise multiplication … and wherein the accumulator circuit is configured to accumulate the summed result values output by the adder tree Appuswamy teaches wherein the plurality of multipliers forms an array of multipliers and the at least one adder tree has an input operably connected to an output of the array of multipliers. ([Appuswamy, col. 9, lines 1-8] “inputs … are distributed to n multipliers… Each multiplier computes a product, with the products being added into a single sum by the adder tree”, AND [Appuswamy, col. 5, lines 25-26] “n parallel multipliers followed by an adder tree,” wherein the examiner interprets Appuswamy’s parallel multipliers whose outputs are supplied to and summed by an adder tree to be the same as an array of multipliers having outputs operably connected to an input of the adder tree.) the array of multipliers being configured to perform elementwise multiplication. ([Appuswamy, col. 7, lines 30-35] describing a vector-matrix multiplier that calculates an output from an input vector and weight matrix, and [Appuswamy, page 1] “Each of the plurality of multipliers is adapted to, in parallel, apply a weight to an input activation to generate an output,” wherein the examiner interprets the parallel application of respective weights to respective input activations to be the same as elementwise multiplication.) and wherein the accumulator circuit is configured to accumulate the summed result values output by the adder tree ([Appuswamy, col. 5, lines 60-67] “a partial sum register can store partial sum vectors of m-elements, and m parallel adders to add the new partial sums (Z) and previously computed ones (Vt-1),” wherein the examiner interprets Appuswamy’s partial-sum register and parallel adders that combine newly generated partial sums with previously stored partial sums to be the same as an accumulator circuit configured to accumulate the summed result values output by the adder tree.) Jouppi, Wang, Appuswamy, and the instant application are analogous art because they are all directed to hardware-based execution of quantized neural-network operations using parallel multiplication, summation, and accumulation circuitry. It would have been obvious to a person of ordinary skill in the art before the effective filing date to modify the hardware accelerator of claim 1 disclosed by Jouppi and Wang to include Appuswamy’s express multiplier-array and adder-tree arrangement. One would have been motivated to do so to efficiently generate, sum, and accumulate parallel neural-network multiplication results using reduced-precision quantized operands, as suggested by Appuswamy ([Appuswamy, page 1] “Each of the plurality of multipliers is adapted to, in parallel, apply a weight to an input activation to generate an output.”) Regarding claim 19, Jouppi and Wang teach The method of claim 14, (see mapping for claim 14.) Jouppi further teaches the adder tree being configured to sum result values of the elementwise multiplication using arithmetic having a bit width corresponding to a number of quantization bits of the quantized artificial neural network. ([Jouppi, page 3] “It contains 256x256 MACs that can perform 8-bit multiply-and-adds on signed or unsigned integers…The 16-bit products are collected in the 4 MiB of 32-bit Accumulators,…The matrix unit produces one 256-element partial sum per clock cycle,” wherein the examiner interprets Jouppi’s summation of products generated from defined reduced-bit integer operands using correspondingly defined arithmetic widths to be the same as summing the elementwise multiplication results using arithmetic having a bit width corresponding to the quantization-bit precision of the quantized artificial neural network.) Wang further teaches the array of multipliers being configured to perform elementwise multiplication using quantized operands including weight kernel of the first layer and the weight kernel of the second layer. ([Wang, [0006] “the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i,” and [Wang, [0074] describing a first layer having a first reduced bit width and a second layer having a second reduced bit width”, wherein the examiner interprets Wang’s quantized weights associated with respective first and second layers to be the same as quantized operands including the weight kernel of the first layer and the weight kernel of the second layer.) Jouppi and Wang do not teach wherein, in performing the convolution, the plurality of multipliers forms an array of multipliers and the at least one adder tree has an input operably connected to an output of the array of multipliers,..., and wherein, in performing the convolution, the accumulator circuit is configured to accumulate the summed result values output by the adder tree. Appuswamy teaches: wherein, in performing the convolution, the plurality of multipliers forms an array of multipliers and the at least one adder tree has an input operably connected to an output of the array of multipliers ([Appuswamy, col. 9, lines 1-8] “inputs … are distributed to n multipliers… Each multiplier computes a product, with the products being added into a single sum by the adder tree,” AND [Appuswamy, col. 5, lines 25-26] “n parallel multipliers followed by an adder tree,” wherein the examiner interprets Appuswamy’s parallel multipliers whose outputs are supplied to and summed by an adder tree to be the same as an array of multipliers having outputs operably connected to an input of the adder tree.) wherein, in performing the convolution, the accumulator circuit is configured to accumulate the summed result values output by the adder tree. ([Appuswamy, col. 5, lines 60–67] “a partial sum register can store partial sum vectors of m-elements, and m parallel adders to add the new partial sums (Z) and previously computed ones (Vt-1),” wherein the examiner interprets Appuswamy’s partial-sum register and parallel adders that combine newly generated partial sums with previously stored partial sums to be the same as an accumulator circuit configured to accumulate the summed result values output by the adder tree.) Jouppi, Wang, Appuswamy, and the instant application are analogous art because they are all directed to hardware-based execution of quantized neural-network operations using parallel multiplication, summation, and accumulation circuitry. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 14 disclosed by Jouppi and Wang to perform convolution using the multiplier-array, adder-tree, and accumulator arrangement disclosed by Appuswamy. One would have been motivated to do so to efficiently generate, sum, and accumulate parallel neural-network multiplication results using reduced-precision quantized operands, as suggested by Appuswamy ([Appuswamy, col. 9, lines 1-8] “Each multiplier computes a product, with the products being added into a single sum by the adder tree.”). Regarding claim 21, Jouppi and Wang teach The hardware accelerator of claim 1 (see rejection for claim 1). Jouppi further teaches wherein the memory is operably connected to the processing unit, and wherein the processing unit is operably connected to the accumulator circuit. ([Jouppi, page 3] “Read_Weights reads weights from Weight Memory into the Weight FIFO as input to the Matrix Unit” and “MatrixMultiply/Convolve causes the Matrix Unit to perform a matrix multiply or a convolution from the Unified Buffer into the Accumulators”, wherein the examiner interprets the Weight Memory supplying weights as input to the Matrix Unit and the Matrix Unit performing convolution into the Accumulators to be the same as the memory being operably connected to the processing unit and the processing unit being operably connected to the accumulator circuit, because they are both connected hardware datapaths that transfer stored neural-network weights from memory to the processing unit and transfer the processing unit’s convolution results to an accumulator.) Jouppi does not teach wherein the accumulator stores a result of the adder and feeds the stored result back to the adder in a feedback loop to accumulate successive outputs of the adder tree. Appuswamy teaches wherein the accumulator stores a result of the adder and feeds the stored result back to the adder in a feedback loop to accumulate successive outputs of the adder tree. ([Appuswamy, col. 10, lines 1–6] “The vector register holds the previously computed partial sum (Vt-1). This partial sum is fed back to the array of adders 1504, which then adds the new partial sum (XBWid). The data flow … consists of an array of adders, registers, feedback paths”, AND [Appuswamy, col. 5, lines 60-67] “a partial sum register can store partial sum vectors of m-elements, and m parallel adders to add the new partial sums (Z) and previously computed ones (Vt-1)”, wherein the examiner interprets the register storing a previously computed partial sum and feeding that stored sum back to adders that combine it with a newly generated partial sum to be the same as the accumulator storing an adder result and feeding the stored result back to the adder in a feedback loop to accumulate successive outputs of the adder tree, because they are both feedback accumulation arrangements that retain an earlier adder-tree partial sum and repeatedly combine it with each newly generated partial sum). Jouppi, Wang, Appuswamy, and the instant application are analogous art because they are all directed to neural-network inference hardware having processing and accumulation circuitry. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware accelerator of claim 1 disclosed by Jouppi and Wang to include the partial sum technique disclosed by Appuswamy. One would be motivated to do so to efficiently accumulate successive partial sums generated during parallel neural-network computations, as suggested by Appuswamy ([Appuswamy, Fig. 20] “Generate Partial Sums in Parallel”). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jouppi in view of Wang in view of Kang further in view of Appuswamy. Regarding claim 13, Jouppi, Wang, and Kang teach The hardware accelerator of claim 12, (see rejection for claim 12.) Jouppi, Wang, and Kang do not teach wherein the processing unit further includes a tree adder configured to sum result values of convolution of each of the plurality of convolution processing units. Appuswamy teaches wherein the processing unit further includes a tree adder configured to sum result values of convolution of each of the plurality of convolution processing units. ([Appuswamy, col. 5, lines 25–26] describing “n parallel multipliers followed by an adder tree,” and [Appuswamy, col. 9, lines 1–8] describing a plurality of multipliers that compute respective products, “with the products being added into a single sum by the adder tree,” wherein the examiner interprets Appuswamy’s adder tree that receives and sums the respective results generated by parallel neural-network processing paths to be the same as a tree adder configured to sum convolution result values produced by a plurality of convolution processing units.) Jouppi, Wang, Kang, Appuswamy, and the instant application are analogous art because they are directed to hardware circuitry for performing parallel neural-network multiplication, convolution, and summation operations. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hardware accelerator of claim 12 disclosed by Jouppi, Wang, and Kang to include the tree adder arrangement disclosed by Appuswamy. One would have been motivated to do so to efficiently combine parallel multiplication results into a summed output for subsequent neural-network processing, as suggested by Appuswamy ([Appuswamy, col. 9, lines 1-8] “Each multiplier computes a product, with the products being added into a single sum by the adder tree.”). Claim(s) 15 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jouppi in view of Jouppi in view of Wang in view of US10262259B2, by Lin et. al. (referred herein as Lin). Regarding claim 15, Jouppi and Wang teach The method of claim 14, (see mapping for claim 14) Jouppi and Wang do not teach further comprising: determining the first bit size as a final number of bits for the weight kernel of the first layer, responsive to the accuracy of the artificial neural network being less than the target value after a further execution of the bit quantization such that the weight kernel of the first layer has a third bit size that is less than the first bit size. Lin teaches further comprising: determining the first bit size as a final number of bits for the weight kernel of the first layer, responsive to the accuracy of the artificial neural network being less than the target value after a further execution of the bit quantization such that the weight kernel of the first layer has a third bit size that is less than the first bit size. ([Lin, col. 4, lines 10-36] “different bit widths may be selected for bias values, activation values, and/or weights of each layer of the neural network. Aspects of the present disclosure are directed to selecting a bit width for different layers and/or different components of each layer of an ANN. Additionally, aspects of the present disclosure are directed to changing the bit widths based on performance specifications and system resources … the effect of quantizing weights and/or activations is the introduction of quantization noise. Similar to other communication systems, when quantization noise increases, the model performance decreases. Accordingly, the SQNR observed at the output may provide an indication of model performance or accuracy” and “every bit in a fixed point representation contributes K dB SQNR. As such, the SQNR may be employed to select an improved or optimized bit width. The bit width may be selected … on a computational stage (e.g., a layer of a deep convolutional network (DCN)) by computational stage basis” AND [Lin, col. 13, lines 1-20] “In block 502, the process injects noise into a computational stage of the machine learning model. In block 504, the process determines a model performance. In some aspects, the model performance may comprise a classification accuracy, classification speed, SQNR, other model performance metric or a combination thereof. The model performance may be evaluated by comparing the performance to a threshold, in block 506. The threshold may comprise a minimally acceptable performance level. If the performance is above the threshold, the process may inject more noise in block 502 and reevaluate the model performance. On the other hand, if the model performance is below the threshold, the bit width may be selected, in block 508, according to the last acceptable noise level,” wherein the examiner interprets the layer-specific bit-width selection, accuracy threshold, further quantization, and selection of the last acceptable bit width to be the same as determining the first-layer kernel’s final bit size after a lower bit size causes accuracy to fall below the target, because they are both accuracy-controlled processes that retain the last acceptable precision.) Jouppi, Wang, Lin, and the instant application are analogous art because they are all directed to processing quantized artificial neural networks (ANNs) while determining appropriate bit widths for neural-network layers based on performance criteria. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 14 disclosed by Jouppi and Wang to include the bit width selection process disclosed by Lin. One would be motivated to do so to effectively determine an acceptable reduced bit width while maintaining neural-network performance, as suggested by Lin ([Lin, col. 13, lines 1–20] “if the model performance is below the threshold, the bit width may be selected, in block 508, according to the last acceptable noise level.”). Regarding claim 16, Jouppi, Wang, and Lin teach The method of claim 15, (see mapping for claim 15.) Wang further teaches: further comprising: selecting a layer in which a final number of bits for a weight kernel is not determined among the plurality of layers. ([Wang, [0005] “each layer of two or more the plurality of layers includes a respective set of quantized parameters, and each quantized parameter is expressed with the preferred values of the respective reduced bit-widths for the layer as determined through the multiple iterations,” and [Wang, [0006] “The result of the above process is the set of quantized weights Wopt,i with the optimal quantization bit-width(s) for each layer i,” wherein the examiner interprets the separate determination of a respective optimal quantization bit width for each layer to be the same as selecting, from among the plurality of layers, a layer whose final number of bits for its weight kernel has not yet been determined, because each layer must be individually subjected to the bit-width-determination process before its respective final bit width is established.) repeatedly executing the bit quantization to determine the final number of bits of the weight kernel of the selected layer. ([Wang, [0005] “preferred values of the respective reduced bit widths are determined through multiple iterations of forward propagation through the first neural network model using a validation data set while each of two or more layers … is expressed with different degrees of quantization … until a predefined information loss threshold is met,” wherein the examiner interprets the multiple iterations using different degrees of quantization to determine a preferred reduced bit width for each selected layer to be the same as repeatedly executing bit quantization to determine the final number of bits of the weight kernel of the selected layer, because both processes iteratively test reduced precisions until an acceptable final layer-specific bit width is identified.) Claim(s) 17 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jouppi in view of Wang in view DOREFA-NET: TRAINING LOW BITWIDTH CONVOLUTIONAL NEURAL NETWORKS WITH LOW BITWIDTH GRADIENTS. by Zhou et. al. (referred herein as Zhou). Regarding claim 17, Jouppi and Wang teach The method of claim 14, (see mapping for claim 14.) Jouppi and Wang do not teach wherein at least one of the first bit size and the second bit size is 1 bit. Zhou teaches wherein at least one of the first bit size and the second bit size is 1 bit. ([Zhou, page 2] “We explore the configuration space of bitwidth for weights, activations and gradients for DoReFa-Net. E.g., training a network using 1-bit weights, 1-bit activations and 2-bit gradients can lead to 93% accuracy on SVHN dataset,” wherein the examiner interprets Zhou’s quantization of neural-network weights to a one-bit representation to be the same as at least one of the first bit size and the second bit size being one bit, because the first and second bit sizes correspond to the quantization precisions of weight kernels associated with respective neural-network layers, and using one-bit quantized weights in the neural network.) Jouppi, Wang, Zhou, and the instant application are analogous art because they are all directed to processing quantized neural networks using reduced-precision weight representations. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 14 disclosed by Jouppi and Wang to include the one-bit weight quantization disclosed by Zhou. One would have been motivated to do so to further reduce the precision and associated storage and computational requirements of the neural-network weights while maintaining acceptable model accuracy, as suggested by Zhou ([Zhou, page 2] “training a network using 1-bit weights, 1-bit activations and 2-bit gradients can lead to 93% accuracy on SVHN dataset.”) Regarding claim 18, Jouppi and Wang teach The method of claim 14, (see mapping for claim 14.) Jouppi and Wang do not teach further comprising: quantizing at least one of feature map data and activation map data of the artificial neural network. Zhou teaches further comprising: quantizing at least one of feature map data and activation map data of the artificial neural network. ([Zhou, pages 1-2, Introduction] “both weights and input activations of convolutional layers are binarized…We generalize the method of binarized neural networks to allow creating DoReFa-Net, a CNN that has arbitrary bitwidth in weights, activations, and gradients,” wherein the examiner interprets Zhou’s binarization and reduced-bit quantization of input activations associated with convolutional layers to be the same as quantizing activation map data of the artificial neural network, because both are directed to reducing the bit-width representation of activation values processed by neural-network layers. Claim 18 recites “at least one of” feature map data and activation map data, Zhou’s express quantization of activation data satisfies the limitation.) Jouppi, Wang, Zhou, and the instant application are analogous art because they are all directed to processing quantized artificial neural networks using reduced-precision neural-network data. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 14 disclosed by Jouppi and Wang to include the activation quantization disclosed by Zhou. One would have been motivated to do so to permit reduced-bit processing of activation data in addition to neural-network weights and thereby provide low-bit-width convolutional neural-network execution, as suggested by Zhou ([Zhou, pages 1-2, Introduction] “a CNN that has arbitrary bitwidth in weights, activations, and gradients.”). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVAN KAPOOR/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Show 7 earlier events
Dec 23, 2025
Request for Continued Examination
Jan 16, 2026
Response after Non-Final Action
Feb 18, 2026
Non-Final Rejection mailed — §103
Apr 28, 2026
Examiner Interview Summary
Apr 28, 2026
Applicant Interview (Telephonic)
May 16, 2026
Response Filed
Aug 05, 2026
Final Rejection mailed — §103
Sep 02, 2026
Response after Non-Final Action

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
7%
Grant Probability
18%
With Interview (+11.1%)
4y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month