DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1-20 are pending examination.
Information Disclosure Statement
The Information Disclosure Statements (IDSs) submitted by Applicant on 2/7/2024, 2/28/2025, 6/20/2025, and 11/17/2025 have been considered.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4, 6, 10, 13, 15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Catanzaro et al., “cuDNN: Efficient Primitives for Deep Learning”, (2014), in view of Son et al. (US 20180032867 A1, filed Jul. 20, 2017 with foreign application priority data dated Jul. 28, 2016), Quinones et al. (US 20080244223 A1 filed Mar. 31, 2007 and published Oct. 2, 2008) and Yao et al. (US 20180046894 A1, filed Aug. 22, 2016 and published Feb. 15, 2018)
Regarding claim 1, Catanzaro teaches:
providing a machine learning framework including a library of machine learning primitives to operations to optimize a machine learning model (Catanzaro, Abstract, teaches we present a library of efficient implementations of deep learning primitives…we have created a library similar in intent to BLAS, with optimized routines for deep learning workloads. Our implementation contains routines for GPUs.; Catanzaro, pg. 8, Section 7, further teaches we proved a set of routines that allow users to train and evaluate complete deep neural networks without the need to write parallel code manually.),
However, Catanzaro does not distinctly disclose:
non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations
the machine learning framework providing a plurality of quantization primitives to perform a plurality of quantization and dequantization operations; and
processing, via the machine learning framework, a trained convolutional neural network (CNN) model having an associated list of instructions to generate a processed CNN model, the trained CNN model having weights in a floating-point format,
wherein processing the trained CNN includes traversing the list of instructions to reduce a count of instructions in the list of instructions, wherein to reduce the count of instructions includes to pruning one or more conditional branches in the list of instructions that descend from a weight having a first threshold value;
quantizing one or more weights in the floating-point format via a quantization primitive provided by the machine learning framework to generate quantized weights, wherein the quantization primitive causes a graphics processor of the one or more processors to execute an instruction to quantize the one or more weights into one or more quantized weights; and
outputting the processed CNN model, the processed CNN model including the reduced count of instructions and the one or more quantized weights.
Nevertheless, Son teaches:
non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations (Son, [0022] teaches in one general aspect, provided is non-transitory computer-readable storage medium storing instructions, which when executed by a processor, cause the processor to implement one or more or all operations described herein.)
the machine learning framework providing a plurality of quantization primitives to perform a plurality of quantization and dequantization operations (Son, [0128] teaches Similar to the n-th layer, the lightening apparatus 1320 performs quantization, regularization, compression and dequantization of an (n-1)-th layer, e.g., respectively toward and through the example first layer. Lightened parameters are stored in the storage 1330 and are used in a recognition process.; Son, [0129] teaches in another example, when quantization is determined to have been applied when generating the lightened parameters, the restoration apparatus 1420 dequantizes the lightened parameters. For example, the restoration apparatus 1420 changes a representation scheme of quantized parameters to a scheme suitable for a system through the dequantization, such as when the lightened parameters are determined to be quantized for 16-bit fixed-point integers, e.g., from a 32-bit floating-point real number scheme of the original trained parameters, the restoration apparatus 1420 dequantizes the parameters to 32-bit floating-point real numbers. Depending on examples, when a fixed-point data type is used for the original plurality of layers 1411, dequantization may not be performed. In addition, though 32-bit floating-point real number schemes are described for a representation scheme for the original trained parameter values, embodiments are not limited thereto, and the original trained parameter values may be represented according to alternate schemes.); and
processing, via the machine learning framework, a trained convolutional neural network (CNN) model having an associated list of instructions to generate a processed CNN model, the trained CNN model having weights in a floating-point format (Son, [0068], teaches FIG. 2 illustrates an example of a quantization process. Quantization refers to a change in a representation scheme to reduce a size of data. Parameters have a predetermined representation scheme based on a type of system or embodiment. For example, the example non-lightened (or ‘original’) parameters of FIG. 1 may be originally represented by decimal floating-point numbers by the corresponding training operation of the neural network.),
quantizing one or more weights in the floating-point format via a quantization primitive provided by the machine learning framework to generate quantized weights (Son, [0069] teaches at least one of quantization, regularization, or compression is applicable to the lightening of the neural network, and accordingly the neural network may be further lightened based on the regularization and/or the compression. For convenience of description, and only as a non-limiting example, an example of quantizing the original parameters to 16-bit fixed-point integers is described below, noting alternate embodiments are also available. In this example, the quantized parameters are represented in an integer range of −2.sup.15 to 2.sup.15-1.), wherein the quantization primitive causes a graphics processor of the one or more processors to execute an instruction to quantize the one or more weights into one or more quantized weights (Son [0101] teaches while original trained parameters may represent connection weights between nodes of neighboring layers of a correspondingly trained original neural network, for example, and accordingly are representative of the trained neural network structure having all of the nodes and weighted connections corresponding to the trained parameters, when lightening of the original training parameter is performed, such as including quantization, truncation and cutoff, distribution range shifting, and/or compression operations discussed above, the weighted connections that existed in the originally neural network may no longer exist or have zero values, then the new neural network according to the lightened parameters would have a different structure without such non-existent weighted connections. Still further, if all previous weighted connections to any original nodes also no longer exist in the lightened parameters, then the new neural network configured according to the lightened parameters may also not include those corresponding original nodes. Thus, with the lightening of originally trained parameters for a particular structured neural network, the resultant lightened parameters may define a different neural network structure than the original neural network structure, and thus, more efficiently and/or with greater performance perform the originally intended recognition, classification, or other operations compared to the efficiency or performance of the original neural network for the same intended recognition, classification, or other operations.; Son, [0142] teaches user adjustments of operations of the lightening operations discussed herein may be provided by UI 2160, which may include a touch screen or other input device/system. In an example, the processor 2120 may be a graphics processor unit (GPU), reconfigurable processor, or have any other type of multi- or single-processor configuration.);
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro, with the neural network lightning features, as taught by Son, in order to improve the processing operation of a recognition apparatus, reduce the space requirements, improve memory access speeds, and/or improve recognition results. (Son, [0110])
However, the combination does not distinctly disclose:
wherein processing the trained CNN includes: traversing the list of instructions to reduce a count of instructions in the list of instructions, wherein to reduce the count of instructions includes to pruning one or more conditional branches in the list of instructions that descend from a weight having a first threshold value;
…
and
outputting the processed CNN model, the processed CNN model including the reduced count of instructions and the one or more quantized weights.
Nevertheless, Quinones teaches:
wherein processing the trained CNN includes: traversing the list of instructions to reduce a count of instructions in the list of instructions (Quinones [0022] teaches the main parts of one example embodiment of an algorithm according to the inventive subject matter are shown in the flow chart 200 of FIG. 2. According to one example embodiment, the first non-pruned version of Opt Region SRC is built by using the data and control dependences among the instructions in the CFG 210. For that, there is first detected all the data dependences where the producer is in Region SRC and the consumer in Region TRG. These dependences are typically referred to as the live-ins of Region TRG. Second, and starting from these initial set of instructions, one may traverse upwards both the data and control dependences in Region SRC. All the instructions traversed are marked as to belong to the initial version of Region SRC.; Quinones, [0023] further teaches with that first step the initial optimized version of Region SRC will contain all the instructions needed to correctly compute the live-ins of Region TRG. The rest will not be included in that initial version. This local optimization is equivalent to applying dead-code elimination. [i.e., as in to reduce a count of instructions in the list of instructions]; Quinones, [0024] further teaches at this point there has to be decided which instructions or groups of instructions are potentially beneficial to remove. To simplify the problem, for each edge, either all or none the instructions contained in it will be considered for removal. The edges that a-priori are beneficial to abort/prune are those that are most biased. The branch instructions that contain these edges (a bias threshold is used) are added to the set of candidates.), wherein to reduce the count of instructions includes to pruning one or more conditional branches in the list of instructions that descend from a weight having a first threshold value (Quinones, [0008] The method and apparatus of inventive subject matter described herein belongs to a family generically known as branch pruning. It focuses on conditional branches that are strongly biased (i.e., they normally follow the same edge) and optimizes the code by eliminating the infrequently taken paths. The benefit of branch pruning is the removal of the instructions that, following data and control dependences, are needed to execute the branch or/and the instructions in the infrequent path.; Quinones [0024] teaches The edges that a-priori are beneficial to abort/prune are those that are most biased. The branch instructions that contain these edges (a bias threshold is used) are added to the set of candidates [Note: the used bias threshold teaching “instructions that descend from a weight having a first threshold value”]);
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, with the branch pruning of Quinones. The method and apparatus of the inventive subject matter may be used to optimize code in a speculative manner, as long as there is a software or hardware mechanism to validate the correctness of the code. One of its important uses is the optimization of speculative threads. In one example embodiment, the existence of an underlying validation mechanism of speculative threads in an architecture makes the architecture very suitable for this kind of optimization. The parallelization of a program can thus be done more effectively. In addition, speculation can be applied to increment the parallelism at thread level, in a similar way to what was done to instructions with branch prediction. Using speculation, a program may be parallelized by a compiler or code analyzer more easily because it does not need to be conservative and can assume the common case. (Quinones, [0038])
Although the combination in view of Son teaches outputting a processed CNN, the processed CNN model including the one or more quantized weights, the combination does not distinctly disclose also including the reduced count of instructions. Nevertheless, Yao teaches the claimed limitation as stated below.
Yao teaches outputting the processed CNN model, the processed CNN model including the reduced count of instructions and the one or more quantized weights (Yao, [0102] teaches As shown in FIG. 6B, the proposed quantization flow mainly consists of two phases: Step 610: the weight quantization phase, and Step 620: the data quantization phase.; Yao [0228] teaches the controller further comprises an instruction granularity transforming module (not shown in FIG. 13) for transforming coarse-granularity instruction into fine-granularity instructions. Said transformation might be based on the number of PE in said computing complex. For example, the 4 phases shown in Table 3 are coarse-granularity instructions. It might be transformed into more fine-granularity instructions so as to improve efficiency.; Yao [0229] further teaches Alternatively, the instruction granularity transforming might be conducted in instruction generating step 720 of FIG. 7, instead of in controller 8210. In this case, the compiling step 415 (e.g. instruction generating step 720) performs instruction granularity transforming in advance. It may simplify the structure of controller 8210 and spare more resources of PL for PEs. Yao, [Note: the data quantization phase understood to read on including a reduced count of instructions]; See also Fig. 6A – outputting Weight and data quantization configuration of the CNN).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son and Quinones, to further include both the quantization of weights and data of the CNN, as taught by Yao, in order to improve efficiency (Yao, [0228] and [0229])
Regarding claim 4, the combination of Cantanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim 1, and the combination further teaches
wherein processing the trained CNN includes bypassing traversal of conditional branches in the list of instructions for conditional branches that descend from a weight having a second threshold value (Yao, Paragraph [0085] teaches pruning said ANN to prune insignificant connections, said insignificant connections decided based on one or more predetermined criteria; Yao, Paragraph [0086] further teaches predetermined criteria includes if a weight connection is zero, said connection is insignificant or if a weight connection is smaller than a threshold said connection is insignificant. Note: if a weight connection is zero is being understood to read on a first threshold value).
Motivation to combine same as stated for claim 1.
Regarding claim 6, the combination of Catanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim 1, and the combination further teaches wherein outputting the processed CNN model includes compressing a representation of instructions associated with the processed CNN model and generating an executable application to perform the instructions associated with the processed CNN model (Yao, Paragraph [0013] teaches compression step and compiling step, for compiling said compressed ANN to generate instructions to be executed by an ANN accelerator, so as to implement said ANN on said ANN accelerator; Note: the generated instructions to be executed reading on executable application as claimed).
Motivation to combine same as stated for claim 1.
Regarding claim 10, the combination of Catanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim 1, and the combination further teaches wherein the floating-point format is a 32-bit floating-point format (Son, [0068] FIG. 2 illustrates an example of a quantization process. Quantization refers to a change in a representation scheme to reduce a size of data. Parameters have a predetermined representation scheme based on a type of system or embodiment. For example, the example non-lightened (or ‘original’) parameters of FIG. 1 may be originally represented by decimal floating-point numbers by the corresponding training operation of the neural network. A lightening apparatus may change a representation scheme of such original parameters to reduce a size of data for lightening the original parameters. For example, the lightening apparatus may change a representation scheme of the original parameters, from the decimal floating-point numbers, to a fixed point representation of an integer. The lightening apparatus may implement a quantization function 2.sup.Q for quantization, for example. As another example, a representation scheme of original parameters may be changed from a 32-bit floating-point representation to a 16-bit fixed-point representation through such a quantization function 2.sup.Q. Additional or alternative quantized approaches are also available.).
Motivation to combine same as stated for claim 1.
Regarding claim 13, Catanzaro teaches:
instructions to provide a machine learning framework including a library of machine learning primitives to operations to optimize a machine learning model (Catanzaro, Abstract, teaches we present a library of efficient implementations of deep learning primitives…we have created a library similar in intent to BLAS, with optimized routines for deep learning workloads. Our implementation contains routines for GPUs.; Catanzaro, pg. 8, Section 7, further teaches we proved a set of routines that allow users to train and evaluate complete deep neural networks without the need to write parallel code manually.),
However, Catanzaro does not distinctly disclose:
a data processing system comprising: one or more processors including one or more graphics multiprocessors; and a memory to store data including data relating to one or more convolutional neural networks (CNNs) and
the instructions to cause the one or more processors to perform operations comprising: processing, via the machine learning framework, a trained convolutional neural network (CNN) model having an associated list of instructions to generate a processed CNN model, the trained CNN model having weights in a floating-point format, wherein processing the trained CNN includes: traversing the list of instructions to reduce a count of instructions, wherein to reduce the count of instructions includes to pruning one or more conditional branches in the list of instructions that descend from a weight having a first threshold value; quantizing one or more weights in the floating-point format via a quantization primitive provided by the machine learning framework to generate quantized weights, wherein the quantization primitive causes a graphics processor of the one or more processors to execute an instruction to quantize the one or more weights into one or more quantized weights and the machine learning framework is to provide a plurality of quantization primitives to perform a plurality of quantization and dequantization operations; and outputting the processed CNN model, the processed CNN model including the reduced count of instructions and the one or more quantized weights.
Nevertheless, Son teaches:
a data processing system comprising: one or more processors including one or more graphics multiprocessors; and a memory to store data including data relating to one or more convolutional neural networks (CNNs) (Son, [0140] teaches FIG. 21 illustrates an example of an electronic system or device 2100. Referring to FIG. 21, the electronic system or device 2100 includes a sensor 2110, a processor 2120, a memory 2130, a display 2150, and a user interface (UI) 2160.; Son, [0143] teaches in an example, the processor 2120 may be a graphics processor unit (GPU), reconfigurable processor, or have any other type of multi- or single-processor configuration.; Son, [0105] teaches as only an example, in one or more embodiments, the trained neural network may be a deep convolutional neural network (DCNN),) and
the instructions to cause the one or more processors to perform operations comprising: processing, via the machine learning framework, a trained convolutional neural network (CNN) model having an associated list of instructions to generate a processed CNN model, the trained CNN model having weights in a floating-point format (Son, [0068], teaches FIG. 2 illustrates an example of a quantization process. Quantization refers to a change in a representation scheme to reduce a size of data. Parameters have a predetermined representation scheme based on a type of system or embodiment. For example, the example non-lightened (or ‘original’) parameters of FIG. 1 may be originally represented by decimal floating-point numbers by the corresponding training operation of the neural network.),
quantizing one or more weights in the floating-point format via a quantization primitive provided by the machine learning framework to generate quantized weights, wherein the quantization primitive causes a graphics processor of the one or more processors to execute an instruction to quantize the one or more weights into one or more quantized weights (Son, [0069] teaches at least one of quantization, regularization, or compression is applicable to the lightening of the neural network, and accordingly the neural network may be further lightened based on the regularization and/or the compression. For convenience of description, and only as a non-limiting example, an example of quantizing the original parameters to 16-bit fixed-point integers is described below, noting alternate embodiments are also available. In this example, the quantized parameters are represented in an integer range of −2.sup.15 to 2.sup.15-1.; Son [0101] teaches while original trained parameters may represent connection weights between nodes of neighboring layers of a correspondingly trained original neural network, for example, and accordingly are representative of the trained neural network structure having all of the nodes and weighted connections corresponding to the trained parameters, when lightening of the original training parameter is performed, such as including quantization, truncation and cutoff, distribution range shifting, and/or compression operations discussed above, the weighted connections that existed in the originally neural network may no longer exist or have zero values, then the new neural network according to the lightened parameters would have a different structure without such non-existent weighted connections. Still further, if all previous weighted connections to any original nodes also no longer exist in the lightened parameters, then the new neural network configured according to the lightened parameters may also not include those corresponding original nodes. Thus, with the lightening of originally trained parameters for a particular structured neural network, the resultant lightened parameters may define a different neural network structure than the original neural network structure, and thus, more efficiently and/or with greater performance perform the originally intended recognition, classification, or other operations compared to the efficiency or performance of the original neural network for the same intended recognition, classification, or other operations.; Son, [0142] teaches user adjustments of operations of the lightening operations discussed herein may be provided by UI 2160, which may include a touch screen or other input device/system. In an example, the processor 2120 may be a graphics processor unit (GPU), reconfigurable processor, or have any other type of multi- or single-processor configuration.) and the machine learning framework is to provide a plurality of quantization primitives to perform a plurality of quantization and dequantization operations (Son, [0128] teaches Similar to the n-th layer, the lightening apparatus 1320 performs quantization, regularization, compression and dequantization of an (n-1)-th layer, e.g., respectively toward and through the example first layer. Lightened parameters are stored in the storage 1330 and are used in a recognition process.; Son, [0129] teaches in another example, when quantization is determined to have been applied when generating the lightened parameters, the restoration apparatus 1420 dequantizes the lightened parameters. For example, the restoration apparatus 1420 changes a representation scheme of quantized parameters to a scheme suitable for a system through the dequantization, such as when the lightened parameters are determined to be quantized for 16-bit fixed-point integers, e.g., from a 32-bit floating-point real number scheme of the original trained parameters, the restoration apparatus 1420 dequantizes the parameters to 32-bit floating-point real numbers. Depending on examples, when a fixed-point data type is used for the original plurality of layers 1411, dequantization may not be performed. In addition, though 32-bit floating-point real number schemes are described for a representation scheme for the original trained parameter values, embodiments are not limited thereto, and the original trained parameter values may be represented according to alternate schemes.);
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro, with the neural network lightning features, as taught by Son, in order to improve the processing operation of a recognition apparatus, reduce the space requirements, improve memory access speeds, and/or improve recognition results. (Son, [0110])
However, the combination does not distinctly disclose:
wherein processing the trained CNN includes: traversing the list of instructions to reduce a count of instructions, wherein to reduce the count of instructions includes to pruning one or more conditional branches in the list of instructions that descend from a weight having a first threshold value
Nevertheless, Quinones teaches:
wherein processing the trained CNN includes: traversing the list of instructions to reduce a count of instructions, wherein to reduce the count of instructions includes to pruning one or more conditional branches in the list of instructions that descend from a weight having a first threshold value (Quinones [0022] teaches the main parts of one example embodiment of an algorithm according to the inventive subject matter are shown in the flow chart 200 of FIG. 2. According to one example embodiment, the first non-pruned version of Opt Region SRC is built by using the data and control dependences among the instructions in the CFG 210. For that, there is first detected all the data dependences where the producer is in Region SRC and the consumer in Region TRG. These dependences are typically referred to as the live-ins of Region TRG. Second, and starting from these initial set of instructions, one may traverse upwards both the data and control dependences in Region SRC. All the instructions traversed are marked as to belong to the initial version of Region SRC.; Quinones, [0023] further teaches with that first step the initial optimized version of Region SRC will contain all the instructions needed to correctly compute the live-ins of Region TRG. The rest will not be included in that initial version. This local optimization is equivalent to applying dead-code elimination. [i.e., as in to reduce a count of instructions in the list of instructions]; Quinones, [0024] further teaches at this point there has to be decided which instructions or groups of instructions are potentially beneficial to remove. To simplify the problem, for each edge, either all or none the instructions contained in it will be considered for removal. The edges that a-priori are beneficial to abort/prune are those that are most biased. The branch instructions that contain these edges (a bias threshold is used) are added to the set of candidates.);
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, with the branch pruning of Quinones. The method and apparatus of the inventive subject matter may be used to optimize code in a speculative manner, as long as there is a software or hardware mechanism to validate the correctness of the code. One of its important uses is the optimization of speculative threads. In one example embodiment, the existence of an underlying validation mechanism of speculative threads in an architecture makes the architecture very suitable for this kind of optimization. The parallelization of a program can thus be done more effectively. In addition, speculation can be applied to increment the parallelism at thread level, in a similar way to what was done to instructions with branch prediction. Using speculation, a program may be parallelized by a compiler or code analyzer more easily because it does not need to be conservative and can assume the common case. (Quinones, [0038])
Although the combination in view of Son teaches outputting a processed CNN, the processed CNN model including the one or more quantized weights, the combination does not distinctly disclose the outputting also including the reduced count of instructions. Nevertheless, Yao teaches the claimed limitation of outputting the processed CNN model, the processed CNN model including the reduced count of instructions and the one or more quantized weights, as stated below.
Yao teaches outputting the processed CNN model, the processed CNN model including the reduced count of instructions and the one or more quantized weights (Yao, [0102] teaches As shown in FIG. 6B, the proposed quantization flow mainly consists of two phases: Step 610: the weight quantization phase, and Step 620: the data quantization phase.; Yao, [0114] further teaches in the above example of data quantization, step 610 is conducted before step 620. That is, it finishes weight quantization of all CONV layers and FC layers of the ANN, and then conducts data quantization for each feature map set on the basis of the quantized CONV layers and FC layers. [Note: data quantization for each feature map understood as reducing the count of instructions]; Yao [0228] also teaches the controller further comprises an instruction granularity transforming module (not shown in FIG. 13) for transforming coarse-granularity instruction into fine-granularity instructions. Said transformation might be based on the number of PE in said computing complex. For example, the 4 phases shown in Table 3 are coarse-granularity instructions. It might be transformed into more fine-granularity instructions so as to improve efficiency.; Yao [0229] further teaches Alternatively, the instruction granularity transforming might be conducted in instruction generating step 720 of FIG. 7, instead of in controller 8210. In this case, the compiling step 415 (e.g. instruction generating step 720) performs instruction granularity transforming in advance. It may simplify the structure of controller 8210 and spare more resources of PL for PEs. Yao, [Note: the data quantization phase understood to read on including a reduced count of instructions]; See also Fig. 6A – outputting Weight and data quantization configuration of the CNN).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son and Quinones, to further include both the quantization of weights and data of the CNN, as taught by Yao, in order to improve efficiency (Yao, [0228] and [0229])
Regarding claim 15,
Claim 15 recites the same and/or analogous limitations as claim 4. Therefore, it is rejected under the same rationale and motivation as claim 4.
Regarding claim 17,
Claim 17 recites the same and/or analogous limitations as claim 6. Therefore, it is rejected under the same rationale and motivation as claim 6.
Claims 2 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Cantanzaro in view of Son, Quinones, and Yao, as applied to claim 1, and further in view of Dally et al. (US 20180046900 A1, filed Jul. 25, 2017- claiming priority to provisional application 62/373, 919 filed Aug. 11, 2016)
Regarding claim 2, the combination of Cantanzaro in view of Son, Quinones, and Yao, however, the combination does not distinctly disclose wherein the graphics processor includes a graphics multiprocessor having a single instruction multiple thread (SIMT) architecture and the instruction to quantize the one or more weights is an instruction provided by an instruction set architecture of the graphics multiprocessor.
Nevertheless, Dally teaches wherein the graphics processor includes a graphics multiprocessor having a single instruction multiple thread (SIMT) architecture and the instruction to quantize the one or more weights is an instruction provided by an instruction set architecture of the graphics multiprocessor (Dally, [0166] teaches the SM 740 comprises a programmable streaming processor that is configured to process tasks represented by a number of threads. Each SM 740 is multi-threaded and configured to execute a plurality of threads (e.g., 32 threads) from a particular group of threads concurrently. In one embodiment, the SM 740 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (i.e., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions. In another embodiment, the SM 740 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution. In other words, when an instruction for the group of threads is dispatched for execution, some threads in the group of threads may be active, thereby executing the instruction, while other threads in the group of threads may be inactive, thereby performing a no-operation (NOP) instead of executing the instruction. The SM 740 may be described in more detail below in conjunction with FIG. 8.).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the SIMT architecture of the sparse CNN architecture, as taught by Dally. A sparse CNN (SCNN) accelerator architecture described herein, exploits weight and/or activation sparsity to reduce energy consumption and improve processing throughput. The SCNN accelerator architecture couples an algorithmic dataflow that eliminates all multiplications with a zero operand while employing a compressed representation of both weights and activations through almost the entire computation. (Dally, [0036])
Regarding claim 14,
Claim 14 recites the same and/or analogous limitations as claim 2. Therefore, it is rejected under the same rationale and motivation as claim 2.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Cantanzaro in view of Son, Quinones, and Yao, as applied to claim 1, and further in view of Khalvati et al., “Window memoization: toward high-performance image processing software”, J Real-Time Image Proc (2015)
Regarding claim 3, the combination of Cantanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim 1, however, the combination does not distinctly disclose wherein processing the trained CNN additionally includes: determining whether pixels of a second convolution window of the trained CNN model differ from a previously stored first convolution window and eliminating the second convolution window from the trained CNN model in response to determination that the second convolution window matches the first convolution window, and wherein eliminating the second convolution window from the trained CNN model in response to determination that the second convolution window matches the first convolution window includes generating a checksum signature for the first convolution window and eliminating the second convolution window upon determination that the checksum signature for the second convolution window matches the checksum signature for the first convolution window.
Nevertheless, Khalvati teaches:
wherein processing the trained CNN additionally includes: determining whether pixels of a second convolution window of the trained CNN model differ from a previously stored first convolution window and eliminating the second convolution window from the trained CNN model in response to determination that the second convolution window matches the first convolution window (Khalvati, Abstract, teaches window memoization minimizes the number of redundant computations performed on an image by identifying similar neighborhoods of pixels in the image and skipping the computations that are not necessary.; Khalvati, pg. 5-6, col. 2 par. 3 and col. 1 par. 1, teaches window memoization minimizes the number of redundant computations performed on an image, and uses a memory array, reuse table, to store the results of previously performed computations. When a set of computations has to be performed for the first time, the computations are performed and the corresponding result is stored in the reuse table. When the same set of computations has to be performed again in the future, the previously calculated result is reused and the actual computations are skipped.; Khalvati, pg. 1, col. 2 teaches the main goal behind this research is to improve the performance of local image processing algorithms in software by reducing the amount of computations that the algorithm must perform.);
, and
wherein eliminating the second convolution window from the trained CNN model in response to determination that the second convolution window matches the first convolution window includes generating a checksum signature for the first convolution window and eliminating the second convolution window upon determination that the checksum signature for the second convolution window matches the checksum signature for the first convolution window (Khalvati, pg. 8, col. 1, par. 2, teaches for each symbol in an image, the mask operations must be performed only once…window memoization reduces the number of redundant operations by identifying similar windows. This is done by determining which windows belong to one particular symbol.; Khalvati, Section 3.3 par. 1, further teaches upon arrival of a window, if the memoization mechanism is able to find the matching symbol in the reuse table, a hit occurs. In this case, the response of the window is looked up from the reuse table and the actual computations are skipped.; Note: symbol as disclosed in Khalvati reading on checksum as claimed.).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the window memorization, as taught by Khalvati, to obtain speedups for different image processing algorithms with different input images across different processors. (Khalvati, Abstract)
Claims 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Catanzaro in view of Son, Quinones, and Yao, as applied to claim 1, and further in view of Anwar et al., “Structured Pruning of Deep Convolutional Neural Networks” (Feb. 2017)
Regarding claim 5, the combination of Catanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim 1, however, the combination does not distinctly disclose wherein processing the trained CNN includes expanding the CNN model into elementary operations.
Nevertheless, Anwar teaches wherein processing the trained CNN includes expanding the CNN model into elementary operations (Anwar, Section 4.5 teaches feature maps and kernels can be unrolled [i.e., expanded]; Anwar, Section 4.5 teaches feature maps and kernels can be unrolled into elementary matrix multiplication on the CPU into elementary matrix multiplication on the CPU).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the include the elementary operation representation, taught by Anwar, in order to provide intra-kernel stride sparsity where convolutions are differently implemented as matrix-matrix multiplications. (Anwar, pg. 32:2)
Regarding claim 16,
Claim 16 recites the same and/or analogous limitations as claim 5. Therefore, it is rejected under the same rationale and motivation as claim 5.
Claims 7 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Catanzaro in view of Son, Quinones, and Yao, as applied to claim 1, and further in view of Shin et al., “14.2 DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks” (Feb. 2017)
Regarding claim 7, the combination of over Catanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim, however, the combination does not distinctly disclose wherein quantizing one or more weights in the floating-point format via a quantization primitive includes generating a quantization table to enable non-uniform quantization of the weights.
Nevertheless Shin teaches wherein quantizing one or more weights in the floating-point format via a quantization primitive includes generating a quantization table to enable non-uniform quantization of the weights (Shin, pg. 239, col. 1, teaches a quantization table (Q-table)-based matrix multiplication to reduce off-chip memory access an remove duplicated multiplications in the FRP; Shin pg. 239, col. 1 further teaches the FRP performs matrix multiplication with the 128-entry Q-table , and 8 16b fixed point-point multipliers are used to update the Q-table; Shin, pg. 239, col. 2, further teaches each entry of the Q-table contains the pre-computed multiplication result between a 16b fixed-point input and a 16b fixed point weight. After the Q-table is constructed once, only quantized indexes are required to compute the product. The Q-table can function as one 7b Q-table to eight 4b Q-tables.[i.e., understood to read on to enable non-uniform quantization of weights]; [Note: the limitation in the claim reading “to enable non-uniform quantization of the weights" has been understood as intended use language carrying no patentable weight]).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the quantization table, as taught by Shin. With the Q-tables, off-chip accesses can be reduced by 75%, and 99% of the 16b fixed-point multiplications can be avoided. (Shin, pg. 239, col. 2)
Regarding claim 18,
Claim 18 recites the same and/or analogous limitations as claim 7. Therefore, it is rejected under the same rationale and motivation as claim 7.
Claims 8 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Catanzaro in view of Son, Quinones, Yao, and Shin, as applied to claim 7, and further in view of Yang et al. (US 20190073582 A1 filed Sep. 23, 2015 and published Mar. 7, 2019)
Regarding claim 8, the combination of Catanzaro in view of Son, Quinones, Yao, and Shin teaches all of the limitations of claim 7, however, the combination does not distinctly disclose quantization table is structured to maintain accuracy of inference by the processed CNN after quantization of the weights of the trained CNN.
Nevertheless, Yang teaches quantization table is structured to maintain accuracy of inference by the processed CNN after quantization of the weights of the trained CNN (Yang, [0121] teaches FIG. 14 illustrates one embodiment in which CNN program code is executed on a graphics processor or CPU 1408. Quantization logic 1401 implements the techniques described herein to convert non-quantized data 1415 to quantized data 1420 in accordance with a specified quantization bias 1410 and quantization factor 1411. The quantization bias 1410 and quantization factor 1411 are calculated by the quantization window logic 1403. In one embodiment, the quantization logic 1401 maintains a quantization factor dictionary 1430 for the convenience of de-quantization operations. (as described in detail below).; Yang [0132] As mentioned with respect to FIG. 14, one embodiment of the quantization logic 1401 uses a quantization factor dictionary 1430 to identify a quantization factor and corresponding bias for each subregion. In one embodiment, an 3×O×O×N (O=ceil((M−K)/S+1)) floating point matrix is used to store these factors and bias for dequantization. More specifically, a 2×O×O×N matrix may be used to store factors and bias of input images, and a O×O×N matrix may be used to store factors of kernels (since kernels don't use bias in one embodiment). The combined data structure is referred to as the quantization factor dictionary 1430.[Note: here, quantization factor dictionary is understood to read on quantization table, as claimed]).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the quantization method, as taught by Yang, in order to speed up performance without the loss of accuracy. (Yang, Paragraph [0138])
Regarding claim 19,
Claim 19 recites the same and/or analogous limitations as claim 8. Therefore, it is rejected under the same rationale and motivation as claim 8.
Claims 9 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Son, Quinones, Yao, Shin, and Yang, as applied to claim 8, and further in view of Brothers et al. (US 20180082181 A1, filed Jan. 31, 2017 and published Mar. 22, 2018)
Regarding claim 9, the combination Catanzaro in view of Son, Quinones, Yao, Shin, and Yang teaches all of the limitations of claim 8, however the combination does not distinctly disclose wherein the quantization of the weights of the trained CNN is performed without retraining.
Nevertheless, Brothers teaches wherein the quantization of the weights of the trained CNN is performed without retraining (Brothers, [0033] teaches Quantization of the weights 625 may be performed with optional retraining).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the neural network weight compression, as taught by Brothers in order to help speed up execution and improve compression at the same time. (Brothers, [0028])
Regarding claim 20,
Claim 20 recites the same and/or analogous limitations as claim 9. Therefore, it is rejected under the same rational and motivation as claim 9.
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Catanzaro in view of Son, Quinones, and Yao, as applied to claim 1, and further in view of Rasmusson et al. (US 20100060629 A1, filed Sep. 9, 2008 and published Mar. 11, 2010)
Regarding claim 11, the combination of Catanzaro in view of Son, Quinones, and Yao teaches all of the limitations of claim 1, however, the combination does not distinctly disclose wherein the floating-point format is a 16-bit floating-point format.
Nevertheless, Rasmusson teaches wherein the floating-point format is a 16-bit floating-point format (Rasmusson, [0028] teaches in some embodiments, one or more of the local error-control units 160 evaluates whether an approximation technique can be performed for a current graphics processing operation and selects an appropriate technique. The selection of a particular technique is based, at least in part, on an error budget assigned by global error-control unit 170, as will be described in more detail below. Based on this error budget, the error-control unit 160 for some stages may further determine the degree of approximation that may be utilized. Thus, a local error-control unit 160 may evaluate whether 16-bit floating-point precision may be utilized instead of 32-bit precision, whether a lossy compression scheme may be employed, or whether other quantization or sub-sampling schemes may be used. Other local error-control units, in these same embodiments or in other embodiments, may instead receive specific instructions, designating a specific approximation technique and/or a degree of approximation to be used, from the global error control unit 170. Those skilled in the art will appreciate that the particular options available for a given processing stage will depend on the operations performed by that stage, since some approximation techniques are more appropriate for some operations than for others. Further, the availability of various approximation schemes may also be limited by the desired complexity of the stage or the overall processing pipeline.).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the error control unit selecting 16-bit float precision as the floating-bit precision format, as taught by Rasmusson. Those skilled in the art will appreciate that the particular options available for a given processing stage will depend on the operations performed by that stage, since some approximation techniques are more appropriate for some operations than for others. Further, the availability of various approximation schemes may also be limited by the desired complexity of the stage or the overall processing pipeline. (Rasmusson, [0028])
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over of Cantanzaro in view of Son, Quinones, and Yao, as applied to claim 1, and further in view of Yang et al.
Regarding claim 12, the combination of Cantanzaro in view of Son, Quinones, and Yao, teaches all of the limitations of claim 1, however, the combination does not distinctly disclose wherein the quantized weights are in an 8-bit integer format.
Nevertheless, Yang teaches wherein the quantized weights are in an 8-bit integer format (Yang, Paragraph [0063] teaches 8-bit data elements; Yang, Paragraph [0138] teaches fixed point CNN algorithm implemented with 128 bit integer vector instructions provided by Intel SSSE3 and SSSE4 ISA, using quantized parameters with 8 bit width).
Before the effective filing date of the claimed invention, it would have been obvious to one or ordinary skill in the art to have modified the optimization machine learning library, as taught by Catanzaro in view of Son, Quinones, and Yao, to further include the quantization method, as taught by Yang, in order to speed up performance without the loss of accuracy. (Yang, Paragraph [0138])
Conclusion
The following prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Chandrasekhar et al., “Compression of Deep Neural Networks for Image Instance Retrieval” (Apr. 4, 2017)
Park et al., “Zero and Data Reuse-aware Fast Convolution for Deep Neural Networks on GPU”, (2016)
Qiu et al., “Going Deeper with Embedded FPGA Platform for Convolutional Neural Network”, (2016)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BEATRIZ RAMIREZ BRAVO whose telephone number is 571-272-2156. The examiner can normally be reached Mon. - Fri. 7:30a.m.-5:00p.m..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, USMAAN SAEED can be reached at 571-272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.R.B./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146