Prosecution Insights
Last updated: October 02, 2026
Application No. 18/717,275

METHOD AND APPARATUS FOR ACCELERATING DEEP LEANING INFERENCE BASED ON HW-AWARE SPARSITY PATTERN

Non-Final OA §103
Filed
Jun 06, 2024
Priority
Mar 04, 2022 — nonprovisional of PCTCN2022079424
Examiner
GURMU, MULUEMEBET
Art Unit
Tech Center
Assignee
Intel Corporation
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
398 granted / 496 resolved
+20.2% vs TC avg
Strong +18% interview lift
Without
With
+17.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
27 currently pending
Career history
520
Total Applications
across all art units

Statute-Specific Performance

§101
18.2%
-21.8% vs TC avg
§103
68.1%
+28.1% vs TC avg
§102
3.3%
-36.7% vs TC avg
§112
1.4%
-38.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 496 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-45 are present in this application. Claims 1-25 are cancelled. Claims 26-45 are pending in this office action. This office action is NON-FINAL. Drawings The Drawings filed on 06/06/24 are acceptable for examination purposes. Specification The Specification filed on 06/06/24 is acceptable for examination purposes. Information Disclosure Statement The information disclosure statements (IDS) filed on 06/06/24 has been considered by the Examiner and made of record in the application file. Claim Rejections 35 U.S.C. §103 6. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: Claims 26-31 and 33-44 are rejected under 35 U.S.C. § 103 as being un patentable over FRUMKIN et al. (US 2020/0342632 A1) in view of CHEN et al. (US 2021/0286860 A1). Regarding claim 26 FRUMKIN teaches an apparatus for a deep neural network (DNN), comprising, (See FRUMKIN paragraph [0064], a deep neural network (DNN) model includes multiple layers of many connected nodes (e.g., perceptrons) interface circuitry, (See FRUMKIN paragraph [0276], well-known interfaces for communicating with external devices); computer readable instructions, (See FRUMKIN paragraph [0328], computer-readable media); and at least one processor circuit to be programmed by the computer readable instructions to, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing), the sparsity pattern specifying a block size and a sparsity ratio for block-wise, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern), sparsification of a weight matrix of an operator in the DNN, (See FRUMKIN paragraph [0067], data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. The operations in forward propagation include multiplying a weight matrix by an activation matrix. The weight matrix models the weights applied to connections between neurons in layer N and layer N+1); perform the block-wise sparsification for the weight matrix based on the sparsity pattern, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern), to obtain a sparse weight matrix, during a training process of the DNN, (See FRUMKIN paragraph [0067], data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. The operations in forward propagation include multiplying a weight matrix by an activation matrix. The weight matrix models the weights applied to connections between neurons in layer N and layer N+1); according to a training dataset received via the interface circuitry, (See FRUMKIN paragraph [0346], the device driver may perform operations, at least in part, by launching operations on the PPU 300 utilizing an input/output interface between the CPU and the PPU 300); compress the sparse weight matrix into a concentrated weight matrix by removing all-zero blocks from the sparse weight matrix, (See FRUMKIN paragraph [0112], A vector D containing the compacted nonzero values along each of a succession of diagonals in turn (i.e., all zero values from the diagonals are removed, so the data representation is “compacted” meaning made more compact by removing zero values). Size of the vector depends on the number of nonzero values); and generate a mask to indicate an index of each row of non-zero blocks in the sparse weight matrix to enable extraction of corresponding elements, (See FRUMKIN paragraph [0187],The task of determining a destination index of each non-zero element in vector D is assigned to a different thread. Each thread performs operations to determine non-zero element coordinates in matrix A (shown in FIG. 10A) or transposed matrix AT (shown in FIG. 10B). Each thread determines Idx(k) indicating row number of the element k in D and J(k) indicating diagonal number of the element k in D), from an activation matrix of the operator during the deep learning inference, (See FRUMKIN paragraph [0305], perform matrix operations, and, in an embodiment, one or more tensor cores are included in the cores 550. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inferencing). FRUMKIN does not explicitly disclose determine a hardware-aware sparsity pattern based on a register width specified by an Instruction Set Architecture (ISA) of a hardware unit for implementing the DNN for deep learning inference. However, CHEN teaches determine a hardware-aware sparsity pattern based on a register width specified by an Instruction Set Architecture (ISA) of a hardware unit, (See CHEN paragraph [0087], the non-zero elements in the first section to one or more vector registers are identified, e.g., by first analyzer 331. In some embodiments, the first distribution pattern is determined as row-dominant, loading the non-zero elements according to the first order can be loading the non-zero elements by row. An exemplary instruction set architecture (e.g., RISC-V) of some embodiments can include multiple vector registers), for implementing the DNN for deep learning inference, (See CHEN paragraph [0015], a neural network system, larger models, such as deep learning models, may require more memory and computational resources). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify determine a hardware-aware sparsity pattern based on a register width specified by an Instruction Set Architecture (ISA) of a hardware unit for implementing the DNN for deep learning inference of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 27, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein the block size and the sparsity ratio are preset, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern), according to a register shape determined by the register width and a data type of elements in the weight matrix, (See FRUMKIN paragraph [0067], The weight matrix models the weights applied to connections between neurons in layer N and layer N+1. The activation matrix for layer N+1 is the output of that operation from layer N), and one or more of the at least one processor circuit, (See FRUMKIN paragraph [0017], a method performed by a least one processor executing instructions stored in memory is provided), is to perform the block-wise sparsification by dividing the weight matrix into blocks of the block size, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern) and setting a subset of the blocks to be the all-zero blocks according to the sparsity ratio, (See FRUMKIN paragraph [0004], a matrix is “sparse” when it contains lots of zero values (e.g., many more zero values than non-zero values). Because sparse matrices contain relatively few non-zero values, sparse matrices are a natural candidate for data compression). Regarding claim 28, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein when the register shape is NxM, the block size is 1 xM and a number of the non-zero blocks in the sparse weight matrix is an integral multiple of N, (See FRUMKIN paragraph [0078], The diagonal storage needs only one copy of non-zeroes and of the mask (e.g., 8 bytes for an 8×8 sparse matrix)…applied on matrices with 4:8 sparsity, which is a special case of 0.5 sparsity. The diagonal storage is not limited to any particular sparsity, sparsity pattern, and/or particular block size). FRUMKIN does not explicitly disclose where N and M are integers determined by the register width and the data type of elements in the weight matrix. However, CHEN teaches where N and M are integers determined by the register width and the data type of elements in the weight matrix, (non-zero weight elements in a pruned weight matrix of a neural network model are represented by vector registers with variable length. An exemplary instruction set architecture (ISA) having variable length vectors can be RISC-V.). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify where N and M are integers determined by the register width and the data type of elements in the weight matrix of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 29, FRUMKIN taught the apparatus according to claim 28 as described above. FRUMKIN further teaches wherein the sparsity pattern is an N-in-L pattern, the sparsity ratio is equal to (L-N) divided by L, (See FRUMKIN paragraph [0124], The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. In some embodiments, interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern). FRUMKIN does not explicitly disclose L is an integer determined according to a length of a cache line specified by the ISA. However, CHEN teaches L is an integer determined according to a length of a cache line specified by the ISA, (non-zero weight elements in a pruned weight matrix of a neural network model are represented by vector registers with variable length. An exemplary instruction set architecture (ISA) having variable length vectors can be RISC-V.). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify L is an integer determined according to a length of a cache line specified by the ISA of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 30, FRUMKIN taught the apparatus according to claim 29 as described above. FRUMKIN further teaches wherein the register shape is 4x16, and the sparisty pattern is a 4-in-64 pattern, (See FRUMKIN paragraph [0086], provide a look up table…2:4 4×4 blocks, for which transposition information can be stored in a look up table. The downside is that this approach does not scale to larger blocks (there are millions of possibilities for a 2D 4:8 8×8 block)). Regarding claim 31, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein one or more of the at least one processor circuit, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing): is to perform the block-wise sparsification by, (See FRUMKIN paragraph [0205], For 16×16 matrix with 8×8 blocks of 0.5 sparsity, all operations will be done by 4 warps independently working on each block): performing group-lasso regularization on the weight matrix based on the hardware-aware sparsity pattern to generate a preliminary sparse weight matrix, (See FRUMKIN paragraph [0065], The neural network shown may generate sparse matrices during training (e.g., weights of coefficients) and may also generate sparse matrices during operation); and pruning elements in the preliminary sparse weight matrix according to a block-based magnitude so as to obtain the sparse weight matrix, (See FRUMKIN paragraph [0065], The training process adjusts weights of neuron connections in the neural network such that correct decisions are made at the output layer 46 for data provided to the input layer 42. The neural network shown may generate sparse matrices during training (e.g., weights of coefficients): Regarding claim 33, FRUMKIN taught the apparatus according to claim 31 as described above. FRUMKIN further teaches wherein one or more of the at least one processor circuit, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing): is to perform the pruning of the elements in the preliminary sparse weight matrix according to the block-based magnitude by, (See FRUMKIN paragraph [0065], The training process adjusts weights of neuron connections in the neural network such that correct decisions are made at the output layer 46 for data provided to the input layer 42. The neural network shown may generate sparse matrices during training (e.g., weights of coefficients): determining a number K of all-zero blocks in the sparse weight matrix according to the sparse pattern and a shape of the weight matrix, (See CHEN paragraph [0015], setting individual weight elements in a weight matrix to zero). FRUMKIN does not explicitly disclose selecting K blocks containing elements with smallest values among all blocks in the weight matrix and setting the K blocks as the all-zero blocks, where K is an integer greater than 0. However, CHEN teaches selecting K blocks containing elements with smallest values among all blocks in the weight matrix, (See CHEN paragraph [0015], setting individual weight elements in a weight matrix to zero); and setting the K blocks as the all-zero blocks, (See CHEN paragraph [0015], setting individual weight elements in a weight matrix to zero), where K is an integer greater than 0, (See CHEN paragraph [0018], non-zero weight elements in a pruned weight matrix of a neural network model are represented by vector registers with variable length). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify L is an integer determined according to a length of a cache line specified by the ISA of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 34, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches FRUMKIN does not explicitly disclose wherein when a register shape specified by the ISA is NxM, the block size is 1 xM, the sparsity pattern is an N-in-L pattern, the sparsity ratio is equal to (L-N) divided by L, and one or more of the at least one processor circuit is to perform the pruning of the elements in the preliminary sparse weight matrix according to the block-based magnitude by selecting (L-N) blocks containing elements with smallest values among every L blocks in the weight matrix, and setting the (L-N) blocks as the all-zero blocks, wherein N and M are integers determined by the register width and the data type of elements in the weight matrix, and L is an integer determined according to a length of a cache line specified by the ISA. However, CHEN teaches wherein when a register shape specified by the ISA is NxM, the block size is 1 xM, (See CHEN paragraph [0033], The hardware registers can include a memory address register, a byte-count register, one or more control registers, and other types of registers. These registers can specify some combination of source, destination, direction of transfer (reading from the input/output (I/O) device or writing to the I/O device), size of a transfer unit,), the sparsity pattern is an N-in-L pattern, the sparsity ratio is equal to (L-N) divided by L, (See CHEN paragraph [0062], These non-zero elements are sparsely positioned in the weight matrix Y and form different distribution patterns in different areas of the matrix. For example, elements Y.sub.12, Y.sub.13, Y.sub.14, Y.sub.15, and Y.sub.16), and one or more of the at least one processor circuit is to perform the pruning of the elements in the preliminary sparse weight matrix according to the block-based magnitude by, (See CHEN paragraph [0015], pruning includes setting individual weight elements in a weight matrix to zero. As the number of the individual weight elements increases, sparsity of the weight elements of the weight matrix can also increase. In other words, fewer elements are present in the weight matrix such that accuracy is decreased by pruning): selecting (L-N) blocks containing elements with smallest values among every L blocks in the weight matrix, (See CHEN paragraph [0087], section B is determined as a row-dominant section because the total number of rows occupied by non-zero elements (e.g., Y.sub.12, Y.sub.13, Y.sub.14, Y.sub.15, and Y.sub.16) in section B is smaller than the total number of columns occupied by non-zero elements in section B. The value of the output element Z.sub.02 is the dot product of the first row of the input matrix X and the third column of the weight matrix Y); and setting the (L-N) blocks as the all-zero blocks, (See CHEN paragraph [0015], setting individual weight elements in a weight matrix to zero), wherein N and M are integers determined by the register width and the data type of elements in the weight matrix, (See CHEN paragraph [0045], non-zero weight elements in a pruned weight matrix of a neural network model are represented by vector registers having variable length), and L is an integer determined according to a length of a cache line specified by the ISA, (See CHEN paragraph [0058], System 300 employs an exemplary instruction set architecture (ISA) having variable-length vectors that can include multiple vector register). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify wherein when a register shape specified by the ISA is NxM, the block size is 1 xM, the sparsity pattern is an N-in-L pattern, the sparsity ratio is equal to (L-N) divided by L, and one or more of the at least one processor circuit is to perform the pruning of the elements in the preliminary sparse weight matrix according to the block-based magnitude by selecting (L-N) blocks containing elements with smallest values among every L blocks in the weight matrix, and setting the (L-N) blocks as the all-zero blocks, wherein N and M are integers determined by the register width and the data type of elements in the weight matrix, and L is an integer determined according to a length of a cache line specified by the ISA. of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 35, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein one or more of the at least one processor circuit, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing): is to prune gradients of elements in respective blocks of the weight matrix during the training process to set gradients of elements in a block of the weight matrix to be zeroes, (See FRUMKIN paragraph [0069], The higher the gradient, the steeper the slope of the function and the faster a network can learn. When the slope is zero, the network stops learning. Generating the gradients in backward propagation includes multiplying the activations by the transposed weight matrix), when values of the elements in the block have been trained to be zeroes, (See FRUMKIN paragraph [0075], sparse data meaning that many values in the matrix are zero. Similarly, matrices representing characteristics of the neural network during training and inference may include sparse matrix data). Regarding claim 36, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein one or more of the at least one processor circuit is to, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing): extract elements corresponding to the non-zero blocks of the sparse weight matrix, (See FRUMKIN paragraph [0049], information extracted from the compressed data (e.g., non-zero values and/or matrix indices of the non-zero values), without reconstructing the sparse matrix array 104), from the activation matrix based on the mask; broadcast the extracted elements to generate a concatenated activation matrix aligned with the register width; (See FRUMKIN paragraph [0067], The operations in forward propagation include multiplying a weight matrix by an activation matrix. The weight matrix models the weights applied to connections between neurons in layer N and layer N+1) and apply the concatenated activation matrix and the concentrated weight matrix to the operator, (See FRUMKIN paragraph [0067], The operations in forward propagation include multiplying a weight matrix by an activation matrix. The weight matrix models the weights applied to connections between neurons in layer N and layer N+1), so as to perform the deep learning inference based on the DNN, (See FRUMKIN paragraph [0067], generating, storing and/or using compressed and decompressed sparse matrix data in deep neural networks (DNNs). Regarding claim 37, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein the ISA includes any Single Instruction Multiple Data (SIMD) ISA, (See FRUMKIN paragraph [0308], provide the original matrix, transposed matrix, compacted original matrix, and/or compacted transposed matrix. Up until the very last storage prior to instruction, the single matrix data stored by diagonals may be maintained), available for implementing the DNN to perform the deep learning inference, (See FRUMKIN paragraph [0305], the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inferencing). Regarding claim 38, FRUMKIN taught the apparatus according to claim 26 as described above. FRUMKIN further teaches wherein the operator includes a General Matrix Multiplication (GEMM) or a convolution, (See FRUMKIN paragraph [0305], perform deep learning matrix arithmetic, such as convolution operations for neural network training and inferencing). Regarding claim 39 FRUMKIN teaches a non-transitory computer-readable medium comprising instructions to cause at least one processor circuit to, (See FRUMKIN paragraph [0127], The operations shown in FIG. 8 may be performed by one or more processors (e.g., CPUs and/or GPUs) executing software instructions stored in non-transitory memory and/or a hardware-based processing circuit(s)): the sparsity pattern specifying a block size and a sparsity ratio for block-wise, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern), sparsification of a weight matrix of an operator in the DNN, (See FRUMKIN paragraph [0067], data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. The operations in forward propagation include multiplying a weight matrix by an activation matrix. The weight matrix models the weights applied to connections between neurons in layer N and layer N+1); perform the block-wise sparsification for the weight matrix based on the sparsity pattern, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern), to obtain a sparse weight matrix, during a training process of the DNN, (See FRUMKIN paragraph [0067], data flows through the DNN in a forward propagation phase until a prediction is produced that indicates a label corresponding to the input. The operations in forward propagation include multiplying a weight matrix by an activation matrix. The weight matrix models the weights applied to connections between neurons in layer N and layer N+1); compress the sparse weight matrix into a concentrated weight matrix by removing all-zero blocks from the sparse weight matrix, (See FRUMKIN paragraph [0112], A vector D containing the compacted nonzero values along each of a succession of diagonals in turn (i.e., all zero values from the diagonals are removed, so the data representation is “compacted” meaning made more compact by removing zero values). Size of the vector depends on the number of nonzero values); and generate a mask to indicate an index of each row of non-zero blocks in the sparse weight matrix to enable extraction of corresponding elements, (See FRUMKIN paragraph [0187],The task of determining a destination index of each non-zero element in vector D is assigned to a different thread. Each thread performs operations to determine non-zero element coordinates in matrix A (shown in FIG. 10A) or transposed matrix AT (shown in FIG. 10B). Each thread determines Idx(k) indicating row number of the element k in D and J(k) indicating diagonal number of the element k in D), from an activation matrix of the operator during the deep learning inference, (See FRUMKIN paragraph [0305], perform matrix operations, and, in an embodiment, one or more tensor cores are included in the cores 550. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inferencing). FRUMKIN does not explicitly disclose determine a hardware-aware sparsity pattern based on a register width specified by an Instruction Set Architecture (ISA) of a hardware unit for implementing a deep neural network (DNN) for deep learning inference. However, CHEN teaches determine a hardware-aware sparsity pattern based on a register width specified by an Instruction Set Architecture (ISA) of a hardware unit, (See CHEN paragraph [0087], the non-zero elements in the first section to one or more vector registers are identified, e.g., by first analyzer 331. In some embodiments, the first distribution pattern is determined as row-dominant, loading the non-zero elements according to the first order can be loading the non-zero elements by row. An exemplary instruction set architecture (e.g., RISC-V) of some embodiments can include multiple vector registers), for implementing a deep neural network (DNN) for deep learning inference, (See CHEN paragraph [0015], a neural network system, larger models, such as deep learning models, may require more memory and computational resources). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify determine a hardware-aware sparsity pattern based on a register width specified by an Instruction Set Architecture (ISA) of a hardware unit for implementing a deep neural network (DNN) for deep learning inference of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 40, FRUMKIN taught the non-transitory computer-readable medium according to claim 26 as described above. FRUMKIN further teaches wherein the block size and the sparsity ratio are preset, (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern), according to a register shape determined by the register width and a data type of elements in the weight matrix, (See FRUMKIN paragraph [0067], The weight matrix models the weights applied to connections between neurons in layer N and layer N+1. The activation matrix for layer N+1 is the output of that operation from layer N), and the instructions cause one or more of the at least one processor circuit, (See FRUMKIN paragraph [0017], a method performed by a least one processor executing instructions stored in memory is provided), to perform the block-wise sparsification to divide the weight matrix into blocks of the block size, , (See FRUMKIN paragraph [0124], structured sparsity or any particular block size. The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern) and set a subset of the blocks to be the all-zero blocks according to the sparsity ratio, (See FRUMKIN paragraph [0004], a matrix is “sparse” when it contains lots of zero values (e.g., many more zero values than non-zero values). Because sparse matrices contain relatively few non-zero values, sparse matrices are a natural candidate for data compression). Regarding claim 41, FRUMKIN taught the non-transitory computer-readable medium according to claim 40 as described above. FRUMKIN further teaches wherein when the register shape is NxM, the block size is 1 xM and a number of the non-zero blocks in the sparse weight matrix is an integral multiple of N, (See FRUMKIN paragraph [0078], The diagonal storage needs only one copy of non-zeroes and of the mask (e.g., 8 bytes for an 8×8 sparse matrix)…applied on matrices with 4:8 sparsity, which is a special case of 0.5 sparsity. The diagonal storage is not limited to any particular sparsity, sparsity pattern, and/or particular block size), FRUMKIN does not explicitly disclose where N and M are integers determined by the register width and the data type of elements in the weight matrix. However, CHEN teaches where N and M are integers determined by the register width and the data type of elements in the weight matrix, (See CHEN paragraph [0045], non-zero weight elements in a pruned weight matrix of a neural network model are represented by vector registers having variable length). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify where N and M are integers determined by the register width and the data type of elements in the weight matrix of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 42, FRUMKIN taught the non-transitory computer-readable medium according to claim 41 as described above. FRUMKIN further teaches wherein the sparsity pattern is an N-in-L pattern, the sparsity ratio is equal to (L-N) divided by L, ((See FRUMKIN paragraph [0124], The approaches can work with 2:8 sparsity, 4:16 sparsity, 8:16 sparsity, or even general 25% or 50% density without a structure. In some embodiments, interaction with sparse Matrix Multiply Accumulate may require a compliant sparsity pattern). FRUMKIN does not explicitly disclose L is an integer determined according to a length of a cache line specified by the ISA. However, CHEN teaches where N and M are integers determined by the register width and the data type of elements in the weight matrix, (See CHEN paragraph [0058], System 300 employs an exemplary instruction set architecture (ISA) having variable-length vectors that can include multiple vector register). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify L is an integer determined according to a length of a cache line specified by the ISA of CHEN in order to reduce the size of a model in the neural network system. Regarding claim 43, FRUMKIN taught the non-transitory computer-readable medium according to claim 42 as described above. FRUMKIN further teaches wherein the register shape is 4x16, and the sparisty pattern is a 4-in-64 pattern, (See FRUMKIN paragraph [0086], provide a look up table…2:4 4×4 blocks, for which transposition information can be stored in a look up table. The downside is that this approach does not scale to larger blocks (there are millions of possibilities for a 2D 4:8 8×8 block)). Regarding claim 44, FRUMKIN taught the non-transitory computer-readable medium according to claim 39 as described above. FRUMKIN further teaches wherein the instructions cause one or more of the at least one processor circuit, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing): to perform the block-wise sparsification by, (See FRUMKIN paragraph [0205], For 16×16 matrix with 8×8 blocks of 0.5 sparsity, all operations will be done by 4 warps independently working on each block): : performing group-lasso regularization on the weight matrix based on the hardware-aware sparsity pattern to generate a preliminary sparse weight matrix, (See FRUMKIN paragraph [0065], The neural network shown may generate sparse matrices during training (e.g., weights of coefficients) and may also generate sparse matrices during operation); and pruning elements in the preliminary sparse weight matrix according to a block-based magnitude so as to obtain the sparse weight matrix, (See FRUMKIN paragraph [0065], The training process adjusts weights of neuron connections in the neural network such that correct decisions are made at the output layer 46 for data provided to the input layer 42. The neural network shown may generate sparse matrices during training (e.g., weights of coefficients): Claims 32 and 45 are rejected under 35 U.S.C. § 103 as being un patentable over FRUMKIN et al. (US 2020/0342632 A1) in view of CHEN et al. (US 2021/0286860 A1) and further in view of QIN et al. (US 2022/0076095 A1). Regarding claim 32, FRUMKIN taught the apparatus according to claim 30 as described above. FRUMKIN further teaches wherein one or more of the at least one processor circuit, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing): is to perform the pruning of the elements in the preliminary sparse weight matrix according to the block-based magnitude by, (See FRUMKIN paragraph [0065], The training process adjusts weights of neuron connections in the neural network such that correct decisions are made at the output layer 46 for data provided to the input layer 42. The neural network shown may generate sparse matrices during training (e.g., weights of coefficients): FRUMKIN does not explicitly disclose determining whether values of elements in a block in the preliminary sparse weight matrix are less than a preset threshold; and setting the block as an all-zero block when the values of the elements in the block are less than the preset threshold. However, QIN teaches determining whether values of elements in a block in the preliminary sparse weight matrix are less than a preset threshold, (See QIN paragraph [0005], determining whether an inference status meets a predetermined condition; executing the layer using the first sparse matrix if the inference status does not meet the predetermined condition); and setting the block as an all-zero block when the values of the elements in the block are less than the preset threshold, (See QIN paragraph [0117], determining whether an inference status meets a predetermined condition. The method can further include executing the layer using the first sparse matrix if the inference status does not meet the predetermined condition). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify determining whether values of elements in a block in the preliminary sparse weight matrix are less than a preset threshold; and setting the block as an all-zero block when the values of the elements in the block are less than the preset threshold of QIN in order to increase efficiency and reduce latency. Indeed, sparsity in an artificial neural network more accurately reflects how neurons in a human brain process information. Regarding claim 45, FRUMKIN taught the non-transitory computer-readable medium according to claim 44 as described above. FRUMKIN further teaches wherein the instructions cause one or more of the at least one processor circuit, (See FRUMKIN paragraph [0278], a program executed by the host processor encodes a command stream in a buffer that provides workloads to the PPU 300 for processing), to prune the elements in the preliminary sparse weight matrix according to the block-based magnitude by, (See FRUMKIN paragraph [0065], The training process adjusts weights of neuron connections in the neural network such that correct decisions are made at the output layer 46 for data provided to the input layer 42. The neural network shown may generate sparse matrices during training (e.g., weights of coefficients). FRUMKIN does not explicitly disclose determining whether values of elements in a block in the preliminary sparse weight matrix are less than a preset threshold and setting the block as an all-zero block when the values of the elements in the block are less than the preset threshold. However, QIN teaches determining whether values of elements in a block in the preliminary sparse weight matrix are less than a preset threshold, (See QIN paragraph [0005], determining whether an inference status meets a predetermined condition; executing the layer using the first sparse matrix if the inference status does not meet the predetermined condition); and setting the block as an all-zero block when the values of the elements in the block are less than the preset threshold, (See QIN paragraph [0117], determining whether an inference status meets a predetermined condition. The method can further include executing the layer using the first sparse matrix if the inference status does not meet the predetermined condition). It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention was made to modify setting the block as an all-zero block when the values of the elements in the block are less than the preset threshold of QIN in order to increase efficiency and reduce latency. Indeed, sparsity in an artificial neural network more accurately reflects how neurons in a human brain process information. Conclusions/Points of Contacts The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See form PTO-892. Jain et al. (US 2022/0237454 A1), determining neural network training data from a neural network data set, obtaining inputs and outputs of layers of a neural network based on the neural network training data compressing at least one layer of the neural network using said inputs and outputs to obtain parameters representing weights and biases corresponding to said layers. DENG (US 2020/0097830 A1) The weight masker may mask weights in each group of weights of a plurality of weight groups of the DNN to generate, and the loss determiner may determine a loss of the DNN based on a network loss of the DNN minus a variance of a count of non-zero weights in the plurality of weight groups. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUEMEBET GURMU whose telephone number is (571)270-7095. The examiner can normally be reached M-F 9am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tony Mahmoudi can be reached at 5712724078. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MULUEMEBET GURMU/Primary Examiner, Art Unit 2163
Read full office action

Prosecution Timeline

Jun 06, 2024
Application Filed
Sep 15, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705549
SYSTEM AND METHOD OF SEGMENTING DATA AND FORECASTING BY A COMBINATION OF MODELS TRAINED ON SEGMENTED DATA
3y 8m to grant Granted Aug 11, 2026
Patent 12705213
Use of Disaggregated Storage by a Distributed Storage System to Facilitate Performance of Data Management Features that Operate at Distributed Scale
2y 5m to grant Granted Aug 11, 2026
Patent 12688169
METHOD, DEVICE, AND COMPUTER PROGRAM PRODUCT FOR DATA MIGRATION
2y 1m to grant Granted Jul 21, 2026
Patent 12664471
MACHINE LEARNING MODEL ANALYSIS
3y 6m to grant Granted Jun 23, 2026
Patent 12650955
SYSTEM AND METHOD FOR EDITING A FILE-BACKED DATABASE TABLE
2y 0m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
98%
With Interview (+17.5%)
3y 1m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 496 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month