Prosecution Insights
Last updated: October 01, 2026
Application No. 17/657,912

SYSTEMS AND METHODS FOR SPARSE MATRIX MULTIPLICATION

Non-Final OA §103
Filed
Apr 04, 2022
Examiner
GUDAS, JAKOB OSCAR
Art Unit
2151
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
3 (Non-Final)
63%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
12 granted / 19 resolved
+8.2% vs TC avg
Strong +64% interview lift
Without
With
+64.2%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
17 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
29.7%
-10.3% vs TC avg
§103
38.8%
-1.2% vs TC avg
§102
6.9%
-33.1% vs TC avg
§112
22.4%
-17.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 19 resolved cases

Office Action

§103
Detailed Action The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is non-final and is in response to claims filed on 05/13/2026 via RCE. Claims 1-4, 6-13, and 15-22 are pending examination. Claims 1, 11, and 20 currently amended. Claims 2-4, 6-10, 12-13, 15-19, and 21-22 are as previously filed. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/13/2026 has been entered. Response to Arguments Rejections under 35 U.S.C. 101 Applicant’s arguments, see Remarks 11-15, filed 05/13/2026, with respect to the rejections under 35 U.S.C. 101 have been fully considered and are persuasive. The previous rejections have been withdrawn. Rejections Under 35 U.S.C. 103 Applicant’s arguments with respect to claims 1-4, 6-13, and 15-22 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 4, 6-8, 11, 12, 15-17, and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu et al. (US 20210027156 A1) hereinafter Zhu in view of Mathuriya et al. (US 11836102 B1) hereinafter Mathuriya further in view of Zhuo et al. (US 20220207374 A1) hereinafter Zhuo. With regards to claim 1, Zhu teaches A method performed by a computing system, the method comprising: receiving a first block of elements derived from a first matrix representing an activation matrix or a weight matrix for a layer of a neural network, (Zhu [0040]: For example, generic sparsifying 300 may reduce weight matrix) the first block of elements having M elements in a first dimension, where M is an integer (Zhu [0040]: For example, generic sparsifying 300 may reduce weight matrix 301 to a sparse weight matrix 305 to reduce a number of calculations required for executing the neural network. Although depicted as a 4×4 weight matrix, weight matrix 301 may be any size; Zhu [0045]: Weight matrix 401 is depicted as an M×N matrix) parsing the first block of elements into a first set of B sub-blocks, where B is an integer <= M, (Zhu [0045]: FIG. 4 is a representation of a block-wise division 400 of a weight matrix 401 of a neural network, consistent with embodiments of the present disclosure. For example, division 400 may divide weight matrix 401 into blocks of size B.sub.y×B.sub.x) and where each of the first set of B sub-blocks include M/B elements in the first dimension; (Zhu [0045]: FIG. 4 is a representation of a block-wise division 400 of a weight matrix 401 of a neural network, consistent with embodiments of the present disclosure. For example, division 400 may divide weight matrix 401 into blocks of size B.sub.y×B.sub.x; Zhu Fig. 4: shows that there are N/Bx blocks with Bx elements in the first dimension each) applying a first sparsity mask [having S% sparsity over M elements] to the first block of elements to obtain a sparsified first block [having fine-grained balanced sparsity,] (Zhu [0048]: Accordingly, as depicted in FIG. 5, block-wise sparsifying 500 may include selecting one or more elements, e.g., elements 503a, 503b, 503c, and 503d from block 501. Although depicted as selecting four elements, block-wise sparsifying 500 may use any predetermined number of elements) receiving a second block of elements derived from a second matrix representing an activation matrix or a weight matrix for the layer of the neural network (Zhu [0069]: a full input matrix) applying a second sparsity mask [having S'% sparsity over M elements] to the second block of elements to obtain a sparsified second block [having coarse-grained balanced sparsity,] (Zhu [0069]: extract elements from a full input matrix based on offset matrix 603 to obtain sparse input matrix) matrix multiplying the sparsified first block and the sparsified second block to generate a resulting block of elements as a matrix-matrix multiplication (matmul) product; [having the target level of combined sparsity] (Zhu [0070]: receive sparse weight matrix 601, offset matrix 603, and sparse input matrix 605 for executing the multiply-accumulate operations of the neural network) and generating an output at the layer of the neural network for an input to the layer based on the resulting block of elements (Zhu [0029]: As further depicted in FIG. 1, neural nework 100 may include one or more hidden layers, e.g., hidden layer 130-1, . . . , hidden layer 130-n. Each hidden layer may comprise one or more nodes. For example, in FIG. 1, hidden layer 130-1 comprises node 130-1-1, node 130-1-2, node 130-1-3, . . . , node 130-1-b, and hidden layer 130-n comprises node 130-n-1, node 130-n-2, node 130-n-3, . . . , node 130-n-c. Similar to nodes of input layer 120, nodes of the hidden layers may apply activation functions to output from connected nodes of the previous layer and weight the output from the activation functions by particular weights associated with the nodes). While Zhu teaches the matrices having dimensions, Zhu fails to teach having M elements in a second dimension, different than the first dimension, parsing the second block of elements into a second set of B sub-blocks, and each of the second set of B sub-blocks including M/B elements in the second dimension. However, Mathuriya does teach having M elements in a second dimension, different than the first dimension (Mathuriya Col 18 Lines 59-60: input matrix X has M rows and N columns) parsing the second block of elements into a second set of B sub-blocks (Mathuriya Col 15 Lines 61-63: the input matrix X and weight matrix W.sup.T are blocked or split into chunks of 4 (e.g., C=4). The size of each block is B, where B=N/C) each of the second set of B sub-blocks including M/B elements in the second dimension (Mathuriya Col 15 Lines 61-63: the input matrix X and weight matrix W.sup.T are blocked or split into chunks of 4 (e.g., C=4). The size of each block is B, where B=N/C). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu with the matrix dimensions and parsing the second block as taught by Mathuriya. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mask to be applied in parallel, speeding up calculations. Zhu in view of Mathuriya fails to teach [applying a first sparsity mask] having S% sparsity over M elements to obtain a sparsified first block] having fine-grained balanced sparsity, such that each of the first set of B sub-blocks of the sparsified first block has S% sparsity, wherein S is based on a target level of combined sparsity; [applying a second sparsity mask] having S'% sparsity over M elements [to obtain a sparsified second block] having coarse-grained balanced sparsity, such that S'% of the second set of B sub-blocks of the sparsified second block have 100% sparsity, and (100-S')% of the second set of B sub-blocks have 0% sparsity, wherein S' is based on the target level of combined sparsity. However, Zhuo teaches [applying a first sparsity mask] having S% sparsity over M elements to obtain a sparsified first block] having fine-grained balanced sparsity, (Zhuo [0011]: in the vector-wise fine-grained sparsity, a weight matrix with the number of rows being #row and the number of columns being #col is filled with zero columns at an edge of the matrix, so that the number of columns of a zero-added minimum matrix is exactly divided by K, and the zero-added minimum matrix is divided into several vector rows with the number of rows being 1 and the number of columns being K; for each vector row, amplitude-based pruning is performed on an element in the vector row, and on a pruning mask I, 1 of a corresponding element position is set as 0, so that the number of 0 on the pruning mask I meets the requirements of the vector-wise fine-grained sparsity; Zhuo [0031]: wherein the joint sparse process is specifically a process of obtaining pruning masks having different pruning granularities by presetting a target sparsity and a mixing ratio of granularity by a user, and the joint sparse process comprises independent vector-wise fine-grained sparsity and block-wise coarse-grained sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) such that each of the first set of B sub-blocks of the sparsified first block has S% sparsity, (Zhuo [0011]: in the vector-wise fine-grained sparsity, a weight matrix with the number of rows being #row and the number of columns being #col is filled with zero columns at an edge of the matrix, so that the number of columns of a zero-added minimum matrix is exactly divided by K, and the zero-added minimum matrix is divided into several vector rows with the number of rows being 1 and the number of columns being K; for each vector row, amplitude-based pruning is performed on an element in the vector row, and on a pruning mask I; Zhuo Fig. 1a: shows that each of the blocks have S% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) wherein S is based on a target level of combined sparsity; (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) [applying a second sparsity mask] having S'% sparsity over M elements [to obtain a sparsified second block] having coarse-grained balanced sparsity, (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo [0031]: wherein the joint sparse process is specifically a process of obtaining pruning masks having different pruning granularities by presetting a target sparsity and a mixing ratio of granularity by a user, and the joint sparse process comprises independent vector-wise fine-grained sparsity and block-wise coarse-grained sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) such that S'% of the second set of B sub-blocks of the sparsified second block have 100% sparsity (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo Fig. 1c: shows that S' blocks have 100% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) and (100-S')% of the second set of B sub-blocks have 0% sparsity (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo Fig. 1c: shows that 100-S' blocks have 0% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) wherein S' is based on the target level of combined sparsity; (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya with the block pruning as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). With regards to claim 2, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu fails to teach wherein S = S'. However, Zhuo does teach wherein S = S' (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively; (if p is 0.5 then the fine grain and coarse grain sparsity would be equal)) Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the sparsities as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). With regards to claim 4, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu further teaches wherein one or more of the first sparsity mask and the second sparsity mask are generated based on absolute magnitudes for a respective set of M/B elements Zhu [0041]: generic sparsifying 300 may include selecting one or more elements, e.g., elements 303a, 303b, 303c, and 303d from weight matrix 301. Although depicted as selecting four elements, generic sparsifying 300 may use any predetermined number of elements. Elements 303a, 303b, 303c, and 303d may be selected on account of having the four largest absolute values). With regards to claim 6, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu further teaches matrix multiplying the sparsified first block and the sparsified second block to generate the resulting block of elements occurs during training of the neural network (Zhu [0073]: Additionally with or alternatively to method 750 of FIG. 7B, iterative re-training of the neural network may be performed). With regards to claim 7, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 6 above. Zhu further teaches wherein the first sparsity mask and second sparsity mask are dynamically recomputed for each iteration of training of the neural network (Zhu [0086]: At step 759, the at least one processor may determine if the re-trained neural network has converged. For example, the at least one processor may determine convergence has occurred when a desired sparsity level has been reached, when an accuracy of the neural network has dropped below a threshold (e.g., as performed by the pseudocode described above), or any other value associated with the neural network has reached or crossed a predetermined threshold. If converged, method 750 may end; if not, method 750 may iterate, as depicted in FIG. 7B). With regards to claim 8, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu further teaches wherein the neural network is a trained neural network (Zhu [0073]: Additionally with or alternatively to method 750 of FIG. 7B, iterative re-training of the neural network may be performed using the example pseudocode below; (this shows that the neural network is a trained neural network)). While Zhu teaches matrix multiplying after training the neural network, Zhu fails to teach and wherein matrix multiplying the sparsified first block and the sparsified second block to generate the resulting block of elements occurs during an inference operation of the trained neural network. However, Mathuriya does teach and wherein matrix multiplying the sparsified first block and the sparsified second block to generate the resulting block of elements occurs during an inference operation of the trained neural network (Mathuriya Col. 19 Lines 47-52: In case of inference operation, stationary weights in arrays 401a are also stored in the bottom die 401. Top die 402 includes a plurality of processing elements (PEs) 402a. Each PE 402a may include one or more MMUs. Each MMU includes matrix multiplication logic (MML), logic, temporary buffer, etc). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the inference operation as taught by Mathuriya. One of ordinary skill in the art would be motivated to make this combination for at least the same reasons listed in claim 1 above. With regards to claim 11, Zhu teaches A computing system for implementing a deep neural network, comprising: one or more logic machines; (Zhu [0006]: a system for providing block-wise sparsity in a neural network may comprise at least one memory storing instructions and at least one processor configured to execute the instructions to perform operations) and one or more storage machines, each storage machine holding instructions, that when executed by the one or more logic machines cause the computing system to: (Zhu [0006]: a system for providing block-wise sparsity in a neural network may comprise at least one memory storing instructions and at least one processor configured to execute the instructions to perform operations) receive a first block of elements derived from a first matrix representing an activation matrix or a weight matrix for a laver of the deep neural network, (Zhu [0040]: For example, generic sparsifying 300 may reduce weight matrix) the first block of elements having M elements in a first dimension, where M is an integer (Zhu [0040]: For example, generic sparsifying 300 may reduce weight matrix 301 to a sparse weight matrix 305 to reduce a number of calculations required for executing the neural network. Although depicted as a 4×4 weight matrix, weight matrix 301 may be any size; Zhu [0045]: Weight matrix 401 is depicted as an M×N matrix) parse the first block of elements into a first set of B sub-blocks, where B is an integer <= M, (Zhu [0045]: FIG. 4 is a representation of a block-wise division 400 of a weight matrix 401 of a neural network, consistent with embodiments of the present disclosure. For example, division 400 may divide weight matrix 401 into blocks of size B.sub.y×B.sub.x) and where each of the first set of B sub-blocks include M/B elements in the first dimension; (Zhu [0045]: FIG. 4 is a representation of a block-wise division 400 of a weight matrix 401 of a neural network, consistent with embodiments of the present disclosure. For example, division 400 may divide weight matrix 401 into blocks of size B.sub.y×B.sub.x; Zhu Fig. 4: shows that there are N/Bx blocks with Bx elements in the first dimension each) apply a first sparsity mask [having S% sparsity over M elements] to the first block of elements to obtain a sparsified first block [having fine-grained balanced sparsity,] (Zhu [0048]: Accordingly, as depicted in FIG. 5, block-wise sparsifying 500 may include selecting one or more elements, e.g., elements 503a, 503b, 503c, and 503d from block 501. Although depicted as selecting four elements, block-wise sparsifying 500 may use any predetermined number of elements) receive a second block of elements derived from a second matrix representing an activation matrix or a weight matrix for the laver of the deep neural network, (Zhu [0069]: a full input matrix) apply a second sparsity mask [having S'% sparsity over M elements] to the second block of elements to obtain a sparsified second block [having coarse-grained balanced sparsity,] (Zhu [0069]: extract elements from a full input matrix based on offset matrix 603 to obtain sparse input matrix) perform sparse matrix multiplication of the sparsified first block and the sparsified second block to generate a resulting block of elements as a matrix- matrix multiplication (matmul) product [having the target level of combined sparsity;] (Zhu [0070]: receive sparse weight matrix 601, offset matrix 603, and sparse input matrix 605 for executing the multiply-accumulate operations of the neural network) and generate an output at the layer of the deep neural network for an input to the layer based on the resulting block of elements (Zhu [0029]: As further depicted in FIG. 1, neural nework 100 may include one or more hidden layers, e.g., hidden layer 130-1, . . . , hidden layer 130-n. Each hidden layer may comprise one or more nodes. For example, in FIG. 1, hidden layer 130-1 comprises node 130-1-1, node 130-1-2, node 130-1-3, . . . , node 130-1-b, and hidden layer 130-n comprises node 130-n-1, node 130-n-2, node 130-n-3, . . . , node 130-n-c. Similar to nodes of input layer 120, nodes of the hidden layers may apply activation functions to output from connected nodes of the previous layer and weight the output from the activation functions by particular weights associated with the nodes). While Zhu teaches the matrices having dimensions, Zhu fails to teach having M elements in a second dimension, different than the first dimension, parse the second block of elements into a second set of B sub-blocks, and each of the second set of B sub-blocks including M/B elements in the second dimension. However, Mathuriya does teach having M elements in a second dimension, different than the first dimension (Mathuriya Col 18 Lines 59-60: input matrix X has M rows and N columns) parse the second block of elements into a second set of B sub-blocks (Mathuriya Col 15 Lines 61-63: the input matrix X and weight matrix W.sup.T are blocked or split into chunks of 4 (e.g., C=4). The size of each block is B, where B=N/C) each of the second set of B sub-blocks including M/B elements in the second dimension (Mathuriya Col 15 Lines 61-63: the input matrix X and weight matrix W.sup.T are blocked or split into chunks of 4 (e.g., C=4). The size of each block is B, where B=N/C). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu with the matrix dimensions and parsing the second block as taught by Mathuriya. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mask to be applied in parallel, speeding up calculations. Zhu in view of Mathuriya fails to teach [apply a first sparsity mask] having S% sparsity over M elements [to obtain a sparsified first block] having fine-grained balanced sparsity, such that each of the first set of B sub-blocks of the sparsified first block has S% sparsity, wherein S is based on a target level of combined sparsity; [apply a second sparsity mask] having S'% sparsity over M elements [to obtain a sparsified second block] having coarse-grained balanced sparsity, such that S'% of the second set of B sub-blocks of the sparsified second block have 100% sparsity, and (100-S')% of the second set of B sub-blocks have 0% sparsity, wherein S' is based on the target level of combined sparsity. However, Zhuo teaches [apply a first sparsity mask] having S% sparsity over M elements to obtain a sparsified first block] having fine-grained balanced sparsity, (Zhuo [0011]: in the vector-wise fine-grained sparsity, a weight matrix with the number of rows being #row and the number of columns being #col is filled with zero columns at an edge of the matrix, so that the number of columns of a zero-added minimum matrix is exactly divided by K, and the zero-added minimum matrix is divided into several vector rows with the number of rows being 1 and the number of columns being K; for each vector row, amplitude-based pruning is performed on an element in the vector row, and on a pruning mask I, 1 of a corresponding element position is set as 0, so that the number of 0 on the pruning mask I meets the requirements of the vector-wise fine-grained sparsity; Zhuo [0031]: wherein the joint sparse process is specifically a process of obtaining pruning masks having different pruning granularities by presetting a target sparsity and a mixing ratio of granularity by a user, and the joint sparse process comprises independent vector-wise fine-grained sparsity and block-wise coarse-grained sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) such that each of the first set of B sub-blocks of the sparsified first block has S% sparsity, (Zhuo [0011]: in the vector-wise fine-grained sparsity, a weight matrix with the number of rows being #row and the number of columns being #col is filled with zero columns at an edge of the matrix, so that the number of columns of a zero-added minimum matrix is exactly divided by K, and the zero-added minimum matrix is divided into several vector rows with the number of rows being 1 and the number of columns being K; for each vector row, amplitude-based pruning is performed on an element in the vector row, and on a pruning mask I; Zhuo Fig. 1a: shows that each of the blocks have S% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) wherein S is based on a target level of combined sparsity; (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) [apply a second sparsity mask] having S'% sparsity over M elements [to obtain a sparsified second block] having coarse-grained balanced sparsity, (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo [0031]: wherein the joint sparse process is specifically a process of obtaining pruning masks having different pruning granularities by presetting a target sparsity and a mixing ratio of granularity by a user, and the joint sparse process comprises independent vector-wise fine-grained sparsity and block-wise coarse-grained sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) such that S'% of the second set of B sub-blocks of the sparsified second block have 100% sparsity (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo Fig. 1c: shows that S' blocks have 100% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) and (100-S')% of the second set of B sub-blocks have 0% sparsity (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo Fig. 1c: shows that 100-S' blocks have 0% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) wherein S' is based on the target level of combined sparsity; (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya with the block pruning as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). With regards to claim 12, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 11 above. Zhu fails to teach wherein S = S'. However, Zhuo does teach wherein S = S' (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively; (if p is 0.5 then the fine grain and coarse grain sparsity would be equal)) Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the sparsities as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). With regards to claim 15, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 11 above. Zhu further teaches wherein the sparse matrix multiplication occurs during training of the deep neural network (Zhu [0073]: Additionally with or alternatively to method 750 of FIG. 7B, iterative re-training of the neural network may be performed). With regards to claim 16, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 15 above. Zhu further teaches wherein the first sparsity mask and second sparsity mask are dynamically recomputed for each iteration of training of the deep neural network (Zhu [0086]: At step 759, the at least one processor may determine if the re-trained neural network has converged. For example, the at least one processor may determine convergence has occurred when a desired sparsity level has been reached, when an accuracy of the neural network has dropped below a threshold (e.g., as performed by the pseudocode described above), or any other value associated with the neural network has reached or crossed a predetermined threshold. If converged, method 750 may end; if not, method 750 may iterate, as depicted in FIG. 7B). With regards to claim 17, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 11 above. hu further teaches wherein the deep neural network is a trained neural network (Zhu [0073]: Additionally with or alternatively to method 750 of FIG. 7B, iterative re-training of the neural network may be performed using the example pseudocode below; (this shows that the neural network is a trained neural network)). While Zhu teaches matrix multiplying after training the neural network, Zhu fails to teach and wherein the sparse matrix multiplication occurs during an inference operation of the trained neural network. However, Mathuriya does teach and wherein the sparse matrix multiplication occurs during an inference operation of the trained neural network (Mathuriya Col. 19 Lines 47-52: In case of inference operation, stationary weights in arrays 401a are also stored in the bottom die 401. Top die 402 includes a plurality of processing elements (PEs) 402a. Each PE 402a may include one or more MMUs. Each MMU includes matrix multiplication logic (MML), logic, temporary buffer, etc). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the inference operation as taught by Mathuriya. One of ordinary skill in the art would be motivated to make this combination for at least the same reasons listed in claim 11 above. With regards to claim 20, Zhu teaches A method performed by a computing system for training a deep neural network, the method comprising, at the computing system: receiving a first block of elements derived from a weight matrix for a layer of the deep neural network, the first block of elements having M elements in a first dimension, where M is an integer (Zhu [0006]: In some embodiments, a system for providing block-wise sparsity in a neural network; Zhu [0040]: For example, generic sparsifying 300 may reduce weight matrix 301 to a sparse weight matrix 305 to reduce a number of calculations required for executing the neural network. Although depicted as a 4×4 weight matrix, weight matrix 301 may be any size; Zhu [0045]: Weight matrix 401 is depicted as an M×N matrix) parsing the first block of elements into a first set of B sub-blocks, where B is an integer <= M, (Zhu [0045]: FIG. 4 is a representation of a block-wise division 400 of a weight matrix 401 of a neural network, consistent with embodiments of the present disclosure. For example, division 400 may divide weight matrix 401 into blocks of size B.sub.y×B.sub.x) and where each of the first set of B sub-blocks include M/B elements in the first dimension; (Zhu [0045]: FIG. 4 is a representation of a block-wise division 400 of a weight matrix 401 of a neural network, consistent with embodiments of the present disclosure. For example, division 400 may divide weight matrix 401 into blocks of size B.sub.y×B.sub.x; Zhu Fig. 4: shows that there are N/Bx blocks with Bx elements in the first dimension each) applying a first sparsity mask [having S% sparsity over M elements] to the first block of elements to obtain a sparsified first block [having fine-grained balanced sparsity,] (Zhu [0048]: Accordingly, as depicted in FIG. 5, block-wise sparsifying 500 may include selecting one or more elements, e.g., elements 503a, 503b, 503c, and 503d from block 501. Although depicted as selecting four elements, block-wise sparsifying 500 may use any predetermined number of elements) [wherein S is based on a target level of combined sparsity] for a hardware configuration of the computing system (Zhu [0025]: The disclosed embodiments relate to computer-implemented systems and methods for providing block-wise sparse neural networks. Advantageously, the exemplary embodiments can provide improved speed and power efficiency by reducing both mathematical operations and memory transfers required to execute the neural network) receiving a second block of elements derived from an activation matrix, (Zhu [0069]: a full input matrix) applying a second sparsity mask [having S'% sparsity over M elements] to the second block of elements to obtain a sparsified second block [having coarse-grained balanced sparsity,] (Zhu [0069]: extract elements from a full input matrix based on offset matrix 603 to obtain sparse input matrix) matrix multiplying the sparsified first block and the sparsified second block to generate a resulting block of elements as a matrix-matrix multiplication (matmul) product [having the target level of combined sparsity] (Zhu [0070]: receive sparse weight matrix 601, offset matrix 603, and sparse input matrix 605 for executing the multiply-accumulate operations of the neural network) and dynamically recomputing the first sparsity mask and the second sparsity mask for each iteration of training of the neural network (Zhu [0086]: At step 759, the at least one processor may determine if the re-trained neural network has converged. For example, the at least one processor may determine convergence has occurred when a desired sparsity level has been reached, when an accuracy of the neural network has dropped below a threshold (e.g., as performed by the pseudocode described above), or any other value associated with the neural network has reached or crossed a predetermined threshold. If converged, method 750 may end; if not, method 750 may iterate, as depicted in FIG. 7B.). While Zhu teaches the matrices having dimensions, Zhu fails to teach having M elements in a second dimension, different than the first dimension, parsing the second block of elements into a second set of B sub-blocks, and each of the second set of B sub-blocks including M/B elements in the second dimension. However, Mathuriya does teach having M elements in a second dimension, different than the first dimension (Mathuriya Col 18 Lines 59-60: input matrix X has M rows and N columns) parsing the second block of elements into a second set of B sub-blocks (Mathuriya Col 15 Lines 61-63: the input matrix X and weight matrix W.sup.T are blocked or split into chunks of 4 (e.g., C=4). The size of each block is B, where B=N/C) each of the second set of B sub-blocks including M/B elements in the second dimension (Mathuriya Col 15 Lines 61-63: the input matrix X and weight matrix W.sup.T are blocked or split into chunks of 4 (e.g., C=4). The size of each block is B, where B=N/C). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu with the matrix dimensions and parsing the second block as taught by Mathuriya. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mask to be applied in parallel, speeding up calculations. Zhu in view of Mathuriya fails to teach [applying a first sparsity mask] having S% sparsity over M elements to obtain a sparsified first block] having fine-grained balanced sparsity, such that each of the first set of B sub-blocks of the sparsified first block has S% sparsity, wherein S is based on a target level of combined sparsity [for a hardware configuration of the computing system;] [applying a second sparsity mask] having S'% sparsity over M elements [to obtain a sparsified second block] having coarse-grained balanced sparsity, such that S'% of the second set of B sub-blocks of the sparsified second block have 100% sparsity, and (100-S')% of the second set of B sub-blocks have 0% sparsity, wherein S' is based on the target level of combined sparsity. However, Zhuo teaches [applying a first sparsity mask] having S% sparsity over M elements to obtain a sparsified first block] having fine-grained balanced sparsity, (Zhuo [0011]: in the vector-wise fine-grained sparsity, a weight matrix with the number of rows being #row and the number of columns being #col is filled with zero columns at an edge of the matrix, so that the number of columns of a zero-added minimum matrix is exactly divided by K, and the zero-added minimum matrix is divided into several vector rows with the number of rows being 1 and the number of columns being K; for each vector row, amplitude-based pruning is performed on an element in the vector row, and on a pruning mask I, 1 of a corresponding element position is set as 0, so that the number of 0 on the pruning mask I meets the requirements of the vector-wise fine-grained sparsity; Zhuo [0031]: wherein the joint sparse process is specifically a process of obtaining pruning masks having different pruning granularities by presetting a target sparsity and a mixing ratio of granularity by a user, and the joint sparse process comprises independent vector-wise fine-grained sparsity and block-wise coarse-grained sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) such that each of the first set of B sub-blocks of the sparsified first block has S% sparsity, (Zhuo [0011]: in the vector-wise fine-grained sparsity, a weight matrix with the number of rows being #row and the number of columns being #col is filled with zero columns at an edge of the matrix, so that the number of columns of a zero-added minimum matrix is exactly divided by K, and the zero-added minimum matrix is divided into several vector rows with the number of rows being 1 and the number of columns being K; for each vector row, amplitude-based pruning is performed on an element in the vector row, and on a pruning mask I; Zhuo Fig. 1a: shows that each of the blocks have S% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) wherein S is based on a target level of combined sparsity [for a hardware configuration of the computing system;] (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) [applying a second sparsity mask] having S'% sparsity over M elements [to obtain a sparsified second block] having coarse-grained balanced sparsity, (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo [0031]: wherein the joint sparse process is specifically a process of obtaining pruning masks having different pruning granularities by presetting a target sparsity and a mixing ratio of granularity by a user, and the joint sparse process comprises independent vector-wise fine-grained sparsity and block-wise coarse-grained sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) such that S'% of the second set of B sub-blocks of the sparsified second block have 100% sparsity (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo Fig. 1c: shows that S' blocks have 100% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) and (100-S')% of the second set of B sub-blocks have 0% sparsity (Zhuo [0012]: in the block-wise coarse-grained sparsity, a weight matrix with a row number being #row and a column number being #col is filled with zero rows and/or zero columns at an edge of the matrix, so that a zero-added minimum matrix is exactly divided by blocks with sizes of R rows and S columns, and is divided into several vector blocks with the number of rows being R and the number of columns being S; an importance psum of each vector block not containing zero-filled rows or zero columns are calculated; amplitude-based pruning is performed on all vector blocks participating in the calculation of the importance psum according to the importance psum and size; and 1 of the corresponding element position of the vector block participating in the calculation of the importance psum on a pruning mask II is set to 0, so that the number of 0 on the pruning mask II meets the requirements of sparsity of the block-wise coarse-grained sparsity; Zhuo Fig. 1c: shows that 100-S' blocks have 0% sparsity; Zhuo [0041]: if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively) wherein S' is based on the target level of combined sparsity; (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya with the block pruning as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). With regards to claim 21, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu further teaches [wherein the target level of combined sparsity] is based on a hardware configuration of the computing system (Zhu [0025]: The disclosed embodiments relate to computer-implemented systems and methods for providing block-wise sparse neural networks. Advantageously, the exemplary embodiments can provide improved speed and power efficiency by reducing both mathematical operations and memory transfers required to execute the neural network). Zhu fails to teach wherein the target level of combined sparsity [is based on a hardware configuration of the computing system]. However, Zhuo teaches wherein the target level of combined sparsity [is based on a hardware configuration of the computing system] (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the combined sparsity as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). With regards to claim 22, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 11 above. Zhu further teaches [wherein the target level of combined sparsity] is based on a hardware configuration of the computing system (Zhu [0025]: The disclosed embodiments relate to computer-implemented systems and methods for providing block-wise sparse neural networks. Advantageously, the exemplary embodiments can provide improved speed and power efficiency by reducing both mathematical operations and memory transfers required to execute the neural network). Zhu fails to teach wherein the target level of combined sparsity [is based on a hardware configuration of the computing system]. However, Zhuo teaches wherein the target level of combined sparsity [is based on a hardware configuration of the computing system] (Zhuo [0041]: In order to obtain the mixed sparse granularity of the joint sparse method, an artificially set hyperparameter is set in the present invention, and represented as a granularity mixing ratio p, so as to control the sparsity ratio of a target sparsity contribution of vector-wise fine-grained sparsity. For example, if the target sparsity of the convolutional layer is 0.7 (i.e. the ratio of zeros in the weight matrix of the pruned convolutional layer reaches 70%), and the mixing ratio p of the granularity is 0.8, then the sparsities contributed by the fine-grained sparsity and the block-wise coarse sparsity should be 0.56 and 0.14, respectively). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the combined sparsity as taught by Zhuo. One of ordinary skill in the art would be motivated to make this combination because it can realize hybrid sparse granularity, thereby reducing reasoning overheads and ensuring the accuracy of a model as taught by Zhuo (Zhuo [0022]). Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu in view of Mathuriya further in view of Zhuo further in view of Kumar et al. (“Pruning filters with L1-norm and capped L1-norm for CNN compression”) hereinafter Kumar. With regards to claim 3, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu fails to teach wherein one or more of the first sparsity mask and the second sparsity mask are generated based on a set of lowest one-norms for a respective set of M/B elements. However, Kumar does teach wherein one or more of the first sparsity mask and the second sparsity mask are generated based on a set of lowest one-norms for a respective set of M/B elements (Kumar Page 1153 Section 1 Right Column: we used an L1-norm and Capped L1-norm based filter pruning to tackle the aforementioned issues). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the lowest 1 norm masks as taught by Kumar. One of ordinary skill in the art would be motivated to make this combination because right after the pruning process, the result in a slimmer network is far more compact concerning runtime memory, model size, and computational cost with comparison to the original wide architecture as taught by Kumar (Kumar Page 1153 Section 1 Right Column). With regards to claim 13, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 11 above. Zhu fails to teach wherein one or more of the first sparsity mask and the second sparsity mask are generated based on a set of lowest one-norms for a respective set of M/B elements. However, Kumar does teach wherein one or more of the first sparsity mask and the second sparsity mask are generated based on a set of lowest one-norms for a respective set of M/B elements (Kumar Page 1153 Section 1 Right Column: we used an L1-norm and Capped L1-norm based filter pruning to tackle the aforementioned issues). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the lowest 1 norm masks as taught by Kumar. One of ordinary skill in the art would be motivated to make this combination because right after the pruning process, the result in a slimmer network is far more compact concerning runtime memory, model size, and computational cost with comparison to the original wide architecture as taught by Kumar (Kumar Page 1153 Section 1 Right Column). Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu in view of Mathuriya further in view of Zhuo further in view of Schlicht et al. (US 20220222537 A1) hereinafter Schlicht. With regards to claim 9, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 8 above. Zhu further teaches and wherein the second sparsity mask is dynamically recomputed for each forward phase of the inference operation (Zhu [0086]: At step 759, the at least one processor may determine if the re-trained neural network has converged. For example, the at least one processor may determine convergence has occurred when a desired sparsity level has been reached, when an accuracy of the neural network has dropped below a threshold (e.g., as performed by the pseudocode described above), or any other value associated with the neural network has reached or crossed a predetermined threshold. If converged, method 750 may end; if not, method 750 may iterate, as depicted in FIG. 7B). Zhu fails to teach wherein the first sparsity mask is maintained during each iteration of the inference operation. However, Schlicht teaches wherein the first sparsity mask is maintained during each iteration of the inference operation (Schlicht [0023]: the classic filter is initialized with fixed filter parameters when the deep neural network is initialized). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with maintaining the first sparsity mask as taught by Schlicht. One of ordinary skill in the art would be motivated to make this combination because overall, this may increase the robustness of the deep neural network with respect to interference as taught by Schlicht (Schlicht [0018]). With regards to claim 18, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 17 above. Zhu further teaches and wherein the second sparsity mask is dynamically recomputed for each forward phase of the inference operation (Zhu [0086]: At step 759, the at least one processor may determine if the re-trained neural network has converged. For example, the at least one processor may determine convergence has occurred when a desired sparsity level has been reached, when an accuracy of the neural network has dropped below a threshold (e.g., as performed by the pseudocode described above), or any other value associated with the neural network has reached or crossed a predetermined threshold. If converged, method 750 may end; if not, method 750 may iterate, as depicted in FIG. 7B). Zhu fails to teach wherein the first sparsity mask is maintained during each iteration of the inference operation. However, Schlicht teaches wherein the first sparsity mask is maintained during each iteration of the inference operation (Schlicht [0023]: the classic filter is initialized with fixed filter parameters when the deep neural network is initialized). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with maintaining the first sparsity mask as taught by Schlicht. One of ordinary skill in the art would be motivated to make this combination because overall, this may increase the robustness of the deep neural network with respect to interference as taught by Schlicht (Schlicht [0018]). Claims 10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu in view of Mathuriya further in view of Zhuo further in view of Li et al. (“Comparing BERT and XLNet from the Perspective of Computational Characteristics”) hereinafter Li further in view of Ahmad et al. (“Optimizing Hardware Accelerated General Matrix-Matrix Multiplication for CNNs on FPGAs”) hereinafter Ahmad. With regards to claim 10, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 1 above. Zhu fails to teach wherein matrix multiplying the sparsified first block and the sparsified second block to generate the resulting block of elements occurs within a self-attention layer of a transformer language model. However, Li does teach wherein matrix multiplying the sparsified first block and the sparsified second block to generate the resulting block of elements occurs within a self-attention layer of a transformer language model, (Li Page 1 Section 2: BERT is a popular attention-based Transformer language model and auto-encoding model... The encoder uses a self attention layer as a unit structure. As shown in Figure 1, it calculates the final attention output by using a query/key/value (feature), which is obtained by matrix multiplication). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the self-attention layer as taught by Li. One of ordinary skill in the art would be motivated to make this combination because recently, Transformer [6], which is based solely on attention mechanism, shows superior performance in training time and accuracy compared to CNN and RNN on various NLP (Natural Language Processing) tasks as taught by Li (Li Page 1 Section 1). Zhu in view of Mathuriya further in view of Zhuo further in view of Li fails to teach and wherein the first block of elements and second block of elements are both derived from activation matrices. However, Ahmad does teach and wherein the first block of elements and second block of elements are both derived from activation matrices (Ahmad Page 2695 Section V A: GeMM operations between activation matrices). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo further in view of Li with the activation matrices as taught by Ahmad. One of ordinary skill in the art would be motivated to make this combination because it improves the efficiency of the layers of the neural network as taught by Ahmad (Ahmad Page 2692 Abstract). With regards to claim 19, Zhu in view of Mathuriya further in view of Zhuo teaches all of the limitations of claim 11 above. Zhu fails to teach wherein the sparse matrix multiplication occurs within a self-attention layer of a transformer language model. However, Li does teach wherein the sparse matrix multiplication occurs within a self-attention layer of a transformer language model, (Li Page 1 Section 2: BERT is a popular attention-based Transformer language model and auto-encoding model... The encoder uses a self attention layer as a unit structure. As shown in Figure 1, it calculates the final attention output by using a query/key/value (feature), which is obtained by matrix multiplication). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo with the self-attention layer as taught by Li. One of ordinary skill in the art would be motivated to make this combination because recently, Transformer [6], which is based solely on attention mechanism, shows superior performance in training time and accuracy compared to CNN and RNN on various NLP (Natural Language Processing) tasks as taught by Li (Li Page 1 Section 1). Zhu in view of Mathuriya further in view of Zhuo further in view of Li fails to teach and wherein the first block of elements and second block of elements are both derived from activation matrices. However, Ahmad does teach and wherein the first block of elements and second block of elements are both derived from activation matrices (Ahmad Page 2695 Section V A: GeMM operations between activation matrices). Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Zhu in view of Mathuriya further in view of Zhuo further in view of Li with the activation matrices as taught by Ahmad. One of ordinary skill in the art would be motivated to make this combination because it improves the efficiency of the layers of the neural network as taught by Ahmad (Ahmad Page 2692 Abstract). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jakob O Gudas whose telephone number is (571)272-0695. The examiner can normally be reached Monday-Thursday: 7:30AM-5:00PM Friday: 7:30AM-4:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.O.G./Examiner, Art Unit 2151 /James Trujillo/Supervisory Patent Examiner, Art Unit 2151
Read full office action

Prosecution Timeline

Show 3 earlier events
Sep 18, 2025
Examiner Interview Summary
Oct 15, 2025
Response Filed
Feb 13, 2026
Final Rejection mailed — §103
Mar 31, 2026
Examiner Interview Summary
Mar 31, 2026
Applicant Interview (Telephonic)
May 13, 2026
Request for Continued Examination
May 17, 2026
Response after Non-Final Action
Aug 18, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744548
PERFORMING COMPARISON OPERATIONS USING VECTOR FLOATING POINT VALUES
4y 6m to grant Granted Sep 22, 2026
Patent 12730610
INTEGRATED CIRCUIT AND METHOD OF OPERATING SAME
4y 7m to grant Granted Sep 08, 2026
Patent 12675257
Hybrid Compute-in-Memory
4y 3m to grant Granted Jul 07, 2026
Patent 12645426
System and Method for Accelerating Neural Networks
4y 5m to grant Granted Jun 02, 2026
Patent 12602200
ANALOG MULTIPLY-ACCUMULATE UNIT FOR MULTIBIT IN-MEMORY CELL COMPUTING
4y 6m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+64.2%)
4y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 19 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month