DETAILED ACTION
Claims 1-18 and 20-21 are pending.
The office acknowledges the following papers:
Claims and remarks filed on 4/22/2026,
IDS filed on 5/4/2026.
Withdrawn objections and rejections
The specification objection has been withdrawn.
The 35 U.S.C. 112(a) rejections for claims 1-20 have been withdrawn due to amendment.
The 35 U.S.C. 112(b) rejections for claims 1-20 have been withdrawn due to amendment.
New Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-2 are rejected under 35 U.S.C. 102(a)(1 & 2) as being anticipated by Gunnam et al. (U.S. 2021/0191733).
As per claim 1:
Gunnam disclosed a neural network inference accelerator, comprising:
a memory configured to store an activation tensor and a weight tensor (Gunnam: Figures 2 and 4 elements 215, 220, and 430, paragraphs 32 and 47)(Input features and weights are compressed and stored in either DRAM or SRAM.);
a first neural processing unit configured to receive the activation tensor and the weight tensor from the memory based on an activation sparsity density of the activation tensor and a weight sparsity density of the weight tensor corresponding to the activation tensor both being greater than a predetermined sparsity density (Gunnam: Figures 2 and 9 elements 245 and 900, paragraphs 34, 36, and 64)(The sparse tensor compute unit (i.e. first NPU) receives input feature maps and weights from the SRAM/DRAM based on sparsity exceeding a predetermined threshold (e.g. 50%).);
a second neural processing unit configured to receive the activation tensor and the weight tensor from the memory based on at least one of the activation sparsity density of the activation tensor or the weight sparsity density of the weight tensor corresponding to the activation tensor being less than or equal to the predetermined sparsity density (Gunnam: Figures 2 and 5 elements 240 and 500, paragraphs 34, 36, and 54)(The dense tensor compute cluster (i.e. second NPU) receives input feature maps and weights from the SRAM/DRAM based on a sparsity being less than a predetermined threshold (e.g. 50%).); and
sparsity management circuitry configured to control transfer of the activation tensor and the weight tensor corresponding to the activation tensor from the memory to the first neural processing unit or to the second neural processing unit based on the activation sparsity density of the activation tensor and the weight sparsity density of the weight tensor with respect to the predetermined sparsity density (Gunnam: Figures 2 and 10 elements 235 and 1025-1030, paragraphs 34, 36-37, and 73-75)(The scheduling engine (i.e. sparsity management circuitry) assigns input and weight tensors to dense and sparse tensor compute clusters based on the sparsity analysis. Gunnam gives the example that input feature maps and weights that have a number of zero elements exceeding 50% are allocated to the sparse tensor compute cluster. Additionally, Gunnam gives the example that input feature maps and weights that have a number of zero elements less than 50% are allocated to the dense tensor compute cluster.).
As per claim 2:
Gunnam disclosed the neural network inference accelerator of claim 1, wherein the first neural processing unit is configured to compute a first result for the activation tensor and the weight tensor (Gunnam: Figures 2, 6-7, and 9 elements 245, 615, 700, and 900, paragraphs 34, 36, 56, 58, and 64)(The sparse tensor compute unit (i.e. first NPU) receives input feature maps and weights from the SRAM/DRAM based on sparsity exceeding a predetermined threshold (e.g. 50%). The systolic array performs matrix dot-product calculations.), and
wherein the second neural processing unit is configured to compute a second result for the activation tensor and the weight tensor (Gunnam: Figures 2 and 5-7 elements 240, 500, 615, and 700, paragraphs 34, 36, 54, 56, and 58)(The dense tensor compute cluster (i.e. second NPU) receives input feature maps and weights from the SRAM/DRAM based on a sparsity being less than a predetermined threshold (e.g. 50%). The systolic array performs matrix dot-product calculations.).
New Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3-4 are rejected under 35 U.S.C. 103 as being unpatentable over Gunnam et al. (U.S. 2021/0191733), further in view of Official Notice.
As per claim 3:
Gunnam disclosed the neural network inference accelerator of claim 2, further comprising a compressor unit configured to receive and compress the first result computed by the first neural processing unit, and to receive and compress the second result computed by the second neural processing unit (Gunnam: Figures 2 and 4 elements 215, 220, and 430, paragraphs 32 and 47)(Gunnam disclosed compressing input features and weights into either DRAM or SRAM, but doesn’t discuss compressing execution results. Official notice is given that execution results can be compressed for storage into memory for the advantage of reducing storage requirements. Thus, it would have been obvious to one of ordinary skill in the art to implement compressing execution results from the tensor compute clusters.), and
wherein the memory is further configured to store the first result compressed by the compressor unit and store the second result compressed by the compressor unit (Gunnam: Figures 2 and 4 elements 215, 220, and 430, paragraphs 32 and 47)(In view of the above official notice, tensor compute results are compressed and stored in the DRAM or SRAM.).
As per claim 4:
Gunnam disclosed the neural network inference accelerator of claim 3, wherein the compressor unit is further configured to generate first metadata associated with the first result and to generate second metadata associated with the second result, and wherein the memory is further configured to store the first metadata and the second metadata (Gunnam: Figures 2 and 4 elements 215, 220, and 430, paragraphs 31-33, 42-43, 47, and 52-53)(Gunnam disclosed compressing input features and weights into either DRAM or SRAM, as well as information indicating compression information (i.e. metadata). In view of the above official notice, tensor compute results are compressed and stored in the DRAM or SRAM with corresponding compression information (i.e. first/second metadata.).
Claims 5-13, 17-18, and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Gunnam et al. (U.S. 2021/0191733), in view of Fishel et al. (U.S. 2019/0340488).
As per claim 5:
The additional limitation(s) of claim 5 basically recite the additional limitation(s) of claim 11. Therefore, claim 5 is rejected for the same reason(s) as claim 11.
As per claim 6:
Gunnam and Fishel disclosed the neural network inference accelerator of claim 5, wherein the decompressor unit is further configured to decompress the activation tensor to the activation sparsity density using third metadata associated with the activation tensor based on the activation tensor being compressed, and to decompress the weight tensor to the weight sparsity density using fourth metadata associated with the weight tensor based on the weight tensor being compressed (Fishel: Figures 3-4 elements 314A-N and 432, paragraph 48, 81)(Gunnam: Figures 2, 4, and 6 elements 215, 220, 430, and 615, paragraphs 31-33, 42-43, 47, 52-53, and 56)(Gunnam disclosed that input features and weights are compressed and stored in either DRAM or SRAM, as well as information indicating compression information (i.e. metadata). Fishel disclosed decompressing compressed kernel data prior to processing. The combination implements decompression elements for the input features and weights such that these compressed inputs are decompressed before systolic array processing using the corresponding compression information for both input features and weights (i.e. third and fourth metadata).).
As per claim 7:
Gunnam and Fishel disclosed the neural network inference accelerator of claim 5, wherein the activation sparsity density is based on a structured-sparsity arrangement or a random-sparsity arrangement (Gunnam: Figures 2 and 10 elements 235 and 1025-1030, paragraphs 34, 36-37, and 73-75)(The scheduling engine (i.e. sparsity management unit) assigns input and weight tensors to dense and sparse tensor compute clusters based on the sparsity analysis. Gunnam gives examples of sparse input feature tensor having 80% zeros and greater than 50% zeros (i.e. structured-sparsity).).
As per claim 8:
Gunnam and Fishel disclosed the neural network inference accelerator of claim 7, wherein the activation sparsity density is based on a 1:4 structured-sparsity arrangement or a 2:8 structured-sparsity arrangement (Gunnam: Figures 2 and 10 elements 235 and 1025-1030, paragraphs 34, 36-37, and 73-75)(The scheduling engine (i.e. sparsity management unit) assigns input and weight tensors to dense and sparse tensor compute clusters based on the sparsity analysis. Gunnam gives examples of sparse input feature tensors having 80% zeros and greater than 50% zeros (i.e. structured-sparsity). It would have been obvious to one of ordinary skill in the art that the sparsity threshold can be 75% zeros. In addition, according to “In re Rose” (105 USPQ 237 (CCPA 1955)), changes in size or range doesn’t give patentability over prior art.).
As per claim 9:
Gunnam and Fishel disclosed the neural network inference accelerator of claim 5, wherein the weight sparsity density is based on a structured-sparsity arrangement or a random-sparsity arrangement (Gunnam: Figures 2 and 10 elements 235 and 1025-1030, paragraphs 34, 36-37, and 73-75)(The scheduling engine (i.e. sparsity management unit) assigns input and weight tensors to dense and sparse tensor compute clusters based on the sparsity analysis. Gunnam gives examples of sparse weight tensors having 80% zeros and greater than 50% zeros (i.e. structured-sparsity).).
As per claim 10:
Gunnam and Fishel disclosed the neural network inference accelerator of claim 9, wherein the weight sparsity density is based on a 1:4 structured-sparsity arrangement or a 2:8 structured-sparsity arrangement (Gunnam: Figures 2 and 10 elements 235 and 1025-1030, paragraphs 34, 36-37, and 73-75)(The scheduling engine (i.e. sparsity management unit) assigns input and weight tensors to dense and sparse tensor compute clusters based on the sparsity analysis. Gunnam gives examples of sparse weight tensors having 80% zeros and greater than 50% zeros (i.e. structured-sparsity). It would have been obvious to one of ordinary skill in the art that the sparsity threshold can be 75% zeros. In addition, according to “In re Rose” (105 USPQ 237 (CCPA 1955)), changes in size or range doesn’t give patentability over prior art.).
As per claim 11:
Claim 11 essentially recites the same limitations of claim 1. Claim 11 additionally recites the following limitations:
a decompressor unit configured to decompress an activation tensor to a first predetermined sparsity density based on the activation tensor being compressed, and to decompress a weight tensor to a second predetermined sparsity density based on the weight tensor being compressed (Fishel: Figures 3-4 elements 314A-N, 432, paragraph 48, 81)(Gunnam: Figures 2, 4, and 6 elements 215, 220, 430, and 615, paragraphs 32, 47, and 56)(Gunnam disclosed that input features and weights are compressed and stored in either DRAM or SRAM. Fishel disclosed decompressing compressed kernel data prior to processing. The combination implements decompression elements for the input features and weights such that these compressed inputs are decompressed before systolic array processing.).
Gunnam disclosed compressing input features and weights into memory, but doesn’t explicitly state how the tensor compute units handle compressed inputs. One of ordinary skill in the art would have been motivated by this lack of teaching in Gunnam to find the Fishel reference that disclosed neural engines including decompression logic to decompress compressed weight inputs. Thus, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date to implement the decompression logic of Fishel into the system of Gunnam for the advantage of accurately handling incoming compressed input features and weights that are to be processed.
As per claim 12:
Claim 12 essentially recites the same limitations of claim 1. Claim 12 additionally recites the following limitations:
wherein the decompressor unit receives the activation tensor and weight tensor from the memory (Fishel: Figures 3-4 elements 314A-N, 432, paragraph 48, 81)(Gunnam: Figures 2, 4, and 6 elements 215, 220, 430, and 615, paragraphs 32, 47, and 56)(The combination implements decompression elements for the input features and weights such that these compressed inputs are decompressed before systolic array processing. The decompression elements receive compressed inputs from the SRAM/DRAM.).
As per claim 13:
The additional limitation(s) of claim 13 basically recite the additional limitation(s) of claim 2. Therefore, claim 13 is rejected for the same reason(s) as claim 2.
As per claim 17:
The additional limitation(s) of claim 17 basically recite the additional limitation(s) of claim 7. Therefore, claim 17 is rejected for the same reason(s) as claim 7.
As per claim 18:
The additional limitation(s) of claim 18 basically recite the additional limitation(s) of claim 8. Therefore, claim 18 is rejected for the same reason(s) as claim 8.
As per claim 20:
The additional limitation(s) of claim 20 basically recite the additional limitation(s) of claim 10. Therefore, claim 20 is rejected for the same reason(s) as claim 10.
As per claim 21:
Gunnam and Fishel disclosed the neural network inference accelerator of claim 11, wherein a compression mask generated after training a dense neural network model is used to decompress the weight tensor to the second predetermined sparsity density based on the weight tensor being compressed and wherein the compression mask is used as metadata for output to a metadata buffer of a memory configured to store the activation tensor and the weight tensor (Fishel: Figures 3-4 elements 314A-N, 432, paragraph 48, 81)(Gunnam: Figures 2, 4, and 6 elements 215, 220, 430, and 615, paragraphs 32, 47, and 56)(Gunnam disclosed that input features and weights are compressed and stored in either DRAM or SRAM. Fishel disclosed decompressing compressed kernel data prior to processing using compression masks that indicate non-zero locations. The combination implements decompression elements for the input features and weights such that these compressed inputs are decompressed before systolic array processing. The combination stores the compression masks in memory as metadata.).
Claims 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Gunnam et al. (U.S. 2021/0191733), in view of Fishel et al. (U.S. 2019/0340488), in view of Official Notice.
As per claim 14:
The additional limitation(s) of claim 14 basically recite the additional limitation(s) of claim 3. Therefore, claim 14 is rejected for the same reason(s) as claim 3.
As per claim 15:
The additional limitation(s) of claim 15 basically recite the additional limitation(s) of claim 4. Therefore, claim 15 is rejected for the same reason(s) as claim 4.
As per claim 16:
The additional limitation(s) of claim 16 basically recite the additional limitation(s) of claim 6. Therefore, claim 16 is rejected for the same reason(s) as claim 6.
Response to Arguments
The arguments presented by Applicant in the response, received on 4/22/2026 are not considered persuasive.
Applicant argues regarding claims 1 and 11:
“Applicant submits that analyzing the sparsity of an input feature tensor and assigning the input feature tensor to a dense cluster, a sparse cluster, or an accelerator based the sparsity does not correspond to "a second neural processing unit configured to receive the activation tensor and the weight tensor . .. based on at least one of the activation sparsity density . .. or the weight sparsity density ... being less than or equal to the predetermined sparsity density" as recited in amended claim 1. This is because an input feature tensor is not an activation tensor and a weight tensor.”
This argument is not found to be persuasive for the following reason. The claims don’t require the input feature tensor of Gunnam to be mapped to both of the claimed activation tensor and weight tensor. Instead, Gunnam’s input feature tensor maps to the claimed activation tensor. In this type of matrix convolutions, input feature and activations are different terms for the same input matrix data provided to be multiplied with input weights. Thus, Gunnam reads upon the claimed limitation at issue.
Applicant argues regarding claims 1 and 11:
“Applicant submits that a dense tensor compute cluster configured to process an input feature tensor having a number of zeros below a threshold and a sparse tensor compute cluster configured to process an input feature tensor having a number of zeros above a threshold does not correspond to "a first neural processing unit configured to receive the activation tensor and the weight tensor . .. based on an activation sparsity density . .. and a weight sparsity density ... both being greater than a predetermined sparsity density; [and] a second neural processing unit configured to receive the activation tensor and the weight tensor . .. based on at least one of the activation sparsity density . .. or the weight sparsity density . .. being less than or equal to the predetermined sparsity density" as recited in amended claim 1. This is because a first compute cluster configured to process a "dense" input feature tensor is not a first neural processing unit configured to receive an activation tensor and a weight tensor if an activation sparsity density of the activation tensor and a weight sparsity density of the weight tensor are both greater than a predetermined sparsity density. This is also because a second compute cluster configured to process a "sparse" input feature tensor is not a second neural processing unit configured to receive the activation tensor and the weight tensor if at least one of the activation sparsity density of the activation tensor or the weight sparsity density of the weight tensor is less than or equal to the predetermined sparsity density. This is additionally because an input feature tensor is not a weight tensor and an activation tensor.”
This argument is not found to be persuasive for the following reason. The dense and sparse compute clusters reads upon the 1st and 2nd NPUs, respectfully. The dense compute cluster receives input activation tensors and weight tensors when the sparsity is less than a threshold. The sparse compute cluster receives input activation tensors and weight tensors when the sparsity exceeds a threshold. The claimed limitations include no further details that further differentiate or prevent the dense and sparse compute clusters from reading upon the 1st and 2nd NPUs. Thus, Gunnam reads upon the amended claim limitations.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The following is text cited from 37 CFR 1.111(c): In amending in reply to a rejection of claims in an application or patent under reexamination, the applicant or patent owner must clearly point out the patentable novelty which he or she thinks the claims present in view of the state of the art disclosed by the references cited or the objections made. The applicant or patent owner must also show how the amendments avoid such references or objections.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB A. PETRANEK whose telephone number is (571)272-5988. The examiner can normally be reached on M-F 8:00-4:30.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached on (571) 270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB PETRANEK/Primary Examiner, Art Unit 2183