DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-9, 12, 15-14 and 27 are pending in this application. Claims 1-9, 12, 15-24 and 27 are currently amended; claims 10-11, 13-14, 25-26 and 28-29 are canceled.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 07/27/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 4, 15-24 and 27 are objected to under 37 C.F.R. 1.71(a) which requires “full, clear, concise, and exact terms” as to enable any person skilled in the art or science to which the invention or discovery appertains, or with which it is most nearly connected, to make and use the same. The following should be corrected.
A. In claim 4 line 1, “the threshold value” should read “the non-zero threshold value” instead for consistency of claim terminologies. Claim 19 recites a similar limitation and is objected to for the same reason.
B. In claim 15 line 4, “processing circuitry” should read “the processing circuitry” instead because processing circuitry is already introduced in line 2.
B. In claim 16 lines 5-6, “processing circuitry” should read “the processing circuitry” instead because processing circuitry is already introduced in line 2. Claims 17-24 and 27 inherit the same deficiency as claim 16 by reason of dependence.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 4-8, 12, 15-17, 19-23 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 20220004385 A1), hereinafter Zhang, in view of Bleiweiss et al. (US 20190205737 A1), hereinafter Bleiweiss, and Gunnam et al. (US 20210191733 A1), hereinafter Gunnam.
Regarding claim 16, Zhang teaches a system for thread reduction in tensor processing, comprising:
processing circuitry (Zhang Fig. 1 and paragraph [0020] processing circuitry - graphics processing unit 100); and
memory, (Zhang paragraph [0020] memory – storage):
(Zhang paragraph [0023 and Table 1; plurality of blocks – matrices 1-5) , the processing circuitry configured to process in parallel a plurality of threads (Zhang Fig. 3 and paragraph [0003] “GPU has achieved the first-mover advantage in performing convolution neural network calculation with its own mature parallel computing hardware architecture and software applications”; paragraph [0025]);
perform a first depth test a first (Zhang Fig. 3 and paragraph [0029] “determining whether the matrices are zero matrices or non-zero matrices (step S304)”; depth test - determining whether a matrix is a zero or a non-zero matrix; first block – one of the zero matrices);
generate a first indicator bit value for the first (Zhang Fig. 3 and paragraph [0029] “In step S310, the method of the present invention marks the zero matrices”; paragraphs [0022-0024] “when a matrix in the sparse matrix detection unit 102 meets a first condition (for example, the matrix is a zero matrix), the assertion register 106 marks the matrix as "0" … As shown in Table 1, all elements in the matrices 1 and 4 are 0, thus the sparse matrix detection unit 102 determines the matrices 1 and 4 are zero matrices, and the assertion register 106 marks the matrices 1 and 4 as "0"”; indicator bit – mark; first indicator bit value – “0”);
generate a second indicator bit value for a second (Zhang Fig. 3 and paragraph [0029] “In step S306, the method of the present invention marks the nonzero matrices”; paragraphs [0022-0024] “When a matrix in the sparse matrix detection unit 102 meets a second condition (for example, the matrix is a non-zero matrix), the assertion register 106 marks the matrix as "1" … The matrices 2, 3, and 5 all have at least one non-zero element, thus the sparse matrix detection unit 102 determines the matrices 2, 3, and 5 are non-zero matrices, and the assertion register 106 marks the matrices 2, 3, and 5 as "1"”; second block - one of the non-zero matrices; second indicator bit value – “1”); and
spawn a thread for processing the second (Zhang paragraphs [0025-0026] “In FIG. 1, the thread scheduling and instruction distribution unit 110 selects a thread to be processed from a plurality of threads, fetches the corresponding instructions, and sends the corresponding instructions to the integer calculation unit 112 … For example, the thread scheduling and instruction distribution unit 110 … sends an integer calculation instruction and a matrix calculation instruction to the integer calculation unit 112 … The integer calculation unit 112 reads the integer calculation instruction from the instruction set, and passes the matrix calculation instruction in the instruction set to the matrix calculation unit 108 … the calculation sub-units 202 of the matrix calculation unit 108 perform matrix calculations on the non-zero matrices, and ignore the zero matrices:”; paragraph [0029] “in step S308, the method of the present invention reads and performs matrix calculations on the non-zero matrices”); and
refrain from spawning another thread for processing the first (Zhang Fig. 3 step S312 and paragraph [0029] “In step S310, the method of the present invention marks the zero matrices. Then, in step S312, the method of the present invention directly ignores the zero matrices, and does not perform matrix calculations on the zero matrices”).
Zhang does not explicitly teach the memory containing instructions; generate a plurality of result blocks by processing a tensor input with a predefined kernel on processing circuitry; perform a first depth test a first result block of the plurality of result blocks; generate a first indicator bit value for the first result block based on the first depth test, the first indicator bit value indicating that the first result block is sparse, comprising comparing one or more values of the first result block with a non-zero threshold value; generate a second indicator bit value for a second result block of the plurality of result blocks based on a second depth test, the second indicator bit value indicating that the second result block is not sparse; and spawn a thread for processing the second result block and a second predefined kernel, based on the indicator bit having the second indicator bit value; and refrain from spawning another thread for processing the first result block based on the first indicator bit value.
However, on the same field of endeavor, Bleiweiss discloses a memory containing instructions (Bleiweiss Fig. 1 and paragraph [0054] “the memory device 120 can operate as system memory for the system 100, to store data 122 and instructions 121 for use when the one or more processors 102 executes an application or process). Further, Bleiweiss discloses generating a plurality of result blocks by processing a tensor input with a predefined kernel on a processing circuitry and processing the plurality of result blocks and a second predefined kernel (Bleiweiss Figs. 16A and 25 and paragraphs [0165, 0171-0172, 0232]; plurality of result blocks – output of a convolution layer; a tensor input – input of the convolution layer; predefined kernel – convolution kernel; second predefined kernel - convolution kernel of a subsequent convolution layer).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Zhang using Bleiweiss and configure the system to store the instructions in the memory because memory devices are normally used for storing the instructions used by a processing circuitry when the processing circuitry executes an application or process (Bleiweiss paragraph [0054]). Further, configure the system such that the matrices 1-5 in Table 1 are output matrices of a layer of a neural network generated by processing a tensor input with a predefined kernel on the processing circuitry which are input to a next layer to be processed with a second predefined kernel in order to implement a system for processing layers of a convolutional neural network (CNN) (Bleiweiss paragraphs [0165, 0171]).
Therefore, the combination of Zhang as modified in view of Bleiweiss teaches the memory containing instructions; generate a plurality of result blocks by processing a tensor input with a predefined kernel on processing circuitry; perform a first depth test a first result block of the plurality of result blocks; generate a first indicator bit value for the first result block based on the first depth test, the first indicator bit value indicating that the first result block is sparse; generate a second indicator bit value for a second result block of the plurality of result blocks based on a second depth test, the second indicator bit value indicating that the second result block is not sparse; and spawn a thread for processing the second result block and a second predefined kernel, based on the indicator bit having the second indicator bit value; and refrain from spawning another thread for processing the first result block based on the first indicator bit value.
Zhang does not explicitly teach generate a first indicator bit value for the first result block based on the first depth test, the first indicator bit value indicating that the first result block is sparse, comprising comparing one or more values of the first result block with a non-zero threshold value.
However, on the same field of endeavor, Gunnam discloses determining whether a matrix is sparse by comprising comparing one or more values of the matrix with a non-zero threshold value (Gunnam paragraph [0073] “To perform the sparsity analysis, the accelerator 200 determines the number or percentage of non-zeroes in each of the input feature tensors, and compares the number or percentage against a predetermined threshold programmed within the accelerator 200 (e.g., within the scheduling engine 235). For example, any input feature tensor that has the number or percentage of zeroes that is greater than or equal to the predetermined threshold may be considered a sparse tensor. For example, if a predetermined threshold is 50%, a tensor is determined to be a sparse tensor when it has more zero values than non-zero values. Similarly, any input feature tensor that has the number or percentage of zeroes that is less than the predetermined threshold may be considered a dense tensor. Following this same 50% threshold example, a dense tensor has more non-zero values than zero values”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Zhang using Gunnam and configure the system generate the first indicator bit value for the first result block indicating that the first result block is sparse by comparing one or more values of the first result block with a non-zero threshold value in order to further reduce the amount of data that needs to be computed with negligible accuracy loss (Gunnam paragraph [0021] by ignoring matrices that have non-zero values below a threshold. For example, marking the second matrix as zero since it only contains two non-zero matrices.
Therefore, the combination of Zhang as modified in view of Bleiweiss and Gunnam teaches generate a first indicator bit value for the first result block based on the first depth test, the first indicator bit value indicating that the first result block is sparse, comprising comparing one or more values of the first result block with a non-zero threshold value.
Regarding claim 17, Zhang as modified in view of Bleiweiss and Gunnam teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: store the second result block and the second indicator bit value in the memory (Zhang paragraphs [0022-0024] memory – register file and assertion register).
Regarding claim 19, Zhang as modified in view of Bleiweiss and Gunnam teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the threshold value is any one of: a dynamic value, a static non-zero value, and or an adaptive value (Zhang paragraphs [0021-0024] the threshold value is static as it is the same condition used for all five matrices; Gunnam paragraph [0034] “the sparsity analysis performed by the scheduling engine is a dynamic sparsity analysis”).
Regarding claim 20, Zhang as modified in view of Bleiweiss and Gunnam teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: perform a third depth test on two or more result blocks of the plurality of result blocks (Zhang paragraphs [0023, 0029] two or more result blocks – two or more other matrices of the plurality of matrices).
Regarding claim 21, Zhang as modified in view of Bleiweiss and Gunnam all the limitations of claim 20 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: generate the first indicator bit value for the two or more result blocks of the plurality of result blocks based on the third depth test, the first indicator bit value indicating that the two or more result blocks are sparse (Zhang Fig. 3 step S310 and paragraph [0029] “when the matrices are zero matrices, the method continues to perform step S310. In step S310, the method of the present invention marks the zero matrices”).
Regarding claim 22, Zhang as modified in view of Bleiweiss and Gunnam teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the first result block is a block representing a plurality of pixels (Bleiweiss Fig. 16A and paragraphs [0165, 0171] “ CNN is a specialized feedforward neural network for processing data having a known, grid-like topology, such as image data”).
Regarding claim 23, Zhang as modified in view of Bleiweiss and Gunnam teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: process the thread on a core of the processing circuitry, wherein the processing circuitry is a multi-core processing circuitry; and generate a third result block based on an output of the core (Zhang paragraphs [0003, 0026-0027, 0032] “(for example, the matrix calculation unit 108) is added to a single core of the graphics processing unit”; second result - matrix calculation result; Bleiweiss Figs. 1-2 and 14A-14B and paragraphs [0150, , 0152-0155]).
Regarding claim 27, Zhang as modified in view of Bleiweiss and Gunnam teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss and Gunnam teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: generate the first indicator bit value (Zhang paragraphs [0022-0024]).
Zhang does not explicitly teach generate the first indicator bit value based on a weight associated with the predefined kernel.
However, on the same field of endeavor, Bleiweiss discloses a weight associated with a predefined kernel (Bleiweiss paragraph [0232] “The forward pass compute multiplies an input matrix with a weight (or model) matrix”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Zhang using Bleiweiss and configure the system to also generate an indicator bit based on a weight of the predefined kernel in order to determine whether the kernel is a zero matrix or an non-zero matrix and to ignore matrix calculations in response to the kernel being a zero matrix to reduce the amount of calculation and/or number of data access (Zhang paragraph [0031]).
Therefore, the combination of Zhang as modified in view of Bleiweiss teaches generate the first indicator bit value based on a weight associated with the predefined kernel.
Regarding claims 1-2, 4-8 and 12, they are directed to a method practiced by the system of claims 16-17, 19-23 and 27 respectively. All steps performed by the method of claims 1-2, 4-8 and 12 would be practiced by the system of claims 16-17, 19-23 and 27 respectively. Claims 16-17, 19-23 and 27 analysis applies equally to claims 1-2, 4-8 and 12 respectively.
Regarding claim 15, it is directed to a non-transitory computer readable medium having stored thereon instructions for causing the processing circuitry of claim 16 to execute the process recited in claim 1 or 16. All the instructions stored in the non-transitory computer readable medium of claim 15 would be executed by the processing circuitry of claim 16. Claim 16 analysis applies equally to claims 15.
Claims 3 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Bleiweiss and Gunnam as applied to claims 1 and 16 above, and further in view of Baraniuk et al. (US 20140279727 A1), hereinafter Baraniuk.
Regarding claim 18, Zhang as modified in view of Bleiweiss teaches all the limitations of claim 16 as stated above. Further, Zhang as modified in view of Bleiweiss teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: generate the first indicator bit value indicating that the first result block is sparse, based on determining that the first depth test indicates that (Zhang paragraphs [0021-0024]; Gunnam paragraph [0073])).
Zhang does not explicitly teach generate the first indicator bit value indicating that the first result block is sparse, based on determining that the first depth test indicates that an average of the one or more values of the block first result block is within the non-zero threshold value.
However, on the same field of endeavor, Baraniuk discloses sparsifying a matrix based on an average value (Baraniuk paragraph [0421] average value – mean value).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Zhang using Baraniuk and generate the first indicator bit value indicating that the first result block is sparse, based on an average of the one or more values of the block first result block being within the non-zero threshold value in order to further reduce the amount of data that needs to be computed with negligible accuracy loss (Gunnam paragraph [0021]) by ignoring matrices that have an average non-zero values within a threshold (i.e., matrices with values close to zero).
Therefore, the combination of Zhang as modified in view of Bleiweiss, Zhang and Baraniuk teaches generate the first indicator bit value indicating that the first result block is sparse, based on determining that the first depth test indicates that an average of the one or more values of the block first result block is within the non-zero threshold value.
Regarding claim 3, it is directed to a method practiced by the system of claim 18. All steps performed by the method of claim 3 would be practiced by the system of claim 18. Claim 18 analysis applies equally to claim 3.
Claims 9 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Bleiweiss as applied to claims 8 and 23 above, and further in view of Ulrich et al. (US 20210349694 A1), hereinafter Ulrich.
Regarding claim 24, Zhang as modified in view of Bleiweiss teaches all the limitations of claim 23 as stated above.
Zhang does not explicitly teach wherein the instructions, when executed by the processing circuitry, further configure the system to: generate a fourth result block based on a plurality of zero values, in response to determining that a fifth result block has the first indicator bit value.
However, on the same field of endeavor, Ulrich discloses generating a result based on a zero value (Ulrich paragraph [0034, 0045]).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Zhang using Ulrich and configure the system to generate a result matrix based on a plurality of zero values in a zero matrix, in response to determining that the indicator bit has the first indicator bit value by outputting a corresponding zero matrix because the matrix multiplication result is already known to be zero and can be bypassed which reduces power consumption (Ulrich paragraphs [0039, 0045]). As discussed, Zhang discloses ignoring matrix calculation of zero matrices, therefore, it is obvious to directly output a zero result matrix as the matrix multiplication result is already known to be zero matrix.
Therefore, the combination of Zhang as modified in view of Bleiweiss and Ulrich teaches wherein the instructions, when executed by the processing circuitry, further configure the system to: generate a fourth result block based on a plurality of zero values, in response to determining that a fifth result block has the first indicator bit value.
Regarding claim 9, it is directed to a method practiced by the system of claim 24. All steps performed by the method of claim 9 would be practiced by the system of claim 24. Claim 24 analysis applies equally to claim 9.
Response to Arguments
In view of amendments made and Applicant’s arguments, the objection to the drawings and specification has been withdrawn.
The amendments made has not addressed all the objection to the claims. Further, the amendments made raises new objection to the claims as discussed above.
Applicant’s arguments, see remarks page 11-16, filed 08/26/2026, with respect to the 35 U.S.C. 112(b) and 35 U.S.C. 101 rejection of the claims have been fully considered and are persuasive. The 35 U.S.C. 112(b) and 35 U.S.C. 101 rejection of the claims has been withdrawn.
Applicant’s arguments, see remarks page 17-18, filed 08/26/2026, with respect to the rejection(s) of claim(s) 1-9, 12, 15-24 and 27 under 35 U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of amendments made and newly found prior art reference.
Applicant amended independent claims 1, 15 and 16 to recite the features of “generate a first indicator bit value for the first result block based on the first depth test, the first indicator bit value indicating that the first result block is sparse, comprising comparing one or more values of the first result block with a non-zero threshold value” and argues that Zhang fails to explicitly teach or suggest generating the indicator bit value by “comparing one or more values of the first result block with a non-zero threshold value”.
Response: Examiner agrees. However, the concept of determining the sparsity of a matrix by comparing one or more values of the matrix with a non-zero threshold value is disclosed by Gunnam in at least paragraph [0073]. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to combine the teaching of Zhang and Gunnam and generate the first indicator bit value for the first result block indicating that the first result block is sparse by comparing one or more values of the first result block with a non-zero threshold value in order to further reduce the amount of data that needs to be computed with negligible accuracy loss (Gunnam paragraph [0021] by ignoring matrices that have non-zero values below a threshold. For example, setting 80% as predetermined threshold value to determine that a matrix or sparse (i.e., 20% non-zero threshold value) as disclosed in paragraph [0034] of Gunnam would further result in determining the second matrix in paragraph [0023] of Zhang is sparse and would be ignored in the subsequent matrix calculation as it only contains 2/16 or 12.5% non-zero values.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Carlo Waje whose telephone number is (571)272-5767. The examiner can normally be reached 9:00-6:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Carlo Waje/Examiner, Art Unit 2151 (571)272-5767