Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 28 August 2026 have been fully considered but they are not persuasive.
Applicant contends that Wang further in view of Vengullur does not teach “instructing two processing elements to start computation at different times ‘in response to determining that the workload of the first processing element is greater than the workload of the second processing element,’” because Vengullur merely teaches staggering computations to avoid memory transaction conflicts. Remarks at 10-11.
The Examiner respectfully disagrees. Wang teaches the recited “determining that the workload of the first processing element is greater than the workload of the second processing element,” as recited in claim 1 (Wang, ¶ 120, “the controller can determine whether a workload imbalance exists between the first PE and the second PE based on the first number of the non-zero weights in the first weights and the second number of the non-zero weights in the second weights”). Wang additionally teaches that allocation is performed to balance the workload (¶ 119, “one or more non-zero weights of the first weights can be allocated by a controller in the neural network processor to the second PE for computing the PSUM of the output activation of the first OC to balance a workload”).
Vengallur provides the “instructing the first processing element to start a first computation at a first time” and “instructing the second processing element to start a second computation at a second time, the second time later than the first time” (see Non-Final Rejection at pages 5-6.
One cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. MPEP § 2145(IV).
It would have been obvious to apply Vengallur’s staggered PE starts as the action taken when Wang determines that the first PE’s workload is larger than the second PE’s.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 4, 6-9, 12, 14-17, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang (US 2023/0103750) and further in view of Vengallur (US 2020/0233803).
Regarding claim 1, Wang teaches: A method of scheduling computations in a deep neural network (DNN) (¶ 3, “A deep learning accelerator (DLA) as customized hardware can be used to accelerate the processing of deep neural networks (DNNs)”), comprising:
determining a workload for each respective processing element in a group of processing elements (¶ 105, “the controller 1130 can determine the workload of each OC (OC0 or OC1) based on a number of respective non-zero weights and detect workload imbalance between the pair of OCs (OC0 and OC1)”) based on an input operand (¶ 5, “the first weights and the second weights correspond to a same set of input channels (ICs) of the convolutional layer”) and a weight operand (¶ 4, “a partial sum (PSUM) of an output activation of the first OC based on the non-zero weights in the first weights”), the respective processing element configured to perform a computation on the input operand and weight operand (¶ 4, “computing, by a first PE in the neural network processor, a partial sum (PSUM) of an output activation of the first OC based on the non-zero weights in the first weights”), the input operand comprising one or more activations of a convolution (¶ 44, “the corresponding 36 input activations in the ICs 101-104”), the weight operand comprising one or more weights of the convolution (¶ 72, “Each set of the OC weights corresponding to an OC in a sequence of OCs of a convolutional layer”);
determining whether a workload of a first processing element in the group of processing elements is greater than a workload of a second processing element in the group of processing elements (¶ 119, “when a first number of the non-zero weights in the first weights is larger than a second number of the non-zero weights in the second weights”); and
in response to determining that the workload of the first processing element is greater than the workload of the second processing element (¶ 120, “the controller can determine whether a workload imbalance exists between the first PE and the second PE based on the first number of the non-zero weights in the first weights and the second number of the non-zero weights in the second weights”): performing an action (¶ 119, “At S1340, one or more non-zero weights of the first weights can be allocated by a controller in the neural network processor to the second PE for computing the PSUM of the output activation of the first OC to balance a workload”).
Wang does not teach as clearly as Vengallur discloses: instructing the first processing element to start a first computation at a first time (¶ 9, “A control finite state machine (FSM) may exploit the convolutional reuse of input feature maps to schedule the threads/grids in a staggered manner and avoid memory conflict”), and
instructing the second processing element to start a second computation at a second time (¶ 27, “The GCNN controller 204 generally triggers the computations in the tile groups 206 in a staggered manner such that the memory reads and/or writes from the tile groups 206 do not overlap and/or conflict”), the second time later than the first time (¶ 29, “In cycle 1, therefore, the PEs 210 of the tiles 207 of tile group 206-1 may compute a MAC operation 401-1. The MAC operation 401-1 may be based on the first row of the IFM 104 and the kernel 108. Doing so may produce an intermediate output pixel, which may be stored in the ORAM 212. In cycle 2, the data in the shift register 208-1 (e.g., the first row of the IFM 104) is shifted (e.g., a left shift) and the tile group 206-1 may compute a second MAC operation 401-2”).
It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the invention, to have applied the known technique of instructing the first processing element to start a first computation at a first time, and instructing the second processing element to start a second computation at a second time, the second time later than the first time, as taught by Vengallur, in the same way to the action performed responsive to detecting an imbalance, as taught by Wang. Both inventions are in the field of accelerating DNN processing, and combining them would have predictably resulted in “avoid[ing] memory conflict,” as indicated by Vengallur (¶ 9).
Regarding claim 4, Wang teaches: The method of claim 1, wherein determining whether the workload of the first processing element in the group of processing elements is greater than the workload of the second processing element in the group of processing elements comprises: determining whether a number of ones in a combined sparsity bitmap associated with the first processing element is greater than a number of ones in a combined sparsity bitmap associated with the second processing element (¶ 5, “the controller in the neural network processor, whether a workload imbalance exists between the first PE and the second PE based on the first number of the non-zero weights in the first weights and the second number of the non-zero weights in the second weights”).
Regarding claim 6, Vengallur teaches: The method of claim 1, wherein instructing the first processing element to start the first computation at the first time comprises: determining that the workload of the first processing element is greater than at least one workload of another processing element in the group of processing elements (¶ 84, “when the weight buffer 305 within any PE 210 becomes full, broadcasting of the weight values into the weight buffer 305 is stalled. If the weight buffer 305 is big enough to hold a few input channels of a tile, some PEs 210 can move ahead to the next input channel while one or more other PEs 210 are a few channels behind—smoothing out load imbalance between PEs 210”); and instructing the first processing element to start the first computation in a first clock cycle in a sequence of clock cycles, wherein the other processing element having less workload than the first processing element start in one or more clock cycles that are subsequent to the first clock cycle in the sequence of clock cycles (¶ 23, “The tile groups 206 may further share an ORAM 212 for storing intermediate OFMs 106 (e.g., a convolution operation requires several compute cycles, and the intermediate output may correspond to the output of one or more such compute cycles)”).
Regarding claim 7, Vengallur teaches: The method of claim 6, wherein instructing the second processing element to start the second computation at the second time comprises: associating a sequence of numbers with the sequence of clock cycles, each respective clock cycle associated with a greater number than another clock cycle subsequent to the respective clock cycle in the sequence of clock cycles (¶ 27, “shift register 208-1 serves tile groups 206-1 and 206-2, while shift register 208-2 serves tile groups 206-3 and 206-N. Doing so allows the input features to be reused over K cycles, where K is the dimensionality of the kernels 108”), the first clock cycle associated with a first number representing the workload of the first processing element (¶ 29, “operations performed by each tile group 206-1 through 206-4 of FIG. 2A is illustrated over N cycles”); determining a second number representing the workload of the second processing element (¶ 29, “In cycle 2, the data in the shift register 208-1 (e.g., the first row of the IFM 104) is shifted (e.g., a left shift) and the tile group 206-1 may compute a second MAC operation 401-2”); identifying, from the sequence of clock cycles, a second clock cycle associated with the second number (¶ 30, “Returning to cycle 2, as shown, a third row of the IFM 104 is read from the IRAM 202 and stored in the buffer 208-2. In cycle 2, therefore, the PEs 210 of the tiles 207 of tile group 206-3 may compute a MAC operation 403-1”); and instructing the second processing element to start the second computation in the second clock cycle (¶ 60, “performing the MAC operation on the first shifted first row of the IFM and the kernel in a second cycle”).
Regarding claim 8, Vengallur teaches: The method of claim 1, wherein the group of processing elements is at least part of an array of processing elements (¶ 23, “the tile groups 206-1 through 206-N reflect a 3D compute grid organized as an array of 32 tiles 207”), wherein the array of processing elements is configured to perform at least part of the convolution (¶ 29, “In cycle 1, therefore, the PEs 210 of the tiles 207 of tile group 206-1 may compute a MAC operation 401-1. The MAC operation 401-1 may be based on the first row of the IFM 104 and the kernel 108”), and wherein the array of processing elements comprises rows and columns, and the group of processing elements is arranged in one of the columns (¶ 34, “the number of OFMs 106 are equal to 1, the 16×8×8 grid may be transformed into eight (1,1,16), or 1×1×16, grids for depth wise/2D convolutions, where each of the 8 grids operates in parallel”).
Claims 9, 12, 14-17, and 19 recite commensurate subject matter as claims 1, 4, and 6-8. Therefore, they are rejected for the same reasons.
Claim(s) 2, 3, 5, 10, 11, 13, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang and Vengallur, as applied above, and further in view of Moshovos (US 2021/0004668).
Regarding claim 2, Wang and Vengallur do not teach; however, Moshovos discloses: Moshovos teaches: determining the workload of the respective processing element based on an input sparsity bitmap and a weight sparsity bitmap (¶ 52, “Embodiments of the present invention exploit ineffective activation bits, either separately or in combination with exploiting weight sparsity;” ¶ 33, “accelerator 3000 uses a 16-bit fixed point format to represent activations and weights, as do many embodiments of inference accelerators;” and ¶ 64, “a bitmap where each bit represents whether a value within the group is equal to or different from zero as shown in Table 3”), wherein the input sparsity bitmap comprises a sequence of bits, each of which indicates whether a value of a respective activation in the input operand is zero (¶ 71, “zero values are not stored and instead a bit vector per group identifies the position of the non-zero values”), and the weight sparsity bitmap comprises another sequence of bits, each of which indicates whether a value of a respective weight in the weight operand is zero (¶ 71, “In some embodiments, a group of 16 activations or weights may be used as offering a good balance between compression rate and metadata overhead. For each group, he [sic] precision is stored in bits and the zero-value bit-vector”).
It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the invention, to have applied the known technique of determining the workload of the respective processing element based on an input sparsity bitmap and a weight sparsity bitmap, wherein the input sparsity bitmap comprises a sequence of bits, each of which indicates whether a value of a respective activation in the input operand is zero, and the weight sparsity bitmap comprises another sequence of bits, each of which indicates whether a value of a respective weight in the weight operand is zero, as taught by Moshovos, in the same way to the determining the workload of the respective processing element based on the input operand and the weight operand, as taught by Wang and Vengallur. Both inventions are in the field of accelerating processing for DNN workload, and combining them would have predictably resulted in “neural network hardware accelerators,” as indicated by Moshovos (¶ 1).
Regarding claim 3, Moshovos teaches: The method of claim 2, wherein determining the workload based on the input sparsity bitmap and the weight sparsity bitmap comprises: generating a combined sparsity bitmap based on the input sparsity bitmap and the weight sparsity bitmap (¶ 52, “Embodiments of the present invention exploit ineffective activation bits, either separately or in combination with exploiting weight sparsity;” ¶ 33, “accelerator 3000 uses a 16-bit fixed point format to represent activations and weights, as do many embodiments of inference accelerators;” and ¶ 64, “a bitmap where each bit represents whether a value within the group is equal to or different from zero as shown in Table 3”), the combined sparsity bitmap comprising a plurality of bits, each of which is a result of a bit in the input sparsity bitmap multiplying a bit in the weight sparsity bitmap (¶ 32, “Each input activation is multiplied with k weights, one per filter of the set of filters 1200 as follows: each IPU 3100 accepts a vector of N weights per cycle, one per input activation, calculates N products, reduces them via an adder tree, and accumulates the result into an output register”); and determining the workload based on a number of ones in the combined sparsity bitmap (¶ 33, “accelerator 3000 uses a 16-bit fixed point format to represent activations and weights, as do many embodiments of inference accelerators”).
Regarding claim 5, Wang and Vengallur do not teach; however, Moshovos teaches: determining the workload of the respective processing element based on an input sparsity bitmap and a weight sparsity bitmap (¶ 52, “Embodiments of the present invention exploit ineffective activation bits, either separately or in combination with exploiting weight sparsity;” ¶ 33, “accelerator 3000 uses a 16-bit fixed point format to represent activations and weights, as do many embodiments of inference accelerators;” and ¶ 64, “a bitmap where each bit represents whether a value within the group is equal to or different from zero as shown in Table 3”), wherein the input sparsity bitmap comprises a sequence of bits, each of which indicates whether a value of a respective activation in the input operand is zero (¶ 71, “zero values are not stored and instead a bit vector per group identifies the position of the non-zero values”), and the weight sparsity bitmap comprises another sequence of bits, each of which indicates whether a value of a respective weight in the weight operand is zero (¶ 71, “In some embodiments, a group of 16 activations or weights may be used as offering a good balance between compression rate and metadata overhead. For each group, he [sic] precision is stored in bits and the zero-value bit-vector”).
It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the invention, to have applied the known technique of determining the workload of the respective processing element based on an input sparsity bitmap and a weight sparsity bitmap, wherein the input sparsity bitmap comprises a sequence of bits, each of which indicates whether a value of a respective activation in the input operand is zero, and the weight sparsity bitmap comprises another sequence of bits, each of which indicates whether a value of a respective weight in the weight operand is zero, as taught by Moshovos, in the same way to the first and second time, as taught by Wang and Vengallur. Both inventions are in the field of accelerating processing for DNN workload, and combining them would have predictably resulted in “neural network hardware accelerators,” as indicated by Moshovos (¶ 1).
Claims 10, 11, 13, 18, and 20 recite commensurate subject matter as claims 2, 3, and 5. Therefore, they are rejected for the same reasons.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB D DASCOMB whose telephone number is (571)272-9993. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Vital can be reached at (571) 272-4215. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB D DASCOMB/ Primary Examiner, Art Unit 2198