DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The following action is in response to the preliminary amendment of 10/06/2023.
By the amendment, claims 3, 5, 6, 8-13 and 15 have been amended. Claim 14 has been canceled. Claims 16-21 have been newly added.
Claims 1-13 and 15-21 are pending and have been considered below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, 6, 8-18 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Datta et al., US 2020/0024856 A1 (DATTA) in view of Li et al., US 2022/0188613 A1 effective filing 12/15/2020 (LI).
Regarding claim 1, DATTA discloses a method performed by one or more computers (pp. 2, Fig. 6, Fig. 20), the method comprising:
obtaining data specifying a neural network to be deployed on a hardware accelerator computer chip (pp. 2: determining relationships and other data specifying a neural network for deployment onto a chip comprising neural cores), wherein:
the hardware accelerator comprises on-chip memory and has a particular hardware datapath (pp. 39-45: inference processing unit comprising on-chip memory, pp. 46-47: having a datapath for the hardware)
the neural network comprises a plurality of neural network layers (pp. 25),
each neural network layer has an associated set of tensors that comprises (i) a weight tensor for the neural network layer, (ii) an input activation tensor for the neural network layer, and (iii) an output activation tensor for the neural network layer (pp. 27), and
each neural network layer is configured to receive the respective input activation tensor for the neural network layer and to process the respective input activation tensor for the neural network layer in accordance with the respective weight tensor for the neural network layer to generate the respective output activation tensor for the neural network layer (pp. 28-23, Equation 1, Equation 2);
providing, as input to a computer chip performance simulator, the data specifying the neural network and data specifying the hardware datapath of the hardware accelerator (pp. 56-57: providing scheme/data specifying the architecture parameters of the neural network, pp. 58-67: schedule utilizes scheme to provide a schedule or execution plan, pp. 68-69: execution plan is provided to a simulator for simulating the schedule);
obtaining, as output from the computer chip performance simulator and for each of the plurality of neural network layers, a respective initial estimate of performance statistics for the layer when the neural network is executed on the hardware accelerator computer chip (pp. 90: simulator can compute performance statistics, ex. number of cycles to process each layer, in order to dynamically update the schedule); and
determining, from the respective initial estimates for the neural network layers and for each tensor that is associated with each of the plurality of neural networks layers, whether the tensor is stored in the on-chip memory while performing inference for the neural network while the neural network is deployed on the hardware accelerator computer chip (pp. 87-90: dynamically updating the schedule for storing/executing each tensor in activation memory during run-time of the inference processing unit).
DATTA fails to explicitly disclose wherein the determination, from respective initial estimates, regarding the storing of a tensor in on-chip memory further considers whether to store the tensor off-chip memory.
LI discloses methods for performing memory efficient execution of neural networks on hardware accelerator chips (pp. 7), an analogous art. In particular, LI discloses determining whether to store tensors in either on-chip memory or off-chip memory based on constraints (pp. 27-28: determining storage and access between global buffer and off-chip utilizing loop optimization techniques, pp. 30: determining where to read and store tensor/matrix between off-chip and global buffer, pp. 40) based on initial estimates regarding each layer (pp. 8). Therefore it would have been obvious to one having ordinary skill in the art and the teachings of DATTA and LI before them before the effective filing of the claimed invention to combine the determination on whether to store a tensor in on-chip or off-chip memory, as taught by LI, with the determination of storing the tensor in on-chip memory while performing inference using the deployed neural network on the hardware accelerator chip of DATTA. One would have been motivated to make this combination in order to better utilize limited hardware resources effectively, resulting in increased performance, as suggested by LI (pp. 6-7).
Regarding claim 2, DATTA and LI disclose the method of claim 1, and DATTA further discloses:
determining a schedule for executing the operations of the plurality of neural network layers (pp. 58-67: schedule utilizes scheme to provide a schedule or execution plan), wherein providing, as input to a computer chip performance simulator, the data specifying the neural network and data specifying the hardware datapath of the hardware accelerator comprises:
providing, as input to a computer chip performance simulator, the data specifying the neural network, the data specifying the hardware datapath of the hardware accelerator, and data specifying the schedule (pp. 56-57: providing scheme/data specifying the architecture parameters of the neural network, pp. 58-67: schedule utilizes scheme to provide a schedule or execution plan, pp. 68-69: execution plan is provided to a simulator for simulating the schedule), and wherein the respective initial estimates are estimates of the performance statistics when the operations are performed according to the schedule during execution of the neural network on the hardware accelerator computer chip (pp. 90: simulator can compute performance statistics, ex. exact number of cycles to process each layer, in order to dynamically update the schedule).
Regarding claim 3, DATTA and LI disclose the method of claim 1, and LI further discloses wherein
determining, from the respective initial estimates and for each tensor that is associated with each of the plurality of neural networks layers, whether the tensor is stored in the on-chip memory or in off-chip memory while performing inference for the neural network while the neural network is deployed on the hardware accelerator computer chip (pp. 8, pp. 53, pp. 71: initially estimating via simulator, pp. 48) comprises:
determining, from the respective initial estimates, a fusion strategy that minimizes a sum of respective execution times for each of the plurality of neural network layers subject to one or more constraints, the fusion strategy assigning each tensor that is associated with each of the plurality of neural network layers to either the on-chip memory or the off-chip memory (pp. 51-53: estimating design variables to minimize objective, pp. 55: determining a loop fusion strategy from simulation results, pp. 48: Loop fusion details fusion strategy for reducing data transfer between off-chip and on-chip memory of the tensor/matrix by matrix chunk assignment, pp. 53: constraints across dataflow includes at least 4 loop optimization techniques).
It would have been obvious to one having ordinary skill in the art and the teachings of DATTA and LI before them before the effective filing of the claimed invention to further combine the determined fusion strategy for assigning a tensor in on-chip or off-chip memory utilizing one or more constraints, as taught by LI, with the determination of storing the tensor in on-chip memory or off-chip memory of DATTA and LI. One would have been motivated to make this further combination in order to better utilize limited hardware resources effectively, resulting in increased performance, as suggested by LI (pp. 6-7, pp. 30, pp. 48).
Regarding claim 6, DATTA and LI disclose the method of claim 3, and LI further discloses wherein the one or more constraints include:
a third constraint specifying that a respective usage of the on-chip memory during the processing of each neural network layer does not exceed a capacity of the on-chip memory (pp. 46: loop tiling determines capacity of on-chip memory and selects correct tile size).
Regarding claim 8, DATTA and LI disclose the method of claim 3, wherein the one or more constraints include a fourth constraint specifying that, for each particular neural network layer of the plurality of neural network layers, if the respective output activation tensor for the particular neural network layer is assigned to the on-chip memory by the fusion strategy, the respective input activation tensor for each neural network layer that receives as input the output of the particular neural network layer is also assigned to the on-chip memory by the fusion strategy (pp. 48: loop fusion causes execution of both tensor elements to stay within the current chip rather than transferring elements other chip).
Regarding claim 9, DATTA and LI disclose the method of claim 3, and LI further discloses wherein the one or more constraints include a fifth constraint specifying that, for each particular neural network layer of the plurality of neural network layers:
if the respective output activation tensor for the particular layer is assigned to the off-chip memory by the fusion strategy, the respective input activation tensor for each layer that receives as input the output of the particular neural network layer is also assigned to the off-chip memory by the fusion strategy (pp. 48: loop fusion causes execution of both tensor elements to stay within the current chip rather than transferring elements other chip).
Regarding claim 10, DATTA and LI disclose the method of claim 3, and LI further discloses wherein the one or more constraints include a sixth constraint specifying that, for each particular neural network layer of the plurality of neural network layers:
if the respective output activation tensor for the particular neural network layer is assigned to the on-chip memory by the fusion strategy, the respective input activation tensor for at least one neural network layer that receives as input the output of the particular neural network layer is also assigned to the on-chip memory by the fusion strategy (pp. 48: loop fusion causes execution of both tensor elements to stay on-chip rather than sending elements to off-chip DRAM).
Regarding claim 11, DATTA and LI disclose the method of claim 3, wherein the one or more constraints include a seventh constraint specifying that, for each particular neural network layer of the plurality of neural network layers:
if the respective input activation tensor for the particular neural network layer is assigned to the on-chip memory by the fusion strategy, the neural network layer that generates the input to the particular neural network layer must be executed immediately before the particular neural network layer (pp. 48: loop fusion causes execution of sequential tensor elements to stay within the current chip rather than transferring elements other chip, ex. SpMM1 must execute before SpMM2).
Regarding claim 12, DATTA and LI disclose method of claim 3, and LI further discloses wherein:
determining, from the respective initial estimates, a fusion strategy that minimizes a sum of respective execution times for each of the plurality of neural network layers subject to one or more constraints comprises determining the fusion strategy through integer linear programming using the respective initial estimates (pp. 52-53, Equation 2).
Regarding claim 13, DATTA and LI disclose the method of claim 1, and DATTA further discloses:
performing inference for the neural network while the neural network is deployed on the hardware accelerator computer chip, comprising:
while performing inference for the neural network while the neural network is deployed on the hardware accelerator computer chip:
for each tensor that was determined to be stored in the on-chip memory, storing the tensor in on-chip memory and
for each tensor that was determined to be stored in the off-chip memory, storing the tensor in off-chip memory (pp. 87-90: dynamically updating the schedule for storing/executing each tensor in activation memory during run-time of the inference processing unit).
Regarding claim 15, claim 15 recites limitations similar to claim 1 and is similarly rejected.
Regarding claims 16-18 and 21, claims 16-18 and 21 recite limitations similar to claims 1-3 and 6, respectively, and are similarly rejected.
Claim 4 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over DATTA in view of LI and in further view of Sun et al., US 2020/0192803 A1 (SUN).
Regarding claim 4, DATTA and LI disclose the method of claim 3, and LI further discloses wherein:
the initial estimates include, for each of the plurality of neural network layers, a respective minimum execution time for the neural network layer when the input and output activation tensors for the layer are assigned to the on-chip memory (pp. 52-53: estimating computation latency for a given layer assigned to on-chip access).
DATTA and LI fail to disclose wherein the one or more constraints include a first constraint specifying that, for each of the plurality of neural network layers, the respective execution time is greater than or equal to the respective minimum execution time for the neural network layer (pp. 51: latency depends on loop un-rolling factors).
SUN discloses methods for optimizing access and control of tensors in memory (pp. 5), an analogous art. In particular, SUN discloses calculating the time needed to execute input and output tensors of a layer in memory and performing a constraint including a respective execution time that is greater than or equal to a respective minimum execution time needed to execute the layer (pp. 132: combining time of tensors to perform and comparing them to a maximum). Therefore it would have been obvious to one having ordinary skill in the art and the teachings of DATTA, LI and SUN before them before the effective filing of the claimed invention to combine the use of one or more constraints specifying respective execution time greater than or equal to a minimum execution time for a neural network layer, as taught by SUN, with the use of constraints when determining initial estimates including execution time of the tensors assigned to on-chip memory of DATTA and LI. One would have been motivated to make this combination in order to improve performance and power consumption of the system, as suggested by SUN (pp. 132).
Regarding claim 19, claim 19 recites limitations similar to claim 4 and is similarly rejected.
Allowable Subject Matter
Claims 5, 7 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Nurvitadhi; Eriko et al.
US 20200279349 A1
MACHINE LEARNING SPARSE COMPUTATION MECHANISM
Sui; Lingzhi et al.
US 20190303762 A1
METHODS OF OPTIMIZATION OF COMPUTATIONAL GRAPHS OF NEURAL NETWORKS
Bleiweiss; Amit et al.
US 20190205736 A1
COMPUTE OPTIMIZATION MECHANISM FOR DEEP NEURAL NETWORKS
Baum; Avi et al.
US 20180285254 A1
SYSTEM AND METHOD OF MEMORY ACCESS OF MULTI-DIMENSIONAL DATA
Wei, Xuechao, Yun Liang, and Jason Cong. "Overcoming data transfer bottlenecks in FPGA-based DNN accelerators via layer conscious memory management." Proceedings of the 56th Annual Design Automation Conference 2019. 2019.
Capra, Maurizio, et al. "Hardware and software optimizations for accelerating deep neural networks: Survey of current trends, challenges, and the road ahead." IEEE Access 8 (2020): 225134-225180.
Xing, Yu, et al. "DNNVM: End-to-end compiler leveraging heterogeneous optimizations on FPGA-based CNN accelerators." IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 39.10 (2019): 2668-2681.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW L TANK whose telephone number is (571)270-1692. The examiner can normally be reached Monday-Thursday 9a-6p.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Ell can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW L TANK/Primary Examiner, Art Unit 2141