Prosecution Insights
Last updated: October 01, 2026
Application No. 18/285,578

FULL-STACK HARDWARE ACCELERATOR SEARCH

Non-Final OA §103
Filed
Oct 04, 2023
Priority
Apr 06, 2021 — provisional 63/171,526 +2 more
Examiner
RAMESH, TIRUMALE K
Art Unit
2121
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
1 (Non-Final)
26%
Grant Probability
At Risk
1-2
OA Rounds
1y 9m
Est. Remaining
50%
With Interview

Examiner Intelligence

Grants only 26% of cases
26%
Career Allowance Rate
13 granted / 49 resolved
-28.5% vs TC avg
Strong +24% interview lift
Without
With
+23.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 9m
Avg Prosecution
19 currently pending
Career history
85
Total Applications
across all art units

Statute-Specific Performance

§101
26.8%
-13.2% vs TC avg
§103
63.9%
+23.9% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
4.6%
-35.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§103
CTNF 18/285,578 CTNF 97128 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 6-9, 11, 15, and 19-20 are rejected under 35 U.S.C. 103 unpatentable over in view of Animesh Jain et.al. (hereinafter Jain ) US 11809981 B1, in view of Kristof Denolf et.al. (hereinafter Denolf ) US 2020/0104715 A1. in view of Eriko Nurvitadhi et.al. (hereinafter Nurvi ) US 11216722 B2. In regard to claim 1: (Original) Jain discloses: - A method performed by one or more computers, the method comprising: obtaining data specifying a target set of one or more neural networks; [Col 1, lines 6-10]: Artificial neural networks are computing systems with an architecture based on biological neural networks. Artificial neural networks can be trained, using training data, to learn about how to perform a certain computing task for an application. [Col 6, lines 45-49]: It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 6, lines 33-36]: Prediction model 103 can be in the form of an artificial neural network . The artificial neural network may include a plurality of processing nodes, with each processing node configured to process part of the input pixel data, [Col 6, lines 38-42]: FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. (i) one or more weight tensors each representing weights of a respective layer of the neural network, [Col 6, lines 50-59]: Layer 207 may process pixel data representing different portions of image 104. For example, in the example of FIG. 2A, layer 207 may process the pixel data of image 204. Each processing node of layer 207 is assigned to receive a pixel value (e.g., x.sub.0, x.sub.1, x.sub.2, . . . x.sub.n) corresponding to a predetermined pixel within image 104, and transmit one or more weights with the received pixel value to layer 209. In a case where prediction model 203 is a DNN, each processing node of layer 207 can be assigned a set of weights defined based on a matrix W1 (BRI: A set of weights defined on a matrix does represent a weight tensor) [Col 6, lines 38-45]: FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. Prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer (not shown in FIG. 2A). [Col 6, lines 66-67]: Different neural network models may include different topologies (e.g., including a different number of layers, [Col 7, lines 1-2]: different connections between layers, etc.), and/or include a different set of weights for each layer. (ii) one or more input activation tensors each representing an input to a respective layer of the neural network, or [Col 1, lines 58-61]: An artificial neural network typically consist of a sequence of layers/operators. The operations involved in each operator/layer may involve, for example, multiplication and summation operations, activation function processing [Col 9, lines 22-28]: Referring back to FIG. 2A, one processing node of layer 209 may be configured to generate the convolution output elements of one convolution output array, and a set M of processing nodes of layer 209 can correspond to a set M of convolution output arrays. The processing node of layer 209 can also process each convolution output with an activation function to generate an activation output. [Col 6, lines 34-48]: The artificial neural network may include a plurality of processing nodes, with each processing node configured to process part of the input pixel data, or to further process the intermediate outputs from other processing nodes. FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. Prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer (not shown in FIG. 2A). It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 3, lines 53-59]: Specifically, as part of the splitting operation, the compiler can determine, for each read instruction at the virtual data node for an input data element to the second operator, one or more corresponding write instructions that supply the output data element(s) of the first operator included in the input data element, based on the tensor addresses included in the read and write instructions - and (ii) a respective fusion strategy for each of the one or more neural networks when deployed on the hardware accelerator computer chip having the determined architecture, [Col 1, lines 6-10]: Artificial neural networks are computing systems with an architecture based on biological neural networks. Artificial neural networks can be trained, using training data, to learn about how to perform a certain computing task for an application. [Col 6, lines 45-49]: It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 1, lines 58-67]: An artificial neural network typically consist of a sequence of layers/operators . The operations involved in each operator/layer may involve, for example, multiplication and summation operations, activation function processing, pooling, etc . Depending on the topology of the neural network (e.g., convolutional neural network, fully-connected deep neural network, etc.), the connectivity between layers, which describes how the outputs of the first operator/layer connect with the inputs of the second operator/layer, can also be different. [Col 1, line 67]: An artificial neural network can be [Col 2, lines 1-3]: implemented in various computing systems, such as a general purpose central processing unit (CPU), a hardware accelerator, etc. [Col 24, lines 52-58]: In some examples, the accelerator 702 can implement a neural network processing engine. In these examples, the accelerator 702, for a set of input data 750, can execute a neural network to perform a task for which the neural network was trained. Executing a neural network on a set of input data can be referred to as inference or performing inference . [Col 2, lines 57-67]: creating fused kernels for a hardware accelerator can present extra challenges. Specifically, the on-chip memory of the hardware accelerator, which can act as a scratchpad memory, is typically managed by the kernel. As part of the management, the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators . A fuse d kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from the [Col 3, lines 1-6]: on-chip memory, to control the data movement between the operators. Orchestrating such data movements and compute operations for these slices across operators is tedious and error-prone, which can further increase the engineering effort required to support different fused operators and for different neural network topologies. - the respective fusion strategy for each of the one or more neural networks specifying, for each tensor in the set of associated tensors for the neural network, whether or not the tensor is stored in on-chip memory of the hardware accelerator computer chip during processing of inputs using the neural network, the determining comprising: [Col 2, lines 57-67]: creating fused kernels for a hardware accelerator can present extra challenges. Specifically, the on-chip memory of the hardware accelerator, which can act as a scratchpad memory, is typically managed by the kernel. As part of the management, the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators . A fuse d kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from the [Col 3, lines 1-6]: on-chip memory, to control the data movement between the operators. Orchestrating such data movements and compute operations for these slices across operators is tedious and error-prone, which can further increase the engineering effort required to support different fused operators and for different neural network topologies. [Col 2, lines 4-20]: A computing system can be programmed to implement a multi-layer artificial neural network. The instruction file may include a plurality of kernels, with each kernel including instructions that define the operations involved in a layer/operator. The computing system can execute a kernel of a first operator, generate a first intermediate tensor, and store the first intermediate tensor at an off-chip memory (e.g., DRAM). The computing system can then execute a kernel of a second operator. To execute the second kernel, the computing system can fetch the first intermediate tensor from the memory, map the first intermediate tensor to the inputs of the second operator based on the connectivity between the two operators/layers, and execute the second kernel to generate a third intermediate tensor. The computing system can store the third intermediate tensor back to the memory, and then repeat the execution and memory access operations for subsequent layers/operators. [Col 2, lines 61-67]: the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators. A fused kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from [Col 2, lines 1-2]: the on-chip memory , to control the data movement between the operators. [Col 3, lines 25-31]: the virtual data node can represent a logical tensor of the output data elements by the first operator. Each output data element of the first operator may be associated with a tensor address (e.g., coordinates) within the logical tensor represented by the virtual data node. The write instructions in the first kernel can include the tensor addresses to which the output data elements are to be stored. - and (ii) a respective optimized fusion strategy for each of the one or more neural networks from a search space of possible optimized fusion strategies for the neural network when deployed on a hardware accelerator chip [Col 19, lines 4-9]: For each virtual data node (which does not have corresponding read/write instructions), vertical fusion module 516 can convert the read or write instructions into memory read or memory write instructions for an on-chip memory (e.g., a CPU cache, a scratchpad in a hardware accelerator, etc.) [Col 17, lines 33-46]: Virtual data node splitting module 514 can then process read instruction 526 which accesses tensor addresses 0, 1, and 2 for input data element b1. For read instruction 526, virtual data node splitting module 514 can identify the corresponding write instructions 524, 532, and 534 which writes output data elements a.sub.0, a.sub.1, and a.sub.2 to tensor addresses 0, 1, and 2. Virtual data node splitting module 514 can look for an access group which includes write instructions 524, 532, and 534 (or a superset including write instructions 528a-528c). As only access group 520 is created at this point and access group 520 includes only write instruction 524, virtual data node splitting module 514 can create an access group 540 which includes read instructions 526, 528, 530 and corresponding write instructions 524, 532, and 534) [Col 21, lines 40-52]: In various implementations, the memory subsystem 704 can include multiple memory banks 714. In these implementations, each memory bank 714 can be independently accessible, meaning that the read of one memory bank is not dependent on the read of another memory bank. Similarly, writing to one memory bank does not affect or limit writing to a different memory bank. In some cases, each memory bank can be read and written at the same time. Various techniques can be used to have independently accessible memory banks 714. For example, each memory bank can be a physically separate memory component that has an address space that is separate and in depende nt of the address spaces of each other memory bank. [Col 21, lines 57-67]: the memory subsystem 704 can include arbitration logic such that arbitration between, for example, the outputs of multiple memory banks 714 can result in more than one memory bank's output being used . In these and other examples, though globally managed by the memory subsystem 704, each memory bank can be operated independently of any other. In some examples, accelerator 702 can be programmed by instructions generated based on the disclosed techniques to perform data transfer between fused operators using memory subsystem 704. (BRI: a fusion model is an accelerator framework and the fusion module uses a cost model to decide which fusion to apply a fusion search space without violating the data dependency and memory constraints) [Col 27, lines 48-56]: perform various steps before producing the instructions that are to be executed by the acceleration engine 812. These steps can include, for example, removing redundant dependencies, resolving or handling dependencies between nodes by inserting synchronization instructions into the code, identifying possibly optimi zations in memory usage or memory bandwidth usage, and other operations. [Col 15, lines 39-49]: In some examples, the compiler can generate the sequence of executable instructions to maximize the number of operators to be fuse d, but exclude an operator from the fuse d operators when, for example, limitations from the computing system prevent the fusion of that operator with other operators . Such limitations may arise from various sources. For example, the data can be of multi-channel and a particular channel of data needs to be accessed at a particular time, but the on-chip memory cannot provide such access at that time. As another example, the on-chip memory is simply too small to fit the multi-channel data. In all these cases, the compiler can remove an operator from the fused operators and generate off-chip memory read/write instructions to handle data transfer between that operator and the other fused operators. (BRI: maximizing the number of operators to be fused can represent an fusing optimization strategy) - having an architecture specified by the candidate hardware datapath; [Col 10, lines 5-26]: a neural network performs a sequence of computation operations to generate a decision . The sequence of computation operations can be represented by a computational graph . The left of FIG. 3A illustrates an example of a simplified computational graph 300 representing a sequence of operators. As shown in FIG. 3A, computation graph 300 includes a set of nodes 302, 304, and 306, as well as edges 308 and 310 connecting between the nodes. Each node in computation graph 300 can represent an operator, which can represent a neural network layer in, for example, prediction model 103. For example, node 302 can correspond to operator Op1 which can represent input layer 207 of FIG. 2A, node 304 can correspond to intermediate layer 209 of FIG. 2A, whereas node 306 can correspond to intermediate layer 211 of FIG. 2A. Moreover, edge 308 can represent flow of data from another node (not shown in FIG. 3A) to node 302 (input layer 207), edge 310 can represent flow of data from node 302 to node 304 (intermediate layer 209), edge 312 can represent flow of data from node 304 (intermediate layer 209) to node 306 (intermediate layer 211), whereas edge 314 can represent flow of data from node 306 to another node (not shown in FIG. 3A). PNG media_image1.png 798 646 media_image1.png Greyscale [Col 11, lines 39-43]: A computing system, such as a hardware accelerator, a CPU, etc., can execute the operators Op1, Op2, and Op3, as well as the data transfer operations, by executing the kernel instructions 322, 324, and 326 as well as the read/write operations represented in blocks 328, 340, and 342. [Col 19, lines 54-67]: compiler receives a first set of instructions including a kernel of a first operator and a kernel of a second operator. The kernel of the first operator can include instructions of the first operator and write instructions to a virtual data node , whereas the kernel of the second operator can include instructions of the second operator and read instructions to the virtual data node . The first set of instructions can be in the form of a computational graph instruction represented by block diagram 400 of FIG. 4A. The first operator and the second operator can correspond to, respectively, a first layer and a second layer of a neural network and each can include various operations such as multiplication and summation operation, activation function processing, pooling operation, [Col 20, lines 1-6]: etc. The output data of the first operator/layer can be fetched to the second operator/layer as input. As shown in FIG. 4C, the virtual data node can correspond to a logical tensor to store the output data of the first operator and from which the second operator fetches input data , to provide the data transfer from the first operator to the second operator. (BRI: data transfer operations, multiplication and summation operations indeed represent by the datapath in a CPU) Jain does not explicitly disclose: - obtaining data specifying an objective function that measures a performance of a hardware accelerator computer chip when performing inference for the target set of one or more neural networks, each of the one or more neural networks having a respective set of associated tensors that includes one or more of - and determining (i) an architecture for the hardware accelerator computer chip - Repeatedly performing operations comprising: - determining (i) a candidate set of hyperparameters that define a candidate hardware datapath for the hardware accelerator computer chip - from a search space of possible hardware datapaths for the hardware accelerator computer chip - and determining a value of the objective function for the candidate hardware datapath by, for each neural network in the set, However, Denolf discloses: - obtaining data specifying an objective function that measures a performance of a hardware accelerator computer chip when performing inference for the target set of one or more neural networks, each of the one or more neural networks having a respective set of associated tensors that includes one or more of [0030]: FIG. 2 is a block diagram depicting a computing system (“computer 200”) according to an example. The computer 200 includes a software platform 204 executing on a hardware platform 202. The hardware platform 202 includes a central processing unit (CPU) 206, a system memory 208, storage devices 210, support circuits 211, a training platform 212, and a hardware acce lerator 214 . [0023]: Reinforcement learning provides for multi-objective optimization, but without adding the implementation cost of the neural network itself as an objective. - and determining (i) an architecture for the hardware accelerator computer chip [0030]: FIG. 2 is a block diagram depicting a computing system (“computer 200”) according to an example. The computer 200 includes a software platform 204 executing on a hardware platform 202. The hardware platform 202 includes a central processing unit (CPU) 206 , a system memory 208, storage devices 210, support circuits 211, a training platform 212, and a hardware acce lerator 214 . PNG media_image2.png 710 648 media_image2.png Greyscale [0017] : FIG. 7 is a block diagram depicting a programmable integrated circuit (IC) according to an example. PNG media_image3.png 767 402 media_image3.png Greyscale [0018]: FIG. 8 is a block diagram depicting a System-on- Chip (SoC) implementation of the programmable IC of FIG. 7 PNG media_image4.png 796 648 media_image4.png Greyscale (BRI: the entire FIG 8 as an SoC embodies the hardware accelerator) [0031]: in some examples, the CPU 206 can be a System-in-Package (SiP), System-on- Chip (SoC), or the like, which absorbs all or a substantial portion of the functionality of the chip set Multi-Objective Optimization [ 0037 ]: The inclusion of inference implementation cost when evaluating the performance of networks means there are at least two objectives that are to be optimized . As such, multiple objectives should be balanced in a meaningful way . For example, assume the accuracy of the network is given by classification error, C.sub.E, and the estimated implementation cost is given by the time taken to process a new input, C.sub.T. If minimizing C.sub.T is given too much importance, then it is possible an optimizer will produce a network with zero layers, zero operations, and zero memory requirements . This could yield a network that has C.sub.T=0, despite incurring a significantly high C.sub.E. Multi-objective optimization aims to balance C.sub.E and C.sub.T to give a desirable solution. [0038] : A general formulation of multi-objective optimization is as follows: PNG media_image5.png 72 481 media_image5.png Greyscale where f.sub.1, . . . , f.sub.x are functions that define the cost of each objective that is being optimized, x is a vector representing the current solution, and X is the search space of all possible solutions . In the examples described herein, x represents a neural network topology and its associated hyperparameters (i.e., the model-capacity hyperparameters 108). The functions represent metrics of interest of the current neural network topology in relation to its accuracy and implementation/hardware cost. For accuracy, these functions include mean squares error (MSE), classification error, l.sub.p norm, hingle loss, or a similar metric suitable for the target domain. For implementation/hardware cost, these functions include memory requirements, bandwidth requirements, clock cycles, datapath width, quantization scheme, arithmetic style, number formats, silicon area, and energy consumption, and error tolerance. - Repeatedly performing operations comprising: [0047]: The basic methodology of evolutionary algorithms is to generate N random strings of genes (which correspond to neural network architectures ) (step 402). These architectures are then evaluated using a fitness function, which may require training each network architecture individually (step 404). At this point, a subset of the architectures are selected, randomly combined and mutated to generate the next N architectures (step 406). Over time, this results in architectures which are highly optimized for the given cost functions , which in this case means high accuracy and low implementation/hardware cost. At step 408, a determination is made whether to end. If not, the method 400 proceeds to step 404 and repeats. Otherwise, the method 400 proceeds to step 410, where the training platform outputs the trained neural network. - determining (i) a candidate set of hyperparameters that define a candidate hardware datapath for the hardware accelerator computer chip [0035]: the applications 236 include software that trains neural networks on the training platform 212 and implements neural networks on the hardware accelerator 214. [0059] : FIG. 8 is a block diagram depicting a System-on- Chip (SoC ) implementation of the programmable IC 1 according to an example. [0006]: a method of implementing a neural network includes: selecting a first neural network architecture from a search space ; training the neural network having the first neural network architecture to obtain an accuracy and an implementation cost , the implementation cost based on a programmable device of an inference platform ; selecting a second neural network architecture from the search space based on the accuracy and the implementation cost; and outputting weights and hyperparameter s for the neural network having the second neural network architecture. [0052]: A Bayesian hyperparameter search is a more sophisticated technique which attempts to develop a statistical model which maps the hyperpar ameter values to our cost function . Usually, this statistical model is a Gaussian Process (GP) which generates functions which closely approximates the observed data. GPs provide a prediction for the chosen cost function in the hyperpar ameter space , along with the uncertainty of such predictions, this has the following benefits over random search and grid search: (BRI: a hyperparameter space represents a set of hyperparameters) - from a search space of possible hardware datapaths for the hardware accelerator computer chip [0008]: computer system includes: a memory having program code stored therein; and a processor, configured to execute the program code, to implement a neural network by: selecting a first neural network architecture from a search space ; [0028] : The topology 120 generally includes an arrangement of neurons. For example, the topology 120 can include a plurality of layers of neurons. The layers generally include an input layer, an output layer , and zero or more hidden layers . Each neuron includes a plurality of inputs and an output. The plurality of inputs for each neuron are associated with a plurality of weights. Each neuron further includes a bias associated with its output. The weights and biases of the neural network 106 are referred to as trained network weights 114. For a given layer, the inputs of its neurons are referred to as input feature maps and the outputs of its neurons are referred to as output feature maps. Input feature maps and output feature maps are generally referred to as “feature maps.” (BRI: the above represents “ neural network architecture”) [0031]: In an example, the CPU 206 can be any type of general-purpose central processing unit (CPU), such as an x86-based processor, ARM®-based processor, or the like. The CPU 206 can include one or more cores and associated circuitry (e.g., cache memories, memory management units (MMUs), interrupt controllers, etc.). The CPU 206 is configured to execute program code that perform one or more operations described herein and which can be stored in the system memory 208 and/or the storage devices 210. The support circuits 211 include various devices that cooperate with the CPU 206 to manage data flow between the CPU 206 , the system memory 208, the storage devices 210, the training platform 212, the hardware accelerator 214 , or any other peripheral device. [0053]: In the methods above, the size/complexity of the neural architecture search space can be reduced by only making certain aspects of the network variable. For instance, making only the bit width of the feature map elements and the number of channels of the feature maps variable enables training for their optimum setting. Typically, reducing the bit width of the feature map elements results in less accuracy while allowing a more efficient implementation. (BRI: the bit width is directly tied to the datapath width as a result of the hardware processing unit handling the ALU or data transfer operation based on that bit width) - and determining a value of the objective function for the candidate hardware datapath by, for each neural network in the set, [0047]: The basic methodology of evolutionary algorithms is to generate N random strings of genes (which correspond to neural network architectures) (step 402). These architectures are then evaluated using a fitness function, which may require training each network architecture individually (step 404). At this point, a subset of the architectures are selected, randomly combined and mutated to generate the next N architectures (step 406). Over time, this results in architectures which are highly optimized for the given cost function s, which in this case means high accuracy and low implementation/hardware cost. ( BRI: the cost function is the value of the objective function) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). Jain and Denolf do not explicitly disclose: - that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; - simulating a performance of a candidate hardware accelerator computer chip that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; - and selecting a final hardware datapath for the hardware accelerator computer chip from the candidate hardware datapaths based on the respective values of the objective functions for the candidate hardware datapaths. However, Nurvi discloses: - that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; [Col 18, lines 46-54]: One implementation of the accelerator 900 can be programmed through a software library (similar to Intel® Math Kernel Library). Such library prepares the matrix data in memory, sets control registers in the accelerator 900 with information about the computation (e.g., computation type, memory pointer to matrix data), and starts the accelerator. Then, the accelerator independently accesses matrix data in memory, performs the computation, and writes the results back to memory for the software to consume. (BRI: Perhaps as known to a POSITA, the software library of kernel library provide fusing [Col 7, lines 28-33]: some embodiments can produce an accelerator instance optimiz ed for a target FPGA chip with a particular number of hardware multiply and on-chip RAM resources, and some embodiments can produce an accelerator instance optimized for an ASIC for a particular market segment, programmable to support all RNN applications in this segment . [Col 7, lines 3-11]: the RNN variants may dictate the size s and types of the matrix and vector operations, as well as their data dependencies . For example, matrix and vector size s can be related to the number of hidden units in the RNN, and the activation function (AF) type (tan h, sigmoid, etc.) can relate to the type of vector operations . Thus, many matrix and vector operations in RNNs make them computationally intensive, so being able to execute RNNs as efficient as possible is of critical importance. (BRI: the RNN variant that dictate the matrix operations is a “hyperparameter”) - simulating a performance of a candidate hardware accelerator computer chip that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; [Col 9, lines 16-20]: The hardware accelerator template 450 also includes one or more vector processing units 470A-470N (VPUs), which includes one or more FMAs 472 and/or one or more activation function blocks (for performing needed activation functions efficiently in hardware) [Col 22, lines 33-45]: the accelerator architecture instance (synthesizable RTL) produced by the template mapping is then automatically validated . To do this, one implementation of the framework derives a functional model of the vertex program to be used as the “golden” reference. Test benches are generated to compare the execution of this golden reference against simula tions of the RTL implementation of the architecture instance . The framework also performs performance validation by comparing RTL simulations against analytical performance model and cycle-accurate software simulator. It reports runtime breakdown and pinpoint the bottlenecks of the design that affect performance. [Col 7, lines 28-37]: some embodiments can produce an accelerator instance optimized for a target FPGA chip with a particular number of hardware multiply and on-chip RAM resources, and some embodiments can produce an accelerator instance optimized for an ASIC for a particular market segment, programmable to support all RNN applications in this segment. For example , in embodiments where the accelerator instance comprises RTL code, the RTL code can be used as an input for a standard ASIC developmental tool (e.g., a logic synthesis tool) to generate an ASIC design . - and selecting a final hardware datapath for the hardware accelerator computer chip from the candidate hardware datapaths based on the respective values of the objective functions for the candidate hardware datapaths. [Col 25, lines 58-67]: The accelerator logic chip 2305 at the bottom of the accelerator stack is customized to the needs of sparse-matrix computations, and is able to consume the bandwidth offered by a DRAM stack 2301-2304 while only expending 2-4 Watts of power, with energy consumption proportional to the bandwidth of the stack. To be conservative, a stack bandwidth of 273 GB/sec is assumed (the expected bandwidth of WIO3 stacks) for the remainder of this application. Designs based on higher-bandwidth stacks would incorporate more parallelism in order to consume the memory bandwidth. [Col 3, lines 42-44]: FIG. 23 illustrates one implementation of an accelerator includes an accelerator logic die and one of more stacks of DRAM die according to some embodiments. PNG media_image6.png 476 658 media_image6.png Greyscale [Col 8, lines 52-63]: Turning back to the FIG. 4, framework 400 module can include a template mapping module 404 that produces a customized accelerator instance (e.g., such as synthesizable register transfer language (RTL) utilizing a hardware description language (HDL) such as Verilog, VHDL, etc.) of the hardware accelerator that best meets the input constraints 402 and optimization goals 408 . Alongside the RTL, in some embodiments the framework 400 module also generates a compiler to program the accelerator , e.g., via providing micro-code executed by control units , as described further herein. (BRI: the above represents the synthesis of the hardware and the arrangement of matrix processing unit may indeed represent selection of the final hardware as the matrix processing (datapath) directly impact the overall performance of the accelerator) [Col 11, lines 22-31]: Flow 800 also includes, at block 815, mapping the plurality of operations of the flow graph to an accelerator hardware template to yield the accelerator instance comprising register transfer language code t hat describes how one or more matrix processing units (MPUs) and one or more vector processing units (VPUs) are to be arranged to perform the RNN algorithm . At least one of the one or more MPUs, as part of implementing the RNN algorithm, is to directly provide or directly receive a value from one of the one or more VPUs. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) In regard to claim 2: (Original) Jain does not explicitly disclose: - providing the value of the objective function to the optimizer for use in generating a new candidate set of hyperparameters. However, Denolf discloses: - providing the value of the objective function to the optimizer for use in generating a new candidate set of hyperparameters. [0048]: FIG. 5 is a method 500 of training a neural network according to an example. The method 500 begins at step 502, where a tuning agent 105 selects a set of hyperparameters. As noted above, the model-capacity hyperparameters allow definition/description of the architecture of the neural network. The model-capacity hyperparameters define both the topology parameters (e.g., the number of layers, number of channels per layer, etc.) and the related implementation attributes. The tuning agent 105 collects knowledge about the relation between the hyperparameters (both algorithm behavior and model-capacity [0050]: Examples of hyperparameter optimization techniques include grid search, random search, and Bayesian optimization. A grid search involves selecting a set of candidate values for each hyperparameter within a neural network. A grid search is then performed by training a network for each permutation of hyper parameters. The best model is then chosen as the one which performs desirably with respect to our cost functions , described above in the multi-objective optimization section (BRI: A grid search space of hyperparameters and selection does represent a new hyperparameter) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). In regard to claim 3: (Original) Jain and Denolf do not explicitly disclose: - obtaining, from the optimizer, data specifying the respective optimized fusion strategy for each of the one or more neural networks. However , Nurvi discloses: - obtaining, from the optimizer, data specifying the respective optimized fusion strategy for each of the one or more neural networks. [Col 18, lines 46-54]: One implementation of the accelerator 900 can be programmed through a software library (similar to Intel® Math Kernel Library). Such library prepares the matrix data in memory, sets control registers in the accelerator 900 with information about the computation (e.g., computation type, memory pointer to matrix data), and starts the accelerator. Then, the accelerator independently accesses matrix data in memory, performs the computation, and writes the results back to memory for the software to consume. (BRI: Perhaps as known to a POSITA, the software library of kernel library provide fusing [Col 7, lines 28-33]: some embodiments can produce an accelerator instance optimiz ed for a target FPGA chip with a particular number of hardware multiply and on-chip RAM resources, and some embodiments can produce an accelerator instance optimized for an ASIC for a particular market segment, programmable to support all RNN applications in this segment . [Col 7, lines 3-11]: the RNN variants may dictate the size s and types of the matrix and vector operations, as well as their data dependencies . For example, matrix and vector size s can be related to the number of hidden units in the RNN, and the activation function (AF) type (tan h, sigmoid, etc.) can relate to the type of vector operations . Thus, many matrix and vector operations in RNNs make them computationally intensive, so being able to execute RNNs as efficient as possible is of critical importance. [Col 8, lines 8-12]: The design framework 400 , in some embodiments, utilizes optimization goal inputs 408, such as latency, throughput, power use, required layout area, etc., as inputs, which can be used when making design instances for the accelerator instance to meet the goals of the particular user. (BRI: the design framework includes “optimizer”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) In regard to claim 4: (Original) Jain discloses: - and determining a respective fusion strategy for the neural network that optimizes an execution of the neural network on the hardware accelerator computer chip having the candidate hardware datapath based on the initial estimates for each of the layers. However, Jain and Denolf do not explicitly disclose: - for each neural network: simulating, using a computer chip performance simulator, However, Nurvi discloses: - for each neural network: simulating, using a computer chip performance simulator, [Col 22, lines 33-45]: In one implementation, the accelerator architecture instance (synthesizable RTL) produced by the template mapping is then automatically validated. To do this, one implementation of the framework derives a functional model of the vertex program to be used as the “golden” reference. Test benches are generated to compare the execution of this golden reference against simulations of the RTL implementation of the architecture instance. The framework also performs performance validation by comparing RTL simulations against analytical performance model and cycle-accurate software simulator. It reports runtime breakdown and pinpoint the bottlenecks of the design that affect performance. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) In regard to claim 6: (Original) Jain discloses: - include one or more of: execution time for the layer when activation inputs for the layer are stored in off-chip memory, execution time for the layer when activation inputs for the layer are stored in on-chip memory, time required to access activation inputs for the layer from off-chip memory, or time required to access weights for the layer from off-chip memory [Col 24, lines 52-58]: the accelerator 702 can implement a neural network processing engine . In these examples, the accelerator 702, for a set of input data 750, can execute a neural network to perform a task for which the neural network was trained . Executing a neural network on a set of input data can be referred to as inference or performing inference. [Col 4, lines 62-67]: the data can be of multi-channel and a particular channel of data needs to be accessed at a particular time, but the on-chip memory cannot provide such access at that time. As another example, the on-chip memory is simply too small to fit the multi-channel data . In all these cases, the compiler can exclude an operator [Col 5, lines 1-5]: from the fused operators and generate off-chip memory read/write instructions to handle data transfer between that operator and the other fused operators. [Col 19, lines 61-67]: The first set of instructions can be in the form of a computational graph instruction represented by block diagram 400 of FIG. 4A. The first operator and the second operator can correspond to, respectively, a first layer and a second layer of a neural network and each can include various operations such as multiplication and summation operation, activation function processing , pooling operation, etc [Col 21, lines 30-35]: The example of FIG. 7 illustrates an accelerator 702. In various examples, the accelerator 702, for a set of input data (e.g., input data 750 ), can execute computations using a processing engine array 710, an activation engine 716 , and/or a pooling engine 718 [Col 12, lines 36-40]: Block 380 can include access instructions to an on-chip memory, such as a scratchpad memory of a hardware accelerator, a SRAM cache of a CPU, etc ., which is typically much faster than off-chip memory. [Col 12, lines 41-44]: FIG. 3D illustrates an example of fused kernel instructions 372. As shown in FIG. 3D, kernel instructions 372 can include read instructions 382 to fetch input data elements i.sub.0 . . . i.sub.n from an off-chip memory, (BRI: accessing memory during execution is a measurable performance statistic) In regard to claim 7: (Original) Jain and Denolf do not explicitly disclose: - performing pre-processing to optimize compute-intensive operations performed by one or more of the layers of the neural network model. However, Nurvi discloses: - performing pre-processing to optimize compute-intensive operations performed by one or more of the layers of the neural network model. [Abstract]: Hardware accelerator templates and design frameworks for implementing recurrent neural networks (RNNs) and variants thereof [Col 14, lines 1-6]: a first means for obtaining a flow graph for a recurrent neural network ( RNN ) algorithm , the flow graph identifying a plurality of operations to be performed to implement the RNN algorithm and further identifying data dependencies between ones of the plurality of operations, [Col 27, lines 44-50]: In a sparse matrix-dense vector multiplication the location of each element in the vector is determined by its index, making it feasible to gather the vector elements that correspond to the non-zero values in a region of the matrix and to pre- compute the set of vector elements that need to be gathered for any dense vector that the matrix will be multiplied by. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) In regard to claim 8: (Currently Amended) Jain discloses: - determining an optimized execution schedule for the layers of the neural network to optimize at least an execution time of the neural network on the computer chip, wherein the initial estimates of the performance statistics are determined when the model is executed in accordance with the optimized schedule [Col 24, lines 52-58]: the accelerator 702 can implement a neural network processing engine . In these examples, the accelerator 702, for a set of input data 750, can execute a neural network to perform a task for which the neural network was trained . Executing a neural network on a set of input data can be referred to as inference or perform ing inference. [Col 4, lines 62-67]: the data can be of multi-channel and a particular channel of data needs to be accessed at a particular time, but the on-chip memory cannot provide such access at that time. As another example, the on-chip memory is simply too small to fit the multi-channel data . In all these cases, the compiler can exclude an operator [Col 5, lines 1-5]: from the fused operators and generate off-chip memory read/write instructions to handle data transfer between that operator and the other fused operators. (BRI: the ability to choose an on-chip or off-chip memory for execution is a “scheduling” operation In regard to claim 9: (Currently Amended) Jain and Denolf do not explicitly disclose: - selecting, as a final hardware datapath for the hardware accelerator computer chip, the candidate hardware datapath specified that resulted in the optimized value of the objective function while repeatedly performing the operations. However, Nurvi discloses: - selecting, as a final hardware datapath for the hardware accelerator computer chip, the candidate hardware datapath specified that resulted in the optimized value of the objective function while repeatedly performing the operations. [Col 25, lines 58-67]: The accelerator logic chip 2305 at the bottom of the accelerator stack is customized to the needs of sparse-matrix computations, and is able to consume the bandwidth offered by a DRAM stack 2301-2304 while only expending 2-4 Watts of power, with energy consumption proportional to the bandwidth of the stack. To be conservative, a stack bandwidth of 273 GB/sec is assumed (the expected bandwidth of WIO3 stacks) for the remainder of this application. Designs based on higher-bandwidth stacks would incorporate more parallelism in order to consume the memory bandwidth. [Col 3, lines 42-44]: FIG. 23 illustrates one implementation of an accelerator includes an accelerator logic die and one of more stacks of DRAM die according to some embodiments. PNG media_image6.png 476 658 media_image6.png Greyscale [Col 8, lines 52-63]: Turning back to the FIG. 4, framework 400 module can include a template mapping module 404 that produces a customized accelerator instance (e.g., such as synthesizable register transfer language (RTL) utilizing a hardware description language (HDL) such as Verilog, VHDL, etc.) of the hardware accelerator that best meets the input constraints 402 and optimization goals 408 . Alongside the RTL, in some embodiments the framework 400 module also generates a compiler to program the accelerator , e.g., via providing micro-code executed by control units , as described further herein. (BRI: the above represents the synthesis of the hardware and the arrangement of matrix processing unit may indeed represent selection of the final hardware as the matrix processing (datapath) directly impact the overall performance of the accelerator) [Col 24, lines 54-59]: repeat ed use of the same matrix makes it practical to transfer matrices to/from an accelerator during program execution and/or to re-format the matrix in a way that simplifies the hardware's task, since the cost of data transfers/transformations can be amortized across many operations on each matrix. [Col 28, lines 34-46]: The second challenge is that fetching the entire vector from stack DRAM for each block of the matrix has the potential to waste significant amounts of bandwidth ( i.e., fetching vector elements for which there is no corresponding non-zero in the block). This is particularly an issue for sparse matrix-dense vector multiplication, where the vector can be a significant fraction of the size of the sparse matrix. To address this, one implementation constructs a fetch list 2611-2612 for each block 2601-2602 in the matrix, which lists the set of vector 2610 elements that correspond to non-zero values in the block, and only fetch those elements when processing the block. [Col 28, lines 51-63]: Thus, a matrix-vector multiplication on Accelerator will involve the following sequence of operations: 1. Fetch a block of matrix data from the DRAM stack and distribute it across the dot-product engines; 2. Generate fetch list based on non-zero elements in the matrix data; 3. Fetch each vector element in the fetch list from stack DRAM and distribute it to the dot-product engines; 4. Compute the dot-product of the rows in the block with the vector and write the results out to stack DRAM ; and 5. In parallel with the computation, fetch the next block of matrix data and repeat until the entire matrix has been processed. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) In regard to claim 11: (Currently Amended) Jain does not explicitly disclose: - i) different configurations of processing elements included in the computer chip and (ii) different memory configurations for on-chip memory included in the computer chip. However, Denolf discloses: - i) different configurations of processing elements included in the computer chip and [0058]: FIG. 7 is a block diagram depicting a programmable IC 1 according to an example that can be used to implement the inference platform and/or training platform. The programmable IC 1 can be used as the IC 220 in FIG. 2. The programmable IC 1 includes programmable logic 3, configura tion logic 25, and configura tion memory 26 . [0058]: The logic cells 30 include circuits that can be configured to implement general logic functions of a plurality of inputs. The support circuits 31 include dedicated circuits, such as transceivers, input/output blocks, digital signal processors, memories, and the like. The logic cells and the support circuits 31 can be interconnected using the programmable interconnect 32. Information for programming the logic cells 30, for setting parameters of the support circuits 31, and for programming the programmable interconnect 32 is stored in the configuration memory 26 by the configuration logic 25. The configuration logic 25 can obtain the configura tion data from the nonvolatile memory 27 or any other source (e.g., the DRAM 28 or from the other circuits 29). In some examples, the programmable IC 1 includes a processing system 2. The processing system 2 can include microprocessor(s), memory, support circuits, IO circuits, and the like. [0062]: FIG. 9 illustrates a field programmable gate array (FPGA) implementation of the programmable IC 1 that includes a large number of different programmable tiles including transceivers 37, configura ble logic blocks (“CLBs”) 33 - (ii) different memory configurations for on-chip memory included in the computer chip. [0058]: FIG. 7 is a block diagram depicting a programmable IC 1 according to an example that can be used to implement the inference platform and/or training platform. The programmable IC 1 can be used as the IC 220 in FIG. 2. The programmable IC 1 includes programmable logic 3, configura tion logic 25, and configuration memory 26. The programmable IC 1 can be coupled to external circuits, such as nonvolatile memory 27, DRAM 28, and other circuits 29. [0058]: The configuration logic 25 can obtain the configuration data from the nonvolatile memory 27 or any other source (e.g., the DRAM 28 or from the other circuits 29). In some examples, the programmable IC 1 includes a processing system 2. The processing system 2 can include microprocessor(s), memory, support circuits, IO circuits, and the like. [0062]: FIG. 9 illustrates a field programmable gate array (FPGA) implementation of the programmable IC 1 that includes a large number of different programmable tiles including transceivers 37, configura ble logic blocks (“CLBs”) 33, random access memory blocks (“BRAMs”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). In regard to claim 15: (Currently Amended) Jain does not explicitly disclose: - wherein the optimizer is configured to generate candidate sets of hyperparameters that satisfy one or more constraints on an area of the computer chip. However, Denolf discloses: - wherein the optimizer is configured to generate candidate sets of hyperparameters that satisfy one or more constraints on an area of the computer chip. [0023]: the techniques described herein for training using implementation cost as an objective are complementary to those techniques. These and further aspects of optimizing network parameters and/or feature maps based on architecture constraints of the inference platform [0038]: A general formulation of multi-objective optimization is as follows: PNG media_image7.png 43 282 media_image7.png Greyscale where f.sub.1, . . . , f.sub.x are functions that define the cost of each objective that is being optimized, x is a vector representing the current solution, and X is the search space of all possible solutions. In the examples described herein, x represents a neural network topology and its associated hyperparameters (i.e., the model-capacity hyperparameters 108). The functions represent metrics of interest of the current neural network topology in relation to its accuracy and implementation/hardware cost . For accuracy, these functions include mean squares error (MSE), classification error, l.sub.p norm, hingle loss, or a similar metric suitable for the target domain. For implementation/hardware cost, these functions include memory requirements, bandwidth requirements, clock cycles, datapath width, quantization scheme, arithmetic style, number formats, silicon area, and energy consumption, and error tolerance. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). In regard to claim 19: (Currently Amended) - A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising: [Abstract]: A method of generating executable instructions for a computing system is provided [Col 1, lines 34-37]: FIG. 3A-FIG. 3D illustrate examples of computational graph instructions that can be converted into executable instructions to be executed by a computing system to implement a neural network; [Col 21, lines 28-67], - obtaining data specifying a target set of one or more neural networks; obtaining data specifying an objective function that measures a performance of a hardware accelerator computer chip when performing inference for the target set of one or more neural networks, each of the one or more neural networks having a respective set of associated tensors that includes one or more of [Col 1, lines 6-10]: Artificial neural networks are computing systems with an architecture based on biological neural networks. Artificial neural networks can be trained, using training data, to learn about how to perform a certain computing task for an application. [Col 6, lines 45-49]: It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 6, lines 33-36]: Prediction model 103 can be in the form of an artificial neural network . The artificial neural network may include a plurality of processing nodes, with each processing node configured to process part of the input pixel data, [Col 6, lines 38-42]: FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. - (i) one or more weight tensors each representing weights of a respective layer of the neural network, [Col 6, lines 50-59]: Layer 207 may process pixel data representing different portions of image 104. For example, in the example of FIG. 2A, layer 207 may process the pixel data of image 204. Each processing node of layer 207 is assigned to receive a pixel value (e.g., x.sub.0, x.sub.1, x.sub.2, . . . x.sub.n) corresponding to a predetermined pixel within image 104, and transmit one or more weights with the received pixel value to layer 209. In a case where prediction model 203 is a DNN, each processing node of layer 207 can be assigned a set of weights defined based on a matrix W1 (BRI: A set of weights defined on a matrix does represent a weight tensor) [Col 6, lines 38-45]: FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. Prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer (not shown in FIG. 2A). [Col 6, lines 66-67]: Different neural network models may include different topologies (e.g., including a different number of layers, [Col 7, lines 1-2]: different connections between layers, etc.), and/or include a different set of weights for each layer. - (ii) one or more input activation tensors each representing an input to a respective layer of the neural network, or (ii) one or more output activation tensors each representing an output of a respective layer of the neural network; Col 1, lines 58-61]: An artificial neural network typically consist of a sequence of layers/operators. The operations involved in each operator/layer may involve, for example, multiplication and summation operations, activation function processing [Col 9, lines 22-28]: Referring back to FIG. 2A, one processing node of layer 209 may be configured to generate the convolution output elements of one convolution output array, and a set M of processing nodes of layer 209 can correspond to a set M of convolution output arrays. The processing node of layer 209 can also process each convolution output with an activation function to generate an activation output. [Col 6, lines 34-48]: The artificial neural network may include a plurality of processing nodes, with each processing node configured to process part of the input pixel data, or to further process the intermediate outputs from other processing nodes. FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. Prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer (not shown in FIG. 2A). It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 3, lines 53-59]: Specifically, as part of the splitting operation, the compiler can determine, for each read instruction at the virtual data node for an input data element to the second operator, one or more corresponding write instructions that supply the output data element(s) of the first operator included in the input data element, based on the tensor addresses included in the read and write instructions - (ii) a respective fusion strategy for each of the one or more neural networks when deployed on the hardware accelerator computer chip having the determined architecture, the respective fusion strategy for each of the one or more neural networks specifying, for each tensor in the set of associated tensors for the neural network, whether or not the tensor is stored in on-chip memory of the hardware accelerator computer chip during processing of inputs using the neural network, the determining comprising: [Col 2, lines 57-67]: creating fused kernels for a hardware accelerator can present extra challenges. Specifically, the on-chip memory of the hardware accelerator, which can act as a scratchpad memory, is typically managed by the kernel. As part of the management, the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators . A fuse d kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from the [Col 3, lines 1-6]: on-chip memory, to control the data movement between the operators. Orchestrating such data movements and compute operations for these slices across operators is tedious and error-prone, which can further increase the engineering effort required to support different fused operators and for different neural network topologies. [Col 2, lines 4-20]: A computing system can be programmed to implement a multi-layer artificial neural network. The instruction file may include a plurality of kernels, with each kernel including instructions that define the operations involved in a layer/operator. The computing system can execute a kernel of a first operator, generate a first intermediate tensor, and store the first intermediate tensor at an off-chip memory (e.g., DRAM). The computing system can then execute a kernel of a second operator. To execute the second kernel, the computing system can fetch the first intermediate tensor from the memory, map the first intermediate tensor to the inputs of the second operator based on the connectivity between the two operators/layers, and execute the second kernel to generate a third intermediate tensor. The computing system can store the third intermediate tensor back to the memory , and then repeat the execution and memory access operations for subsequent layers/operators. [Col 2, lines 61-67]: the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators. A fused kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from [Col 2, lines 1-2]: the on-chip memory , to control the data movement between the operators. [Col 3, lines 25-31]: the virtual data node can represent a logical tensor of the output data elements by the first operator. Each output data element of the first operator may be associated with a tensor address (e.g., coordinates) within the logical tensor represented by the virtual data node. The write instructions in the first kernel can include the tensor addresses to which the output data elements are to be stored. - and (ii) a respective optimized fusion strategy for each of the one or more neural networks from a search space of possible optimized fusion strategies for the neural network when deployed on a hardware accelerator chip [Col 19, lines 4-9]: For each virtual data node (which does not have corresponding read/write instructions), vertical fusion module 516 can convert the read or write instructions into memory read or memory write instructions for an on-chip memory (e.g., a CPU cache, a scratchpad in a hardware accelerator, etc.) [Col 17, lines 33-46]: Virtual data node splitting module 514 can then process read instruction 526 which accesses tensor addresses 0, 1, and 2 for input data element b1. For read instruction 526, virtual data node splitting module 514 can identify the corresponding write instructions 524, 532, and 534 which writes output data elements a.sub.0, a.sub.1, and a.sub.2 to tensor addresses 0, 1, and 2. Virtual data node splitting module 514 can look for an access group which includes write instructions 524, 532, and 534 (or a superset including write instructions 528a-528c). As only access group 520 is created at this point and access group 520 includes only write instruction 524, virtual data node splitting module 514 can create an access group 540 which includes read instructions 526, 528, 530 and corresponding write instructions 524, 532, and 534) [Col 21, lines 40-52]: In various implementations, the memory subsystem 704 can include multiple memory banks 714. In these implementations, each memory bank 714 can be independently accessible, meaning that the read of one memory bank is not dependent on the read of another memory bank. Similarly, writing to one memory bank does not affect or limit writing to a different memory bank. In some cases, each memory bank can be read and written at the same time. Various techniques can be used to have independently accessible memory banks 714. For example, each memory bank can be a physically separate memory component that has an address space that is separate and in depende nt of the address spaces of each other memory bank. [Col 21, lines 57-67]: the memory subsystem 704 can include arbitration logic such that arbitration between, for example, the outputs of multiple memory banks 714 can result in more than one memory bank's output being used . In these and other examples, though globally managed by the memory subsystem 704, each memory bank can be operated independently of any other. In some examples, accelerator 702 can be programmed by instructions generated based on the disclosed techniques to perform data transfer between fused operators using memory subsystem 704. (BRI: a fusion model is an accelerator framework and the fusion module uses a cost model to decide which fusion to apply a fusion search space without violating the data dependency and memory constraints) [Col 27, lines 48-56]: perform various steps before producing the instructions that are to be executed by the acceleration engine 812. These steps can include, for example, removing redundant dependencies, resolving or handling dependencies between nodes by inserting synchronization instructions into the code, identifying possibly optimi zations in memory usage or memory bandwidth usage, and other operations. [Col 15, lines 39-49]: In some examples, the compiler can generate the sequence of executable instructions to maximize the number of operators to be fuse d, but exclude an operator from the fuse d operators when, for example, limitations from the computing system prevent the fusion of that operator with other operators . Such limitations may arise from various sources. For example, the data can be of multi-channel and a particular channel of data needs to be accessed at a particular time, but the on-chip memory cannot provide such access at that time. As another example, the on-chip memory is simply too small to fit the multi-channel data. In all these cases, the compiler can remove an operator from the fused operators and generate off-chip memory read/write instructions to handle data transfer between that operator and the other fused operators. (BRI: maximizing the number of operators to be fused can represent an fusing optimization strategy) - having an architecture specified by the candidate hardware datapath; Col 10, lines 5-26]: a neural network performs a sequence of computation operations to generate a decision . The sequence of computation operations can be represented by a computational graph . The left of FIG. 3A illustrates an example of a simplified computational graph 300 representing a sequence of operators. As shown in FIG. 3A, computation graph 300 includes a set of nodes 302, 304, and 306, as well as edges 308 and 310 connecting between the nodes. Each node in computation graph 300 can represent an operator, which can represent a neural network layer in, for example, prediction model 103. For example, node 302 can correspond to operator Op1 which can represent input layer 207 of FIG. 2A, node 304 can correspond to intermediate layer 209 of FIG. 2A, whereas node 306 can correspond to intermediate layer 211 of FIG. 2A. Moreover, edge 308 can represent flow of data from another node (not shown in FIG. 3A) to node 302 (input layer 207), edge 310 can represent flow of data from node 302 to node 304 (intermediate layer 209), edge 312 can represent flow of data from node 304 (intermediate layer 209) to node 306 (intermediate layer 211), whereas edge 314 can represent flow of data from node 306 to another node (not shown in FIG. 3A). PNG media_image1.png 798 646 media_image1.png Greyscale [Col 11, lines 39-43]: A computing system, such as a hardware accelerator, a CPU, etc., can execute the operators Op1, Op2, and Op3, as well as the data transfer operations, by executing the kernel instructions 322, 324, and 326 as well as the read/write operations represented in blocks 328, 340, and 342. [Col 19, lines 54-67]: compiler receives a first set of instructions including a kernel of a first operator and a kernel of a second operator. The kernel of the first operator can include instructions of the first operator and write instructions to a virtual data node , whereas the kernel of the second operator can include instructions of the second operator and read instructions to the virtual data node . The first set of instructions can be in the form of a computational graph instruction represented by block diagram 400 of FIG. 4A. The first operator and the second operator can correspond to, respectively, a first layer and a second layer of a neural network and each can include various operations such as multiplication and summation operation, activation function processing, pooling operation, [Col 20, lines 1-6]: etc. The output data of the first operator/layer can be fetched to the second operator/layer as input. As shown in FIG. 4C, the virtual data node can correspond to a logical tensor to store the output data of the first operator and from which the second operator fetches input data , to provide the data transfer from the first operator to the second operator. (BRI: data transfer operations, multiplication and summation operations indeed represent by the datapath in a CPU) However, Denolf discloses: - obtaining data specifying an objective function that measures a performance of a hardware accelerator computer chip when performing inference for the target set of one or more neural networks, each of the one or more neural networks having a respective set of associated tensors that includes one or more of [0030]: FIG. 2 is a block diagram depicting a computing system (“computer 200”) according to an example. The computer 200 includes a software platform 204 executing on a hardware platform 202. The hardware platform 202 includes a central processing unit (CPU) 206, a system memory 208, storage devices 210, support circuits 211, a training platform 212, and a hardware acce lerator 214 . [0023]: Reinforcement learning provides for multi-objective optimization, but without adding the implementation cost of the neural network itself as an objective. - and determining (i) an architecture for the hardware accelerator computer chip [0030]: FIG. 2 is a block diagram depicting a computing system (“computer 200”) according to an example. The computer 200 includes a software platform 204 executing on a hardware platform 202. The hardware platform 202 includes a central processing unit (CPU) 206 , a system memory 208, storage devices 210, support circuits 211, a training platform 212, and a hardware acce lerator 214 . PNG media_image2.png 710 648 media_image2.png Greyscale [0017] : FIG. 7 is a block diagram depicting a programmable integrated circuit (IC) according to an example. PNG media_image3.png 767 402 media_image3.png Greyscale [0018]: FIG. 8 is a block diagram depicting a System-on- Chip (SoC) implementation of the programmable IC of FIG. 7 PNG media_image4.png 796 648 media_image4.png Greyscale (BRI: the entire FIG 8 as an SoC embodies the hardware accelerator) [0031]: in some examples, the CPU 206 can be a System-in-Package (SiP), System-on- Chip (SoC), or the like, which absorbs all or a substantial portion of the functionality of the chip set Multi-Objective Optimization [ 0037 ]: The inclusion of inference implementation cost when evaluating the performance of networks means there are at least two objectives that are to be optimized . As such, multiple objectives should be balanced in a meaningful way . For example, assume the accuracy of the network is given by classification error, C.sub.E, and the estimated implementation cost is given by the time taken to process a new input, C.sub.T. If minimizing C.sub.T is given too much importance, then it is possible an optimizer will produce a network with zero layers, zero operations, and zero memory requirements . This could yield a network that has C.sub.T=0, despite incurring a significantly high C.sub.E. Multi-objective optimization aims to balance C.sub.E and C.sub.T to give a desirable solution. [0038] : A general formulation of multi-objective optimization is as follows: PNG media_image5.png 72 481 media_image5.png Greyscale where f.sub.1, . . . , f.sub.x are functions that define the cost of each objective that is being optimized, x is a vector representing the current solution, and X is the search space of all possible solutions . In the examples described herein, x represents a neural network topology and its associated hyperparameters (i.e., the model-capacity hyperparameters 108). The functions represent metrics of interest of the current neural network topology in relation to its accuracy and implementation/hardware cost. For accuracy, these functions include mean squares error (MSE), classification error, l.sub.p norm, hingle loss, or a similar metric suitable for the target domain. For implementation/hardware cost, these functions include memory requirements, bandwidth requirements, clock cycles, datapath width, quantization scheme, arithmetic style, number formats, silicon area, and energy consumption, and error tolerance. - Repeatedly performing operations comprising: [0047]: The basic methodology of evolutionary algorithms is to generate N random strings of genes (which correspond to neural network architectures ) (step 402). These architectures are then evaluated using a fitness function, which may require training each network architecture individually (step 404). At this point, a subset of the architectures are selected, randomly combined and mutated to generate the next N architectures (step 406). Over time, this results in architectures which are highly optimized for the given cost functions , which in this case means high accuracy and low implementation/hardware cost. At step 408, a determination is made whether to end. If not, the method 400 proceeds to step 404 and repeats. Otherwise, the method 400 proceeds to step 410, where the training platform outputs the trained neural network. - determining (i) a candidate set of hyperparameters that define a candidate hardware datapath for the hardware accelerator computer chip [0035]: the applications 236 include software that trains neural networks on the training platform 212 and implements neural networks on the hardware accelerator 214. [0059] : FIG. 8 is a block diagram depicting a System-on- Chip (SoC ) implementation of the programmable IC 1 according to an example. [0006]: a method of implementing a neural network includes: selecting a first neural network architecture from a search space ; training the neural network having the first neural network architecture to obtain an accuracy and an implementation cost , the implementation cost based on a programmable device of an inference platform ; selecting a second neural network architecture from the search space based on the accuracy and the implementation cost; and outputting weights and hyperparameter s for the neural network having the second neural network architecture. [0052]: A Bayesian hyperparameter search is a more sophisticated technique which attempts to develop a statistical model which maps the hyperpar ameter values to our cost function . Usually, this statistical model is a Gaussian Process (GP) which generates functions which closely approximates the observed data. GPs provide a prediction for the chosen cost function in the hyperpar ameter space , along with the uncertainty of such predictions, this has the following benefits over random search and grid search: (BRI: a hyperparameter space represents a set of hyperparameters) - from a search space of possible hardware datapaths for the hardware accelerator computer chip [0008]: computer system includes: a memory having program code stored therein; and a processor, configured to execute the program code, to implement a neural network by: selecting a first neural network architecture from a search space ; [0028] : The topology 120 generally includes an arrangement of neurons. For example, the topology 120 can include a plurality of layers of neurons. The layers generally include an input layer, an output layer , and zero or more hidden layers . Each neuron includes a plurality of inputs and an output. The plurality of inputs for each neuron are associated with a plurality of weights. Each neuron further includes a bias associated with its output. The weights and biases of the neural network 106 are referred to as trained network weights 114. For a given layer, the inputs of its neurons are referred to as input feature maps and the outputs of its neurons are referred to as output feature maps. Input feature maps and output feature maps are generally referred to as “feature maps.” (BRI: the above represents “ neural network architecture”) [0031]: In an example, the CPU 206 can be any type of general-purpose central processing unit (CPU), such as an x86-based processor, ARM®-based processor, or the like. The CPU 206 can include one or more cores and associated circuitry (e.g., cache memories, memory management units (MMUs), interrupt controllers, etc.). The CPU 206 is configured to execute program code that perform one or more operations described herein and which can be stored in the system memory 208 and/or the storage devices 210. The support circuits 211 include various devices that cooperate with the CPU 206 to manage data flow between the CPU 206 , the system memory 208, the storage devices 210, the training platform 212, the hardware accelerator 214 , or any other peripheral device. [0053]: In the methods above, the size/complexity of the neural architecture search space can be reduced by only making certain aspects of the network variable. For instance, making only the bit width of the feature map elements and the number of channels of the feature maps variable enables training for their optimum setting. Typically, reducing the bit width of the feature map elements results in less accuracy while allowing a more efficient implementation. (BRI: the bit width is directly tied to the datapath width as a result of the hardware processing unit handling the ALU or data transfer operation based on that bit width) - and determining a value of the objective function for the candidate hardware datapath by, for each neural network in the set, [0047]: The basic methodology of evolutionary algorithms is to generate N random strings of genes (which correspond to neural network architectures) (step 402). These architectures are then evaluated using a fitness function, which may require training each network architecture individually (step 404). At this point, a subset of the architectures are selected, randomly combined and mutated to generate the next N architectures (step 406). Over time, this results in architectures which are highly optimized for the given cost function s, which in this case means high accuracy and low implementation/hardware cost. ( BRI: the cost function is the value of the objective function) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). Jain and Denolf do not explicitly disclose: - simulating a performance of a candidate hardware accelerator computer chip that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; - and selecting a final hardware datapath for the hardware accelerator computer chip from the candidate hardware datapaths based on the respective values of the objective functions for the candidate hardware datapaths. However, Nurvi discloses: - simulating a performance of a candidate hardware accelerator computer chip that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; [Col 9, lines 16-20]: The hardware accelerator template 450 also includes one or more vector processing units 470A-470N (VPUs), which includes one or more FMAs 472 and/or one or more activation function blocks (for performing needed activation functions efficiently in hardware) [Col 22, lines 33-45]: the accelerator architecture instance (synthesizable RTL) produced by the template mapping is then automatically validated . To do this, one implementation of the framework derives a functional model of the vertex program to be used as the “golden” reference. Test benches are generated to compare the execution of this golden reference against simula tions of the RTL implementation of the architecture instance . The framework also performs performance validation by comparing RTL simulations against analytical performance model and cycle-accurate software simula tor . It reports runtime breakdown and pinpoint the bottlenecks of the design that affect performance. [Col 7, lines 28-37]: some embodiments can produce an accelerator instance optimized for a target FPGA chip with a particular number of hardware multiply and on-chip RAM resources, and some embodiments can produce an accelerator instance optimized for an ASIC for a particular market segment, programmable to support all RNN applications in this segment. For example , in embodiments where the accelerator instance comprises RTL code, the RTL code can be used as an input for a standard ASIC developmental tool (e.g., a logic synthesis tool) to generate an ASIC design . [Col 18, lines 46-54]: One implementation of the accelerator 900 can be programmed through a software library (similar to Intel® Math Kernel Library). Such library prepares the matrix data in memory, sets control registers in the accelerator 900 with information about the computation (e.g., computation type, memory pointer to matrix data), and starts the accelerator. Then, the accelerator independently accesses matrix data in memory, performs the computation, and writes the results back to memory for the software to consume. (BRI: Perhaps as known to a POSITA, the software library of kernel library provide fusing [Col 7, lines 3-11]: the RNN variants may dictate the size s and types of the matrix and vector operations, as well as their data dependencies . For example, matrix and vector size s can be related to the number of hidden units in the RNN, and the activation function (AF) type (tan h, sigmoid, etc.) can relate to the type of vector operations . Thus, many matrix and vector operations in RNNs make them computationally intensive, so being able to execute RNNs as efficient as possible is of critical importance. [Col 18, lines 55-61]: The accelerator handles the different compute patterns by setting its PEs to the proper datapat h configuration , as depicted in FIGS. 15a-15b . In particular, FIG. 15a highlights paths (using dotted lines) for spMspV_csc and scale_update operations and FIG. 15b illustrates paths for a spMdV_csr operation. The accelerator operation to perform each compute pattern is detailed below. [Col 18, lines 62-67]: For spMspV_csc, the initial y vector subset is loaded in to PE's RAM 1421 by the DMU 905. It then reads x vector elements from memory. For each x element, the DMU 905 streams the elements of the corresponding matrix column from memory and supplies them to the PE 901. Each matrix element contains a value (A.val) and an index (A.idx) which [Col 19, lines 1-9]: points to the y element to read from PE's RAM 1421. The DMU 1005 also provides the x vector element (x.val) that is multiplied against A.val by the multiply-accumulate (FMA) unit . The result is used to update the y element in the PE's RAM pointed to by A.idx. Note that even though not used by our workloads, the accelerator also supports column-wise multiplication against a dense x vector (spMdV_csc) by processing all matrix columns instead of only a subset (since x is dense). (BRI:Perhaps as known to a POSITA, the sizes of the matrices and hidden state in RNN variant does represent hyperparameters. The matrix operations can represent datpath operation) - and selecting a final hardware datapath for the hardware accelerator computer chip from the candidate hardware datapaths based on the respective values of the objective functions for the candidate hardware datapaths. 208) The accelerator logic chip 2305 at the bottom of the accelerator stack is customized to the needs of sparse-matrix computations, and is able to consume the bandwidth offered by a DRAM stack 2301-2304 while only expending 2-4 Watts of power, with energy consumption proportional to the bandwidth of the stack. To be conservative, a stack bandwidth of 273 GB/sec is assumed (the expected bandwidth of WIO3 stacks) for the remainder of this application. Designs based on higher-bandwidth stacks would incorporate more parallelism in order to consume the memory bandwidth. [Col 3, lines 42-44]: FIG. 23 illustrates one implementation of an accelerator includes an accelerator logic die and one of more stacks of DRAM die according to some embodiments. PNG media_image6.png 476 658 media_image6.png Greyscale [Col 8, lines 52-63]: Turning back to the FIG. 4, framework 400 module can include a template mapping module 404 that produces a customized accelerator instance (e.g., such as synthesizable register transfer language (RTL) utilizing a hardware description language (HDL) such as Verilog, VHDL, etc.) of the hardware accelerator that best meets the input constraints 402 and optimization goals 408 . Alongside the RTL, in some embodiments the framework 400 module also generates a compiler to program the accelerator , e.g., via providing micro-code executed by control units , as described further herein. (BRI: the above represents the synthesis of the hardware and the arrangement of matrix processing unit may indeed represent selection of the final hardware as the matrix processing (datapath) directly impact the overall performance of the accelerator) [Col 11, lines 22-31]: Flow 800 also includes, at block 815, mapping the plurality of operations of the flow graph to an accelerator hardware template to yield the accelerator instance comprising register transfer language code t hat describes how one or more matrix processing units (MPUs) and one or more vector processing units (VPUs) are to be arranged to perform the RNN algorithm . At least one of the one or more MPUs, as part of implementing the RNN algorithm, is to directly provide or directly receive a value from one of the one or more VPUs. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) In regard to claim 20: (Currently Amended) - One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising: [Col 25, lines 48-54], [Col 21, lines 28-67] - obtaining data specifying a target set of one or more neural networks; [Col 1, lines 6-10]: Artificial neural networks are computing systems with an architecture based on biological neural networks. Artificial neural networks can be trained, using training data, to learn about how to perform a certain computing task for an application. [Col 6, lines 45-49]: It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 6, lines 33-36]: Prediction model 103 can be in the form of an artificial neural network . The artificial neural network may include a plurality of processing nodes, with each processing node configured to process part of the input pixel data, [Col 6, lines 38-42]: FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. - obtaining data specifying an objective function that measures a performance of a hardware accelerator computer chip when performing inference for the target set of one or more neural networks, each of the one or more neural networks having a respective set of associated tensors that includes one or more of (i) one or more weight tensors each representing weights of a respective layer of the neural network, [Col 6, lines 50-59]: Layer 207 may process pixel data representing different portions of image 104. For example, in the example of FIG. 2A, layer 207 may process the pixel data of image 204. Each processing node of layer 207 is assigned to receive a pixel value (e.g., x.sub.0, x.sub.1, x.sub.2, . . . x.sub.n) corresponding to a predetermined pixel within image 104, and transmit one or more weights with the received pixel value to layer 209. In a case where prediction model 203 is a DNN, each processing node of layer 207 can be assigned a set of weights defined based on a matrix W1 (BRI: A set of weights defined on a matrix does represent a weight tensor) [Col 6, lines 38-45]: FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. Prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer (not shown in FIG. 2A). [Col 6, lines 66-67]: Different neural network models may include different topologies (e.g., including a different number of layers, [Col 7, lines 1-2]: different connections between layers, etc.), and/or include a different set of weights for each layer. (ii) one or more input activation tensors each representing an input to a respective layer of the neural network, or [Col 1, lines 58-61]: An artificial neural network typically consist of a sequence of layers/operators. The operations involved in each operator/layer may involve, for example, multiplication and summation operations, activation function processing [Col 9, lines 22-28]: Referring back to FIG. 2A, one processing node of layer 209 may be configured to generate the convolution output elements of one convolution output array, and a set M of processing nodes of layer 209 can correspond to a set M of convolution output arrays. The processing node of layer 209 can also process each convolution output with an activation function to generate an activation output. [Col 6, lines 34-48]: The artificial neural network may include a plurality of processing nodes, with each processing node configured to process part of the input pixel data, or to further process the intermediate outputs from other processing nodes. FIG. 1 illustrates an example of prediction model 103 that uses techniques disclosed herein. In FIG. 1, prediction model 103 may be a multi-layer neural network such as a deep neural network (DNN), a convolutional neural network (CNN), etc. Prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer (not shown in FIG. 2A). It is understood that prediction model 103 can also include other different types of neural networks including, for example, long short-term memory (LSTM), multilayer perception (MTP), multiscale densenet (MSDNET), etc. [Col 3, lines 53-59]: Specifically, as part of the splitting operation, the compiler can determine, for each read instruction at the virtual data node for an input data element to the second operator, one or more corresponding write instructions that supply the output data element(s) of the first operator included in the input data element, based on the tensor addresses included in the read and write instructions - and (ii) a respective fusion strategy for each of the one or more neural networks when deployed on the hardware accelerator computer chip having the determined architecture, [Col 2, lines 57-67]: creating fused kernels for a hardware accelerator can present extra challenges. Specifically, the on-chip memory of the hardware accelerator, which can act as a scratchpad memory, is typically managed by the kernel. As part of the management, the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators . A fuse d kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from the [Col 3, lines 1-6]: on-chip memory, to control the data movement between the operators. Orchestrating such data movements and compute operations for these slices across operators is tedious and error-prone, which can further increase the engineering effort required to support different fused operators and for different neural network topologies. [Col 2, lines 4-20]: A computing system can be programmed to implement a multi-layer artificial neural network. The instruction file may include a plurality of kernels, with each kernel including instructions that define the operations involved in a layer/operator. The computing system can execute a kernel of a first operator, generate a first intermediate tensor, and store the first intermediate tensor at an off-chip memory (e.g., DRAM). The computing system can then execute a kernel of a second operator. To execute the second kernel, the computing system can fetch the first intermediate tensor from the memory, map the first intermediate tensor to the inputs of the second operator based on the connectivity between the two operators/layers, and execute the second kernel to generate a third intermediate tensor. The computing system can store the third intermediate tensor back to the memory, and then repeat the execution and memory access operations for subsequent layers/operators. [Col 2, lines 61-67]: the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators. A fused kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from [Col 2, lines 1-2]: the on-chip memory , to control the data movement between the operators. [Col 3, lines 25-31]: the virtual data node can represent a logical tensor of the output data elements by the first operator. Each output data element of the first operator may be associated with a tensor address (e.g., coordinates) within the logical tensor represented by the virtual data node. The write instructions in the first kernel can include the tensor addresses to which the output data elements are to be stored. - the respective fusion strategy for each of the one or more neural networks specifying, for each tensor in the set of associated tensors for the neural network, whether or not the tensor is stored in on-chip memory of the hardware accelerator computer chip during processing of inputs using the neural network, the determining comprising: [Col 2, lines 57-67]: creating fused kernels for a hardware accelerator can present extra challenges. Specifically, the on-chip memory of the hardware accelerator, which can act as a scratchpad memory, is typically managed by the kernel. As part of the management, the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators . A fuse d kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from the [Col 3, lines 1-6]: on-chip memory, to control the data movement between the operators. Orchestrating such data movements and compute operations for these slices across operators is tedious and error-prone, which can further increase the engineering effort required to support different fused operators and for different neural network topologies. [Col 2, lines 4-20]: A computing system can be programmed to implement a multi-layer artificial neural network. The instruction file may include a plurality of kernels, with each kernel including instructions that define the operations involved in a layer/operator. The computing system can execute a kernel of a first operator, generate a first intermediate tensor, and store the first intermediate tensor at an off-chip memory (e.g., DRAM). The computing system can then execute a kernel of a second operator. To execute the second kernel, the computing system can fetch the first intermediate tensor from the memory, map the first intermediate tensor to the inputs of the second operator based on the connectivity between the two operators/layers, and execute the second kernel to generate a third intermediate tensor. The computing system can store the third intermediate tensor back to the memory, and then repeat the execution and memory access operations for subsequent layers/operators. [Col 2, lines 61-67]: the kernel needs to divide the data to be stored into the on-chip memory by one operator into data slices to fit into the memory. The data slices are then fetched to the next operator based on the connectivity between the operators. A fused kernel needs to include instructions to indicate where the data slices are stored in the on-chip memory and how the data slices are fetched from [Col 2, lines 1-2]: the on-chip memory , to control the data movement between the operators. [Col 3, lines 25-31]: the virtual data node can represent a logical tensor of the output data elements by the first operator. Each output data element of the first operator may be associated with a tensor address (e.g., coordinates) within the logical tensor represented by the virtual data node. The write instructions in the first kernel can include the tensor addresses to which the output data elements are to be stored. - determining (i) a candidate set of hyperparameters that define a candidate hardware datapath for the hardware accelerator computer chip from a search space of possible hardware datapaths for the hardware accelerator computer chip [0035]: the applications 236 include software that trains neural networks on the training platform 212 and implements neural networks on the hardware accelerator 214. [0059] : FIG. 8 is a block diagram depicting a System-on- Chip (SoC ) implementation of the programmable IC 1 according to an example. [0006]: a method of implementing a neural network includes: selecting a first neural network architecture from a search space ; training the neural network having the first neural network architecture to obtain an accuracy and an implementation cost , the implementation cost based on a programmable device of an inference platform ; selecting a second neural network architecture from the search space based on the accuracy and the implementation cost; and outputting weights and hyperparameter s for the neural network having the second neural network architecture. [0052]: A Bayesian hyperparameter search is a more sophisticated technique which attempts to develop a statistical model which maps the hyperpar ameter values to our cost function . Usually, this statistical model is a Gaussian Process (GP) which generates functions which closely approximates the observed data. GPs provide a prediction for the chosen cost function in the hyperpar ameter space , along with the uncertainty of such predictions, this has the following benefits over random search and grid search: (BRI: a hyperparameter space represents a set of hyperparameters) - from a search space of possible hardware datapaths for the hardware accelerator computer chip [0008]: computer system includes: a memory having program code stored therein; and a processor, configured to execute the program code, to implement a neural network by: selecting a first neural network architecture from a search space ; [0028] : The topology 120 generally includes an arrangement of neurons. For example, the topology 120 can include a plurality of layers of neurons. The layers generally include an input layer, an output layer , and zero or more hidden layers . Each neuron includes a plurality of inputs and an output. The plurality of inputs for each neuron are associated with a plurality of weights. Each neuron further includes a bias associated with its output. The weights and biases of the neural network 106 are referred to as trained network weights 114. For a given layer, the inputs of its neurons are referred to as input feature maps and the outputs of its neurons are referred to as output feature maps. Input feature maps and output feature maps are generally referred to as “feature maps.” (BRI: the above represents “ neural network architecture”) [0031]: In an example, the CPU 206 can be any type of general-purpose central processing unit (CPU), such as an x86-based processor, ARM®-based processor, or the like. The CPU 206 can include one or more cores and associated circuitry (e.g., cache memories, memory management units (MMUs), interrupt controllers, etc.). The CPU 206 is configured to execute program code that perform one or more operations described herein and which can be stored in the system memory 208 and/or the storage devices 210. The support circuits 211 include various devices that cooperate with the CPU 206 to manage data flow between the CPU 206 , the system memory 208, the storage devices 210, the training platform 212, the hardware accelerator 214 , or any other peripheral device. [0053]: In the methods above, the size/complexity of the neural architecture search space can be reduced by only making certain aspects of the network variable. For instance, making only the bit width of the feature map elements and the number of channels of the feature maps variable enables training for their optimum setting. Typically, reducing the bit width of the feature map elements results in less accuracy while allowing a more efficient implementation. (BRI: the bit width is directly tied to the datapath width as a result of the hardware processing unit handling the ALU or data transfer operation based on that bit width) - and (ii) a respective optimized fusion strategy for each of the one or more neural networks from a search space of possible optimized fusion strategies for the neural network when deployed on a hardware accelerator chip [Col 19, lines 4-9]: For each virtual data node (which does not have corresponding read/write instructions), vertical fusion module 516 can convert the read or write instructions into memory read or memory write instructions for an on-chip memory (e.g., a CPU cache, a scratchpad in a hardware accelerator, etc.) [Col 17, lines 33-46]: Virtual data node splitting module 514 can then process read instruction 526 which accesses tensor addresses 0, 1, and 2 for input data element b1. For read instruction 526, virtual data node splitting module 514 can identify the corresponding write instructions 524, 532, and 534 which writes output data elements a.sub.0, a.sub.1, and a.sub.2 to tensor addresses 0, 1, and 2. Virtual data node splitting module 514 can look for an access group which includes write instructions 524, 532, and 534 (or a superset including write instructions 528a-528c). As only access group 520 is created at this point and access group 520 includes only write instruction 524, virtual data node splitting module 514 can create an access group 540 which includes read instructions 526, 528, 530 and corresponding write instructions 524, 532, and 534) [Col 21, lines 40-52]: In various implementations, the memory subsystem 704 can include multiple memory banks 714. In these implementations, each memory bank 714 can be independently accessible, meaning that the read of one memory bank is not dependent on the read of another memory bank. Similarly, writing to one memory bank does not affect or limit writing to a different memory bank. In some cases, each memory bank can be read and written at the same time. Various techniques can be used to have independently accessible memory banks 714. For example, each memory bank can be a physically separate memory component that has an address space that is separate and in depende nt of the address spaces of each other memory bank. [Col 21, lines 57-67]: the memory subsystem 704 can include arbitration logic such that arbitration between, for example, the outputs of multiple memory banks 714 can result in more than one memory bank's output being used . In these and other examples, though globally managed by the memory subsystem 704, each memory bank can be operated independently of any other. In some examples, accelerator 702 can be programmed by instructions generated based on the disclosed techniques to perform data transfer between fused operators using memory subsystem 704. (BRI: a fusion model is an accelerator framework and the fusion module uses a cost model to decide which fusion to apply a fusion search space without violating the data dependency and memory constraints) [Col 27, lines 48-56]: perform various steps before producing the instructions that are to be executed by the acceleration engine 812. These steps can include, for example, removing redundant dependencies, resolving or handling dependencies between nodes by inserting synchronization instructions into the code, identifying possibly optimi zations in memory usage or memory bandwidth usage, and other operations. [Col 15, lines 39-49]: In some examples, the compiler can generate the sequence of executable instructions to maximize the number of operators to be fuse d, but exclude an operator from the fuse d operators when, for example, limitations from the computing system prevent the fusion of that operator with other operators . Such limitations may arise from various sources. For example, the data can be of multi-channel and a particular channel of data needs to be accessed at a particular time, but the on-chip memory cannot provide such access at that time. As another example, the on-chip memory is simply too small to fit the multi-channel data. In all these cases, the compiler can remove an operator from the fused operators and generate off-chip memory read/write instructions to handle data transfer between that operator and the other fused operators. (BRI: maximizing the number of operators to be fused can represent an fusing optimization strategy) - having an architecture specified by the candidate hardware datapath; [Col 10, lines 5-26]: a neural network performs a sequence of computation operations to generate a decision . The sequence of computation operations can be represented by a computational graph . The left of FIG. 3A illustrates an example of a simplified computational graph 300 representing a sequence of operators. As shown in FIG. 3A, computation graph 300 includes a set of nodes 302, 304, and 306, as well as edges 308 and 310 connecting between the nodes. Each node in computation graph 300 can represent an operator, which can represent a neural network layer in, for example, prediction model 103. For example, node 302 can correspond to operator Op1 which can represent input layer 207 of FIG. 2A, node 304 can correspond to intermediate layer 209 of FIG. 2A, whereas node 306 can correspond to intermediate layer 211 of FIG. 2A. Moreover, edge 308 can represent flow of data from another node (not shown in FIG. 3A) to node 302 (input layer 207), edge 310 can represent flow of data from node 302 to node 304 (intermediate layer 209), edge 312 can represent flow of data from node 304 (intermediate layer 209) to node 306 (intermediate layer 211), whereas edge 314 can represent flow of data from node 306 to another node (not shown in FIG. 3A). PNG media_image1.png 798 646 media_image1.png Greyscale [Col 11, lines 39-43]: A computing system, such as a hardware accelerator, a CPU, etc., can execute the operators Op1, Op2, and Op3, as well as the data transfer operations, by executing the kernel instructions 322, 324, and 326 as well as the read/write operations represented in blocks 328, 340, and 342. [Col 19, lines 54-67]: compiler receives a first set of instructions including a kernel of a first operator and a kernel of a second operator. The kernel of the first operator can include instructions of the first operator and write instructions to a virtual data node , whereas the kernel of the second operator can include instructions of the second operator and read instructions to the virtual data node . The first set of instructions can be in the form of a computational graph instruction represented by block diagram 400 of FIG. 4A. The first operator and the second operator can correspond to, respectively, a first layer and a second layer of a neural network and each can include various operations such as multiplication and summation operation, activation function processing, pooling operation, [Col 20, lines 1-6]: etc. The output data of the first operator/layer can be fetched to the second operator/layer as input. As shown in FIG. 4C, the virtual data node can correspond to a logical tensor to store the output data of the first operator and from which the second operator fetches input data , to provide the data transfer from the first operator to the second operator. (BRI: data transfer operations, multiplication and summation operations indeed represent by the datapath in a CPU) - and determining a value of the objective function for the candidate hardware datapath by, for each neural network in the set, [0047]: The basic methodology of evolutionary algorithms is to generate N random strings of genes (which correspond to neural network architectures) (step 402). These architectures are then evaluated using a fitness function, which may require training each network architecture individually (step 404). At this point, a subset of the architectures are selected, randomly combined and mutated to generate the next N architectures (step 406). Over time, this results in architectures which are highly optimized for the given cost function s, which in this case means high accuracy and low implementation/hardware cost. ( BRI: the cost function is the value of the objective function) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain and Denolf. Jain teaches training transfer models on a training data set related to users from a source domain and target domain and using a transfer score from users in the source domain to make recommendations to the users in the target domain. Jagmohan teaches generating a transfer score for each user. One of ordinary skill would have motivation to combine Jain and Denolf that use an optimized search space that can exploit the resources of a chip ( Denolf [0054]) . optimize for the given cost function (Denolf[0046]). Jain and Denolf do not explicitly disclose: - simulating a performance of a candidate hardware accelerator computer chip that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; - and selecting a final hardware datapath for the hardware accelerator computer chip from the candidate hardware datapaths based on the respective values of the objective functions for the candidate hardware datapaths. However, Nurvi discloses: - simulating a performance of a candidate hardware accelerator computer chip that has the hardware datapath defined by the candidate set of hyperparameters when performing inference for the neural network in accordance with the respective optimized fusion strategy for the neural network; [Col 9, lines 16-20]: The hardware accelerator template 450 also includes one or more vector processing units 470A-470N (VPUs), which includes one or more FMAs 472 and/or one or more activation function blocks (for performing needed activation functions efficiently in hardware) [Col 22, lines 33-45]: the accelerator architecture instance (synthesizable RTL) produced by the template mapping is then automatically validated . To do this, one implementation of the framework derives a functional model of the vertex program to be used as the “golden” reference. Test benches are generated to compare the execution of this golden reference against simula tions of the RTL implementation of the architecture instance . The framework also performs performance validation by comparing RTL simulations against analytical performance model and cycle-accurate software simula tor . It reports runtime breakdown and pinpoint the bottlenecks of the design that affect performance. [Col 7, lines 28-37]: some embodiments can produce an accelerator instance optimized for a target FPGA chip with a particular number of hardware multiply and on-chip RAM resources, and some embodiments can produce an accelerator instance optimized for an ASIC for a particular market segment, programmable to support all RNN applications in this segment. For example , in embodiments where the accelerator instance comprises RTL code, the RTL code can be used as an input for a standard ASIC developmental tool (e.g., a logic synthesis tool) to generate an ASIC design . [Col 18, lines 46-54]: One implementation of the accelerator 900 can be programmed through a software library (similar to Intel® Math Kernel Library). Such library prepares the matrix data in memory, sets control registers in the accelerator 900 with information about the computation (e.g., computation type, memory pointer to matrix data), and starts the accelerator. Then, the accelerator independently accesses matrix data in memory, performs the computation, and writes the results back to memory for the software to consume. (BRI: Perhaps as known to a POSITA, the software library of kernel library provide fusing [Col 7, lines 3-11]: the RNN variants may dictate the size s and types of the matrix and vector operations, as well as their data dependencies . For example, matrix and vector size s can be related to the number of hidden units in the RNN, and the activation function (AF) type (tan h, sigmoid, etc.) can relate to the type of vector operations . Thus, many matrix and vector operations in RNNs make them computationally intensive, so being able to execute RNNs as efficient as possible is of critical importance. [Col 18, lines 55-61]: The accelerator handles the different compute patterns by setting its PEs to the proper datapat h configuration , as depicted in FIGS. 15a-15b . In particular, FIG. 15a highlights paths (using dotted lines) for spMspV_csc and scale_update operations and FIG. 15b illustrates paths for a spMdV_csr operation. The accelerator operation to perform each compute pattern is detailed below. [Col 18, lines 62-67]: For spMspV_csc, the initial y vector subset is loaded in to PE's RAM 1421 by the DMU 905. It then reads x vector elements from memory. For each x element, the DMU 905 streams the elements of the corresponding matrix column from memory and supplies them to the PE 901. Each matrix element contains a value (A.val) and an index (A.idx) which [Col 19, lines 1-9]: points to the y element to read from PE's RAM 1421. The DMU 1005 also provides the x vector element (x.val) that is multiplied against A.val by the multiply-accumulate (FMA) unit . The result is used to update the y element in the PE's RAM pointed to by A.idx. Note that even though not used by our workloads, the accelerator also supports column-wise multiplication against a dense x vector (spMdV_csc) by processing all matrix columns instead of only a subset (since x is dense). (BRI:Perhaps as known to a POSITA, the sizes of the matrices and hidden state in RNN variant does represent hyperparameters. The matrix operations can represent datpath operation) - and selecting a final hardware datapath for the hardware accelerator computer chip from the candidate hardware datapaths based on the respective values of the objective functions for the candidate hardware datapaths. [Col 25, lines 58-63]: The accelerator logic chip 2305 at the bottom of the accelerator stack is customized to the needs of sparse-matrix computations, and is able to consume the bandwidth offered by a DRAM stack 2301-2304 while only expending 2-4 Watts of power, with energy consumption proportional to the bandwidth of the stack. To be conservative, a stack bandwidth of 273 GB/sec is assumed (the expected bandwidth of WIO3 stacks) for the remainder of this application. Designs based on higher-bandwidth stacks would incorporate more parallelism in order to consume the memory bandwidth. [Col 3, lines 42-44]: FIG. 23 illustrates one implementation of an accelerator includes an accelerator logic die and one of more stacks of DRAM die according to some embodiments. PNG media_image6.png 476 658 media_image6.png Greyscale [Col 8, lines 52-63]: Turning back to the FIG. 4, framework 400 module can include a template mapping module 404 that produces a customized accelerator instance (e.g., such as synthesizable register transfer language (RTL) utilizing a hardware description language (HDL) such as Verilog, VHDL, etc.) of the hardware accelerator that best meets the input constraints 402 and optimization goals 408 . Alongside the RTL, in some embodiments the framework 400 module also generates a compiler to program the accelerator , e.g., via providing micro-code executed by control units , as described further herein. (BRI: the above represents the synthesis of the hardware and the arrangement of matrix processing unit may indeed represent selection of the final hardware as the matrix processing (datapath) directly impact the overall performance of the accelerator) [Col 11, lines 22-31]: Flow 800 also includes, at block 815, mapping the plurality of operations of the flow graph to an accelerator hardware template to yield the accelerator instance comprising register transfer language code t hat describes how one or more matrix processing units (MPUs) and one or more vector processing units (VPUs) are to be arranged to perform the RNN algorithm . At least one of the one or more MPUs, as part of implementing the RNN algorithm, is to directly provide or directly receive a value from one of the one or more VPUs. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf and Nurvi. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. One of ordinary skill would have motivation to combine Jain, Denolf and Nurvi that can provide accelerator optimized in various types of chip implementation such as FPGA and ASIC( Nurvi [Col 7, lines 28-37]) Claim 5 is rejected under 35 U.S.C. 103 unpatentable over Animesh Jain et.al. (hereinafter Jain ) US 11809981 B1, in view of Kristof Denolf et.al. (hereinafter Denolf ) US 2020/0104715 A1. in view of Eriko Nurvitadhi et.al. (hereinafter Nurvi ) US 11216722 B2, further in view of Srinivasa PRASANNA et.al. (hereinafter Pras ) US 2022/0188241 A1. In regard to claim 5: (Original) However , Pras discloses: - solving a constrained optimization using integer linear programming. [0150]: Partly computed operator results (especially for subsets/information content) can be stored for later access. [0151]: The graphical model can be derived from data, or from known apriori constraints (e.g. flow conservation), exemplarily following the techniques in Petitjean below (and similar ones), where an increasingly complex model is tried in sequence (forward selection). [0231]: An exemplary embodiment in FIG. 14, FIG. 15, and FIG. 16, above shows a Bayesian Network, and its associated (non-convex) polyhedron. In the Figure above y and z are conditionally independent given x. Hence in the y-z plane they satisfy one or more linear box constraint (s) Y min(x, . . . ) <=y <=Y max(x, . . . ) Z min(x, . . . ) <=z <=Z max(x, . . . ) [0237]: The hypercube dimensions are dependent on the separator (conditioning) variable. The boundaries can be analytically characterized, given the rate of change of the dimensions, with the separator variable. Based on this hypercube cross section, instead of using linear programming (LP) , we can use analytical methods to perform set theoretic/information theoretic operations (volume, . . . ). [0395]: FIG. 31, engine based on the computer system presented in FIG. 12 and FIG. 133 Exemplary Operation of IE in Ganaka”, depicts some of the facilities offered by this database, including the modeling module which includes nonlinear transformations to obtain features, algorithms to extract features (SVM, Neural Networks), Linear Classifiers (LP, German Tank), I-structures, multiple data model classes, including graphical models, inference engines beyond linear and integer linear programming , and summaries of the data. A namespace manager is also shown. All interactions between these are not depicted for clarity. The CmdB interacts with embodiments like the constraint manager , It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Pras. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. Pras teaches constrained optimization using integer linear programming. One of ordinary skill would have motivation to combine Jain, Denolf, Nurvi and Pras that can provide improved performance speed, memory I/O using based on ALU and structural inference engine (Pras [0221]). Claim 16 is rejected under 35 U.S.C. 103 unpatentable over Animesh Jain et.al. (hereinafter Jain ) US 11809981 B1, in view of Kristof Denolf et.al. (hereinafter Denolf ) US 2020/0104715 A1. in view of Eriko Nurvitadhi et.al. (hereinafter Nurvi ) US 11216722 B2, further in view of Santiago Pagani et.al. (hereinafter Pagani ) Machine Learning for Power, Energy, and Thermal Management on Multicore Processors: A Survey, IEEE TRANSACTIONS ON COMPUTER-AIDED DESIGN OF INTEGRATED CIRCUITS AND SYSTEMS, VOL. 39, NO. 1, JANUARY 2020. In regard to claim 16: (Currently Amended) Jain, Denolf and Nurvi do not explicitly disclose: - wherein the optimizer is configured to generate candidate sets of hyperparameters that satisfy one or more constraints on a thermal design power of the computer chip. However, Pagani discloses: - wherein the optimizer is configured to generate candidate sets of hyperparameters that satisfy one or more constraints on a thermal design power of the computer chip. [I. Introduction, Page 101]: HIGH power densities and temperatures on many-core systems are the result of ever-increasing transistor integration coupled with the observed limits on voltage scaling for next-generation technology nodes, [I. Introduction, Page 101]: Particularly, if we wish to keep cooling costs constant (by using a common cooling solution for several scaling generations) without violating the chip’s thermal constraints, [Abstract, Page 101]: This paper presents an overview of several research efforts that propose to use machine learning (ML) techniques for power and thermal management on single-core and multicore processors. Traditional power and thermal management techniques rely on a certain a-priori knowledge of the chip’s thermal model, as well as information of the workloads/applications to be executed (e.g., transient and average power consumption ). [I. Introduction, Page 101]: Moreover, to prolong battery lifetime of embedded systems or to cut power bills of servers, energy management for minimizing overall energy consumption while satisfying performance (or real-time) constraints is another relevant (almost dual) problem . [II, Background, Page 102]: 1) Computational Performance: Computational performance refers to how quickly a system can execute an application or a given set of applications. An application’s [II, Background, Page 103]: performance can be measured using multiple metrics, such as execution time, throughput, instructions per cycle (IPC), instructions per second (IPS), speed-up factor (normalized to a known reference), and so on. [II, Background, Page 103]: To aid designing the cooling solution for a certain chip , the common industry practice is to provide the thermal design power (TDP) of a particular chip , defined as “the highest expected sustainable power while running known power intensive real applications” [49]. Hence, given that the system should be able to safely consume TDP power, manu facturers normally recommend to design the cooling solution to dissipate TDP, avoiding the cooling solution from being over-dimensioned. [IV. TECHNIQUES FOR HOMOGENEOUSMULTICORES, Page 111]: a hierarchical power management technique that attempts to deliver maximum the energy efficiency (particularly, to maximize the total executed instructions pe r joule) under a fixed per-chip power budget (e.g., TDP). [III B, Page 108]: The work in [28] presents an online adaptive DPM technique based on model-free reinforcement learning, where the tradeoff between power consumption and latency can be controlled by a user-defined parameter. [III B, Page 107]: Given that most multimedia applications are naturally arranged into frames and the computation load of processing each frame has high temporal correlation, the environment observation and policy adaptation is performed at a single frame gran ularity. The technique learns the temperature change and workload switching patterns by observing the thermal sensor and event counters of the processor, finding a management policy that provides good performance-thermal tradeoff during runtime . (BRI: Perhaps as known to the POSITA particularly in the area of thermal management in VLSI chips , the use of RL to control the power dissipation and TDP may indeed can be framed as a hyperparameter optimization under a policy learning for adaptive control) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Pagani. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. Pagani teaches thermal power management for chips. One of ordinary skill would have motivation to combine Jain, Denolf, Nurvi and Pagani that can provide adaptation to environment changes as the system workload changes [I, Page 101]) Claims 10 and 17-18 are rejected under 35 U.S.C. 103 unpatentable over Animesh Jain et.al. (hereinafter Jain ) US 11809981 B1, in view of Kristof Denolf et.al. (hereinafter Denolf ) US 2020/0104715 A1. in view of Eriko Nurvitadhi et.al. (hereinafter Nurvi ) US 11216722 B2, further in view of David Roberts et.al. (hereinafter Roberts ) US 2020/0174748 A1. In regard to claim 10: (Currently Amended) Jain, Denolf and Nurvi do not explicitly disclose: - wherein the objective function measures one or more of: a latency of the computer chip when performing inference for each of the neural networks, or a throughput of the computer chip when performing inference for each of the neural networks. However, Roberts discloses: - wherein the objective function measures one or more of: a latency of the computer chip when performing inference for each of the neural networks, or a throughput of the computer chip when performing inference for each of the neural networks. [0033] : Processor 202 is a functional block that performs computational operations in electronic device 200. For example, processor 202 may be or include one or more central processing unit (CPU) cores, graphics processing unit (GPU) cores, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc . In some embodiments, processor 202 includes circuit elements or functional blocks such as pipelines, execut ion units, compute units, etc. that execut e program code that causes the circuit elements/functional blocks to perform associated operations (BRI: execution to perform the operations is “inferencing”) [0063]: In the described embodiments, a presorter (e.g., presorter 304) performs operations for determining sorted orders in which instances of input data are processed through a neural network by a neural network processor. In an attempt to ensure that the benefits of using the presorter to determine the sorted orders (e.g., more effective reuse of result values when processing instances of input data through the neural network) are not outweighed by the detriments of using the presorter (e.g ., electrical power consumption, delay, etc.), the described embodiments monitor the efficiency of using the presorter and dynamically adjust operations of the presorter. [0018]: As described above, processing similar instances of input data through a neural network can result in similar results from nodes and outputs from the neural network. For example, when processing neighboring frames of video such as those captured by a security camera or a vehicle's forward-facing camera, there may be only small differences in the images, and thus the results of nodes and output of the neural network for the neighboring frames can be similar. The described embodiments take advantage of the responses of neural networks to similar instances of input data in order to simplify the computation of node results when processing instances of input data. In the described embodiments, when processing instances of input data in a neural network, a neural network processor uses stored previous result values from nodes in the neural network for computing current result values . Because the computations when using the stored previous result values are typically simpler, faster, and require less memory accesses than the computations and memory accesses associated with the normal computation of result values, the described embodiments can process instances of input data through the neural network while consuming less electrical power and with reduced latency . [0075]: In some embodiments, a data structure representative of some or all of the structures and mechanisms described herein (e.g., electronic device 200, neural network processor 206, presorter 304, controller 308, and/or some portion thereof) is stored on a non-transitory computer-readable storage medium that includes a database or other data structure which can be read by an electronic device and used, directly or indirectly, to fabricate hardware including the structures and mechanisms. For example, the data structure may be a behavioral-level description or register-transfer level (RTL) description of the hardware functionality in a high level design language (HDL) such as Verilog or VHDL. The description may be read by a synthesis tool which may synthesize the description to produce a netlist including a list of gates/circuit elements from a synthesis library that represent the functionality of the hardware including the above-described structures and mechanisms. The netlist may then be placed and routed to produce a data set describing geometric shapes to be applied to masks. The masks may then be used in various semiconductor fabrication steps to produce a semiconductor circuit or circuits (e.g., integrated circuits) corresponding to the above-described structures and mechanisms . (BRI: a subsequent chip layout processing can provide the area of the chip) [0072]: In some embodiments, the controller automatically adjusts the one or more parameters that control how the presorter determines subsequent sorted orders in order to seek out a particular balance between an efficiency of the presorter and an efficiency of the neural network processor. In these embodiments, the controller may increase the accuracy or quality of the sorting operations using the above-described adjustments to the difference value threshold, the specified number of sorting passes, and/or the number of available locations in the presort buffer, etc. until the particular balance is reached—and decrease the accuracy or quality if the particular balance is overshot. In other words, the controller may “try” different accuracies or qualities of the sorting operations in an attempt to find a “best” sorting operation for which the efficiencies meet or exceed the particular balance. As described above, the particular balance may be expressed in terms of electrical energy consumed/saved, latency , rate or operations/ throughput, memory accesses, etc. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Roberts. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. Roberts teaches fabricating a accelerator computer chip such as GPU, ASICS, FPGAs. One of ordinary skill would have motivation to combine Jain, Denolf, Nurvi and Roberts that can provide the overall efficiency of the controller in the electronic device using presorter ( Roberts[0028]). In regard to claim 17: (Currently Amended) Jain, Denolf and Nurvi do not explicitly disclose: - further comprising: providing data specifying the determined architecture for use in fabricating a computer chip. However, Roberts discloses: - further comprising: providing data specifying the determined architecture for use in fabricating a computer chip. [ 0024 ]: In the described embodiments, an electronic device that includes a neural network processor performs operations for processing instances of input data through a neural network. As described above, when processing instances of input data through the neural network, the neural network processor stores and reuses node result values for processing the instances of input data, [0019]: Generally, in the described embodiments, results of nodes in a neural network —which function as input values to downstream nodes— are stored/saved and used when processing one or more subsequent instances of input data in order to avoid certain computations and memory accesses . The use of the stored result values is based on the observation that the current result of a node in a neural network, r.sup.2, can be computed by updating/adjusting a previous result of the node, r.sup.1, using associated input values—which, again, are result values from nodes in a previous layer in the neural network. For example, consider a node of a fully-connected layer with three inputs, such as the topmost intermediate node in layer 110 in FIG. 1. When processing a first instance of input data through the neural network, the result r.sup.1 of such a node is computed [ 0048 ]: In describing FIG. 4 and elsewhere herein, “ normal” processing of instances of input data through the neural network is described . Generally, the normal processing of instances of input data involves computing values for internal elements of the neural network using common neural network operations [0048]: “normally” processing an instance of input data through the neural network, input values are first determined for the instance of input data and the input values are provided as inputs to input nodes in the neural network . For instance, when the instance of input data is an image (i.e., for a neural network that performs image recognition), each input node may be provided with values representing the colors of one or more respective pixels from the image. A result of an activation function (e.g., a rectified linear unit (ReLU) function, a hyperbolic tan (tanh) function, a soft-step/logistic function, etc.) for each input node is computed using the respective input values and the results are forwarded to one or more intermediate nodes in the neural network. The weighted input values for intermediate nodes are then computed based on the results from the input nodes and weights associated with corresponding directed edges . The weighted input values are next summed to determine an internal value for each of the intermediate nodes . [0035]: Neural network processor 206 is a functional block that performs operations for processing instances of input data through a neural network and other operations. In some embodiments, neural network processor 206 includes circuit elements and functional blocks that are dedicated to and optimized for performing neural network processing operations—and may not perform general computing operations such as executing program code for an operating system or application program, etc. For example, neural network processor 206 may be included in a neural network accelerator functional block. In some embodiments, neural network processor 206 is or is included in a general-purpose processor such as a CPU or a GPU, which may execute program code to cause the general-purpose processor to perform operations of neural network processor 206. [0075]: In some embodiments, a data structure representative of some or all of the structures and mechanisms described herein (e.g., electronic device 200, neural network processor 206 , presorter 304, controller 308, and/or some portion thereof) is stored on a non-transitory computer-readable storage medium that includes a database or other data structure which can be read by an electronic device and used, directly or indirectly, to fabrica te hardware including the structures and mechanisms . [0020] : By computing the result for the node for processing the second instance of input data using stored results as described, operations such as fetching weight values from memory and mathematical computations can be reduced or avoided. More specifically, continuing the example above, one weight has to be acquired from memory instead of three, the bias is not needed, and three computations are performed instead of six (only one of which is a multiplication operation ). ( BRI: a MAC function ( is a significant part of the datapath) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Roberts. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. Roberts teaches fabricating a accelerator computer chip such as GPU, ASICS, FPGAs. One of ordinary skill would have motivation to combine Jain, Denolf, Nurvi and Roberts that can provide the overall efficiency of the controller in the electronic device using presorter ( Roberts[0028]). In regard to claim 18: (Currently Amended) Jain, Denolf and Nurvi do not explicitly disclose: - f urther comprising: fabricating a computer chip having the determined architecture. [0032]: FIG. 2 presents a block diagram illustrating an electronic device 200 in accordance with some embodiments. As can be seen in FIG. 2, electronic device 200 includes processor 202, memory 204, and neural network processor 206 . [0032]: Generally, processor 202, memory 204, and neural network processor 206 are implemented in hardware, i.e., using various circuit elements and devices. For example, processor 202, memory 204, and neural network processor 206 can be entirely fabricated on one or more semiconductor chip s , including on one or more separate semiconductor chips, can be fashioned from semiconductor chips in combination with discrete circuit elements, can be fabricated from discrete circuit elements alone, etc. As described herein, some or all of processor 202, memory 204, and neural network processor 206 perform operations associated with determining the sorted orders for processing instances of input data. [0035]: neural network processor 206 may be included in a neural network accelerator functional block. In some embodiments, neural network processor 206 is or is included in a general-purpose processor such as a CPU or a GPU, which may execute program code to cause the general-purpose processor to perform operations of neural network processor 206. ( BRI: the neural network processor 206 which can also be a NN accelerator is being fabricated on a chip) [0075]: In some embodiments, a data structure representative of some or all of the structures and mechanisms described herein (e.g., electronic device 200, neural network processor 206, presorter 304, controller 308, and/or some portion thereof) is stored on a non-transitory computer-readable storage medium that includes a database or other data structure which can be read by an electronic device and used, directly or indirectly, to fabricate hardware including the structures and mechanisms. For example, the data structure may be a behavioral-level description or register-transfer level (RTL) description of the hardware functionality in a high level design language (HDL) such as Verilog or VHDL. The description may be read by a synthesis tool which may synthesize the description to produce a netlist including a list of gates/circuit elements from a synthesis library that represent the functionality of the hardware including the above-described structures and mechanisms. The netlist may then be placed and routed to produce a data set describing geometric shapes to be applied to masks. The masks may then be used in various semiconductor fabrication steps to produce a semiconductor circuit or circuits (e.g., integrated circuits) corresponding to the above-described structures and mechanisms . (BRI: a subsequent chip layout processing can provide the area of the chip) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Roberts. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Nurvi teaches simulation and final datapath selection. Roberts teaches fabricating a accelerator computer chip such as GPU, ASICS, FPGAs. One of ordinary skill would have motivation to combine Jain, Denolf, Nurvi and Roberts that can provide the overall efficiency of the controller in the electronic device using presorter ( Roberts[0028]). Claims 12 -14 are rejected under 35 U.S.C. 103 unpatentable over Animesh Jain et.al. (hereinafter Jain ) US 11809981 B1, in view of Kristof Denolf et.al. (hereinafter Denolf ) US 2020/0104715 A1. in view of Eriko Nurvitadhi et.al. (hereinafter Nurvi ) US 11216722 B2, further in view of Sarah Verhulst et.al. (hereinafter Verhuslt) US 2022/0248148 A1. In regard to claim 12: (Original) Jain, Denolf and Nurvi do not explicitly disclose: wherein the candidate set of hyperparameters comprises - (i) one or more hyperparameters that specify a dimensionality of an array of processing elements included in the computer chip from a set of a plurality of possible dimensionalities and - (ii) one or more hyperparameters that specify a respective configuration of each of one or more memory buffers included in the computer chip from a set of a plurality of possible configurations. However, Verhulst discloses: (i) one or more hyperparameters that specify a dimensionality of an array of processing elements included in the computer chip from a set of a plurality of possible dimensionalities and [0119]: Sharing a number of weights in a layer-to-layer mapping is saving memory and computation requirements and allows for an efficient hardware implementation on dedicated accelerator chips. [0069]: In some embodiments, the different set of neural network hyperparameters include one or more of: a different nonlinear transformation applied by the at least one nonlinear unit, a different number of convolutional layers in the encoder and/or decoder, a different number of convolutional filters in any one convolutional layer of the neural network, a different length as the predetermined length for the input sequence, a different configuration of shortcut connections, or optionally a different size of the convolutional filters in any one convolutional layer of the neural network. [0036]: The processing device may be a specially designed processing unit such as an ASIC, or may be a dedicated, energy-efficient machine learning hardware module , for instance a convolution accelerator chip, suitable for portable and embedded applications, e.g. battery-powered applications. The processing device may comprise a systolic array of processing elements for a distributed computation of convolutions via a systolic data flow on the array. Such data flow on the array of processing elements may be row-stationary and the data flow mapping may be flexible , f or instance layer-size dependent. [0104]: a neural network is considered a convolutional neural network if it comprises at least one convolutional layer. A convolutional layer comprises one or more filters, or kernels, which operate on the layer input by convolution, wherein a convolution direction is along one or several dimens ions defined by an input map. [0104]: Performing a convolution operation along a convolution direction generally involves multiple input dimensions or even input batches, and multiple output dimensions (filter depths), and therefore is most often carried out as a generalized tensor convolution between layer input (map) and layer filters to produce a corresponding (nonlinearly activated) layer output (map). - (ii) one or more hyperparameters that specify a respective configuration of each of one or more memory buffers included in the computer chip from a set of a plurality of possible configurations. [0151]: With reference to FIG. 7, an example of a processing device according to some embodiments is now described briefly. The processing device 700, which may be an integrated semiconductor device (e.g. chip), may comprise a control unit 701, a global buffer 702 , an array 703 of processing elements 704 [0151]: The network weight parameters required for performing an inference pass of the neural network implementation may be stored on an external memory unit 706 with large storage capacity (e.g. DRAM, SRAM, non-volatile storage device) and can be transferred to the processing device 700 on request , [0151]: The array 703 may be configured as a systolic array for enabling massively parallelized, distributed computations, e.g. the convolutions between layer inputs and convolutional filters of the layer are parallelized on the systolic array. [0151]: the processing elements 704 may store a neural network weight parameter during many computation cycles so that it is efficiently reused , which leads to less memory access cycles and latency, and also to energy savings that are important for battery-driven, portable devices. Known data flow mappings may be provided to further improve the usage efficiency of the array 703 and further reduce the energy cost, e.g. provide a row-stationary data flow. The accumulated, processed data flows are read out at an edge of the array 703 and are stored in the global buffer 702 , or, if further processing is not required, may be sent to the external memory unit 706 or applied to the connector 705. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Verhulst. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Verhulst teaches processing array dimension and memory configurations. One of ordinary skill would have motivation to combine Jain, Denolf, Nurvi and Verhulst that can provide user efficient array to reduce the energy cost (Verhulst [0151]). In regard to claim 13: (Original) Jain does not explicitly disclose: - wherein the candidate set of hyperparameters further comprises (iii) one or more hyperparameters that specify a configuration of on-chip global memory included in the computer chip from a plurality of possible configurations. However, Denolf discloses: - wherein the candidate set of hyperparameters further comprises (iii) one or more hyperparameters that specify a configuration of on-chip global memory included in the computer chip from a plurality of possible configurations. PNG media_image8.png 631 365 media_image8.png Greyscale [0058]: FIG. 7 is a block diagram depicting a programmable IC 1 according to an example that can be used to implement the inference platform and/or training platform. The programmable IC 1 can be used as the IC 220 in FIG. 2. The programmable IC 1 includes programmable logic 3, configuration logic 25, and config uration memory 26. The programmable IC 1 can be coupled to external circuits, such as nonvolatile memory 27, DRAM 28, and other circuits 29. The programmable logic 3 includes logic cells 30, support circuits 31, and programmable interconnect 32. The logic cells 30 include circuits that can be configured to implement general logic functions of a plurality of inputs. The support circuits 31 include dedicated circuits, such as transceivers, input/output blocks, digital signal processors, memories, and the like. The logic cells and the support circuits 31 can be interconnected using the programmable interconnect 32. Information for programming the logic cells 30, for setting parameters of the support circuits 31, and for programming the programmable interconnect 32 is stored in the config uration memory 26 by the config uration logic 25 . The config uration logic 25 can obtain the config uration data from the nonvolatile memory 27 or any other source (e.g., the DRAM 28 or from the other circuits 29 ). In some examples, the programmable IC 1 includes a processing system 2. The processing system 2 can include microprocessor(s), memory, support circuits, IO circuits, and the like. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). In regard to claim 14: (Currently Amended) Jain does not explicitly disclose: However, Denolf discloses: - a configuration of on-chip global memory included in the computer chip; [0059]: FIG. 8 is a block diagram depicting a System- on-Chip (SoC) implementation of the programmable IC 1 according to an example. In the example, the programmable IC 1 includes the processing system 2 [0059]: The processing system 2 includes various processing units, such as a real-time processing unit (RPU) 4, an application processing unit (APU) 5, a graphics processing unit (GPU) 6, a configuration and security unit (CSU) 12, a platform management unit (PMU) 122, and the like. The processing system 2 also includes various support circuits, such as on-chip memory (OCM) 14. Combine It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, and Denolf. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. One of ordinary skill would have motivation to combine Jain, and Denolf that can optimize for the given cost function (Denolf[0046]). Jain, Denolf and Nurvi do not explicitly disclose: - a dimensionality of an array of processing elements included in the computer chip; - a dimensionality of a systolic array included in the computer chip; - a respective configuration of each of one or more memory buffers included in the computer chip; However, Verhulst discloses: - a dimensionality of an array of processing elements included in the computer chip; [0119]: Sharing a number of weights in a layer-to-layer mapping is saving memory and computation requirements and allows for an efficient hardware implementation on dedicated accelerator chips. [0069]: In some embodiments, the different set of neural network hyperparameters include one or more of: a different nonlinear transformation applied by the at least one nonlinear unit, a different number of convolutional layers in the encoder and/or decoder, a different number of convolutional filters in any one convolutional layer of the neural network, a different length as the predetermined length for the input sequence, a different configuration of shortcut connections, or optionally a different size of the convolutional filters in any one convolutional layer of the neural network. [0036]: The processing device may be a specially designed processing unit such as an ASIC, or may be a dedicated, energy-efficient machine learning hardware module , for instance a convolution accelerator chip, suitable for portable and embedded applications, e.g. battery-powered applications. The processing device may comprise a systolic array of processing elements for a distributed computation of convolutions via a systolic data flow on the array. Such data flow on the array of processing elements may be row-stationary and the data flow mapping may be flexible, for instance layer-size dependent. [0104]: a neural network is considered a convolutional neural network if it comprises at least one convolutional layer. A convolutional layer comprises one or more filters, or kernels, which operate on the layer input by convolution, wherein a convolution direction is along one or several dimens ions defined by an input map. [0104]: Performing a convolution operation along a convolution direction generally involves multiple input dimensions or even input batches, and multiple output dimensions (filter depths), and therefore is most often carried out as a generalized tensor convolution between layer input (map) and layer filters to produce a corresponding (nonlinearly activated) layer output (map). - a dimensionality of a systolic array included in the computer chip; [0119]: Sharing a number of weights in a layer-to-layer mapping is saving memory and computation requirements and allows for an efficient hardware implementation on dedicated accelerator chips. [0069]: In some embodiments, the different set of neural network hyperparameters include one or more of: a different nonlinear transformation applied by the at least one nonlinear unit, a different number of convolutional layers in the encoder and/or decoder, a different number of convolutional filters in any one convolutional layer of the neural network, a different length as the predetermined length for the input sequence, a different configuration of shortcut connections, or optionally a different size of the convolutional filters in any one convolutional layer of the neural network. [0036]: The processing device may be a specially designed processing unit such as an ASIC, or may be a dedicated, energy-efficient machine learning hardware module , for instance a convolution accelerator chip, suitable for portable and embedded applications, e.g. battery-powered applications. The processing device may comprise a systolic array of processing elements for a distributed computation of convolutions via a systolic data flow on the array. Such data flow on the array of processing elements may be row-stationary and the data flow mapping may be flexible , f or instance layer-size dependent. [0104]: a neural network is considered a convolutional neural network if it comprises at least one convolutional layer. A convolutional layer comprises one or more filters, or kernels, which operate on the layer input by convolution, wherein a convolution direction is along one or several dimens ions defined by an input map. [0104]: Performing a convolution operation along a convolution direction generally involves multiple input dimensions or even input batches, and multiple output dimensions (filter depths), and therefore is most often carried out as a generalized tensor convolution between layer input (map) and layer filters to produce a corresponding (nonlinearly activated) layer output (map). - a respective configuration of each of one or more memory buffers included in the computer chip; [0151]: With reference to FIG. 7, an example of a processing device according to some embodiments is now described briefly. The processing device 700, which may be an integrated semiconductor device (e.g. chip), may comprise a control unit 701, a global buffer 702 , an array 703 of processing elements 704 [0151]: The network weight parameters required for performing an inference pass of the neural network implementation may be stored on an external memory unit 706 with large storage capacity (e.g. DRAM, SRAM, non-volatile storage device) and can be transferred to the processing device 700 on request , [0151]: The array 703 may be configured as a systolic array for enabling massively parallelized, distributed computations, e.g. the convolutions between layer inputs and convolutional filters of the layer are parallelized on the systolic array. [0151]: the processing elements 704 may store a neural network weight parameter during many computation cycles so that it is efficiently reused , which leads to less memory access cycles and latency, and also to energy savings that are important for battery-driven, portable devices. Known data flow mappings may be provided to further improve the usage efficiency of the array 703 and further reduce the energy cost, e.g. provide a row-stationary data flow. The accumulated, processed data flows are read out at an edge of the array 703 and are stored in the global buffer 702 , or, if further processing is not required, may be sent to the external memory unit 706 or applied to the connector 705. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Jain, Denolf, Nurvi and Verhulst. Jain teaches hardware accelerator, fusing strategy Denolf teaches datapath and hyperparameters that define datapath. Verhulst teaches processing array dimension and memory configurations. One of ordinary skill would have motivation to combine Jain, and Verhulst that can provide user efficient array to reduce the energy cost (Verhulst [0151]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to TIRUMALE KRISHNASWAMY RAMESH whose telephone number is (571)272-4605. The examiner can normally be reached by phone. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached on phone (571-272-3768). The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TIRUMALE K RAMESH/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121 Application/Control Number: 18/285,578 Page 2 Art Unit: 2121 Application/Control Number: 18/285,578 Page 3 Art Unit: 2121 Application/Control Number: 18/285,578 Page 4 Art Unit: 2121 Application/Control Number: 18/285,578 Page 5 Art Unit: 2121 Application/Control Number: 18/285,578 Page 6 Art Unit: 2121 Application/Control Number: 18/285,578 Page 7 Art Unit: 2121 Application/Control Number: 18/285,578 Page 8 Art Unit: 2121 Application/Control Number: 18/285,578 Page 9 Art Unit: 2121 Application/Control Number: 18/285,578 Page 10 Art Unit: 2121 Application/Control Number: 18/285,578 Page 11 Art Unit: 2121 Application/Control Number: 18/285,578 Page 12 Art Unit: 2121 Application/Control Number: 18/285,578 Page 13 Art Unit: 2121 Application/Control Number: 18/285,578 Page 14 Art Unit: 2121 Application/Control Number: 18/285,578 Page 15 Art Unit: 2121 Application/Control Number: 18/285,578 Page 16 Art Unit: 2121 Application/Control Number: 18/285,578 Page 17 Art Unit: 2121 Application/Control Number: 18/285,578 Page 18 Art Unit: 2121 Application/Control Number: 18/285,578 Page 19 Art Unit: 2121 Application/Control Number: 18/285,578 Page 20 Art Unit: 2121 Application/Control Number: 18/285,578 Page 21 Art Unit: 2121 Application/Control Number: 18/285,578 Page 22 Art Unit: 2121 Application/Control Number: 18/285,578 Page 23 Art Unit: 2121 Application/Control Number: 18/285,578 Page 24 Art Unit: 2121 Application/Control Number: 18/285,578 Page 25 Art Unit: 2121 Application/Control Number: 18/285,578 Page 26 Art Unit: 2121 Application/Control Number: 18/285,578 Page 27 Art Unit: 2121 Application/Control Number: 18/285,578 Page 28 Art Unit: 2121 Application/Control Number: 18/285,578 Page 29 Art Unit: 2121 Application/Control Number: 18/285,578 Page 30 Art Unit: 2121 Application/Control Number: 18/285,578 Page 31 Art Unit: 2121 Application/Control Number: 18/285,578 Page 32 Art Unit: 2121 Application/Control Number: 18/285,578 Page 33 Art Unit: 2121 Application/Control Number: 18/285,578 Page 34 Art Unit: 2121 Application/Control Number: 18/285,578 Page 35 Art Unit: 2121 Application/Control Number: 18/285,578 Page 36 Art Unit: 2121 Application/Control Number: 18/285,578 Page 37 Art Unit: 2121 Application/Control Number: 18/285,578 Page 38 Art Unit: 2121 Application/Control Number: 18/285,578 Page 39 Art Unit: 2121 Application/Control Number: 18/285,578 Page 40 Art Unit: 2121 Application/Control Number: 18/285,578 Page 41 Art Unit: 2121 Application/Control Number: 18/285,578 Page 42 Art Unit: 2121 Application/Control Number: 18/285,578 Page 43 Art Unit: 2121 Application/Control Number: 18/285,578 Page 44 Art Unit: 2121 Application/Control Number: 18/285,578 Page 45 Art Unit: 2121 Application/Control Number: 18/285,578 Page 46 Art Unit: 2121 Application/Control Number: 18/285,578 Page 47 Art Unit: 2121 Application/Control Number: 18/285,578 Page 48 Art Unit: 2121 Application/Control Number: 18/285,578 Page 49 Art Unit: 2121 Application/Control Number: 18/285,578 Page 50 Art Unit: 2121 Application/Control Number: 18/285,578 Page 51 Art Unit: 2121 Application/Control Number: 18/285,578 Page 52 Art Unit: 2121 Application/Control Number: 18/285,578 Page 53 Art Unit: 2121 Application/Control Number: 18/285,578 Page 54 Art Unit: 2121 Application/Control Number: 18/285,578 Page 55 Art Unit: 2121 Application/Control Number: 18/285,578 Page 56 Art Unit: 2121 Application/Control Number: 18/285,578 Page 57 Art Unit: 2121 Application/Control Number: 18/285,578 Page 58 Art Unit: 2121 Application/Control Number: 18/285,578 Page 59 Art Unit: 2121 Application/Control Number: 18/285,578 Page 60 Art Unit: 2121 Application/Control Number: 18/285,578 Page 61 Art Unit: 2121 Application/Control Number: 18/285,578 Page 62 Art Unit: 2121 Application/Control Number: 18/285,578 Page 63 Art Unit: 2121 Application/Control Number: 18/285,578 Page 64 Art Unit: 2121 Application/Control Number: 18/285,578 Page 65 Art Unit: 2121 Application/Control Number: 18/285,578 Page 66 Art Unit: 2121 Application/Control Number: 18/285,578 Page 67 Art Unit: 2121 Application/Control Number: 18/285,578 Page 68 Art Unit: 2121 Application/Control Number: 18/285,578 Page 69 Art Unit: 2121 Application/Control Number: 18/285,578 Page 70 Art Unit: 2121 Application/Control Number: 18/285,578 Page 71 Art Unit: 2121 Application/Control Number: 18/285,578 Page 72 Art Unit: 2121 Application/Control Number: 18/285,578 Page 73 Art Unit: 2121 Application/Control Number: 18/285,578 Page 74 Art Unit: 2121 Application/Control Number: 18/285,578 Page 75 Art Unit: 2121 Application/Control Number: 18/285,578 Page 76 Art Unit: 2121 Application/Control Number: 18/285,578 Page 77 Art Unit: 2121
Read full office action

Prosecution Timeline

Oct 04, 2023
Application Filed
May 19, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699873
NEURAL NETWORK PROCESSING USING MIXED-PRECISION DATA REPRESENTATION
6y 6m to grant Granted Aug 04, 2026
Patent 12688395
Neural Network Processor with On-Chip Convolution Kernel Storage
8y 5m to grant Granted Jul 21, 2026
Patent 12518153
TRAINING MACHINE LEARNING SYSTEMS
5y 12m to grant Granted Jan 06, 2026
Patent 12293284
META COOPERATIVE TRAINING PARADIGMS
4y 4m to grant Granted May 06, 2025
Patent 12229651
BLOCK-BASED INFERENCE METHOD FOR MEMORY-EFFICIENT CONVOLUTIONAL NEURAL NETWORK IMPLEMENTATION AND SYSTEM THEREOF
4y 4m to grant Granted Feb 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
26%
Grant Probability
50%
With Interview (+23.7%)
4y 9m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month