Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or
composition of matter, or any new and useful improvement thereof, may obtain a patent
therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
1. A computer program product for balancing utilization of tiles in an analog in-memory computing system, the computer program product comprising:
one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media,
the program instructions comprising: identifying, by a computer processor, a plurality of tiles in the analog in-memory computing system;
receiving, by the computer processor, a plurality of layers in a neural network being processed by the analog in-memory computing system;
defining, by the computer processor, a number of operations for each layer in the plurality of layers;
assigning a utilization rate to each layer;
determining, by the computer processor, a target equalized utilization rate based on the identified plurality of tiles;
and assigning the layers to the plurality of tiles, by the computer processor, wherein a first utilization rate of a first tile is balanced relative to a second utilization rate of a second tile in the analog in-memory computing system.
by the computer processor
Claim 1 Step 1:
Claim 1 is directed to a UTILIZATION-BASED MAPPING FOR THREE DIMENSIONAL ANALOG IN-MEMORY COMPUTING, the method comprising: a series of steps, and is therefore directed to a process, which is one of the fourstatutory categories.
Claim 1 Step 2A, Prong One:
Limitations 1 b,d,e,f,g can be performed in the human mind through observation, evaluation,
judgement and opinion, with the aid of pen and paper, and is/are therefore reciting a mental process.
Accordingly, claim 1 recites a judicial exception (i.e., an abstract idea).
Claim 1 Step 2A, Prong Two:
Limitations 1 a,h provide mere instructions to implement the limitations which can be
performed in the human mind, i.e., the judicial exception, on a computer, which is not indicative of
integration into a practical application. See MPEP 2106.04(d) and 2106.0S(f).
Limitation 1 c amounts to insignificant extra-solution activity of necessary data outputting, as it
is merely outputting the result of the judicial exception, which is not indicative of integration into a
practical application. See MPEP 2106.04(d) and 2106.0S(g).
Claim 1Step 2B:
The combination of these additional elements amounts to a method comprising steps which can
be performed mentally implemented by generic computing components, and comprising a step of
insignificant extra-solution and well-understood, routine and conventional activity. Therefore, the
additional elements, when considered individually and in combination, fail to add an inventive concept
to the claim. Consequently, claim 1 as a whole does not amount to significantly more than the recited
judicial exceptions and the claim is not eligible.
Claim 2 is dependent on claim 1, and therefore inherits the same judicial exception.
2. The computer program product of claim 1, wherein:
the program instructions further comprise receiving, by the computer processor, a locality constraint;
and assigning the layers to the plurality of tiles is based at least in part on the locality constraint.
by the computer processor
Limitation 2 b recites steps that can performed in the human mind through observation,
evaluation, judgement and opinion, with the aid of pen and paper, and is therefore reciting a mental
process. . In addition, the mention of the steps being performed on a computer in limitation 6 a, amounts to mere instructions to apply the exception for the same reasons presented with respect to claim 1. Limitation 2 a amounts to mere data gathering and outputting, and is therefore
insignificant extra-solution activity. This additional element of insignificant extra-solution activity is not
indicative of integration into a practical application. Even when considered in combination with the additional elements of claim 1, the additional elements comprise mere instructions to apply the
exception and insignificant extra-solution activity, which are not indicative of integration into a practical
application.
Claim 3 is dependent on claim 2, and therefore inherits the same judicial exception.
3. The computer program product of claim 2,
wherein the locality constraint includes using only successive layers in the tiles.
An integration into a practical application is not substantiated from the above. Thus, the claim is not eligible.
Claim 4 is dependent on claim 2, and therefore inherits the same judicial exception.
4. The computer program product of claim 2,
wherein the program instructions further comprise selecting the locality constraint based on reducing latency in an output of the neural network.
An integration into a practical application is not substantiated from the above. Thus, the claim is not eligible.
Claim 5 is dependent on claim 1, and therefore inherits the same judicial exception.
5. The computer program product of claim 1,
wherein: the neural network operates under one or more architecture-specific restraints; and the determination of the equalized utilization rate for the tiles is made under the one or more architecture-specific restraints.
An integration into a practical application is not substantiated from the above. Thus, the claim is not eligible.
Claim 6 is dependent on claim 1, and therefore inherits the same judicial exception.
6. The computer program product of claim 1,
wherein the analog in-memory computing system comprises a plurality of neural network models, and wherein the program instructions further comprise:
receiving an input token of an image for processing by the plurality of neural network models; mapping the plurality of neural network models to a planar space;
dividing the planar space into a plurality of subspaces;
and assigning the plurality of neural network models to the subspaces, wherein an activation rate for the processing of the input token is evenly distributed amongst the subspaces.
Limitation 6 d recites steps that can performed in the human mind through observation,
evaluation, judgement and opinion, with the aid of pen and paper, and is therefore reciting a mental
process. In addition, the mention of the steps being performed on a computer in limitation 6 a, amounts
to mere instructions to apply the exception for the same reasons presented with respect to claim 1.
Furthermore, limitation 6 c,b amounts to mere data gathering and outputting, and is therefore
insignificant extra-solution activity. This additional element of insignificant extra-solution activity is not
indicative of integration into a practical application. Even when considered in combination with the
additional elements of claim 1, the additional elements comprise mere instructions to apply the
exception and insignificant extra-solution activity, which are not indicative of integration into a practical
application. Thus, claim 6 is not eligible.
Claim 7 is dependent on claim 1, and therefore inherits the same judicial exception.
7. The computer program product of claim 1, wherein:
the analog in-memory computing system comprises a first neural network model and a second neural network model;
and the program instructions further comprise: stacking layers of a first tile from the first neural network model with layers of the first tile from the second neural network model into a new tile;
and determining the target equalized utilization rate for the new tile.
Limitation 7 c recites steps that can performed in the human mind through observation,
evaluation, judgement and opinion, with the aid of pen and paper, and is therefore reciting a mental
process. In addition, the mention of the steps being performed on a computer in limitation 7 a, amounts
to mere instructions to apply the exception for the same reasons presented with respect to claim 1.
Furthermore, limitation 7 b amounts to mere data gathering and outputting, and is therefore
insignificant extra-solution activity. This additional element of insignificant extra-solution activity is not
indicative of integration into a practical application. Even when considered in combination with the
additional elements of claim 1, the additional elements comprise mere instructions to apply the
exception and insignificant extra-solution activity, which are not indicative of integration into a practical
application. Thus, claim 7 is not eligible.
Limitations 8 a-g correspond to limitations 1b-h and thus inherit the same judicial exceptions.
8. A computer implemented method for balancing utilization of tiles in an analog in-memory computing system,
comprising: identifying, by a computer processor, a plurality of tiles in the analog in-memory computing system;
receiving, by the computer processor, a plurality of layers in a neural network being processed by the analog in-memory computing system;
defining, by the computer processor, a number of operations for each layer in the plurality of layers;
assigning a utilization rate to each layer;
determining, by the computer processor, a target equalized utilization rate based on the identified plurality of tiles;
and assigning the layers to the plurality of tiles, by the computer processor, wherein a first utilization rate of a first tile is balanced relative to a second utilization rate of a second tile in the analog in-memory computing system.
by the computer processor
Limitations 9 a-c correspond to limitation 2 a-c and thus inherit the same judicial exceptions.
9. The method of claim 8, further comprising:
receiving, by the computer processor, a locality constraint;
and wherein assigning the layers to the plurality of tiles is based at least in part on the locality constraint.
by the computer processor
Limitation 10 a corresponds to limitation 3 a and thus inherits the same judicial exceptions.
10. The method of claim 9,
wherein the locality constraint includes using only successive layers in the tiles.
Limitation 11 a corresponds to limitation 4 a and thus inherits the same judicial exceptions.
11. The method of claim 9,
further comprising selecting the locality constraint based on reducing latency in an output of the neural network.
Limitation 12 a corresponds to limitation 5 a and thus inherits the same judicial exceptions.
12. The method of claim 8, wherein:
the neural network operates under one or more architecture-specific restraints; and the determination of the equalized utilization rate for the tiles is made under the one or more architecture-specific restraints.
Limitations 13 a-d correspond to limitations 6 a-d and thus inherit the same judicial exceptions.
13. The method of claim 8,
wherein the analog in-memory computing system comprises a plurality of neural network models, and wherein the method further comprises:
receiving an input token of an image for processing by the plurality of neural network models; mapping the plurality of neural network models to a planar space;
dividing the planar space into a plurality of subspaces;
and assigning the plurality of neural network models to the subspaces, wherein an activation rate for the processing of the input token is evenly distributed amongst the subspaces.
Limitations 14 a-c correspond to limitations 7 a-c and thus inherit the same judicial exceptions.
14. The method of claim 8, wherein:
the analog in-memory computing system comprises a first neural network model and a second neural network model;
and the method further comprises: stacking layers of a first tile from the first neural network model with layers of the first tile from the second neural network model into a new tile;
and determining the target equalized utilization rate for the new tile.
15. A computing device configured to balance utilization of tiles in an analog in-memory computing system, comprising:
a processor operating an analog in-memory computing engine in the analog in-memory computing system; and a memory coupled to the processor,
Limitations 8 b-h correspond to limitations 1b-h and thus inherit the same judicial exceptions.
the memory storing instructions to cause the processor to perform acts comprising: identifying, by a computer processor, a plurality of tiles in the analog in-memory computing system;
receiving, by the computer processor, a plurality of layers in a neural network being processed by the analog in-memory computing system;
defining, by the computer processor, a number of operations for each layer in the plurality of layers;
assigning a utilization rate to each layer;
determining, by the computer processor, a target equalized utilization rate based on the identified plurality of tiles;
and assigning the layers to the plurality of tiles, by the computer processor, wherein a first utilization rate of a first tile is balanced relative to a second utilization rate of a second tile in the analog in-memory computing system.
by the computer processor,
Limitations 16 a-c correspond to limitation 2 a-c and thus inherit the same judicial exceptions.
16. The computing device of claim 15, wherein the instructions cause the processor to perform further acts comprising:
receiving, by the computer processor, a locality constraint;
and wherein assigning the layers to the plurality of tiles is based at least in part on the locality constraint.
by the computer processor,
Limitation 17 a corresponds to limitation 4 a and thus inherits the same judicial exceptions.
17. The computing device of claim 16,
wherein the instructions cause the processor to perform further acts comprising selecting the locality constraint based on reducing latency in an output of the neural network.
Limitation 18 a corresponds to limitation 5 a and thus inherits the same judicial exceptions.
18. The computing device of claim 15, wherein: t
The neural network operates under one or more architecture-specific restraints; and the determination of the equalized utilization rate for the tiles is made under the one or more architecture-specific restraints.
Limitations 19 a-d correspond to limitations 6 a-d and thus inherit the same judicial exceptions.
19. The computing device of claim 15,
wherein the analog in-memory computing system comprises a plurality of neural network models,
and wherein the memory storing instructions cause the processor to perform acts further comprising: receiving an input token of an image for processing by the plurality of neural network models; mapping the plurality of neural network models to a planar space;
dividing the planar space into a plurality of subspaces;
and assigning the plurality of neural network models to the subspaces, wherein an activation rate for the processing of the input token is evenly distributed amongst the subspaces.
Limitations 20 a-c correspond to limitations 7 a-c and thus inherit the same judicial exceptions.
20. The computing device of claim 15, wherein:
The analog in-memory computing system comprises a first neural network model and a second neural network model; a
and the instructions cause the processor to perform further acts comprising: stacking layers of a first tile from the first neural network model with layers of the first tile from the second neural network model into a new tile;
and determining the target equalized utilization rate for the new tile.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections
set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed
invention is not identically disclosed as set forth in section 102, if the differences between the
claimed invention and the prior art are such that the claimed invention as a whole would have
been obvious before the effective filing date of the claimed invention to a person having
ordinary skill in the art to which the claimed invention pertains. Patentability shall not be
negated by the manner in which the invention was made.
Claim(s) 1,8 is/are rejected under 35 U.S.C. 103 as being unpatentable over 20230074229 A1, Jia et all (Jia hereafter), 2021-02-05. In view of WO2021164752A1, Huawei Technologies Co Ltd (Huawei hereafter), 2021-08-26 and in further view of US 12254399 B2, Intel Corporation (Intel hereafter) 2021-01-27.
Jia teaches the following substantially as claimed.
1. A computer program product for balancing utilization of tiles in an analog in-memory computing system, the computer program product comprising:
one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions comprising:
[0194] It will be appreciated that the functions depicted and described herein may be implemented in hardware or in a combination of software and hardware, e.g., using a general purpose computer, one or more application specific integrated circuits (ASIC), or any other hardware equivalents. In one embodiment, the cooperating process 2405 can be loaded into memory 2404 and executed by processor(s) 2402 to implement the functions as discussed herein. Thus, cooperating process 2405 (including associated data) can be stored on a computer readable storage medium, e.g., RAM memory, magnetic or optical drive or diskette, and the like.
identifying, by a computer processor, a plurality of tiles in the analog in-memory computing system;
[0193] As depicted in FIG. 24, computing device 2400 includes a processor element 2402 (e.g., a central processing unit (CPU) or other suitable processor(s)), a memory 2404 (e.g., random access memory (RAM), read only memory (ROM), and the like), a cooperating module/process 2405, and various input/output devices 2406 (e.g., communications modules, network interface modules, receivers, transmitters and the like).
[0135] FIG. 16 depicts a high-level block diagram of a scalable NN accelerator architecture based on IMC in accordance with some embodiments. Specifically, FIG. 16 depicts a scalable NN accelerator based on IMC wherein integrated microarchitectural supports for application mapping around an IMC bank forms a module that enables architectural scale-up by tiling and interconnection.
receiving, by the computer processor, a plurality of layers in a neural network being processed by the analog in-memory computing system;
[0193] As depicted in FIG. 24, computing device 2400 includes a processor element
[0179] FIG. 22 graphically depicts three stages of mapping software flow to an architecture, illustratively, a NN mapping flow being mapped onto an 8×8 array of CIMUs. FIG. 23A depicts a sample placement of layers from a pipeline segment, and FIG. 23B depicts a sample routing from a pipeline segment.
defining, by the computer processor, a number of operations for each layer in the plurality of layers;
[0193] As depicted in FIG. 24, computing device 2400 includes a processor element
[0103] Regarding throughput, pipelining requires throughput matching across CNN layers. The required operations vary widely across layers, due to both the number of weights and the number of operations per weight. As previously mentioned, IMC intrinsically couples data-storage and compute resources. This provides hardware allocation addressing operation scaling with the number of weights. However, the operations per weight is determined by the number of pixels in the output feature map, which itself varies widely (second row of Table I).
determining (…) based on the identified plurality of tiles;
[0136]. The inventors have determined that benchmark throughput, latency, and energy scale with the number of tiles (throughput/latency should scales proportionally, energy remains substantially constant).
Jia does not teach the following limitation.
assigning a utilization rate to each layer;
Huawei does however.
Summary of the invention, Paragraph 3: In the first aspect, the embodiments of the present application first provide a method for searching neural network channel parameters, which can be used in the field of artificial intelligence. The method includes: First, the training device will obtain a data set, which includes multiple training data and Multiple verification data. After that, the training device will train the initial neural network according to the multiple training data in the data set. The training tasks can be classification, detection, segmentation, etc., and then the trained neural network can be obtained, and the training device can obtain the trained neural network. After that, the use efficiency of any layer in the trained neural network will be determined based on the multiple verification data in the data set. The use efficiency of the computing power is the amount of network performance change caused by the unit computing power. Finally, the training equipment Adjust the neural network channel parameters of the trained neural network according to the utilization efficiency of the computing power, thereby obtaining the first neural network.
Jia does not teach the following limitation.
determining, by the computer processor, a target equalized utilization rate
Huawei does however.
Summary of the invention, Paragraph 26: In a fifth aspect, an embodiment of the present application provides a training device, which may include a memory, a processor, and a bus system.
Summary of the invention, Paragraph 3: After that, the use efficiency of any layer in the trained neural network will be determined based on the multiple verification data in the data set. The use efficiency of the computing power is the amount of network performance change caused by the unit computing power.
Jin does not teach the following limitation.
and assigning the layers to the plurality of tiles,
Intel does however.
(39) In block 502, routine 500 receives, by an analog router 104 of a first supertile 106a of a plurality of supertiles 106a-106n of a NoC 102, a first analog output from a first compute-in-memory tile 108a of a plurality of compute-in-memory tiles 108a-108n of the first supertile 106a. In block 504, routine 500 determines, by the analog router 104 based on a configuration of a neural network executing on the NoC, a destination of the first analog output comprises a second compute-in-memory tile 108b of the plurality of compute-in-memory tiles 108a-108n of the first supertile 106. In block 506, routine 500 generates, by a Gaussian distribution circuit 120 of the analog router 104, a distribution comprising a plurality of weights for the weight of the neural network. The weight may be included in the first analog output received at block 502. The weight may be for a node (or neuron) of the neural network, a layer of the neural network, or any other suitable component of the neural network. In block 508, routine 500 transmits, by the analog router, the first analog output and the distribution to the second compute-in-memory tile without converting the second analog output to the digital domain. Advantageously, doing so reduces the amount of energy used by the NoC 102, as the NoC 102 need not convert the analog output and the distribution from the analog domain to the digital domain, and vice versa.
Jin does not teach the following limitation.
by the computer processor, wherein a first utilization rate of a first tile is balanced relative to a second utilization rate of a second tile in the analog in-memory computing system.
Huawei does however.
Summary of the invention, Paragraph 26: In a fifth aspect, an embodiment of the present application provides a training device, which may include a memory, a processor, and a bus system.
Summary of the invention, Paragraph 3: After that, the use efficiency of any layer in the trained neural network will be determined based on the multiple verification data in the data set. The use efficiency of the computing power is the amount of network performance change caused by the unit computing power.
Intel describes layers being assigned to tiles in via weights of the NN and Huawei provides a determination of a utilization rates. It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to modify Jia’s design to assign determined utilization rates to the layers of NN and to assign layers in consideration of these rates to tiles as it would allow the design to achieve more balanced performance among the tiles.
Limitations 8 a-h correspond to limitations 1 b-i and as such are taught by Jin in view of Huawei and Intel as described above.
8. A computer implemented method for balancing utilization of tiles in an analog in-memory computing system, comprising:
identifying, by a computer processor, a plurality of tiles in the analog in-memory computing system;
receiving, by the computer processor, a plurality of layers in a neural network being processed by the analog in-memory computing system;
defining, by the computer processor, a number of operations for each layer in the plurality of layers;
assigning a utilization rate to each layer;
determining, (…) based on the identified plurality of tiles;
determining, by the computer processor, a target equalized utilization rate
and assigning the layers to the plurality of tiles,
by the computer processor, wherein a first utilization rate of a first tile is balanced relative to a second utilization rate of a second tile in the analog in-memory computing system.
Jia teaches the following substantially as claimed.
15. A computing device configured to balance utilization of tiles in an analog in-memory computing system, comprising:
a processor operating an analog in-memory computing engine in the analog in-memory computing system;
[0003] The present disclosure generally relates to the field of in-memory computing and matrix-vector multiplication.
[0194] It will be appreciated that the functions depicted and described herein may be implemented in hardware or in a combination of software and hardware, e.g., using a general purpose computer, one or more application specific integrated circuits (ASIC), or any other hardware equivalents. In one embodiment, the cooperating process 2405 can be loaded into memory 2404 and executed by processor(s) 2402 to implement the functions as discussed herein. Thus, cooperating process 2405 (including associated data) can be stored on a computer readable storage medium, e.g., RAM memory, magnetic or optical drive or diskette, and the like.
and a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts comprising:
[0194] It will be appreciated that the functions depicted and described herein may be implemented in hardware or in a combination of software and hardware, e.g., using a general purpose computer, one or more application specific integrated circuits (ASIC), or any other hardware equivalents. In one embodiment, the cooperating process 2405 can be loaded into memory 2404 and executed by processor(s) 2402 to implement the functions as discussed herein. Thus, cooperating process 2405 (including associated data) can be stored on a computer readable storage medium, e.g., RAM memory, magnetic or optical drive or diskette, and the like.
Limitations 14 c-j correspond to limitations 1 b-i and as such are taught by Jin in view of Huawei and Intel as described above.
identifying, by a computer processor, a plurality of tiles in the analog in-memory computing system;
receiving, by the computer processor, a plurality of layers in a neural network being processed by the analog in-memory computing system;
defining, by the computer processor, a number of operations for each layer in the plurality of layers;
assigning a utilization rate to each layer;
determining, (…) based on the identified plurality of tiles;
determining, by the computer processor, a target equalized utilization rate
and assigning the layers to the plurality of tiles,
by the computer processor, wherein a first utilization rate of a first tile is balanced relative to a second utilization rate of a second tile in the analog in-memory computing system.
Claim(s) 2,4,9.11,16,17 is/are rejected under 35 U.S.C. 103 as being unpatentable over 20230074229 A1, Jia et all (Jia hereafter), 2021-02-05. In view of WO2021164752A1, Huawei Technologies Co Ltd (Huawei hereafter), 2021-08-26, US 12254399 B2, Intel Corporation (Intel hereafter) 2021-01-27 and in further view of 20240272791 A1, Advanced Micro Devices, Inc. (Micro hereafter) 2023-02-12.
Claim 2 inherits from Claim 1 and therefore inherits the same rejections.
Jin in view of Huawei and Intel does not teach the following limitation.
2. The computer program product of claim 1, wherein:
the program instructions further comprise receiving, by the computer processor, a locality constraint; and assigning the layers to the plurality of tiles is based at least in part on the locality constraint.
Micro does however.
[0025] In these color group implementations, the data layout system and the graph adaptation system generate data layout instructions by mapping data objects involved in a sequence of operations to a common color group. By mapping different data objects involved in a sequence of operations to a common color group, the data layout instructions are generated in a manner that simultaneously satisfies PIM locality considerations and row-buffer locality considerations. Mapping data objects to a common color group satisfies PIM locality considerations, but not necessarily row-buffer locality considerations. Thus, to satisfy row-buffer locality considerations, the data layout instructions are generated to assign different colors of a same color group to the data objects associated with a same PIM operation.
It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to modify Jin’s design in view of Huawei and Intel to utilize locality constraints in assigning layers to tiles as this would confer advantages such as minimizing latency and energy expenditure.
Claim 4 inherits from Claim 2 and therefore inherits the same rejections.
Jin in view of Huawei and Intel does not teach the following limitation.
4. The computer program product of claim 2,
wherein the program instructions further comprise selecting the locality constraint based on reducing latency in an output of the neural network.
Micro does however.
[0117] As a specific example, consider a scenario where X and Y are arrays with data objects that each span four super rows with different colors. To simultaneously optimize the data layout instructions 118 for both PIM locality and row-buffer locality considerations, array X is colored with blue, orange, green, and yellow, whereas array Y is colored with orange, blue, yellow, and green.
Here optimization necessarily entails reducing latency.
It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to modify Jin’s design in view of Huawei and Intel to utilize locality constraints in assigning layers to tiles as this would confer advantages such as minimizing latency and energy expenditure.
Limitation 9 a correspond to limitation 2 a and as such is taught by Jin in view of Huawei and Intel as well as in further view of Micro as described above.
9. The method of claim 8, further comprising:
receiving, by the computer processor, a locality constraint; and wherein assigning the layers to the plurality of tiles is based at least in part on the locality constraint.
Limitation 11 a corresponds to limitation 4 a and as such is taught by Jin in view of Huawei and Intel as well as in further view of Micro as described above.
11. The method of claim 9,
further comprising selecting the locality constraint based on reducing latency in an output of the neural network.
Limitations 16 a corresponds to limitation 2 a and as such is taught by Jin in view of Huawei and Intel as well as in further view of Micro as described above.
16. The computing device of claim 15, wherein the instructions cause the processor to perform further acts comprising:
receiving, by the computer processor, a locality constraint; and wherein assigning the layers to the plurality of tiles is based at least in part on the locality constraint.
Limitations 17 a corresponds to limitation 4 a and as such is taught by Jin in view of Huawei and Intel as well as in further view of Micro as described above.
17. The computing device of claim 16,
wherein the instructions cause the processor to perform further acts comprising selecting the locality constraint based on reducing latency in an output of the neural network.
Claim(s) 3,10 is/are rejected under 35 U.S.C. 103 as being unpatentable over 20230074229 A1, Jia et all (Jia hereafter), 2021-02-05. In view of WO2021164752A1, Huawei Technologies Co Ltd (Huawei hereafter), 2021-08-26, US 12254399 B2, Intel Corporation (Intel hereafter), 2021-01-27 20240272791 A1, Advanced Micro Devices, Inc. (Micro hereafter) 2023-02-12. and in further view of US 20240046099 A, Tata Consultancy Services Limited, (Tata hereafter) 2023-07-18.
Claim 3 inherits from Claim 2 and therefore inherits the same rejections.
Jin in view of Huawei and Intel does not teach the following limitation.
3. The computer program product of claim 2,
wherein the locality constraint
Micro does however.
[0025] the data layout instructions are generated in a manner that simultaneously satisfies PIM locality considerations and row-buffer locality considerations.
Jin in view of Huawei and Intel does not teach the following limitation.
includes using only successive layers in the tiles.
Tata does however.
[0007] Embodiments of the present disclosure present technological improvements as solutions to one or more of the above-mentioned technical problems recognized by the inventors in conventional systems. For example, in one embodiment, a system for jointly pruning and hardware acceleration of pre-trained deep learning models is provided. The processor implemented system is configured by the instructions to receive from a user, a pruning request comprising of (i) a plurality of deep neural network (DNN) models, (ii) a plurality of hardware accelerators comprising of one or more processors, a plurality of target performance indicators comprising of a target accuracy, a target inference latency, a target model size, a target network sparsity, and a target energy, and (iii) a plurality of user options comprising of a first pruning search, and a secondary pruning search. The plurality of DNN models and the plurality of hardware accelerators are transformed into a plurality of pruned hardware accelerated DNN models based on at least one of the user options. The first pruning search option executes a hardware pruning search technique, to perform search on each DNN model and each processor based on at least one of a performance indicator and an optimal pruning ratio. The second pruning search option executes an optimal pruning search technique, to perform search on each layer with corresponding pruning ratio. Further, an optimal layer associated with the pruned hardware accelerated DNN model is identified based on the user option. The layer assignment sequence technique creates a static load distributor by partitioning the optimal layer of the DNN model into a plurality of layer sequences and assigning each layer sequence to corresponding processing element of hardware accelerators.
Micro provides the reference to a locality constraint. Tata provides a technique based on the partition of sequential layers. It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to modify Jin’s design in view of Huawei and Intel to utilize locality constraints in assigning layers to tiles as this would confer advantages such as minimizing latency and energy expenditure.
Limitations 10 a,b corresponds to limitation 3 a,b and as such are taught by Jin in view of Huawei and Intel as well as in further view of Micro and Tata as described above.
10. The method of claim 9,
wherein the locality constraint
includes using only successive layers in the tiles.
Claim(s) 5,12,18 is/are rejected under 35 U.S.C. 103 as being unpatentable over 20230074229 A1, Jia et all (Jia hereafter), 2021-02-05. In view of WO2021164752A1, Huawei Technologies Co Ltd (Huawei hereafter), 2021-08-26, US 12254399 B2, Intel Corporation (Intel hereafter), and in further view of 20200175402 A1, CAMERON et all, (CAMERON hereafter) 2018-11-29.
Claim 5 inherits from Claim 1 and therefore inherits the same rejections.
Jin in view of Huawei and Intel does not teach the following limitation.
5. The computer program product of claim 1, wherein:
the neural network operates under one or more architecture-specific restraints; and the determination of the equalized utilization rate for the tiles is made under the one or more architecture-specific restraints.
Cameron does however.
[0023] Note that predictive modeling can be data intensive. For example, a data preparation phase and a learning (or training) phase can require many sweeps of the same data and many calculations on each individual input parameters. Consider a cross statistics step in an algorithm. Such a step may require that statistics be calculated on every input variable with every target variable.
[0024] Some embodiments described herein may be associated with automatic, in-database predictive modeling. Such modeling may be performed in a Big Data environment to overcome the performance and scalability limitations of modeling within a traditional architecture, such as the limitations described above.
It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to consider for Jin’s design in view of Huawei and Intel, architecture specific constraints in determining tile utilization rates. Doing so would make clear the limitations of the system.
Limitation 12 a corresponds to limitation 5 a and as such is taught by Jin in view of Huawei and Intel as well as in further view of Cameron as described above.
12. The method of claim 8, wherein:
the neural network operates under one or more architecture-specific restraints; and the determination of the equalized utilization rate for the tiles is made under the one or more architecture-specific restraints.
Limitation 18 a corresponds to limitation 5 a and as such is taught by Jin in view of Huawei and Intel as well as in further view of Cameron as described above.
18. The computing device of claim 15, wherein:
the neural network operates under one or more architecture-specific restraints; and the determination of the equalized utilization rate for the tiles is made under the one or more architecture-specific restraints.
Claim(s) 6,13,19 is/are rejected under 35 U.S.C. 103 as being unpatentable over 20230074229 A1, Jia et all (Jia hereafter), 2021-02-05. In view of WO2021164752A1, Huawei Technologies Co Ltd (Huawei hereafter), 2021-08-26, US 12254399 B2, Intel Corporation (Intel hereafter), and in further view of 20220004805 A1, SAMSUNG ELECTRONICS CO., LTD. (Samsung hereafter), WO 2023137922 A1, Zheng et all. (Zheng hereafter), 2023-07-27, 2021-05-24. And CN 114020459 A, GAO, 2021-10-30.
Claim 6 inherits from Claim 1 and therefore inherits the same rejections.
Jin in view of Huawei and Intel does not teach the following limitation.
6. The computer program product of claim 1, wherein the analog in-memory computing system comprises
a plurality of neural network models, and wherein the program instructions further comprise:
Samsung does however.
[0006] Provided is an electronic device, which includes a plurality of recognition models, and which is capable of recognizing an object with high accuracy by using computing resources of the electronic device and through object recognition that is performed by logically dividing a physical space and allocating a recognition model corresponding to spatial characteristics to each space. Also provided is a method of operating the electronic device.
Jin in view of Huawei and Intel does not teach the following limitation.
receiving an input token of an image for processing by the plurality of neural network models;
Samsung does however.
[0058] As used herein, the term “recognition model” may refer to, but is not limited to, an artificial intelligence model including one or more neural networks, which are trained to receive an image of an object as input data and obtain object information by performing object recognition on one or more objects in the image.
Jin in view of Huawei and Intel does not teach the following limitation.
mapping the plurality of neural network models
Samsung does however.
[0008] According to an embodiment of the disclosure, there is provided a method, performed by an electronic device, of performing object recognition, the method including obtaining a spatial map of a space, using a first recognition model, recognizing one or more objects in the space, to obtain first object information of the objects, and dividing the space into a plurality of subset spaces, based on the obtained spatial map and the obtained first object information. The method further includes determining at least one second recognition model to be allocated to each of the plurality of subset spaces into which the space is divided, based on characteristic information of each of the plurality of subset spaces, and using the determined at least one second recognition model allocated to each of the plurality of subset spaces, performing object recognition on each of the plurality of subset spaces, to obtain second object information.
Jin in view of Huawei and Intel does not teach the following limitation.
to a planar space;
Zheng does however,
In step S220 of some embodiments, the time-domain signal and the frequency-domain signal are combined into a two-dimensional space, that is, a planar space.
Jin in view of Huawei and Intel does not teach the following limitation.
dividing the planar space into a plurality of subspaces;
Samsung does however.
[0008] and dividing the space into a plurality of subset spaces,. The method further includes determining at least one second recognition model to be allocated to each of the plurality of subset spaces into which the space is divided,
Jin in view of Huawei and Intel does not teach the following limitation.
and assigning the plurality of neural network models to the subspaces,
Samsung does however.
[0008] and dividing the space into a plurality of subset spaces. The method further includes determining at least one second recognition model to be allocated to each of the plurality of subset spaces into which the space is divided,
Jin in view of Huawei and Intel does not teach the following limitation.
wherein an activation rate for the processing of the input token is evenly distributed amongst the subspaces.
GAO does however.
Contents of the Invention, Paragraph 7, by using the technical solution, obtaining the sending target rate of the token bucket of the token bucket by using RAM, according to the different sending target rate requirement, classifying the token bucket, and respectively the number of the token bucket to be used by the RAM and the register, so as to realize the token bucket, combining the advantages of the register and RAM in the FPGA, it not only can realize higher packet sending rate, but also can balance the use of each resource in the FPGA.
Samsung provides the plurality of network models, receiving image input, mapping the models to a space, dividing these into a plurality of subspaces, and the assignment of the models to the subspaces. Zheng provides the spaces being planar. Gao provides a balancing of resources based on packet sending rate (or an activation rate) for tokens. It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to allow Jin’s design in view of Huawei and Intel to accommodate a plurality of models, to receive input tokens, and to assign said models to a division of planar subspaces in consideration of an activation rate . These elements would allow the design to achieve more balanced performance, helping to prevent idling in the system.
Limitations 13 a-g correspond to limitation 6 a-g and as such are taught by Jin in view of Huawei and Intel as well as in further view of Samsung and Zheng as described above.
13. The method of claim 8, wherein the analog in-memory computing system comprises
a plurality of neural network models, and wherein the method further comprises:
receiving an input token of an image for processing by the plurality of neural network models;
mapping the plurality of neural network models
to a planar space;
dividing the planar space into a plurality of subspaces;
and assigning the plurality of neural network models to the subspaces,
wherein an activation rate for the processing of the input token is evenly distributed amongst the subspaces.
Limitations 19 a-g correspond to limitation 6 a-g and as such are taught by Jin in view of Huawei and Intel as well as in further view of Samsung and Zheng as described above.
19. The computing device of claim 15, wherein the analog in-memory computing system comprises a
plurality of neural network models, and wherein the memory storing instructions cause the processor to perform acts further comprising:
receiving an input token of an image for processing by the plurality of neural network models;
mapping the plurality of neural network models
to a planar space;
dividing the planar space into a plurality of subspaces;
and assigning the plurality of neural network models to the subspaces,
wherein an activation rate for the processing of the input token is evenly distributed amongst the subspaces.
Claim(s) 7,14,20 is/are rejected under 35 U.S.C. 103 as being unpatentable over 20230074229 A1, Jia et all (Jia hereafter), 2021-02-05. In view of WO2021164752A1, Huawei Technologies Co Ltd (Huawei hereafter), 2021-08-26, US 12254399 B2, Intel Corporation (Intel hereafter), and in further view of 20220004805 A1, SAMSUNG ELECTRONICS CO., LTD. (Samsung hereafter), 2021-05-24. and US 20080065835 A1, Sun Microsystems, Inc. (Sun hereafter) 2006-09-11
Claim 7 inherits from Claim 1 and therefore inherits the same rejections.
Jin in view of Huawei and Intel teaches the following:
7. The computer program product of claim 1, wherein:
d. and determining the target equalized utilization rate for the new tile.
Summary of the invention, Paragraph 3: After that, the use efficiency of any layer in the trained neural network will be determined based on the multiple verification data in the data set. The use efficiency of the computing power is the amount of network performance change caused by the unit computing power.
Jin in view of Huawei and Intel does not teach the following limitation.
the analog in-memory computing system comprises a first neural network model and a second neural network model;
Samsung does however
[0006] Provided is an electronic device, which includes a plurality of recognition models,
Jin in view of Huawei and Intel does not teach the following limitation.
and the program instructions further comprise: (…) from the first neural network model (…) from the second neural network
Samsung does however.
[0006] Provided is an electronic device, which includes a plurality of recognition models,
Jin in view of Huawei and Intel does not teach the following limitation.
and the program instructions further comprise: stacking layers (…) with layers (…) into a new tile;
Sun does however.
[0006] It has been discovered that at least a portion of the functionality for implementing data coherence in a shared-cache cluster can be offloaded from primary processing units executing instantiated code that performs the functionality to a data coherence offload engine. The offloading reduces the burden on a node's primary processing unit(s) and frees resources typically consumed by performing the operations for data coherence (e.g., accessing memory, generating messages, examining messages, etc.). Offloading also allows the overhead of protocols at upper layers of a protocol stack to either be shifted off of the primary processing unit(s) to the offload engine or to be at least partially avoided. In addition, offloading may reduce cost of licensing software, since a core is not being used to execute code that performs the functionality being offloaded to the data coherence offload engine. At a requestor node, a data coherence offload engine handles a data request initiated by an instantiated code executed on the requestor node's primary processing unit(s). The data coherence offload engine may locally determine a current owner node for a requested data unit, query another node for the current owner node, etc. Assuming the request data unit is available, a current owner node's data coherence offload engine supplies the requested data unit to the requester node. After receiving the requested data unit, the requestor node's data coherence offload engine writes the data unit to memory and indicates availability of the data unit to the instantiated code.
[0015] Example offload engines may comprise one or more of an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device, etc
Samsung provides the plurality of NN models, Sun provides the stacking of layers from processing units into an offload engine which can be implemented as application specific integrated circuit as one example. Tiles are described in this application’s specification to refer to non-volatile memory cells in a two-dimensional or three-dimensional array that include transistors or other circuit devices that control the reading and writing of the non-volatile memory cells. One of ordinary skill in the art would appreciate that both the processing units and the off-load engine could represent tiles based off of this description.
It would have been obvious to one of ordinary skill in the art at the time of this application’s filing to modify Jin’s design in view of Huawei and Intel to stack layers from a plurality of models into a new tile as it would allow the design to achieve more balanced performance among the tiles.
Limitations 14 a-d correspond to limitation 7 a-d and as such are taught by Jin in view of Huawei and Intel as well as in further view of Samsung and Zheng as described above.
14. The method of claim 8, wherein:
the analog in-memory computing system comprises a first neural network model and a second neural network model;
and the program instructions further comprise: (…) from the first neural network model (…) from the second neural network
and the program instructions further comprise: stacking layers (…) with layers (…) into a new tile;
and determining the target equalized utilization rate for the new tile.
Limitations 20 a-d correspond to limitation 7 a-d and as such are taught by Jin in view of Huawei and Intel as well as in further view of Samsung and Zheng as described above.
20. The computing device of claim 15, wherein:
the analog in-memory computing system comprises a first neural network model and a second neural network model;
and the program instructions further comprise: (…) from the first neural network model (…) from the second neural network
and the program instructions further comprise: stacking layers (…) with layers (…) into a new tile;
and determining the target equalized utilization rate for the new tile.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner
should be directed to Luke Absher whose telephone number is (571) 270-1057. The examiner can
normally be reached M-F: 8:00 am - 4:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the US PTO Automated Interview Request (AIR) at http:/ /www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kevin Young can be reached at 571-270-3180.
The fax phone number for the organization where this application or proceeding is assigned is
571-273-8300. Information regarding the status of published or unpublished applications may be
obtained from Patent Center. Unpublished application information in Patent Center is available to
registered users. To file and manage patent submissions in Patent Center, visit: https:/
/patentcenter.uspto.gov. Visit https:/ /www.uspto.gov/patents/apply/patent-center for more
information about Patent Center and https://www.uspto.gov/patents/docx for information about filing
in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197
(toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786
9199 (IN USA OR CANADA) or 571-272-1000.
/LUCAS DONALD ABSHER/
Examiner, Art Unit 2194
/KEVIN L YOUNG/Supervisory Patent Examiner, Art Unit 2194