DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This Office Action has been issued in response to amendments filed 26 June 2026.
Claims 1, 5 – 11, 13 – 19 and 21 – 25 are pending.
Claim Objections
Claim 21 is objected to because of the following informalities. Appropriate correction is required.
Claim 21 should be amended to i) “the multiplexer is configured to select, at a first time, all blocks of the output buffer to output resultants of vector-matrix multiplications for multiple matrices , at a second time, a subset of blocks of the output buffer to output a resultant of a vector-matrix multiplication for a single matrix, wherein the subset of the blocks is less than all of the blocks see spec Fig. 6 and corresponding paragraphs). In addition, this is so that output buffer output less than all (and not all) of its blocks for single matrix multiplication (see spec Fig. 6 and corresponding paragraphs). There is also no support for outputting all of said output buffer’s blocks for single matrix multiplication. In addition, outputting all of said output buffer’s blocks would also result in erroneous result being outputted from blocks of said output buffer that are not part of single matrix multiplication.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6 – 7, 17 and 22 – 25 are rejected under 35 U.S.C. 103 as being unpatentable over Di Febbo (US 20230059200) in view of Deshpande (US 20230022347) and Huang (US 12007937).
Regarding claim 1, Di Febbo teaches
A hardware compute-in-memory (CIM) module (CIM module = Fig. 4 in-memory compute circuit 101), comprising:
storage sites;
compute logic (compute logic = multiple transconductance + multiple ADCs) coupled with the storage sites for performing, in parallel, operations on data (data = weight) stored in the storage sites; (Di Febbo teaches memory cells and multiple transconductance (compute logic), and input values are fed (in parallel (parallel) to said memory cells) to be (performing) multiplied (operations) with weight (data) (stored in each of said memory cells) using respective transconductance in (coupled to) each of said memory cells (see Fig.7, ¶[65-67]). Di Febbo also teaches said memory cells are in sets of rows (storage sites) that is part of memory circuit 120 in in-memory compute circuit 101 (see Fig.4, ¶[23-24]). Di Febbo also teaches said multiplications can also occur concurrently (parallel) (see ¶[86]).)
wherein the CIM hardware module is configured to store weights in blocks of the storage sites, to utilize the blocks and portions of the compute logic corresponding to the blocks to selectively provide outputs of the operations for the weights stored in the blocks, and (Di Febbo teaches in-memory compute circuit 101 (CIM hardware) is configured to receive weights (w00 to w117) (weights) to be stored (store) in memory cells (blocks) for at least a portion of said sets of rows (storage sites) (see Fig. 4, ¶[37]). Di Febbo also teaches said memory cells, each respectively (portions corresponding to the blocks), has transconductance (compute logic) that multiplies (operations) input value with weight (stored in said respective memory cell) (see Fig. 7, ¶[65-67]) to generate a product (outputs) that is accumulated (read) (see ¶[43]) where there are plural input values (see Fig.4 , ¶[43]). Note that each memory cell provides a product that is multiplication of input and weight specific (selective) to each memory cell.)
As noted in claim 1, Di Febbo teaches memory cells (blocks), each with transconductance (compute logic) that multiplies input value to generate a product (outputs) but does not appear to explicitly teach using demultiplexer and multiplexer to route said input value and said product in the following manner.
an output buffer;
a demultiplexer [coupled to the input buffer];
a multiplexer coupled to the output buffer
selectively read, using the multiplexer, the outputs of the operations corresponding to the blocks at different times,
wherein the demultiplexer is configured for routing an input to a portion of the compute logic corresponding to a block of the blocks, and
wherein the multiplexer is configured to select an output corresponding to the block
However, Deshpande teaches
an output buffer;
a demultiplexer [coupled to the input buffer];
a multiplexer (multiplexer = Fig. 10 switches 1041 and 1042) coupled to the output buffer (Deshpande teaches switches 1041 and 1042 coupled to inverters (output buffer) that passes (buffer) RBLs to adder tree (see Fig. 10).)
selectively read, using the multiplexer, the outputs of the operations corresponding to [the] blocks (blocks = Fig. 10 sub-arrays 701 and 702) at different times, (Deshpande teaches i) switch 1041 is used to couple first RBLs (outputs) in sub-array 701 to adder tree, and ii) switch 1042 is used to couple second RBLs (outputs) in sub-array 702 to said adder tree (see Fig. 10, ¶[61]). Deshpande also teaches only switch 1041 or switch 1042 is on at a time (different times) (see ¶[62]). As such, when (selectively) switch 1041 is on, said first RBLs are coupled (read) to said adder tree and when (selectively) switch 1042 is on, said second RBLs are coupled (read) to said adder tree. Deshpande further teaches that said first and second RBLs are products (operations) of bit cells (see ¶[28]) storing weights (see ¶[7]) in (corresponding) subarrays 701 and 702 (blocks) (see ¶[58]).)
wherein the demultiplexer is configured for routing an input to [a portion of the compute logic corresponding to] a block of [the] blocks, and (Deshpande teaches multiplexer 854 (demultiplexer) routing input activation (input) to sub-array 701 (block) of sub-arrays 701 and 702 (blocks) with weights (see Fig. 9-10, ¶[58], [60-61]).)
wherein the multiplexer is configured to select an output corresponding to the block (Deshpande teaches address signal CIM_ADDR<2>, used to route said input activation to said sub-array 701 (block) (see ¶[58]), is also used to control switch 1041 (multiplexer) to couple (select) MULTB signals (output) from (corresponding) said sub-array 701 to adder tree (see Fig. 10, ¶[61-62]) that accumulate products to generate a MAC value (see ¶[28]).)
In view of Deshpande, Di Febbo is modified such that said memory cells (blocks) are configured as sub-arrays where said input value (used to generate said product (outputs)) of each transconductance (compute logic), in each of said memory cells, would be routed to respective sub-array of said sub-arrays using multiplexer (demultiplexer), and switches (multiplexer) (connected to inverters (output buffer)) would be used to couple said product from said respective sub-array to adder tree where only one of said switches is on at a time (different times) resulting in a subset (selective) of said products being coupled to said adder tree at a time.
Di Febbo and Deshpande are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Di Febbo in the manner described supra because configuring cells into sub-arrays (and routing inputs/outputs to said sub-arrays) would reduce device loading and reduce power consumption (Deshpande, ¶[57]).
A noted in claim 1, modified Di Febbo teaches a base CIM module with multiplexer (demultiplexer) that routes input to compute logic. The claimed invention improves upon said base module by using coupling said multiplexer to an input buffer.
This improvement to said base module is an application of known technique from Huang – multiplexer (demultiplexer) provides output (input) from specific storage elements, in buffer 104a (input buffer) (connected (coupled to) to said multiplexer), to processing element (compute logic) (see Huang Fig. 2A, col 11 ln 6-18).
One of ordinary skill in the art would recognize that this known technique of multiplexer using buffer to pass input to processing element can also be applied to modified Di Febbo’s multiplexer that routes said input to said compute logic, and the result would have been predictable. In this instance, said multiplexer would route said input from specific storage elements, in buffer 104 (connected to said multiplexer), to said compute logic. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Huang’s known technique would have yielded i) predictable result of said multiplexer would route said input from buffer 104 (input buffer) (connected to said multiplexer (coupled to)) to said compute logic, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Regarding claim 6, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Di Febbo also teaches
wherein the hardware CIM is configured to store the weights in the blocks based on an optimization of throughput and utilization of the CIM hardware module (Di Febbo teaches storing weights (weights) in memory cells (blocks) (see ¶[37]). Di Febbo teaches to reduce time (throughput) for processing (utilization) data in memory buffer circuit, said stored weight remains constant (see ¶[58]) wherein in-memory compute circuit 101 (CIM hardware module) retrieve said data from said memory buffer circuit (see ¶[33]).)
Regarding claim 7, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Di Febbo also teaches
wherein the operations comprise vector-matrix multiplication operations (Di Febbo teaches memory cells each includes transconductance that multiples (operations) input value with weight (see Fig. 7, ¶[66-67]) wherein said input values originate from respective row (vector) of memory ranges 265a-c (see Fig. 2, ¶[38]), and said respective row originates from a matrix (matrix) (see Fig. 5).)
Regarding claim 17, Di Febbo teaches
A method, comprising:
performing in parallel a first plurality of operations on a first set of weights (first sets of weights = Fig. 7 w00+w01) stored in a first block (first block = Fig. 7 memory cells 727aa+727ab) of a plurality of blocks (plurality of blocks = Fig. 7 memory cells 727aa-727bb) of storage sites of a hardware compute-in-memory (CIM) module (CIM hardware module = Fig. 4 in-memory compute circuit 101), the CIM hardware module including compute logic coupled with the storage sites, the compute logic configured to selectively perform in parallel, operations for the plurality of blocks and provide outputs for the operations, the operations including the first plurality of operations;
outputting, from the CIM hardware module, a first output for the first plurality of operations corresponding to the first block [at a first time];
performing in parallel a second plurality of operations on a second set of weights (second set of weights = Fig. 7 w10+w11) stored in a second block (second block = Fig. 7 memory cells 727ba+memory cell 727bb) of the plurality of blocks, the operations including the second plurality of operations; and
outputting, from the CIM hardware module, a second output for the second plurality of operations corresponding to the second block [at a second time] (Di Febbo teaches memory cells 727aa-727bb (first and second blocks), respectively, has (coupled with) transconductance (compute logic) that multiplies (operations, first and second plurality of operations) respective input value with respective weight stored in respective memory cell (see Fig. 7, ¶[66-67]). Di Febbo also teaches i) multiplications (first plurality of operations) in memory cell 727aa-727ab is accumulated to generate (provide) accumulated voltage 775a (first output) and ii) multiplications (second plurality of operations) in memory cell 727ba-727b is accumulated to generate (provide) accumulated voltage 775b (second output). Note the parallel (parallel) arrangement of i) memory cell 727aa and 727ab in which each multiplies (first plurality of operations) respective input value with respective weight stored in said respective memory cell and ii) memory cell 727ba and 727bb in which each multiplies (second plurality of operations) respective input value with respective weight stored in said respective memory cell, which results in parallel multiplication. Also note that each memory cell performs a multiplication (operations, first and second plurality of operations) for input and weight specific to (selectively) each memory cell (see Fig. 7, ¶[66-67]). Di Febbo also teaches memory cells 727aa-727bb (plurality of blocks) are in sets of rows (storage sites) that is part of memory circuit 120 within (including) in-memory compute circuit 101 (CIM hardware module) (see Fig.4, ¶[23-24]).)
As noted in claim 17, Di Febbo teaches memory cells (first and second blocks), each with transconductance (compute logic) that multiplies input value (first and second inputs) to generate a product (first and second outputs) but does not appear to explicitly teach using demultiplexer and multiplexer to route said input values and said products in the following manner.
an output buffer;
a demultiplexer [coupled to the input buffer];
a multiplexer coupled to the output buffer
routing, using the demultiplexer, a first input to a first portion of the compute logic corresponding to the first block
selectively outputting, using the multiplexer, from the CIM hardware module, a first output for the first plurality of operations corresponding to the first block at a first time
routing, using the demultiplexer, a second input to a second portion of the compute logic corresponding to the second block and
selectively outputting, using the multiplexer, from the CIM hardware module, a second output for the second plurality of operations corresponding to the second block at a second time different from the first time
However, Deshpande teaches
an output buffer;
a demultiplexer [coupled to the input buffer];
a multiplexer (multiplexer = Fig. 10 switches 1041 and 1042) coupled to the output buffer (Deshpande teaches switches 1041 and 1042 coupled to inverters (output buffer) that passes (buffer) RBLs to adder tree (see Fig. 10).)
selectively outputting, using the multiplexer, from the CIM hardware module, a first output for the first plurality of operations corresponding to the first block (first block = Fig. 10 sub-array 701) at a first time
selectively outputting, using the multiplexer, from the CIM hardware module, a second output for the second plurality of operations corresponding to the second block at a second time different from the first time (blocks = Fig. 10 sub-array 702) (Deshpande teaches i) switch 1041 is used to couple first RBLs (first output) in sub-array 701 to adder tree, and ii) switch 1042 is used to couple second RBLs (second output) in sub-array 702 to said adder tree (see Fig. 10, ¶[61]). Deshpande also teaches only switch 1041 or switch 1042 is on at a time (first time, second time different from first time) (see ¶[62]). As such, when (selectively) switch 1041 is on (first time), said first RBLs are coupled (read) to said adder tree, and when (selectively) switch 1042 is on (second time), said second RBLs are coupled (output) to said adder tree. Deshpande further teaches that said first and second RBLs are products (operations) of bit cells (see ¶[28]) storing weights (see ¶[7]) in (corresponding) subarrays 701 and 702 (blocks) (see ¶[58]).) (Deshpande teaches address signal CIM_ADDR<2>, used to route said input activation i) to said sub-array 701 (first block) (see ¶[58]), is also used to control switch 1041 (multiplexer) to couple (select) MULTB signals (first output) from (corresponding) said sub-array 701 to adder tree, and ii) i) to said sub-array 702 (second block) (see ¶[58]), is also used to control switch 1042 (multiplexer) to couple (select) MULTB signals (second output) from (corresponding) said sub-array 702 to adder tree (see Fig. 10, ¶[61-62]).)
routing, using the demultiplexer, a first input to [a first portion of the compute logic corresponding to] the first block;
routing, using the demultiplexer, a second input to [a second portion of the compute logic corresponding to] the second block; (Deshpande teaches multiplexer 854 (demultiplexer) routing input activation (first input, second input) to one (first block, second block) of sub-arrays 701 and 702 (blocks) with weights (see Fig. 9-10, ¶[58], [60-61]).)
In view of Deshpande, Di Febbo is modified such that said memory cells (blocks) are configured as sub-arrays where said input value (first and second inputs) (used to generate said product (first and second outputs)) of each transconductance (compute logic), in each of said memory cells, would be routed to respective sub-array of said sub-arrays using multiplexer (demultiplexer), and switches (multiplexer) (connected to inverters (output buffer)) would be used to couple said product from said respective sub-array to adder tree where only one of said switches is on at a time (different times) resulting in a subset (selective) of said products being coupled to said adder tree at a time.
Di Febbo and Deshpande are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Di Febbo in the manner described supra because configuring cells into sub-arrays (and routing inputs/outputs to said sub-arrays) would reduce device loading and reduce power consumption (Deshpande, ¶[57]).
A noted in claim 17, modified Di Febbo teaches a base method where CIM module includes multiplexer (demultiplexer) that routes input to compute logic. The claimed invention improves upon said base method by coupling said multiplexer to an input buffer.
This improvement to said base method is an application of known technique from Huang – multiplexer (demultiplexer) provides output (input) from specific storage elements, in buffer 104a (input buffer) (connected (coupled to) to said multiplexer), to processing element (compute logic) (see Huang Fig. 2A, col 11 ln 6-18).
One of ordinary skill in the art would recognize that this known technique of multiplexer using buffer to pass input to processing element can also be applied to modified Di Febbo’s multiplexer that routes said input to said compute logic, and the result would have been predictable. In this instance, said multiplexer would route said input from specific storage elements, in buffer 104 (connected to said multiplexer), to said compute logic. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Huang’s known technique would have yielded i) predictable result of said multiplexer would route said input from buffer 104 (input buffer) (connected to said multiplexer (coupled to)) to said compute logic, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Regarding claim 22, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1.
A noted in claim 1, modified Di Febbo teaches a base CIM module with multiplexer (demultiplexer) that routes input to compute logic. The claimed invention improves upon said base module by using said multiplexer to select blocks, for vector-matrix multiplications, in the following manner.
wherein the demultiplexer is configured to select which blocks of the input buffer are loaded with an input vector for vector-matrix multiplications
This improvement to said base module is an application of known technique from Huang – multiplexer selecting specific storage elements for product of two matrices. In particular, Huang teaches
wherein the demultiplexer is configured to select which blocks of the input buffer are loaded with an input vector for vector-matrix multiplications (Huang teaches multiplexer (demultiplexer) provides (select) specific (which) storage elements (blocks), in a row (vector) of input buffer 104a (input buffer), from (loaded) which output (input vector) is provided to (for) processing element (see Huang Fig. 2A, col 11 ln 6-18) that provides product (multiplication) of two matrices (vector-matrix) (see Huang col9 ln 61 – col 10 ln 10) which is dot product of vectors (vector) (see Huang Fig. 10, col 8 ln 35-42).)
One of ordinary skill in the art would recognize that this known technique of multiplexer selecting specific storage elements in which to provide input to processing element can also be applied to modified Di Febbo’s multiplexer that routes said input to said compute logic, and the result would have been predictable. In this instance, said multiplexer would provide specific storage elements in a row of input buffer 104a, from which output is provided to processing element that performs product of two matrices. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Huang’s known technique would have yielded i) predictable result of said multiplexer providing (select) specific storage elements (blocks) in a row (vector) of input buffer 104a (input buffer), from which output (input vector) is provided to (for) processing element that performs product (multiplication) of two matrices (vector-matrix), and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Regarding claim 23, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Deshpande also teaches
wherein time multiplexing is used to perform vector-matrix multiplications for a first set of matrices stored in a first set of blocks at a first time interval and vector-matrix multiplications for a second set of matrices stored in a second set of blocks at a second time interval (Deshpande teaches sub-arrays 701 (first set of blocks) and 702 (second set of blocks) loaded, in two-dimensional matrix (matrix), with weights (first and second set of matrices) where i) input is route to one (time multiplexing) of said sub-arrays 701 and 702 (see Fig. 10, ¶[58], [60-61]), and ii) said input, for a row (vector), is multiplied (multiplication) with said weights stored in said two-dimensional matrix (matrix) (see ¶[38]). As such, when (first interval) said input is routed to said sub-array 701 (first set of blocks), said input is multiplied with subset (first set of matrices) of said weights, and when (second interval) said input is routed to said sub-array 702 (second set of blocks), said input is multiplied with subset (second set of matrices) of said weights.)
In view of Deshpande, modified Di Febbo is further modified such that memory cells are configured into sub-arrays 701 and 702 where when (first interval) row (vector) input is routed to said sub-array 701 (first set of blocks), said row input is multiplied (multiplications) with two dimensional weights (first set of matrices) within said sub-array 701, and when (second interval) said row (vector) input is routed to said sub-array 702 (second set of blocks), said row input is multiplied (multiplications) with two dimensional weights (first set of matrices) within said sub-array 702.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Di Febbo in the manner described supra because configuring cells into sub-arrays (and routing inputs/outputs to said sub-arrays) would reduce device loading and reduce power consumption (Deshpande, ¶[57]).
Regarding claim 24, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Deshpande already teaches in claim 1
wherein the multiplexer is configured to provide the output [to a general-purpose (GP) processor for an activation function,] to another compute tile, [or to a memory] (Deshpande teaches address signal CIM_ADDR<2>, used to route said input activation to said sub-array 701 (block) (see ¶[58]), is also used to control switch 1041 (multiplexer) to couple (select) MULTB signals (output) from (corresponding) said sub-array 701 to adder tree (compute tile) (see Fig. 10, ¶[61-62]) that accumulate products to generate (compute) a MAC value (see ¶[28]).)
Regarding claim 25, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Deshpande also teaches
wherein the output buffer is divided into a plurality of output buffer blocks, each output buffer block corresponding to a column of blocks of the storage sites, and wherein the multiplexer selects which of the output buffer blocks provides output to be read ((Deshpande teaches switches 1041 and 1042 coupled to inverters (output buffer) that passes (buffer) RBLs to adder tree wherein each inverter (output buffer block) of said inverters is connected to respective RBL which is connected to a column of blocks of respective one of sub-arrays 701 and 702 storing weights (see Fig. 10) in bit cells (storage sites) (see ¶[38]). Deshpande also teaches address signal CIM_ADDR<2> is used to control which of switch 1041 or 1042 (multiplexer) to couple (select) MULTB signals (output) from (read) respective one of said sub-arrays 701 and 702 to specific inverter (which of the output buffer blocks) to adder tree (see Fig. 10, ¶[61-62]).)
In view of Deshpande, modified Di Febbo is further modified such that memory cells (storing weights) are configured into sub-arrays 701 and 702 where each inverter (output buffer block) is coupled to respective RBL which is connected to column of blocks of respective one of said sub-arrays 702 and 702 wherein address signal CIM_ADDR<2> is used to control which of switch 1041 or 1042 (multiplexer) to couple (select) MULTB signals (output) from (read) respective one of said sub-arrays 701 and 702 to specific inverter (which of the output buffer blocks) to adder tree.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Di Febbo in the manner described supra because configuring cells into sub-arrays (and routing inputs/outputs to said sub-arrays) would reduce device loading and reduce power consumption (Deshpande, ¶[57]).
Claims 5 and 18 – 19 are rejected under 35 U.S.C. 103 as being unpatentable over Di Febbo in view of Deshpande and Huang and further in view of Kale (US 20240086696).
Regarding claim 5, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Di Febbo also teaches
wherein the blocks include a first block and a second block, [weight of the first block being replicated in weight of the second block] (Di Febbo teaches memory cells (blocks) comprising memory cell 727aa (first block) and memory cell 727ab (second block) (see Fig. 7).)
As noted in claim 5, modified Di Febbo teaches first and second blocks but does not appear to explicitly teach
the weights of the first block being replicated in the weights of the second block
However, Kale teaches
weight of [the] first block being replicated in weight of [the] second block (Kale teaches layer 371’s memory cells (first block) storing (replicated) the same weight matrices as layer 373’s memory cells (second block) (see ¶[217]).)
In view of Kale, modified Di Febbo is modified such that said first and second blocks store the same weights.
Di Febbo, Deshpande, Huang and Kale are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Di Febbo in the manner described supra because replicating weight matrix in multiple memory cells of different layers allows for error detection that reduces erroneous result of multiplication and accumulation (Kale, ¶[210]).
Regarding claim 18, Di Febbo in view of Deshpande and Huang teach the method of claim 17 where first set of weights is stored in first block, and second set of weights is stored in second block but do not appear to explicitly teach storing said first set of weights in said second block, and storing said second set of weights in said first block (see also limitation below).
storing, in the first block and the second block, the first set of weights and the second set of weights
However, Kale teaches
storing, in [the] first block and [the] second block, the first set of weights and the second set of weights (Kale teaches configuring synapse memory cells 207-227 (first block) and synapse memory cells 206-226 (second block) to store same set of weights (first and second set of weights) (see Fig. 5, ¶[204], [208]).)
In view of Kale, modified Di Febbo is modified such that i) said first block (storing said first set of weights) would store said second set of weights, and ii) said second block (storing said second set of weights) would store said first set of weights.
Di Febbo, Deshpande, Huang and Kale are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify modified Di Febbo in the manner described supra because storing same weights across multiple synapse memory cells would allow for detection of error and correction of memory cells with corrupted weights (Kale, ¶[208-209]).
Regarding claim 19, Di Febbo in view of Deshpande, Huang and Kale teach the method of claim 18 where Kale also teaches
storing of the first set of weights and the second set of weights in the first block and the second block based on an optimization of throughput and utilization of the CIM hardware module (Kale teaches configuring synapse memory cells 207-227 (first block) and synapse memory cells 206-226 (second block) to store same set of weights (first and second set of weights) in order to (based on) i) improve reliability (utilization) of computations results generated by synapse memory cells and ii) to select same result generated (throughput) by most of memory cell sets (see Fig. 5, ¶[204], [208]) wherein memory cells are part of array 273 (see Fig. 5) in integrated circuit die B (CIM hardware module).)
In view of Kale, modified Di Febbo is modified such that that i) said first block (storing said first set of weights) would store said second set of weights, and ii) said second block (storing said second set of weights) would store said first set of weights, in order to i) improve realizability (unitization) of computation results generated by memory cells of said CIM hardware module, and ii) select same result generated (throughput) by most of memory cell sets of said CIM hardware module.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify modified Di Febbo in the manner described supra because storing same weights across multiple synapse memory cells would allow for detection of error and correction of memory cells with corrupted weights (Kale, ¶[208-209]).
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Di Febbo in view of Deshpande and Huang, and further in view of Lee (US 20240177751) and Ting (US 9653148).
Regarding claim 8, Di Febbo in view of Deshpande and Huang teach the CIM hardware module of claim 1 where Di Febbo also teaches
wherein the blocks include a first block and a second block (Di Febbo teaches memory cells (blocks) comprising memory cell 727aa (first block) and memory cell 727ba (second block) (see Fig. 7).)
a first output of the operations on the first block [being output at a first time], and a second output of the operations for the second block [being output at a second time different from the first time] (Di Fabbo teaches memory cells 727aa and memory cells 727ba, each multiplies (operations) input value with weight (see Fig. 7, ¶[67]) that generate a product (first and second outputs) (see ¶[43]).)
Modified Di Febbo teaches a base CIM hardware module with compute logic and first and second blocks (see claim 8). The claimed invention improves upon said base module by having said first and second blocks share a portion of said compute logic.
This improvement to said base module is an application of known technique from Lee – memory cells sharing accumulator (see Lee Fig. 3 and corresponding paragraphs).
One of ordinary skill in the art would recognize that this known technique of sharing accumulator among memory cells can also be applied to first and second blocks of modified Di Febbo, and the result would have been predictable. In this instance, said compute logic is modified to include an accumulator that is shared between said first and second blocks. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Lee’s known technique would have yielded i) predictable result of said compute logic being modified to include an accumulator (portion) that is shared between said first and second blocks, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Modified Di Febbo teaches a base CIM module that outputs i) first output from first block, and ii) second output from second block (see claim 8). The claimed invention improves upon said base module by outputting said first output at first time and said second output at second time different from said first time.
This improvement to said base module is an application of known technique from Ting – reading (being output) data D0-D3 (first output) from memory cell array 106_1 (first block) at time period 5-8 (first time), and reading (being output) data D4-D5 (second output) from memory cell array 106_2 (second block) at time period 9-10 (second time) (see Ting Fig. 2, col 1 ln 36-54).
One of ordinary skill in the art would recognize that this known technique of outputting data at different time periods can also be applied to output first and second outputs of modified Di Febbo, and the result would have been predictable. In this instance, said first and second outputs are outputted at different time periods. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Ting’s known technique would have yielded i) predictable result of said first and second outputs are outputted at different time periods, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Di Febbo in view of Deshpande, Huang, Lee and Ting, and further in view of Chih (US 20220019407).
Regarding claim 9, Di Fabbo in view of Deshpande, Huang, Lee and Ting teach the CIM hardware module of claim 8 where Di Fabbo also teaches
wherein the compute logic includes a first plurality of logic gates corresponding to the first block, a second plurality of logic gates corresponding to the second block (Di Fabbo teaches memory cell 727aa (first block) coupled to ADC 285a (compute logic) and memory cell 727ba (second block) coupled to ADC 285b (compute logic) (see Fig. 7) wherein said ADCs are implemented using circuits (see ¶[87]) that are combinatorial logic circuity such as flops, register and latches (first plurality of logic gates and second plurality of logic gates) (see ¶[116]).)
Modified Di Febbo teaches a base CIM module that has ADCs (first and second plurality of logic gates) each coupled respectively to memory cells 727aa, 727ba (first and second blocks) (see claim 9) where said memory cell 727aa (first block) output first output at a different time from second output of said memory cell 727ba (second block) (see claim 8). The claimed invention improves upon said base module by having adders coupled to said ADCs (first and second plurality of logic gates).
This improvement to said base module is an application of known technique from Lee –coupling adder to ADC that is coupled to memory cells (first and second blocks) wherein outputs (first and second outputs) (from said memory cells) are pass from said memory cells to said ADC to said adder to (providing) said accumulator (see Lee Fig. 3, ¶[83-85]).
One of ordinary skill in the art would recognize that this known technique of coupling adder to ADC can also be applied to couple modified Di Febbo’s ADCs, and the result would have been predictable. In this instance, said ADCs are coupled to an adder wherein said first and second outputs (from said memory cells 727aa and 727ba) are passed from said memory cells 727aa and 727ba to said adder to said accumulator. Note that since said first and second outputs are outputted at different times, said first and second outputs would be passed to from said adder at different times. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Lee’s known technique would have yielded i) predictable result of said ADCs (first and second plurality of logic gates) are coupled to an adder wherein said first and second outputs (outputted at different times (first and second times)) are passed from said memory cells 727aa and 727ba to said adder to (providing) said accumulator, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Modified Di Febbo teaches a base CIM module that has adder connected accumulator (see claim 9). The claimed invention improves upon said base module by having plural adders.
This improvement to said base module is an application of known technique from Chih –adder trees 122 (adders) connected to accumulator M 140 (see Chih Fig. 1A).
One of ordinary skill in the art would recognize that this known technique of using plural adder tress to connect to accumulator can also be applied to adder of modified Di Febbo, and the result would have been predictable. In this instance, said adder is modified to be plural adder tress 122 that are connected to said accumulator. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Chih’s known technique would have yielded i) predictable result of said adder being modified to be plural adder tress 122 that are connected to said accumulator, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Di Febbo in view of Deshpande, Huang, Lee and Ting, and further in view of Edso (US 20170357570).
Regarding claim 10, Di Febbo in view of Deshpande, Huang, Lee and Ting teach the CIM hardware module of claim 8 where first and second block outputting respective output at different times but do not appear to explicitly teach turning off block at said different times (see also limitation below).
wherein the second block is not powered on io during the first time and the first block is not powered on during the second time
However, Edso teaches
wherein [the] second block is not powered on during [the] first time and [the] first block is not powered on during [the] second time (Edso teaches for cycles (first and second times) where row/columns (of a bank (first block or second block) (see ¶[22])) are read/written, unused bank (second block or first block) is turned off (not powered on)) (see ¶[162]).)
In view of Edso, modified Di Febbo is further modified such that when said first block is outputting first output at said first time, said second block is turned off, and when said second block is outputting said second output at said second time, said first block is turned off.
Di Febbo, Deshpande, Huang, Lee, Ting and Edso are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify modified Di Febbo in the manner described supra because it turning off unused bank would save power (Chih, ¶[162]).
Claim 11 and 13 – 14 are rejected under 35 U.S.C. 103 as being unpatentable over Garde (US 20090177867) in view of Di Febbo, Deshpande and Huang.
Regarding claim 11, Garde teaches
A compute tile (compute tile = Fig. 2 processor 20), comprising:
a plurality of compute engines (plurality of compute engines = Fig. 2 compute engines 50-57), each of the plurality of compute engines including a memory module [hardware compute-in-memory (CIM) module, the CIM hardware memory module including a plurality of is storage sites and compute logic coupled with the plurality of storage sites for performing, in parallel, operations on data stored in the storage sites]; and (Garde teaches each compute engine includes memory 78-80 (memory module) (see ¶[45]).)
a general-purpose (GP) processor (GP processor = Fig. 2 control block 30) coupled with the plurality of compute engines and configured to provide control instructions and data to the plurality of compute engines; (Garde teaches instructions (control instructions) issued (provide) by control block 30 (GP processor) (with control logic (processor)) and corresponding data (data) flow through compute engines (compute engines) and are executed by said compute engines (see Fig. 2, ¶[41], [42]). Note that said compute engines and said control block 30 are functionally coupled (coupled with) via flow of said instructions.)
Garde teaches each compute engine includes a memory module but does not appear to explicitly teach said memory module is hardware CIM module performing the functions outlined in the limitations below.
each of the plurality of compute engines including a hardware compute-in-memory (CIM) module, the CIM hardware module including a plurality of storage sites and compute logic coupled with the plurality of storage sites for performing, in parallel, operations on data stored in the storage sites; and
wherein the CIM hardware module is configured to store weights in blocks of the storage sites, to utilize the blocks and portions of the compute logic corresponding to the blocks to selectively provide outputs of the operations for the weights stored in the blocks, and to read the outputs of the operations corresponding to the blocks at different times
However, Di Febbo teaches
[each of the plurality of compute engines including a hardware compute-in-memory (CIM) module, the] CIM hardware module including a plurality of storage sites and compute logic coupled with the plurality of storage sites for performing, in parallel, operations on data stored in the storage sites; and
wherein the CIM hardware module is configured to store weights in blocks of the storage sites, to utilize the blocks and portions of the compute logic corresponding to the blocks to selectively provide outputs of the operations for the weights stored in the blocks, and to read the outputs of the operations corresponding to the blocks at different times (Note that claim 1 teaches a CIM hardware module as described here. As such, Di Febbo’s mapping in claim 1 also applies here.)
In view of Di Febbo, Garde is modified such that said memory module (in each of said plurality of compute engines) is a CIM hardware module as described here.
Garde and Di Febbo are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Garde in the manner described supra because use of im-memory compute arrays optimize execution of MAC operations in CNN accelerators (Di Febbo, ¶[3],[20]).
As noted in claim 11, modified Garde teaches memory cells (blocks), each with transconductance (compute logic) that multiplies input value to generate a product (outputs) but does not appear to explicitly teach using demultiplexer and multiplexer to route said input value and said product in the following manner.
an output buffer;
a demultiplexer [coupled to the input buffer];
a multiplexer coupled to the output buffer
selectively read, using the multiplexer, the outputs of the operations corresponding to the blocks at different times,
wherein the demultiplexer is configured for routing an input to a portion of the compute logic corresponding to a block of the blocks, and
wherein the multiplexer is configured to select an output corresponding to the block
However, Deshpande teaches
an output buffer;
a demultiplexer [coupled to the input buffer];
a multiplexer coupled to the output buffer
selectively read, using the multiplexer, the outputs of the operations corresponding to the blocks at different times,
wherein the demultiplexer is configured for routing an input to [a portion of the compute logic corresponding to] a block of [the] blocks, and
wherein the multiplexer is configured to select an output corresponding to the block (These limitations are the same as claim 1. Therefore, Deshpande’s mappings in claim 1 also applies here.)
In view of Deshpande, modified Garde is further modified such that said memory cells (blocks) are configured as sub-arrays where said input value (used to generate said product (outputs)) of each transconductance (compute logic), in each of said memory cells, would be routed to respective sub-array of said sub-arrays using multiplexer (demultiplexer), and switches (multiplexer) (connected to inverters (output buffer)) would be used to couple said product from said respective sub-array to adder tree where only one of said switches is on at a time (different times) resulting in a subset (selective) of said products being coupled to said adder tree at a time.
Garde, Di Febbo and Deshpande are analogous art to the claimed invention because they are in the same field of endeavor, storage management.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which said subject matter pertains to modify Di Febbo in the manner described supra because configuring cells into sub-arrays (and routing inputs/outputs to said sub-arrays) would reduce device loading and reduce power consumption (Deshpande, ¶[57]).
A noted in claim 11, modified Garde teaches a base CIM module with multiplexer (demultiplexer) that routes input to compute logic. The claimed invention improves upon said base module by using coupling said multiplexer to an input buffer.
This improvement to said base module is an application of known technique from Huang – multiplexer (demultiplexer) provides output (input) from specific storage elements, in buffer 104a (input buffer) (connected (coupled to) to said multiplexer), to processing element (compute logic) (see Huang Fig. 2A, col 11 ln 6-18).
One of ordinary skill in the art would recognize that this known technique of multiplexer using buffer to pass input to processing element can also be applied to modified Garde’s multiplexer that routes said input to said compute logic, and the result would have been predictable. In this instance, said multiplexer would route said input from specific storage elements, in buffer 104 (connected to said multiplexer), to said compute logic. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Huang’s known technique would have yielded i) predictable result of said multiplexer would route said input from buffer 104 (input buffer) (connected to said multiplexer (coupled to)) to said compute logic, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Regarding claim 13, Garde in view of Di Febbo, Deshpande and Huang teach the compute tile of claim 11 where Di Febbo also teaches
wherein the hardware CIM is configured to store the weights in the blocks based on an optimization of throughput and utilization of the CIM hardware module (see Di Febbo mapping in claim 6 supra)
Regarding claim 14, Garde in view of Di Febbo, Deshpande and Huang teach the compute tile of claim 11 where Di Febbo also teaches
wherein the operations comprise vector-matrix multiplication operations (see Di Febbo mapping in claim 7 supra)
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Garde in view of Di Febbo, Deshpande and Huang, and further in view of Lee and Ting.
Regarding claim 15, Garde in view of Di Febbo, Deshpande and Huang teach the compute tile of claim 11 where Di Febbo also teaches
wherein the blocks include a first block and a second block
a first output of the operations on the first block [being output at a first time], and a second output of the operations for the second block [being output at a second time different from the first time] (see Di Febbo mapping in claim 8 supra)
Modified Garde teaches a base compute tile with compute logic and first and second blocks (see claim 15). The claimed invention improves upon said compute tile by having said first and second blocks share a portion of said compute logic.
This improvement to said base compute tile is an application of known technique from Lee – memory cells sharing accumulator (see Lee mapping in claim 8 supra).
One of ordinary skill in the art would recognize that this known technique of sharing accumulator among memory cells can also be applied to first and second blocks of modified Garde, and the result would have been predictable. In this instance, said compute logic is modified to include an accumulator that is shared between said first and second blocks. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Lee’s known technique would have yielded i) predictable result of said compute logic being modified to include an accumulator (portion) that is shared between said first and second blocks, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Modified Garde teaches a base compute tile that outputs i) first output from first block, and ii) second output from second block (see claim 15). The claimed invention improves upon said base compute tile by outputting said first output at first time and said second output at second time different from said first time.
This improvement to said base compute tile is an application of known technique from Ting – reading (being output) data D0-D3 (first output) from memory cell array 106_1 (first block) at time period 5-8 (first time), and reading (being output) data D4-D5 (second output) from memory cell array 106_2 (second block) at time period 9-10 (second time) (see Ting mapping in claim 8).
One of ordinary skill in the art would recognize that this known technique of outputting data at different time periods can also be applied to output first and second outputs of modified Garde, and the result would have been predictable. In this instance, said first and second outputs are outputted at different time periods. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Ting’s known technique would have yielded i) predictable result of said first and second outputs are outputted at different time periods, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Garde in view of Di Febbo, Deshpande, Huang, Lee and Ting, and further in view of Chih.
Regarding claim 16, Garde in view of Di Febbo, Deshpande, Huang, Lee and Ting teach the compute tile of claim 15 where Di Fabbo also teaches
wherein the compute logic includes a first plurality of logic gates corresponding to the first block, a second plurality of logic gates corresponding to the second block (see Di Febbo mapping in claim 9)
Modified Garde teaches a base compute tile that has ADCs (first and second plurality of logic gates) each coupled respectively to memory cells 727aa, 727ba (first and second blocks) (see claim 16) where said memory cell 727aa (first block) output first output at a different time from second output of said memory cell 727ba (second block) (see claim 15). The claimed invention improves upon said base compute tile by having adders coupled to said ADCs (first and second plurality of logic gates).
This improvement to said base compute tile is an application of known technique from Lee –coupling adder to ADC that is coupled to memory cells (first and second blocks) wherein outputs (first and second outputs) (from said memory cells) are pass from said memory cells to said ADC to said adder to (providing) said accumulator (see Lee mapping in claim 9).
One of ordinary skill in the art would recognize that this known technique of coupling adder to ADC can also be applied to couple modified Garde’s ADCs, and the result would have been predictable. In this instance, said ADCs are coupled to an adder wherein said first and second outputs (from said memory cells 727aa and 727ba) are passed from said memory cells 727aa and 727ba to said adder to said accumulator. Note that since said first and second outputs are outputted at different times, said first and second outputs would be passed to from said adder at different times. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Lee’s known technique would have yielded i) predictable result of said ADCs (first and second plurality of logic gates) are coupled to an adder wherein said first and second outputs (outputted at different times (first and second times)) are passed from said memory cells 727aa and 727ba to said adder to (providing) said accumulator, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Modified Garde teaches a base compute tile that has adder connected accumulator (see claim 16). The claimed invention improves upon said base compute tile by having plural adders.
This improvement to said base compute tile is an application of known technique from Chih –adder trees 122 (adders) connected to accumulator M 140 (see Chih mapping in claim 9).
One of ordinary skill in the art would recognize that this known technique of using plural adder tress to connect to accumulator can also be applied to adder of modified Garde, and the result would have been predictable. In this instance, said adder is modified to be plural adder tress 122 that are connected to said accumulator. It would have been obvious to one of ordinary skill in the art at the time of filing to recognize that applying Chih’s known technique would have yielded i) predictable result of said adder being modified to be plural adder tress 122 that are connected to said accumulator, and ii) the improved claimed invention (see MPEP 2143(I)(D)).
Allowable Subject Matter
Claim 21 recites, at least, clearing input buffer and output buffer after each set of operations that are performed in parallel on data stored in storage sites. This subject matter is reflected in the following limitations of claim 1 and 21.
compute logic coupled with the storage sites for performing, in parallel, operations on data stored in the storage sites (claim 1)
the input buffer and the output buffer are cleared after each set of the operations (claim 21)
Shin (US 20240419341) teaches resetting (clear) input/output buffer (input buffer, output buffer) after transferring (operation) of each of data 1 and data 2 (each set) (see Shin ¶[104], [107]). However, Shin does not appear to explicitly teach resetting said input/output buffer after transferring, in parallel, each of said data 1 and said data 2. Therefore, claim 21 is allowable over prior art of record.
Response to Remarks
Claim objections, drawing objections and specification objections are withdrawn in view of Applicant’s amendments to the specification and the claims.
Applicant’s remarks, with respect to prior art rejections, have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of newly identified prior art (see rejection supra).
Additional Remarks
In the interest of compact prosecution, Deshpande also teaches when (first time) address signal CIM_ADDR<2> is set to control switch 1041, coupling (select) first RBLs from sub-array 701 to (output) 4 (all) inverters (blocks of output buffer), and when (second time) address signal CIM_ADDR<2> is set to control switch 1042, coupling (select) second RBL0 from sub-array 702 to (output) inverter 1 (MULT0) (subset of output buffer) (see Deshpande Fig. 10, ¶[61-62]), wherein said first RBLs are output (resultants) products (multiplications) of each bit cells (vector) in columns (matrices) of said sub-array 701, and said second RBL0 are output (resultant) products (multiplication) of each bit cells (vector) in first column (single matrix) of said sub-array 702 (see Deshpande ¶[40]). This reads on the following limitations from claim 21.
the multiplexer is configured to select all blocks of the output buffer to output resultants of vector-matrix multiplications for multiple matrices at a first time
select a subset of blocks of the output buffer to output a resultant of a vector-matrix multiplication for a single matrix at a second time
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHIE YEW whose telephone number is (571)270-5282. The examiner can normally be reached Monday - Thursday and alternate Fridays.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Reginald Bragdon can be reached at (571) 272-4204. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHIE YEW/ Primary Examiner, Art Unit 2139