DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 03/24/2026 has been entered.
Remarks
Claim 25 recites “wherein the memory cell is a static random access memory (SRAM) memory cell”. Applicant may want to recite “wherein the memory cell is a static random access memory (SRAM)
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 3, 18 and 20-29 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites “the result” in line 23. There is insufficient antecedent basis for this limitation in the claim. The amendments made deleted a result recited in line 9. For purposes of examination, this is interpreted as a of the bit-serial multiplication. Claim 22 recites a similar limitation in line 1 and is rejected for the same reason. Claims 3 and 22-26 inherit the same deficiency as claim 1 by reason of dependence.
Claim 18 recites “the second partial-product” in lines 19-20. There is insufficient antecedent basis for this limitation in the claim. For purposes of examination, this is interpreted as a second partial-product. Claims 20-21 and 27-29 inherit the same deficiency as claim 18 by reason of dependence.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 3, 8-10, 12-13, 17-18, 20-23, 25-27 and 29 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chih et al. (NPL – “16.4 An 89TOPS/W and 16.3TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute-In Memory Macro in 22nm for Machine-Learning Edge Applications”), hereinafter Chih.
Regarding claim 1, Chih teaches a computing method configured to perform bit-serial multiplication in a compute-in memory (CIM) device, the computing method comprising (Chih page 252 left col second paragraph):
determining at least one input according to a type of an application (Chih page 252 left col second paragraph);
receiving the at least one input by an input driver circuit (Chih Fig. 16.4.1)
determining at least one weight according to a training result or a configuration of a user (Chih page 252 left col second paragraph; right col second and fourth paragraphs);
storing the at least one weight in a memory cell of a memory array (Chih Fig. 16.4.1; page 252 left col third paragraph);
performing a bit-serial multiplication based on the input and the weight, by a multiply circuit, from a most significant bit (MSB) of the input to a least significant bit (LSB) of the input to obtain a plurality of partial-products (Chih page 252 left col second and fourth paragraphs; Fig. 16.4.1 and 16.4.2; multiply circuit – multipliers; plurality of partial-products – output of the multipliers/adder tree);
storing a first partial-sum based on a first one of the plurality of partial products in a first register (Chih Fig. 16.4.2; first register – register left of the shifter);
left-shifting the first partial-sum one bit by a shifter circuit (Chih Fig. 16.4.2);
storing a second one of the plurality of partial products in a second register (Chih Fig. 16.4.2; second register – register left of the multiplexer);
adding the left-shifted first partial-sum and the second one of the plurality of partial products by an adder circuit to obtain a second partial-sum (Chih Fig. 16.4.2; adder circuit – 20b adder);
storing an output of the adder circuit in the first register having an output terminal operably connected to an input terminal of the shifter circuit (Chih Fig. 16.4.2); and
outputting the result, by the adder circuit (Chih Fig. 16.4.2).
Regarding claim 3, Chih teaches all the limitations of claim 1 as stated above. Further, Chih teaches wherein the input comprises a plurality of inputs, and wherein performing the bit-serial multiplication comprises: determining the plurality of the first partial-products for each of the plurality of inputs (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2 plurality of inputs – input activation).
Regarding claim 22, Chih teaches all the limitations of claim 1 as stated above. Further, Chih teaches further comprising outputting the result by the adder circuit to a third register (Chih Fig. 16.4.2; third register – register right of the adder).
Regarding claim 23, Chih teaches all the limitations of claim 1 as stated above. Further, Chih teaches wherein performing the bit-serial multiplication includes a logic NOR operation (Chih page 252 left col third paragraph; Fig. 16.4.2).
Regarding claim 25, Chih teaches all the limitations of claim 1 as stated above. Further, Chih teaches wherein the memory cell is a static random access memory (SRAM) memory cell (Chih page 252 left col third paragraph; Fig. 16.4.1).
Regarding claim 26, Chih teaches all the limitations of claim 1 as stated above. Further, Chih teaches further comprising:
storing a third one of the plurality of partial products in the second register ((Chih Fig. 16.4.2);
left-shifting the second partial sum one bit by the shifter circuit ((Chih Fig. 16.4.2);
adding the left-shifted second partial-sum and the third one of the plurality of partial products by the adder circuit to obtain a third partial-sum ((Chih Fig. 16.4.2); and
storing the output of the adder circuit in the first register ((Chih Fig. 16.4.2).
Regarding claim 8, Chih teaches a device, comprising:
an adder (Chih Fig. 16.4.2 adder – 20b adder);
a shifter having an output terminal operably connected to a first input terminal of the adder, the shifter configured to left-shift one bit (Chih Fig. 16.4.2 shifter - shifter block left of the adder);
a first register configured to store a first partial-sum and having an output terminal operably connected to an input terminal of the shifter, wherein the shifter is configured to left-shift each first partial sum received from the first register one bit (Chih Fig. 16.4.2 first register - register left of the shifter);
a second register having an output terminal operably connected to a second input terminal of the adder (Chih Fig. 16.4.2 second register - register receiving and outputting PSUM<11:0>);
a multiplier configured to perform a bit-serial multiplication based on an input signal and a weight signal from a most significant bit (MSB) of the input signal to a least significant bit (LSB) of the input signal to obtain a plurality of partial-products (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2; multiply circuit - multipliers; input signal – input activation; weight signal - weights);
wherein an input terminal of the second register is connected to an output terminal of the multiplier and is operable to sequentially receive the plurality of partial-products from the multiplier (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2); and
wherein an input terminal of the first register is connected to an output terminal of the adder and is operable to receive an output of the adder (Chih Fig. 16.4.2);
wherein the adder is configured to add the left-shifted first partial-sum and a first one of the plurality of partial-products to obtain a second partial-sum (Chih Fig. 16.4.2).
Regarding claim 9, Chih teaches all the limitations of claim 8 as stated above. Further, Chih teaches further comprising a third register having an input terminal that is operably connected to the output of the adder (Chih Fig. 16.4.2; third register – register on the right of the adder).
Regarding claim 10, Chih teaches all the limitations of claim 8 as stated above. Further, Chih teaches wherein the multiplier comprises a NOR gate (Chih page 252 left col third paragraph; Figs. 16.4.1-2).
Regarding claim 12, Chih teaches all the limitations of claim 8 as stated above. Further, Chih teaches further comprising a memory array configured to store the weight signal (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2; memory array – array of memory cells).
Regarding claim 13, Chih teaches all the limitations of claim 12 as stated above. Further, Chih teaches wherein the memory array includes a plurality of static random access memory (SRAM cells) (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2).
Regarding claim 17, Chih teaches all the limitations of claim 8 as stated above. Further, Chih teaches wherein: the first register is configured to store the obtained second partial-sum; the shifter is configured to left-shift the obtained second partial sum one bit; the adder is configured to add the obtained left-shifted second partial-sum and a next one of the plurality of partial-products to obtain a third partial sum (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2).
Regarding claim 18, Chih teaches a device, comprising:
a memory array storing a weight signal (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2; memory array – array of memory cells).;
an input driver configured to output an input signal (Chih Fig. 16.4.1 input driver - input driver);
a multiplier configured to perform a bit-serial multiplication of the input signal and the weight signal, from a most significant bit (MSB) of the input signal to a least significant bit (LSB) of the input signal to determine a plurality of partial-products (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2; multiplier – multipliers);
an adder (Chih Fig. 16.4.2; adder – 20b adder);
a first register having an input terminal operably connected to an output of the adder, the first register configured to store a first partial-sum (Chih Fig. 16.4.2; first register – register to the left of the shifter block);
a shifter having an input terminal operably connected to an output of the first register and an output terminal operably connected to a first input terminal of the adder, the shifter being configured to left-shift each first partial sum received from the first register one bit and output the left-shifted first partial sum to the adder (Chih Fig. 16.4.2; shifter - shifter);
a second register having an input terminal operably connected to an output terminal of the multiplier to sequentially receive the plurality of partial products from the multiplier, and an output terminal operably connected to a second input terminal of the adder (Chih Fig. 16.4.2 second register - register receiving and outputting PSUM<11:0>);
wherein the adder is configured to add the left-shifted first partial-sum and a second partial- product to obtain a second partial-sum (Chih page 252 left col fourth paragraph; Fig. 16.4.2).
Regarding claim 20, Chih teaches all the limitations of claim 18 as stated above. Further, Chih teaches further comprising a third register having an input terminal that is operably connected to the output of the adder (Chih Fig. 16.4.2; third register – register on the right of the adder).
Regarding claim 21, Chih teaches all the limitations of claim 18 as stated above. Further, Chih teaches wherein the memory array includes a plurality of static random access memory (SRAM cells) (Chih page 252 left col second-fourth paragraphs; Figs. 16.4.1-2).
Regarding claim 27, Chih teaches all the limitations of claim 18 as stated above. Further, Chih teaches wherein the multiplier comprises a NOR gate (Chih page 252 left col third paragraph; Fig. 16.4.2).
Regarding claim 29, Chih teaches all the limitations of claim 18 as stated above. Further, Chih teaches wherein: the first register is configured to store the obtained second partial-sum; the shifter is configured to left-shift the obtained second partial-sum one bit; the adder is configured to add the obtained left-shifted second partial-sum and a next one of the plurality of partial-products to obtain a third partial sum (Chih Fig. 16.4.2).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 8-9, 11-12, 17-18, 20 and 28-29 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 10853066 B1), hereinafter Liu, in view of Judd et al. (NPL – “Stripes: Bit-Serial Deep Neural Network Computing”), hereinafter Judd, and Ware et al. (US 20210326286 A1), hereinafter Ware.
Regarding claim 18, Liu teaches a device, comprising:
a memory array storing a weight signal (Liu Figs. 1-5; memory array - memory cell arrays 110/ first storage location 505; weight signal - set of multipliers; col 5 lines 5 lines 63-64 “At 605, a set of multipliers can be loaded into the first storage location 505”);
an input driver configured to output an input signal (Liu Figs. 1-5; input driver - input registers 120/second storage location 510; input signal - set of multiplicands; col 5 line 64 to col 6 line 6 “a set of multiplicands can be loaded into the second storage location 510 … A given bit position of the given row in the second storage location 510 … can be output to the logic AND circuitry 520, 525”);
a multiplier configured to perform a bit-serial multiplication of the input signal and the weight signal, from a most significant bit (MSB) of the input signal to a least significant bit (LSB) of the input signal to determine a plurality of partial-products (Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows. After rows of the first and second storage locations 505, 510 are incrementally addressed by the address generator 515, the bit position in the second storage location 510 for input to the logic AND circuitry 520, 525 can be shifted by one bit to the left when processing from the most-significant-bit to the least-significant-bit, and the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630 … The processes at 615-630 can be repeated for each bit position of the second storage location 510. After the process at 615-630 are repeated for each bit position of the second storage location 510, the accumulated value in the one or more accumulators 530, 535 can be output as the matrix dot product of the set of multiplier and the set of multiplicands”; multiplier – AND circuitry 520 and/or 525; plurality of partial-products - output of the bitwise AND each time step 620 is repeated);
an adder (Liu Figs. 5-6 and col 6 lines 10-13 “At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows”; adder – accumulator 520/525);
(Liu Figs. 5-6 and col 6 lines 13-20 “After rows of the first and second storage locations 505, 510 are incrementally addressed by the address generator 515 … the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630”; partial-sum – accumulated value that is shifted at 630);
sequentially receive the plurality of partial products from the multiplier, (Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position … The processes at 615-630 can be repeated for each bit position of the second storage location 510”);
wherein the adder is configured to add the left-shifted first partial-sum and a second partial- product to obtain a second partial-sum (Liu Figs. 5-6 and col 6 lines 10-13 “At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows”; second partial-product - output of the AND circuitry when processing the next bit after the MSB).
Liu does not explicitly teach a first register having an input terminal operably connected to an output of the adder, the first register configured to store a first partial-sum; a shifter having an input terminal operably connected to an output of the first register and an output terminal operably connected to a first input terminal of the adder, the shifter being configured to left-shift each first partial sum received from the first register one bit and output the left-shifted first partial sum to the adder; a second register having an input terminal operably connected to an output terminal of the multiplier to sequentially receive the plurality of partial products from the multiplier, and an output terminal operably connected to a second input terminal of the adder.
However, on the same field of endeavor, Judd discloses a first register having an input terminal operably connected to an output of an adder, the first register configured to store a first partial-sum; and a shifter having an input terminal operably connected to an output of the first register and an output terminal operably connected to a first input terminal of the adder, the shifter being configured to left-shift each first partial sum received from the first register one bit and output the left-shifted first partial sum to the adder (Judd Fig. 4 adder – adder; a first register – register downstream of the adder; shifter – shifter below the register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu using Judd and include a shifter to perform the left shift operation in step 630. As discussed, Liu discloses performing a left shift operation while Judd discloses a shifter which is normally used for performing shifting operations. Therefore, one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination yielded nothing more than predictable results to one of ordinary skill in the art. The predictable result is a device that includes a shifter for performing the shift operations. See MPEP 2141 III (A) for more information. Further, include a register for storing the output of the adder that is provided to the shifter in order to implement a processing pipeline as registers are normally used to separate pipeline stages (Ware paragraph [0039]).
Therefore, the combination of Liu as modified in view of Judd and Ware teaches a first register having an input terminal operably connected to an output of the adder, the first register configured to store a first partial-sum; a shifter having an input terminal operably connected to an output of the first register and an output terminal operably connected to a first input terminal of the adder, the shifter being configured to left-shift each first partial sum received from the first register one bit and output the left-shifted first partial sum to the adder.
Liu as modified in view of Judd and Ware does not explicitly teach a second register having an input terminal operably connected to an output terminal of the multiplier to sequentially receive the plurality of partial products from the multiplier, and an output terminal operably connected to a second input terminal of the adder.
However, on the same field of endeavor, Ware discloses a second register having an input terminal connected to an output of a multiplier and an output terminal connected to an input of an adder (Ware Fig. 2D and paragraph [0062] “In each execution cycle, the filter weight value (Fkl value) in the D_r[p] register is multiplied by the Dijk value in the D_i[p] register, via multiplier circuitry, and the result is output to the MULT_r[p] register. In the next pipeline cycle this product (i.e., D*F value) is added to the Yijl accumulation value in the MAC_r[p−1] register (in the previous multiplier-accumulator circuit) and the result is stored in the MAC_r[p] register”; second register - MULT_r[p] register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd using Ware and configure the memory device to include a second register downstream of the AND circuitry and upstream of the accumulator for storing each partial product output of the AND circuitry that is subsequently added/accumulated in order to implement a processing pipeline as registers are normally used to separate pipeline stages (Ware paragraph [0039]).
Therefore, the combination of Liu as modified in view of Judd and Ware teaches a second register having an input terminal operably connected to an output terminal of the multiplier to sequentially receive the plurality of partial products from the multiplier, and an output terminal operably connected to a second input terminal of the adder.
Regarding claim 20, Liu as modified in view of Judd teaches all the limitations of claim 18 as stated above.
Liu does not explicitly teach further comprising a third register having an input terminal that is operably connected to the output of the adder.
However, on the same field of endeavor, Ware discloses a third register having an input terminal that is operably connected to the output of the adder (Liu Fig. 2D and paragraph [0073] “The output data are parallel loaded from the accumulation registers Y (here, 64) into the MAC_SO registers (here, 64)”; third register - MAC_SO [p] register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective
filling date of the claimed invention, to modify Liu in view of Judd and Ware and configure the memory device to include a third register downstream of the adder for storing the final result (Ware paragraph [0073]) such that it can be output in a next pipeline execution cycle.
Therefore, the combination of Liu as modified in view of Judd and Ware teaches further comprising a third register having an input terminal that is operably connected to the output of the adder.
Regarding claim 28, Liu as modified in view of Judd and Ware teaches all the limitations of claim 18 as stated above. Further, Liu as modified in view of Judd and Ware teaches wherein the multiplier comprises an AND gate (Liu Fig. 5 and col 5 lines 41-42 and col 6 lines 6-10; AND gate – AND circuitry; Figs. 12-14; see also Judd Fig. 4).
Regarding claim 29, Liu as modified in view of Judd and Ware teaches all the limitations of claim 1 as stated above. Further, Liu as modified in view of Judd and Ware teaches wherein:
the first register is configured to store the obtained second partial-sum (Judd Fig. 4);
the shifter is configured to left-shift the obtained second partial-sum one bit (Judd Fig. 4; Liu col 5 lines 19-20);
the adder is configured to add the obtained left-shifted second partial-sum and a next one of the plurality of partial-products to obtain a third partial sum (Liu Figs. 5-6 and col 6 lines 25-26).
Regarding claim 8, Liu teaches a device, comprising:
an adder (Liu Figs. 5-6 and col 6 lines 10-13 adder - accumulator 520/525);
(Liu Fig. 6 and col 6 lines 19-20 “the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630”);
(Liu Fig. 6 and col 6 lines 19-20 “the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630”);
a multiplier configured to perform a bit-serial multiplication based on an input signal and a weight signal from a most significant bit (MSB) of the input signal to a least significant bit (LSB) of the input signal to obtain a plurality of partial-products (Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows. After rows of the first and second storage locations 505, 510 are incrementally addressed by the address generator 515, the bit position in the second storage location 510 for input to the logic AND circuitry 520, 525 can be shifted by one bit to the left when processing from the most-significant-bit to the least-significant-bit, and the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630 … The processes at 615-630 can be repeated for each bit position of the second storage location 510. After the process at 615-630 are repeated for each bit position of the second storage location 510, the accumulated value in the one or more accumulators 530, 535 can be output as the matrix dot product of the set of multiplier and the set of multiplicands”; multiplier – AND circuitry 520 and/or 525; input signal - set of multiplicands; weight signal - set of multipliers; plurality of partial-products - output of the bitwise AND each time step 620 is repeated);
(Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position … The processes at 615-630 can be repeated for each bit position of the second storage location 510”); and
;
wherein the adder is configured to add the left-shifted first partial-sum and a first one of the plurality of partial-products to obtain a second partial-sum (Liu Figs. 5-6 and col 6 lines 10-13 “At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows … The processes at 615-630 can be repeated for each bit position of the second storage location 510”; second partial-product - output of the AND circuitry when processing the next bit after the MSB).
Liu does not explicitly teach a shifter having an output terminal operably connected to a first input terminal of the adder, the shifter configured to left-shift one bit; a first register configured to store a first partial-sum and having an output terminal operably connected to an input terminal of the shifter, wherein the shifter is configured to left-shift each first partial sum received from the first register one bit; a second register having an output terminal operably connected to a second input terminal of the adder; wherein an input terminal of the second register is connected to an output terminal of the multiplier and is operable to sequentially receive the plurality of partial-products from the multiplier; and wherein an input terminal of the first register is connected to an output terminal of the adder and is operable to receive an output of the adder.
However, on the same field of endeavor, Judd discloses a shifter having an output terminal operably connected to a first input terminal of an adder, the shifter configured to left-shift one bit and a first register configured to store a first partial-sum and having an output terminal connected to an input terminal of the shifter, wherein the shifter is configured to left-shift each first partial sum received from the first register one bit, wherein an input terminal of the first register is operably connected to an output terminal of the adder to receive the output of the adder (Judd Fig. 4 adder – adder; a first register – register downstream of the adder; shifter – shifter below the register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu using Judd and include a shifter to perform the left shift operation in step 630. As discussed, Liu discloses performing a left shift operation while Judd discloses a shifter which is normally used for performing shifting operations. Therefore, one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination yielded nothing more than predictable results to one of ordinary skill in the art. The predictable result is a device that includes a shifter for performing the shift operations. See MPEP 2141 III (A) for more information. Further, include a register for storing the output of the adder that is provided to the shifter in order to implement a processing pipeline as registers are normally used to separate pipeline stages (Ware paragraph [0039]).
Therefore, the combination of Liu as modified in view of Judd and Ware teaches a shifter having an output terminal operably connected to a first input terminal of the adder, the shifter configured to left-shift one bit; a first register configured to store a first partial-sum and having an output terminal operably connected to an input terminal of the shifter, wherein the shifter is configured to left-shift each first partial sum received from the first register one bit; and wherein an input terminal of the first register is connected to an output terminal of the adder and is operable to receive an output of the adder.
Liu does not explicitly teach a second register having an output terminal operably connected to a second input terminal of the adder; wherein an input terminal of the second register is connected to an output terminal of the multiplier and is operable to sequentially receive the plurality of partial-products from the multiplier.
However, on the same field of endeavor, Ware discloses a second register having an input terminal connected to an output of a multiplier and an output terminal connected to an input of an adder (Ware Fig. 2D and paragraph [0062] “In each execution cycle, the filter weight value (Fkl value) in the D_r[p] register is multiplied by the Dijk value in the D_i[p] register, via multiplier circuitry, and the result is output to the MULT_r[p] register. In the next pipeline cycle this product (i.e., D*F value) is added to the Yijl accumulation value in the MAC_r[p−1] register (in the previous multiplier-accumulator circuit) and the result is stored in the MAC_r[p] register”; second register - MULT_r[p] register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd using Ware and configure the memory device to include a second register downstream of the AND circuitry and upstream of the accumulator for storing each partial product output of the AND circuitry that is subsequently added/accumulated in order to implement a processing pipeline as registers are normally used to separate pipeline stages (Ware paragraph [0039]).
Therefore, the combination of Liu as modified in view of Judd and Ware teaches a second register having an output terminal operably connected to a second input terminal of the adder; wherein an input terminal of the second register is connected to an output terminal of the multiplier and is operable to sequentially receive the plurality of partial-products from the multiplier.
Regarding claim 9, Liu as modified in view of Judd and Ware teaches all the limitations of claim 8 as stated above.
Liu does not explicitly teach further comprising a third register having an input terminal that is operably connected to the output of the adder.
However, on the same field of endeavor, Ware discloses a third register having an input terminal that is operably connected to the output of the adder (Liu Fig. 2D and paragraph [0073] “The output data are parallel loaded from the accumulation registers Y (here, 64) into the MAC_SO registers (here, 64)”; third register - MAC_SO [p] register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective
filling date of the claimed invention, to modify Liu in view of Judd and Ware and configure the memory device to include a third register downstream of the adder for storing the final result (Ware paragraph [0073]) such that it can be output in a next pipeline execution cycle.
Therefore, the combination of Liu as modified in view of Judd and Ware teaches further comprising a third register having an input terminal that is operably connected to the output of the adder.
Regarding claim 11, Liu as modified in view of Judd and Ware teaches all the limitations of claim 8 as stated above. Further, Liu as modified in view of Judd and Ware teaches wherein the multiplier comprises an AND gate (Liu Fig. 5 and col 5 lines 41-42 and col 6 lines 6-10; AND gate – AND circuitry; Figs. 12-14; see also Judd Fig. 4).
Regarding claim 12, Liu as modified in view of Judd and Ware teaches all the limitations of claim 8 as stated above. Further, Liu as modified in view of Judd and Ware teaches further comprising a memory array configured to store the weight signal (Liu Figs. 1-5; memory array - memory cell arrays 110/ first storage location 505; col 5 lines 5 lines 63-64 “At 605, a set of multipliers can be loaded into the first storage location 505”).
Regarding claim 17, Liu as modified in view of Judd and Ware teaches all the limitations of claim 8 as stated above. Further, Liu as modified in view of Judd and Ware teaches wherein:
the first register is configured to store the obtained second partial-sum (Judd Fig. 4 and section V.C);
the shifter is configured to left-shift the obtained second partial sum partial-sum (Liu Figs. 5-6 and col 6 lines 13-20 “the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630”; Judd Fig. 4 and section V.C);
the adder is configured to add the obtained left-shifted second partial-sum and a next one of the plurality of partial-products to obtain a third partial-sum (Liu Figs. 5-6 and col 6 lines 10-30 “At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position … The processes at 615-630 can be repeated for each bit position of the second storage location 510 … After the process at 615-630 are repeated for each bit position of the second storage location 510, the accumulated value in the one or more accumulators 530, 535 can be output as the matrix dot product of the set of multiplier and the set of multiplicands”; Judd Fig. 4 and section V.C).
Claims 10 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Judd and Ware as applied to claims 8 and 18 above, and further in view of Nakayama et al. (NPL – “A GaAs 16x16 bit parallel multiplier”), hereinafter Nakayama.
Regarding claim 10, Liu as modified in view of Judd and Ware teaches all the limitations of claim 8 as stated above.
Liu does not explicitly teach wherein the multiplier comprises a NOR gate.
However, on the same field of endeavor, Nakayama discloses a multiplier that comprises a NOR gate (Nakayama Fig. 4 and page 601 left col middle “The input signals applied to the X and Y lines are inverted by the input buffer and partial products are produced by the NOR gates”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd and Ware using Nakayama by replacing the AND gates with NOR gates receiving inverted inputs for performing the multiplication and generating the partial products. One of ordinary skill in the art could have substituted the AND gates/circuitry for NOR gates/circuitry receiving inverted inputs, and the results of the substitution would have been predictable because both configuration generates partial products from a multiplication of the input signals. See MPEP 2141 III (B) for more information.
Therefore, the combination of Liu as modified in view of Judd, Ware and Nakayama teaches wherein the multiplier comprises a NOR gate.
Regarding claim 27, Liu as modified in view of Judd and Ware teaches all the limitations of claim 18 as stated above.
Liu does not explicitly teach wherein the multiplier comprises a NOR gate.
However, on the same field of endeavor, Nakayama discloses a multiplier that comprises a NOR gate (Nakayama Fig. 4 and page 601 left col middle “The input signals applied to the X and Y lines are inverted by the input buffer and partial products are produced by the NOR gates”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd and Ware using Nakayama by replacing the AND gates with NOR gates receiving inverted inputs for performing the multiplication and generating the partial products. One of ordinary skill in the art could have substituted the AND gates/circuitry for NOR gates/circuitry receiving inverted inputs, and the results of the substitution would have been predictable because both configuration generates partial products from a multiplication of the input signals. See MPEP 2141 III (B) for more information.
Therefore, the combination of Liu as modified in view of Judd, Ware and Nakayama teaches wherein the multiplier comprises a NOR gate.
Claims 13 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Judd and Ware as applied to claims 12 and 18 above, and further in view of Li et al. (US 20210279036 A1), hereinafter, Li
Regarding claim 13, Liu as modified in view of Judd and Ware teaches all the limitations of claim 12 as stated above.
Liu does not explicitly teach wherein the memory array includes a plurality of static random access memory (SRAM cells).
However, on the same field of endeavor, Li discloses a memory array for storing weights wherein the memory array includes a plurality of SRAM cells (Li Fig. 4A and paragraph [0055] “the multi-bit multiplier circuit 400 includes circuitry ( e.g., AND gates) for multiplying, in a sequential fashion, each of the bits (e.g. bits X0, X1, X2, X3) of the bit stream (provided via X activation) by values (e.g., weight parameters) stored in the SRAM array 402”; Figs. 7A-7C and paragraph [0061]; plurality of SRAM cells – SRAM cells 708 in the SRAM array 420).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd and Ware using Li by replacing the RRAM cells with SRAM cells for storing the set of multipliers. One of ordinary skill in the art could have substituted the RRAM cells for SRAM cells, and the results of the substitution would have been predictable because both memory cells have the same function of storing data used in performing a multiplication operation. See MPEP 2141 III (B) for more information.
Therefore, the combination of Liu as modified in view of Judd, Ware and Li teaches wherein the memory array includes a plurality of static random access memory (SRAM cells).
Regarding claim 21, Liu as modified in view of Judd and Ware teaches all the limitations of claim 18 as stated above.
Liu does not explicitly teach wherein the memory array includes a plurality of static random access memory (SRAM cells).
However, on the same field of endeavor, Li discloses a memory array for storing weights wherein the memory array includes a plurality of SRAM cells (Li Fig. 4A and paragraph [0055] “the multi-bit multiplier circuit 400 includes circuitry ( e.g., AND gates) for multiplying, in a sequential fashion, each of the bits (e.g. bits X0, X1, X2, X3) of the bit stream (provided via X activation) by values (e.g., weight parameters) stored in the SRAM array 402”; Figs. 7A-7C and paragraph [0061]; plurality of SRAM cells – SRAM cells 708 in the SRAM array 420).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd and Ware using Li by replacing the RRAM cells with SRAM cells for storing the set of multipliers. One of ordinary skill in the art could have substituted the RRAM cells for SRAM cells, and the results of the substitution would have been predictable because both memory cells have the same function of storing data used in performing a multiplication operation. See MPEP 2141 III (B) for more information.
Therefore, the combination of Liu as modified in view of Judd, Ware and Li teaches wherein the memory array includes a plurality of static random access memory (SRAM cells).
Claims 11, 24 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Chih as applied to claims 8, 1 and 18 above, and further in view of Liu.
Regarding claim 11, Chih teaches all the limitations of claim 8 as stated above.
Chih does not explicitly teach wherein the multiplier comprises an AND gate.
However, on the same field of endeavor, Liu discloses a multiplier that comprises an AND gate (Liu Figs. 5-6 and col 6 lines 6-10; AND gate – AND circuitry).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Chih using Liu by replacing the NOR gates with AND gates for performing the bit-wise multiplication and generating the partial products. One of ordinary skill in the art could have substituted the NOR gates/circuitry for AND gates/circuitry, and the results of the substitution would have been predictable because both logic gates are able to generate partial products for performing a multiplication of two input operands. See MPEP 2141 III (B) for more information.
Therefore, the combination of Chih as modified in view of Liu teaches wherein the multiplier comprises an AND gate.
Regarding claim 24, Chih teaches all the limitations of claim 1 as stated above.
Chih does not explicitly teach wherein performing the bit-serial multiplication includes a logic AND operation.
However, on the same field of endeavor, Liu discloses performing a bit-serial multiplication using a logic AND operation (Liu Figs. 5-6 and col 6 lines 6-10).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Chih using Liu by replacing the NOR gates with AND gates for performing the bit-wise multiplication and generating the partial products. One of ordinary skill in the art could have substituted the NOR gates/circuitry for AND gates/circuitry, and the results of the substitution would have been predictable because both logic gates are able to generate partial products for performing a multiplication of two input operands. See MPEP 2141 III (B) for more information.
Therefore, the combination of Chih as modified in view of Liu teaches wherein performing the bit-serial multiplication includes a logic AND operation.
Regarding claim 28, Chih teaches all the limitations of claim 18 as stated above.
Chih does not explicitly teach wherein the multiplier comprises an AND gate.
However, on the same field of endeavor, Liu discloses a multiplier that comprises an AND gate (Liu Figs. 5-6 and col 6 lines 6-10; AND gate – AND circuitry).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Chih using Liu by replacing the NOR gates with AND gates for performing the bit-wise multiplication and generating the partial products. One of ordinary skill in the art could have substituted the NOR gates/circuitry for AND gates/circuitry, and the results of the substitution would have been predictable because both logic gates are able to generate partial products for performing a multiplication of two input operands. See MPEP 2141 III (B) for more information.
Therefore, the combination of Chih as modified in view of Liu teaches wherein the multiplier comprises an AND gate.
Claims 1, 3, 22 and 24-26 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Li, Judd and Ware.
Regarding claim 1, Liu teaches a computing method configured to perform bit-serial multiplication in a compute-in memory (CIM) device, the computing method comprising (Liu Fig. 5 and col 5 lines 37-39 “Referring now to FIG. 5, a memory device configured to compute matrix dot products, in accordance with aspects of the present technology, is shown”; compute-in memory (CIM) device – memory device; Fig. 6 step 620 and col 6 lines 6-10 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505”; bit-serial multiplication – AND operation between one bit in a row of 510 and the row of 505; Fig. 4):
determining at least one input according to a type of an application (Liu Figs. 4-6 and col 5 lines 64-65 “At 610, a set of multiplicands can be loaded into the second storage location 510”; at least one input – at least one multiplicand of the set of multiplicands; col 5 lines 1-15 “the second matrix X loaded in the input registers is illustrated … in neural network applications, the matrix elements are commonly 8-bit values”; type of an application - neural network application);
receiving the at least one input by an input driver circuit (Liu Figs. 1-5; input driver - input registers 120/second storage location 510; col 5 line 64 to col 6 line 6 “a set of multiplicands can be loaded into the second storage location 510 … A given bit position of the given row in the second storage location 510 … can be output to the logic AND circuitry 520, 525”);
determining at least one weight (Liu Figs. 4-6 and col 5 lines 63-64 “At 605, a set of multipliers can be loaded into the first storage location 505”; at least one weight – at least one multiplier of the set of multipliers);
performing the bit-serial multiplication based on the at least one input and the at least one weight, by a multiply circuit, from a most significant bit (MSB) of the at least one input to a least significant bit (LSB) of the at least one input to obtain a plurality of partial-products (Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows. After rows of the first and second storage locations 505, 510 are incrementally addressed by the address generator 515, the bit position in the second storage location 510 for input to the logic AND circuitry 520, 525 can be shifted by one bit to the left when processing from the most-significant-bit to the least-significant-bit, and the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630 … The processes at 615-630 can be repeated for each bit position of the second storage location 510. After the process at 615-630 are repeated for each bit position of the second storage location 510, the accumulated value in the one or more accumulators 530, 535 can be output as the matrix dot product of the set of multiplier and the set of multiplicands”; multiply circuit – AND circuitry; plurality of partial-products – output of the bitwise AND each time step 620 is repeated).
(Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows”; first partial-sum – accumulated value at 625);
left-shifting the first partial-sum one bit (Liu Figs. 5-6 and col 6 lines 13-20 “After rows of the first and second storage locations 505, 510 are incrementally addressed by the address generator 515 … the contents of the one or more accumulators 530, 535 can shifted one bit to the left, at 630”; partial-sum – accumulated value that is shifted at 630);
(Liu Figs. 5-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows … The processes at 615-630 can be repeated for each bit position of the second storage location 510”;
adding the left-shifted first partial-sum and the second one of the plurality of partial products by an adder circuit to obtain a second partial-sum (Liu Figs. 5-6 and col 6 lines 10-13 “At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows”; adder circuit – accumulators 530, 535; second partial-product - output of the AND circuitry when processing the next bit after the MSB).
; and
and outputting the result, by the adder circuit (Liu Figs. 5-6 and col 6 lines 28-32 “the accumulated value in the one or more accumulators 530, 535 can be output as the matrix dot product of the set of multiplier and the set of multiplicands, at 635”).
Liu does not explicitly teach determining at least one weight according to a training result or a configuration of a user; storing a first partial-sum based on a first one of the plurality of partial products in a first register; left-shifting the first partial-sum one bit by a shifter circuit; storing a second one of the plurality of partial products in a second register; and storing an output of the adder circuit in the first register having an output terminal operably connected to an input terminal of the shifter circuit.
However, on the same field of endeavor, Li discloses performing a bit-serial multiplication in a CIM device wherein one of the input operands are weight data that are determined according to a training result (Li paragraph [0041] “To adjust the weights, a learning algorithm may compute a gradient vector for the weights. The gradient may indicate an amount that an error would increase or decrease if the weight were adjusted”; paragraph [0051] “the weights may be stored in an SRAM configured for in-memory computations”; paragraph [0055] “the multi-bit multiplier circuit 400 includes circuitry ( e.g., AND gates) for multiplying, in a sequential fashion, each of the bits (e.g. bits XO, X1, X2, X3) of the bit stream (provided via X activation) by values (e.g., weight parameters) stored in the SRAM array 402”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu using Li and configure the values stored in the memory 505 as weights that are obtained according to a result of a learning algorithm in order to implement a deep convolutional network (DCN) that is trained to identify or recognize visual features from an input image (Li paragraph [0036]).
Therefore, the combination of Liu as modified in view of Li teaches determining at least one weight according to a training result or a configuration of a user.
Liu as modified in view of Li does not explicitly teach storing a first partial-sum based on a first one of the plurality of partial products in a first register; left-shifting the first partial-sum one bit by a shifter circuit; storing a second one of the plurality of partial products in a second register; and storing an output of the adder circuit in the first register having an output terminal operably connected to an input terminal of the shifter circuit.
However, on the same field of endeavor, Judd discloses storing a first partial-sum based on a first one of a plurality of partial products in a first register; left-shifting the first partial-sum one bit by a shifter circuit; and storing an output of an adder circuit in the first register having an output terminal operably connected to an input terminal of the shifter circuit (Judd section V.C and Fig. 4 adder – adder; a first register – register downstream of the adder; shifter – shifter below the register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu using Judd and include a shifter to perform the left shift operation in step 630. As discussed, Liu discloses performing a left shift operation while Judd discloses a shifter which is normally used for performing shifting operations. Therefore, one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination yielded nothing more than predictable results to one of ordinary skill in the art. The predictable result is a device that includes a shifter for performing the shift operations. See MPEP 2141 III (A) for more information. Further, include a register for storing the output of the adder that is provided to the shifter in order to implement a processing pipeline as registers are normally used to separate pipeline stages (Ware paragraph [0039]).
Therefore, the combination of Liu as modified in view of Li, Judd and Ware teaches storing a first partial-sum based on a first one of the plurality of partial products in a first register; left-shifting the first partial-sum one bit by a shifter circuit; and storing an output of the adder circuit in the first register having an output terminal operably connected to an input terminal of the shifter circuit.
Liu as modified in view of Li, Judd and Ware does not explicitly teach storing a second one of the plurality of partial products in a second register.
However, on the same field of endeavor, Ware discloses storing an output of a multiplier in a second register (Ware Fig. 2D and paragraph [0062] “In each execution cycle, the filter weight value (Fkl value) in the D_r[p] register is multiplied by the Dijk value in the D_i[p] register, via multiplier circuitry, and the result is output to the MULT_r[p] register. In the next pipeline cycle this product (i.e., D*F value) is added to the Yijl accumulation value in the MAC_r[p−1] register (in the previous multiplier-accumulator circuit) and the result is stored in the MAC_r[p] register”; second register - MULT_r[p] register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Li and Judd and generalize the teaching of Ware by configuring the memory device to include a second register downstream of the AND circuitry and upstream of the accumulator for storing each partial product output of the AND circuitry that is subsequently added/accumulated in order to implement a processing pipeline as registers are normally used to separate pipeline stages (Ware paragraph [0039]).
Therefore, the combination of Liu as modified in view of Li, Judd and Ware teaches storing a second one of the plurality of partial products in a second register.
Regarding claim 3, Liu as modified in view of Li, Judd and Ware teaches all the limitations of claim 1 as stated above. Further, Liu as modified in view of Li, Judd and Ware teaches
wherein the at least one input comprises a plurality of inputs (Liu Figs. 4-5 and col 5 lines 64-64 “a set of multiplicands can be loaded into the second storage location 510”; plurality of inputs – X0 to Xr-1 or set of multiplicands), and
wherein performing the bit-serial multiplication comprises: determining the plurality of first partial-products for each of the plurality of inputs (Liu Figs. 4-6 and col 6 lines 6-31 “At 620, the logic AND circuitry 520, 525 can perform a bitwise AND of the given bit position of the given row of the second storage location 510 and the given row of the first set of storage locations 505. At 625, the accumulators 530, 535 can be configured to accumulate the output of the logic AND circuitry 520, 525 for the given bit position, and the processes of 615-625 can be repeated for the plurality of rows …The processes at 615-630 can be repeated for each bit position of the second storage location 510”).
Regarding claim 22, Liu as modified in view of Li, Judd and Ware teaches all the limitations of claim 1 as stated above.
Liu does not explicitly teach further comprising outputting the result by the adder circuit to a third register.
However, on the same field of endeavor, Ware discloses outputting the result of an adder circuit to a third register (Liu Fig. 2D and paragraph [0073] “The output data are parallel loaded from the accumulation registers Y (here, 64) into the MAC_SO registers (here, 64)”; third register - MAC_SO [p] register).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective
filling date of the claimed invention, to modify Liu in view of Li, Judd and Ware and configure the memory device to include a third register downstream of the adder for storing the final result (Ware paragraph [0073]) such that it can be output in a next pipeline execution cycle.
Therefore, the combination of Liu as modified in view of Judd and Ware teaches further comprising outputting the result by the adder circuit to a third register.
Regarding claim 24, Liu as modified in view of Li, Judd and Ware teaches all the limitations of claim 1 as stated above. Further, Liu as modified in view of Li, Judd and Ware teaches wherein performing the bit-serial multiplication includes a logic AND operation (Liu Fig. 5 and col 5 lines 41-42 and col 6 lines 6-10).
Regarding claim 25, Liu as modified in view of Li, Judd and Ware teaches all the limitations of claim 1 as stated above.
Liu does not explicitly teach wherein the memory cell is a static random access memory (SRAM) memory cell.
However, on the same field of endeavor, Li discloses a memory cell that is a static random access memory (SRAM) cell (Li Fig. 4A and paragraph [0055] “the multi-bit multiplier circuit 400 includes circuitry ( e.g., AND gates) for multiplying, in a sequential fashion, each of the bits (e.g. bits X0, X1, X2, X3) of the bit stream (provided via X activation) by values (e.g., weight parameters) stored in the SRAM array 402”; Figs. 7A-7C and paragraph [0061] “The SRAM cell 708 may be part of the SRAM array 402 including an array of word lines (WLs)”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd and Ware using Li by replacing the RRAM cells with SRAM cells for storing the set of multipliers. One of ordinary skill in the art could have substituted the RRAM cells for SRAM cells, and the results of the substitution would have been predictable because both memory cells have the same function of storing data used in performing a multiplication operation. See MPEP 2141 III (B) for more information.
Therefore, the combination of Liu as modified in view of Judd, Ware and Li teaches wherein the memory cell is a static random access memory (SRAM) memory cell.
Regarding claim 26, Liu as modified in view of Li, Judd and Ware teaches all the limitations of claim 1 as stated above. Further, Liu as modified in view of Li, Judd and Ware teaches further comprising:
storing a third one of the plurality of partial products in the second register (Ware Fig. 2D and paragraph [0062]);
left-shifting the second partial sum one bit by the shifter circuit (Judd Fig. 4; Liu col 5 lines 19-20);
adding the left-shifted second partial-sum and the third one of the plurality of partial products by the adder circuit to obtain a third partial-sum (Liu Figs. 5-6 and col 6 lines 25-26); and
storing the output of the adder circuit in the first register (Judd Fig. 4).
Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Li, Judd and Ware as applied to claim 1 above, and further in view of Nakayama.
Regarding claim 23, Liu as modified in view of Judd and Ware teaches all the limitations of claim 1 as stated above.
Liu does not explicitly teach wherein performing the bit-serial multiplication includes a logic NOR operation.
However, on the same field of endeavor, Nakayama discloses a multiplier that comprises a NOR gate (Nakayama Fig. 4 and page 601 left col middle “The input signals applied to the X and Y lines are inverted by the input buffer and partial products are produced by the NOR gates”).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Liu in view of Judd and Ware using Nakayama by replacing the AND gates with NOR gates receiving inverted inputs for performing the multiplication and generating the partial products. One of ordinary skill in the art could have substituted the AND gates/circuitry for NOR gates/circuitry receiving inverted inputs, and the results of the substitution would have been predictable because both configuration generates partial products from a multiplication of the input signals. See MPEP 2141 III (B) for more information.
Therefore, the combination of Liu as modified in view of Li, Judd, Ware and Nakayama teaches wherein performing the bit-serial multiplication includes a logic NOR operation.
Response to Arguments
In view of amendments made, the objection the specification and the claims has been withdrawn.
In view of amendments made, the 35 U.S.C. 112(b) rejection of claim 17 has been withdrawn. However, the amendments made raises new 35 U.S.C. 112(b) issues as discussed above.
Applicant’s arguments, see remarks page 8-13, filed 03/24/2026, with respect to the 35 U.S.C. 103 rejection of claims 8-13, 17-18 and 20-21have been fully considered but they are not persuasive.
Claim 18 recites “a first register having an input terminal operably connected to an output of the adder, the first register configured to store a first partial-sum; a shifter having an input terminal operably connected to an output of the first register and an output terminal operably connected to a first input terminal of the adder, the shifter being configured to left-shift each first partial sum received from the first register one bit and output the left-shifted first partial sum to the adder”.
Applicant argues the following:
1.) the portion of Liu cited in the office action appears to refer to adding the contents of the accumulators 530 and 535 to one another, rather than adding sequential partial sums, where the preceding partial sum has been left-shifted one bit as recited in claim 18. If the contents of the accumulators are left shifted after the rows of the first and second storage locations 505, 510 have been incrementally addressed, this would not appear to refer to adding partial sums output by a multiplier. If the multiplication results of the multiplicands 510 and the multipliers 505 for all of the rows are added after determining partial products for each row ("After rows of the first and second storage locations 505,510 are incrementally addressed ... ), the resulting partial sums would need to be shifted more than one bit to obtain a total sum from several partial sums. There are no details on this accumulation in operation 625 provided in Liu, and it does not appear to refer to left-shifting by one bit a first partial sum and adding this shifted partial sum to a subsequent partial sum.
Response: Examiner respectfully disagrees. In reference to Fig. 5, the storage locations include 12 rows. Step 615 refers to accessing the first bit (MSB) of the first row of 510 and the first row of 505 and bitwise ANDing the first bit (MSB) of the first row of 510 and the first row of 505 at 620. At 625, the output of the logic AND is accumulated. Since this is the first partial product, the accumulated value is equal to the first partial product. As shown in Fig. 6A, the process goes back to step 615 where the first bit (MSB) of the second row of 510 and the second row of 505 are accessed and bitwise ANDed at 620. Then the first partial product and the second partial product are accumulated at 625. This process repeats until all 12 rows are accessed. After the last partial product for the MSB (first partial sum) is accumulated, the processes proceeds to step 630 to left shift the first partial sum, and the process goes back to 615 and the same process is repeated by accessing the next bit of all rows of 510 and all rows of 505 at 620 as described in col 5 lines 6-31. Accumulators 530 and 535 do not their contents together. The content of accumulators 530 and 535 is the dot product of the multipliers and the multiplicand. This is shown in Fig. 4 where the content of accumulator 530 is the first element of B and the content of accumulator 535 is the first element of B. Further, as evidenced in Fig. 4, none of the values in column a0,1 is added to the values in column a0,0.
2.) Liu teaches away from the claim elements because Liu includes a “bit skipping mechanism” which would not correctly calculate the result.
Response: Examiner respectfully disagrees. Examiner is not relying on the embodiments of Liu that includes a bit skipping mechanism. Further, Liu does not teach away but rather providing another configuration. See MPEP 2123 II “Disclosed examples and preferred embodiments do not constitute a teaching away from a broader disclosure or nonpreferred embodiments.” Furthermore, as discussed above, the bit-shifting in 630 is after all the same bit position in all the rows 510 are accessed. Therefore, even if all the MSB of the multiplicand are zero for example, the process would still perform a left-shift of one bit.
3.) including the shift the shifter of Chiueh would not result in a device that operates in the manner recited in claim 18 because the modified device still would not have a shifter configured to left-shift an output of the first register one bit and an adder configured to add the outputs of the left-shifted first partial-sum and the second partial sum.
Response: Examiner respectfully disagrees. However, the rejection has been modified to replace the Chiueh reference with Judd and Fig. 4 of Judd clearly shows a one bit left shifter configured to left-shift an output of a first register one bit where the first register is connected to an adder and receives an output of the adder and, the adder is configured to add the output of the a-shifted partial-sum and a second partial sum.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Carlo Waje whose telephone number is (571)272-5767. The examiner can normally be reached 9:00-6:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached on (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Carlo Waje/Examiner, Art Unit 2182 (571)272-5767