DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
Regarding Claim 1:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
a first multiplier circuit configured to perform multiplication on a first input data with at least part of a first kernel coefficient in a floating-point mode or an integer mode to generate a first multiplied output. This limitation is directed to the abstract idea of a mathematical concept, as performing multiplication is analogous to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.)
a second multiplier circuit configured to perform multiplication on a second input data with a second kernel coefficient in parallel with the first multiplier circuit in the integer mode to generate a second multiplied output; This limitation is directed to the abstract idea of a mathematical concept, as performing multiplication is analogous to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.)
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application.
A multiply-accumulator circuit in a neural processor circuit, comprising: a plurality of multiplier circuits comprising. This limitation recites generic computer components such as circuit, which invokes a system merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application.
a first multiplier circuit This limitation recites generic computer components such as circuit, which invokes a system merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application.
a second multiplier circuit This limitation recites generic computer components such as circuit, which invokes a system merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application.
a first accumulator configured to store a first accumulator value determined by at least adding the first multiplied output; This limitation recites generic computer components such as accumulator, which invokes a system merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application.
a second accumulator configured to [store] a second accumulator value determined by at least adding the second multiplied output. This limitation recites generic computer components such as accumulator, which invokes a system merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application.
Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception.
A multiply-accumulator circuit in a neural processor circuit, comprising: a plurality of multiplier circuits comprising. This limitation invokes a circuit merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception.
a first multiplier circuit. This limitation invokes a circuit merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception.
a second multiplier circuit. This limitation invokes a circuit merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception.
a first accumulator configured to store a first accumulator value determined by at least adding the first multiplied output; This limitation invokes an accumulator merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception.
a second accumulator configured to [store] a second accumulator value determined by at least adding the second multiplied output. This limitation invokes an accumulator merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception.
Step 2A Prong Two and Step 2B:
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible.
Regarding Claim 2:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Claim 3 does not recite abstract ideas other than the ones recited at claim 1.
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application.
the first input data is of a first bit size, the first kernel coefficient is of a second bit size, the second input data is of a third bit size and the second kernel coefficient is of a fourth bit size. The limitation that the input data or a kernel coefficient is of a bit-size specifying the type of data used in performing the abstract idea constitutes insignificant extra-solution activity and therefore does not integrate the judicial exception into a practical application (MPEP 2106.05(g))
Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception.
the first input data is of a first bit size, the first kernel coefficient is of a second bit size, the second input data is of a third bit size and the second kernel coefficient is of a fourth bit size. This limitation specifying that the input data or a kernel coefficient is of a bit-size does not amount to significantly more than the abstract idea. The limitation merely describes the type of data used and represents well-understand, routine, and conventional activity in the field of data processing (MPEP 2106.05(d))
Step 2A Prong Two and Step 2B:
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible.
Regarding Claim 3:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Claim 3 does not recite abstract ideas other than the ones recited at claim 1.
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application.
the second multiplier circuit is inactive in the floating-point mode. This limitation recites the circuit is inactive in the floating-point mode. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and fails to integrate the judicial exception into a practical application.
Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception.
the second multiplier circuit is inactive in the floating-point mode. This limitation recites description of the circuit is inactive in the floating-point mode. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and therefore fails to amount to significantly more than the judicial exception.
Step 2A Prong Two and Step 2B:
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible.
Regarding Claim 4:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Claim 4 does not recite abstract ideas other than the ones recited at claim 1.
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application.
the first multiplied output and the second multiplied output are generated in a same cycle of the neural processor circuit. This limitation recites generating two outputs in same cycle of neural processor circuit. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and fails to integrate the judicial exception into a practical application.
Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception.
the first multiplied output and the second multiplied output are generated in a same cycle of the neural processor circuit. This limitation recites description of generating two outputs in same cycle of neural processor circuit. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and therefore fails to amount to significantly more than the judicial exception.
Step 2A Prong Two and Step 2B:
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible.
Regarding Claim 5:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
the adder-shifter circuit configured to perform a shift operation on a first added value derived from the first multiplied output. This limitation is directed to the abstract idea of a mathematical concept, as performing a shift operation on a value is analogous to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.)
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application.
an adder-shifter circuit coupled to the first multiplier circuit to receive the first multiplied output, the adder-shifter circuit coupled to the second multiplier circuit to receive the second multiplied output. This limitation recites an adder-shifter circuit coupled to the multiplier circuit to receive the output. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and fails to integrate the judicial exception into a practical application.
Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception.
an adder-shifter circuit coupled to the first multiplier circuit to receive the first multiplied output, the adder-shifter circuit coupled to the second multiplier circuit to receive the second multiplied output. This limitation recites an adder-shifter circuit coupled to the multiplier circuit to receive the output. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and therefore fails to amount to significantly more than the judicial exception.
Step 2A Prong Two and Step 2B:
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible.
Regarding Claim 6:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
the adder-shifter circuit does not perform a shift operation on a second added value derived from the second multiplied output. This limitation is directed to the abstract idea of a mathematical concept, as (not) performing a shift operation on a value is analogous to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.)
Regarding Claim 7:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
a third multiplier circuit configured to generate a third multiplied output and a fourth multiplier circuit configured to generate a fourth multiplied output, the first added value derived by at least adding the first multiplied output and the third multiplied output, and the second added value derived by at least adding the second multiplied output with the fourth multiplied output. This limitation is directed to the abstract idea of a mathematical concept, as generating multiplied output and deriving added value by multiplied outputs is analogous to mathematical calculation (see MPEP 2106.04(a)(2) I. C.)
Regarding Claim 8:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Claim 8 does not recite abstract ideas other than the ones recited at claim 1.
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application.
the first input data includes the most significant bits of an input data and the second input data includes bits of the input data other than the most significant bits. The limitation that the first input data includes the most significant bits of an input data specifying the type of data used in performing the abstract idea constitutes insignificant extra-solution activity and therefore does not integrate the judicial exception into a practical application (MPEP 2106.05(g))
Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception.
the first input data includes the most significant bits of an input data and the second input data includes bits of the input data other than the most significant bits. This limitation specifying that the first input data includes the most significant bits of an input data does not amount to significantly more than the abstract idea. The limitation merely describes the type of data used and represents well-understand, routine, and conventional activity in the field of data processing (MPEP 2106.05(d))
Step 2A Prong Two and Step 2B:
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible.
Regarding Claim 9:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
the third bit size and the fourth bit size are different. This limitation is directed to the abstract idea of a mathematical concept, as two bit-size are different is analogous to mathematical relationships (see MPEP 2106.04(a)(2) I. A.).
Regarding Claim 10:
Step 1 – Is the claim to a process, machine, manufacture, or composition of matter?
Yes
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites the abstract ideas of:
the first bit size is larger than the third bit size. This limitation is directed to the abstract idea of a mathematical concept, as one bit-size is lager than another bit-size is analogous to mathematical relationships (see MPEP 2106.04(a)(2) I. A.).
Regarding claims 11- 15
Claims 11 - 15 recites analogous limitations to claims 1 - 5 (respectively) and therefore they are rejected on the same grounds as claims 1 - 5.
Regarding claims 16- 19
Claims 16 - 19 recites analogous limitations to claims 7 - 10 (respectively) and therefore they are rejected on the same grounds as claims 7 - 10.
Regarding claim 20
Claim 20 recites analogous limitations to claims 1 (respectively) and therefore they are rejected on the same grounds as claims 1.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claim(s) 1 – 8, 10 – 17 and 19 - 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Elmer (US20210157549A1 Pub. Date 05/27/2021, by Elmer et al - hereinafter Elmer) in view of Pugh (US20210042087A1, Pub. Date 02/11/2021, by Pugh et al - hereinafter Pugh).
Referring to Claim 1, Elmer teaches:
A multiply-accumulator circuit in a neural processor circuit, comprising. See Elmer at [0041]:” For example, the systolic array 100 may be part of a neural network processor in a computer system.” And See Elmer at [0049]:” The multiplier 208, products 250 and 251, multiplexer 213, first multiplexer product 255, shared adder 210, sums 237 and 238, and multiplexer 215 form an example multiply accumulate datapath 209.” Examiner interprets a neural network processor in a computer system as equivalent to a neural processor circuit, and example multiply accumulate as the multiply-accumulator circuit as claimed.
a plurality of multiplier circuits comprising. See Elmer at [0038]:” For example, a PE may utilize a shared multiplier and separate adders, a PE may utilize separate multipliers and a shared adder, or a PE may utilize any combination of a shared/separate multiplier(s) with shared/separate adder(s).” Examiner interprets separate multipliers as equivalent to a plurality of multiplier circuits.
a first multiplier circuit configured to perform multiplication on a first input data with at least part of a first kernel coefficient in a floating-point mode or an integer mode to generate a first multiplied output. See Elemer at [0051]:” In some implementations, the weight 224 may belong to a set of weight values corresponding to a convolution filter.” Examiner interprets the weight of a set of weight values corresponding to a convolution filter as equivalent to the kernel coefficient since it is well known in the art that the weight values of a convolution filter are equivalent to kernel coefficients, as both represent the multiplicative parameters used in convolution operations. And see Elemer at [0065]:” The integer multiplier sub-circuit 1002 is configured to multiply an integer input data element 244 by an integer weight 246 to produce the integer product 250.” Also see Elemer:” The floating-point multiplier sub-circuit 1004 is configured to multiply a floating-point input data element 244 by a floating-point weight 246 to produce the floating-point product 251.” Examiner interprets sub-circuit 1002 configured to multiply an integer input data by an integer weight to produce the integer product as equivalent as performing multiplication on a first input data with at least part of a first kernel coefficient in an integer mode to generate a first multiplied output, and the operation of sub-circuit 1004 as equivalent as the operation of a floating-point mode as claimed. Thus, Elemer teaches the limitation.
a second multiplier circuit configured to perform multiplication on a second input data with a second kernel coefficient in parallel with the first multiplier circuit in the integer mode to generate a second multiplied output; As discussed above, Elmer teaches sub-circuit 1002 configured to multiply an integer input data by an integer weight to produce the integer product. Also, See Elmer at [0065]: “the integer multiplier sub-circuit 1002 can perform parallel multiplications on shorter integers, such as performing two parallel signed 9-bit integer multiplications.” Examiner interprets multiplier 1002 performing two different multiplications in parallel in the integer mode, which is as equivalent as a first multiplier and a second multiplier as claimed.
However, Elmer fails to teach: a first accumulator configured to store a first accumulator value determined by at least adding the first multiplied output; and a second accumulator configured to a second accumulator value determined by at least adding the second multiplied output.
Pugh teaches a first accumulator configured to store a first accumulator value determined by at least adding the first multiplied output; See Pugh at [0030]:” A typical MAC multiplies two or more products and adds the results to an accumulator. The MACs 130 and 135, in some example embodiments, provide additional functionality by allowing partial products to be summed and provided as an output before being added to the accumulator. Thus, the individual partial products, sums of partial products for a current multiplication, and an accumulation result across multiple multiplication cycles may all be accessed by use of the MACs 130 and 135. Though a single box is shown in FIG. 1 for MACs 130 and 135, in some example embodiments, multiple MACs of each type are used (e.g., two integer MACs and two floating - point MACs).” Pugh expressly teaches an embodiment having two integer MACs and explains that a MAC adds multiplication results to an accumulator. Pugh further discloses a registered accumulation path in which adders 760A-760B add an input value to an accumulated value, and multiplexer 780 supplies the selected addition result to register 790, identified in FIG. 7 as ACCUM_AB_REG. A person of ordinary skill in the art would have implemented each of Pugh’s two disclosed integer-MAC instances with this known registered accumulation path. In the Elmer-Pugh combination, the first instance receives the first parallel multiplied output and stores the resulting first accumulator value, while the second instance receives the second parallel multiplied output and stores the resulting second accumulator value. This predictably preserves independently accumulated results for Elmer’s two parallel integer multiplication lanes.
a second accumulator configured to a second accumulator value determined by at least adding the second multiplied output. As discussed above, Pugh teaches an embodiment regarding two MACs, which satisfies description of the second accumulator as the limitation claimed.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elmer with the above teachings of Pugh by the first and the second multipliers and the configurations, as taught by Elmer; the first and the second accumulators and the configurations, as taught by Pugh. The modification would have been obvious because one of ordinary skill in art would be motivated to treat multiple tiles as larger circuits and increase the input and output bandwidth, as suggested by Pugh. See Pugh at [Abstract]: “The tile may also fuse a memory circuit with the arithmetic circuits. Connections directly between multiple instances of the tile are also available, allowing multiple tiles to be treated as larger memories or arithmetic circuits. By using these connections, referred to as cascade inputs and outputs, the input and output bandwidth of the arithmetic circuit is further increased.”
Referring to Claim 2, Elmer-Pugh teaches the circuit of claim 1. However, Elmer fails to teach:
the first input data is of a first bit size, the first kernel coefficient is of a second bit size, the second input data is of a third bit size and the second kernel coefficient is of a fourth bit size.
Pugh teaches:
the first input data is of a first bit size, the first kernel coefficient is of a second bit size, the second input data is of a third bit size and the second kernel coefficient is of a fourth bit size. See Pugh at [0052]:” Each of the multipliers 520A-520H accepts eight bits of the A operand and eight bits of the B operand.” The findings and rationale for claim 1 are incorporated. Pugh's A-operand portions correspond to the input data, and its B-operand portions correspond to the kernel-coefficient data. Multipliers 520D and 520A therefore receive respective first and second input-data portions and respective first and second kernel-coefficient portions, each eight bits wide. Claim 2 assigns four bit-size designations but does not require any one size to differ from another. Pugh's four eight-bit operand sizes satisfy the additional limitation.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 2.
Referring to Claim 3, Elmer-Pugh teaches the circuit of claim 1. Elmer also teaches:
the second multiplier circuit is inactive in the floating-point mode. See Elmer at [0065]:” Non-shared parts of integer multiplier sub-circuit 1002 can be selectively disabled while the data type control signal 235 indicates that a non-integer data type is selected.” Examiner interprets sub-circuit 1002 can be selectively disabled while data type is non-integer as equivalent as 2nd multiplier circuit is inactive in the floating-point mode.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 3.
Referring to Claim 4, Elmer-Pugh teaches the circuit of claim 1. Elmer also teaches:
the first multiplied output and the second multiplied output are generated in a same cycle of the neural processor circuit. See Elemer at [0090]:” …For example, during a first systolic interval, the shared multiplier 208 can generate a first value for the integer product 250 and first value for the floating – point product 251, and the first multiplexer 213 can select one of the product values to be stored in the delay register 263.” Elemer discloses a shared multiplier 208, which can generate one integer product 250 and one floating point-product 251 during a first systolic interval. Also see Elemer at [0050]:” For example, a systolic interval may correspond to a clock cycle” Elemer also discloses a systolic interval may correspond to a clock cycle. Thus, Elemer teaches the shared multiplier 208 preforming the function of two multipliers and generating two products during the same cycle.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 4.
Referring to Claim 5, Elmer-Pugh teaches the circuit of claim 1. Elmer also teaches:
an adder-shifter circuit coupled to the first multiplier circuit to receive the first multiplied output, the adder-shifter circuit coupled to the second multiplier circuit to receive the second multiplied output, the adder-shifter circuit configured to perform a shift operation on a first added value derived from the first multiplied output. See Elemer at [0073]:” As shown in FIG. 2A and shown with additional detail in FIG. 10A, a shared adder 210 includes an integer adder sub - circuit 1012 and a floating - point adder sub – circuit 1014 that share a shared part 1016. The floating - point adder sub - circuit 1014 also includes adder exponent logic 1018.” See Elemer at [0076]:” The integer adder sub - circuit 1012 is configured to add multiplexer product 255 with the stored input partial sum 236 to produce the integer partial sum 237.” And see Elemer [0077]:” The floating - point adder sub - circuit 1014 is con figured to add multiplexer product 255 with the stored input partial sum 236 to produce the floating - point partial sum 238. ... The floating - point multiplication can also include adder exponent logic 1018 for computing an exponent of the floating - point partial sum and for shifting the significands to align.” Elmer discloses the shared adder 210 including an integer adder sub circuit 1012 and a floating - point adder sub circuit 1014, that both sub circuits configured to add multiplexer product 255, which is as equivalent as the adder-shifter circuit coupled to the first and second multiplier circuits as claimed. Also, see Elmer at [Fig. 10A]: in the integer adder 1012, no shifting operation; in the float-point adder 1014, adder exponent logic 1018 computing the shift operation. Thus, Elmer teaches both claim 5 and claim 6.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 5.
PNG
media_image1.png
200
400
media_image1.png
Greyscale
Referring to Claim 6, Elmer-Pugh teaches the circuit of claim 1. Elmer also teaches:
the adder-shifter circuit does not perform a shift operation on a second added value derived from the second multiplied output. As discussed in claim 5, Elmer teaches one adder-shifter circuit performing a shift operation and another one does not.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 6.
Referring to Claim 7, Elmer-Pugh teaches the circuit of claim 1. Elmer fails to teach:
a third multiplier circuit configured to generate a third multiplied output and a fourth multiplier circuit configured to generate a fourth multiplied output, the first added value derived by at least adding the first multiplied output and the third multiplied output, and the second added value derived by at least adding the second multiplied output with the fourth multiplied output.
However, Pugh teaches:
a third multiplier circuit configured to generate a third multiplied output and a fourth multiplier circuit configured to generate a fourth multiplied output, the first added value derived by at least adding the first multiplied output and the third multiplied output, and the second added value derived by at least adding the second multiplied output with the fourth multiplied output. See Pugh at [Fig. 5] and [0054]:” With respect to the first operation mode, each of the eight multipliers 520A - 520H performs an eight - bit multiplication using a different portion of the operands A and B as inputs. The results of the eight multiplications are pair wise summed by the adders 530A - 530D.” Pugh discloses eight multiplier circuits configured to generate eight multiplied outputs, pairing and summing by four adders, which satisfies the configurations including two adders and four multipliers as claimed.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 7.
PNG
media_image2.png
200
400
media_image2.png
Greyscale
Referring to Claim 8, Elmer-Pugh teaches the circuit of claim 1. Elmer fails to teach:
the first input data includes the most significant bits of an input data and the second input data includes bits of the input data other than the most significant bits.
Pugh teaches:
the first input data includes the most significant bits of an input data and the second input data includes bits of the input data other than the most significant bits. See Pugh at [0055]:” The larger operands are divided into two portions, high and low, and organized as follows, wherein AH represents the high portion of the A operand, AL represents the low portion of the A operand, BH represents the high portion of the B operand, and BL represents the low portion of the B operand.” Also see Pugh at [0056]:” Thus, in the second operation mode, the multiplier 520D multiplies BL with AH and the multiplier 520C multiplies BH with AL. … The multiplier 520A multiples BL with AL.” Pual discloses dividing a large operand into two portions, AH as the high portion and AL as the low portion, which is as equivalent as the first input data includes the most significant bits and the second input data includes bits of the data other than the most significant bits. Further, BL and AL will be inputted into multiplier 520D as the input data. Thus, Pual teaches the limitation.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 8.
Referring to Claim 10, Elmer-Pugh teaches the circuit of claim 1. Elmer fails to teach:
the first bit size is larger than the third bit size.
However, Pugh teaches:
the first bit size is larger than the third bit size. See Pual at [0051]:” FIG. 5 is a diagrammatic view of a portion 500 of a multiple mode arithmetic circuit, according to some example embodiments . The portion 500 comprises registers 510A , 510B , 510C , 510D , 510E , 510F , 510G , 510H , 5101 , 510J , 510K , 510L , 510M , 510N , 5100 , and 510P ; multipliers 520A , 520B , 5200 , 520D , 520E , 520F , 520G , and 520H ; adders 530A , 530B , 530C , 530D, 550A , 550B , and 560 : multiplexers 540A , 540B , 540C , 540D , and 570 ; and stage 2 delay registers 580.” Also, see Pual at [0056]:” The particular value of “SIZE” is implementation - dependent. 8 bits is used by way of example herein. Additionally, the process of combining four multipliers to operate as a larger multiplier may be repeated, so that four larger multipliers (16 original multipliers) are used to form an even larger multiplier. Thus, in an example embodiment, the MLP 115 provides 64 4 - bit multipliers, 16 8 - bit multipliers, 4 16 - bit multipliers, or one 32 - bit multiplier.” Pual discloses portion 500 comprises registers, multipliers and adders, which can be combined as different number of bit size sub-multipliers. Examiner interprets the first bit size is 32-bit, and second bit size is 16-bit, and 32-bit is larger than 16-bit size. Also, both 32-bit and 16-bit input are satisfied with the input rule of portion 500. Thus, Pual teaches the limitation.
The same motivation that was utilized for combining Elmer with Pugh as set forth in claim 1 is equally applicable to claim 10.
Referring to Claims 11 - 15, these claims are rejected on the same basis as claims 1 - 5, mutatis mutandis, since they are analogous claims.
Referring to Claims 16 - 17, these claims are rejected on the same basis as claims 7 - 8, mutatis mutandis, since they are analogous claims.
Referring to Claim 19, these claims are rejected on the same basis as claim 10, mutatis mutandis, since they are analogous claims.
Referring to Claim 20, these claims are rejected on the same basis as claim 1, mutatis mutandis, since they are analogous claims.
Claim(s) 9 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Elmer - Pugh in view of Yoon (US20230015148A1, Pub. Date 01/19/2023, by Yoon et al - hereinafter Yoon).
Referring to Claim 9, Elmer-Pugh teaches the circuit of claim 1. Elmer-Pugh fails to teach:
the third bit size and the fourth bit size are different.
However, Yoon teaches:
wherein the third bit size and the fourth bit size are different. See Yoon at [0023]:” … In particular, the multiplier 158 may receive as inputs the scalar values, a and b, from the flip/ flops 154 and 156, respectively, and may multiply these scalar values.” See Yoon at [0046]:” The precision of what is multiplied by the multiplicand, the multiplier, such as the scalar value b of the matrix B, may affect the height of the carry save adder tree reduction. … Therefore, in such examples, the precision of the scalar value, a, of the matrix A may not impact the latency of computations as significantly as the precision of the scalar value, b, of the matrix B. Thus, in some examples, asymmetric precision may be used in the values/numbers input to a systolic array, and the higher precision may be used for the matrix A rather than the matrix B. In these examples, the matrix that includes higher precision numbers may be defined to be the matrix A. For example, 16-bit or 32-bit integers may be used for the matrix A, while the matrix B may include 8-bit integers.” Yoon discloses the multiplier 158 receiving as inputs the scalar values a of matrix A and b of matrix B with different precision such as 16-bit, 32-bit, or 8-bit. Thus, Yoon teaches the bit size of two input to the multiplier are different, which is as equivalent as the third bit size and the fourth bit size are different as claimed.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elmer-Pugh with the above teachings of Yoon by the multiply-accumulator circuit in a neural processor circuit, as taught by Elmer-Pugh; the third bit size and the fourth bit size are different, as taught by Yoon. The modification would have been obvious because one of ordinary skill in art would be motivated to be more efficient and optimized for performing matrix multiplication include less hardware and more energy efficient by Yoon. See Yoon at [0044]: “In the MAC unit 450, the multiplier and adder conventionally found in a MAC unit may be fused together to produce an enhanced MAC unit design. The enhanced MAC unit design may be more efficient and optimized for performing matrix multiplication, may include less hardware, and may be more energy efficient when compared to a conventional MAC unit design. The enhanced MAC unit design may include these and other advantages when it is used for matrix multiplication in a systolic array, such as those used for accelerators for DNNs.”
Referring to Claim 18, these claims are rejected on the same basis as claim 9, mutatis mutandis, since they are analogous claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIAYUE MA whose telephone number is (571)272-9658. The examiner can normally be reached between 9 am to 5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Jiayue Ma/
Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126