DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
The present application, 18111846, filed 02/20/2023 Claims Priority from Provisional Application 63312141, filed 02/21/2022.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/04/2024 and 09/12/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 501 mentioned in paragraph [0058].
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The specification is objected to as failing to provide proper antecedent basis for the claimed subject matter. See 37 CFR 1.75(d)(1) and MPEP § 608.01(o). Correction of the following is required:
A. The “means for” receiving …; “means for” generating …; and “means for” accumulating in claim 20 does not follow the nomenclature of the specification. Therefore, there is no clear support or explicit antecedent basis in the description for the corresponding structure, material, or acts for performing the claimed functions.
Claim Objections
Claim 12 is objected to under 37 C.F.R. 1.71(a) which requires “full, clear, concise, and exact terms” as to enable any person skilled in the art or science to which the invention or discovery appertains, or with which it is most nearly connected, to make and use the same. The following should be corrected.
In claim 12 line 2, “the during” should read “
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5 and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 5 recites “wherein the multiplier circuit within at least one of the MAC circuits within the given one of the MAC units is coupled to output a carry value to the multiplier circuit of at least one other of the MAC circuits within the given one of the MAC units”. An adder normally outputs a sum and carry value while a multiplier normally outputs a product value. Therefore, it is unclear how the multiplier circuit outputs a carry value because this is not a result that is normally generated by a multiplier circuit. Further clarification is required. Claim 15 recites substantially the same limitations and is rejected for the same reason.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 11 and 20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1 and 22 of copending Application No. 18215993 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because claims 1, 11 and 20 under examination are anticipated by claims 1 and 22 of the reference application. Every limitation in the application under examination claims are recited in the conflicting reference application claim as shown in the table below.
18111846
18215993
1. An integrated circuit device comprising:
1. An integrated circuit device comprising: first, second and third tensor processing units (TPUs), each TPU having:
a plurality of broadcast data paths;
a plurality of broadcast data paths;
a weighting-value memory; and
a weighting-value memory;
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:
a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;
a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;
a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;
a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products.
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products
Mapping for the other claims are not being shown for purposes of brevity.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
Claims 1, 2, 5, 11, 12, 15 and 20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 9 and 10 of copending Application No. 18144288 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because claims 1, 2, 5, 11, 12, 15 and 20 under examination are anticipated by claims 1, 9, 10, 1, 9, 10 and 1 respectively of the reference application. Every limitation in the application under examination claims are recited in the conflicting reference application claims respectively as shown in the table below.
18111846
18144288
1. An integrated circuit device comprising:
1. An integrated circuit device comprising:
a plurality of broadcast data paths;
a plurality of broadcast data paths;
a weighting-value memory; and
a weighting-value memory;
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:
a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;
a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;
a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;
a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products.
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products;
Mapping for the other claims are not being shown for purposes of brevity.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
Claims 1-4, 6, 11-14, 16 and 20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 4-7, 9-13, and 15-19 of copending Application No. 18371247 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because claims 1-4, 6, 11-14, 16 and 20 under examination are anticipated by claims 1, 4-7, 9-13, and 15-19 of the reference application. Every limitation in the application under examination claims are recited in the conflicting reference application claims respectively as shown in the table below.
18111846
18371247
1. An integrated circuit device comprising:
An integrated circuit device comprising: a plurality of tensor processing units (TPUs) to multiply, over a first plurality of timing cycles, an input data matrix having at least first and second dimensions with a filter-weight matrix having at least the first dimension and a third dimension to produce an output data matrix having at least the second and third dimensions, each TPU having:
a plurality of broadcast data paths;
a plurality of broadcast data paths;
a weighting-value memory; and
a weighting-value memory;
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths, each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having:
a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths;
a data input coupled to receive, during each timing cycle of the first plurality of timing cycles, an input data value via a respective one of the broadcast data paths;
a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;
a weighting-value input coupled to receive, during each timing cycle of the first plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths;
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles; and
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each timing cycle of the first plurality of timing cycles; and
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products.
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products;
Mapping for the other claims are not being shown for purposes of brevity.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-2, 9, 11-12 and 19-20 are rejected under 35 U.S.C. 102(a)(1) and (a)(2) as being anticipated by Fang et al. (WO 2019205617 A1), hereinafter Fang.
Regarding claim 1, Fang teaches
a plurality of broadcast data paths (Fang Fig. 4 and page 8 top; plurality of broadcast data paths – data lines from 410-413 to the computing units);
a weighting-value memory (Fang Fig. 4 and page 7 bottom; weighting-value memory - second repository set 414-417); and
a plurality of multiply-accumulate (MAC) units coupled in common to each of the broadcast data paths and coupled to receive respective weighting values from the weighting-value memory via respective weighting-value paths (Fang Fig. 4, page 7 middle and page 8 middle; plurality of multiply-accumulate (MAC) units – columns of the computing units; respective weighting values – multipliers; respective weighting-value paths - data lines from 414-417 to the computing units), each of the MAC units having a plurality of MAC circuits coupled respectively to the broadcast data paths, each of the MAC circuits within a given one of the MAC units having (Fang Fig. 4, page 7 middle and page 8 middle; plurality of MAC circuits – computing units in each column):
a data input coupled to receive, during each of a plurality of timing cycles, an input data value via a respective one of the broadcast data paths (Fang Fig. 4-5 data input – multiplicand input terminal; input data value – multiplicand; Figs. 7-10 and page 8 bottom paragraph; plurality of timing cycles – first to fourth clock cycles; page 8 top “the first repository set may broadcast N data to the N*N computing units in the first clock cycle”; page 3 second paragraph “in each clock cycle, all the computing units in the same row in the matrix of N*N receive the same first input data”);
a weighting-value input coupled to receive, during each of the plurality of timing cycles, a shared one of the weighting values via a shared one of the respective weighting-value paths (Fang Fig. 4-5 weighting-value input – multiplier input terminal; shared one of the weighting values – multiplier; Figs. 7-10; page 8 top “the second repository set may also broadcast N to the N*N computing units in the first clock cycle”; page 3 second paragraph “All computing units in the same column receive the same second input data”);
a multiplier circuit to generate a sequence of multiplication products by multiplying the input data value received during each of the plurality of timing cycles with the shared one of the weighting values received during each of the plurality of timing cycles (Fang Figs. 5 and 7-11 and page 8 bottom paragraph; multiplier circuit – multiplication unit 506; sequence of multiplication products – products); and
an accumulator circuit to accumulate a sum of constituent multiplication products within the sequence of multiplication products (Fang Fig. 5 and 7-11 and page 8 bottom paragraph; accumulator circuit – addition unit 507, 508; sum – addition result).
Regarding claim 2, Fang teaches all the limitations of claim 1 as stated above. Further, Fang teaches wherein each of the MAC circuits further comprises a data operand register, coupled between the data input and the multiplier circuit, to store the input data value received during each of the plurality of timing cycles and to output the data input value received during each of the plurality of timing cycles to the multiplier circuit (Fang Fig. 5 and page 8 bottom; data operand register – register 501/502).
Regarding claim 9, Fang teaches all the limitations of claim 1 as stated above. Further, Fang teaches wherein the plurality of MAC units each having a plurality of MAC circuits constitute a collective array of MAC circuits in which each column of the MAC circuits within the array constitutes a respective one of the MAC units (Fang Fig. 4 and page 8 middle “Maintaining a column connection with the computing unit matrix, and if the third repository set and the computing unit matrix maintain a column connection, each of the third repository sets and each column of the computing unit matrix The computing units are connected”).
Regarding claims 11-12 and 19, they are directed to a method practiced by the device of claims 1-2 and 9 respectively. All steps performed by the method of claims 11-12 and 19 would be practiced by the device of claims 1-2 and 9 respectively. Claims 1-2 and 9 analysis applies equally to claims 11-12 and 19 respectively.
Regarding claim 20, it is directed to an integrated circuit component comprising substantially the same limitations as the device of claim 1. Claim 1 analysis applies equally to claim 20.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3-4, 10 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Fang as applied to claims 1 and 9 above, and further in view of Dasgupta et al. (US 20200410038 A1) hereinafter Dasgupta.
Regarding claim 3, Fang teaches all the limitations of claim 1 as stated above.
Fang does not explicitly teach wherein the given one of the MAC units comprises a weighting-value register to store a respective one of the weighting values received via a respective one of the weighting-value paths.
However, on the same field of endeavor, Dasgupta discloses a weighting-value register to store a respective one of the weighting values received via a respective one of the weighting-value paths (Dasgupta Fig. 23 and paragraph [0169] weighting-value register – 2314, 2316 and/or 2318]).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Fang using Dasgupta and configure the circuit by removing the register 502 in each computing element and replacing them by a single register upstream to each computing element such that each computing element in a column is coupled to an output of a respective single register in order to reduce the circuit area replacing a plurality of input storage circuits (registers) with a single input storage circuit, and to reduce power in the clock tree by significantly decreasing the clock load (Dasgupta paragraph [0169]).
Therefore, the combination of Fang as modified in view of Dasgupta teaches wherein the given one of the MAC units comprises a weighting-value register to store a respective one of the weighting values received via a respective one of the weighting-value paths.
Regarding claim 4, Fang as modified in view of Dasgupta teaches all the limitations of claim 3 as stated above. Further, Fang as modified in view of Dasgupta teaches wherein the weighting-value input of each of the MAC circuits within the given one of the MAC units is coupled in common to the weighting-value register to receive, as the shared one of the weighting values, the respective one of the weighting values stored within the weighting-value register (Dasgupta Fig. 23 and paragraph [0169]).
Regarding claim 10, Fang teaches all the limitations of claim 9 as stated above. Further, Fang teaches
wherein: the plurality of MAC units comprises respective weighting-value registers to store the respective weighting values received from the weighting-value memory (Fang Fig. 5 and page 8 bottom paragraph; respective weighting-value registers – register 502 of each computing unit);
each row of the MAC circuits within the array is coupled in common to a respective one of the broadcast data paths (Fang Fig. 4 and page 7 bottom);
each of the columns of the MAC circuits is coupled in common, via the shared one of the respective weighting-value paths, to (Fang Fig. 4 and page 7 bottom).
Fang does not explicitly teach each of the columns of the MAC circuits is coupled in common, via the shared one of the respective weighting-value paths, to an output of a respective one of the weighting- value registers.
However, on the same field of endeavor, Dasgupta discloses a subset of a matrix multiplier array comprising MAC circuits is coupled in common, via a shared one of a respective weighting-value paths, to an output of a respective one of a plurality of weighting- value registers (Dasgupta Fig. 23 and paragraph [0169] weighting-value registers – 2314, 2316 or 2318]).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Fang using Dasgupta and configure the circuit by removing the register 502 in each computing element and replacing them by a single register upstream to each computing element such that each computing element in a column is coupled to an output of a respective single register in order to reduce the circuit area replacing a plurality of input storage circuits (registers) with a single input storage circuit, and to reduce power in the clock tree by significantly decreasing the clock load (Dasgupta paragraph [0169]).
Therefore, the combination of Fang as modified in view of Dasgupta teaches each of the columns of the MAC circuits is coupled in common, via the shared one of the respective weighting-value paths, to an output of a respective one of the weighting- value registers.
Regarding claims 13-14, they are directed to a method practiced by the device of claims 3-4 respectively. All steps performed by the method of claims 13-14 would be practiced by the device of claims 3-4 respectively. Claims 3-4 analysis applies equally to claims 13-14 respectively.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Fang as applied to claim 1 and 11 above, and further in view of Langhammer (US 8959137 B1).
Regarding claim 5, Fang teaches all the limitations of claim 1 as stated above.
Fang does not explicitly teach wherein the multiplier circuit within at least one of the MAC circuits within the given one of the MAC units is coupled to output a carry value to the multiplier circuit of at least one other of the MAC circuits within the given one of the MAC units.
However, on the same field of endeavor, Langhammer discloses a multiplier circuit within a larger multiplier block is coupled to output a carry value to another multiplier circuit within the larger multiplier block (Langhammer Figs. 3-4 and col 7 lines 6-37).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Fang using Dasgupta and configure the circuit such that
the multiplier unit within at least one of the computing units within a column is coupled to output a carry value to the multiplier unit of at least one other of computing units with the column in order to implement multiply-accumulate operations on higher-precision inputs using lower-precision circuitry by decomposing the multiplication operation into smaller sub-operations (Langhammer col 1 line 57 to col 2 line 3; col 3 line 60 to col 4 line 23).
Therefore, the combination of Fang as modified in view of Langhammer teaches wherein the multiplier circuit within at least one of the MAC circuits within the given one of the MAC units is coupled to output a carry value to the multiplier circuit of at least one other of the MAC circuits within the given one of the MAC units.
Regarding clam 15, it is directed to a method practiced by the device of claim 5. All steps performed by the method of claim 15 would be practiced by the device of claim 5. Claim 5 analysis applies equally to claim 15.
Claims 6-7 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Fang as applied to claim 1 above, and further in view of Temam et al. (US 20190050717 A1), hereinafter Temam.
Regarding claim 6, Fang teaches all the limitations of claim 1 as stated above.
Fang does not explicitly teach further comprising a plurality of shift registers coupled in common to the MAC units, each of the shift registers being coupled to an output of the accumulator circuit within a respective one of the MAC circuits within each of the MAC units.
However, on the same field of endeavor, Temam discloses a shift register coupled in common to a set of MAC circuits, the shift register being coupled to an output of respective one of the MAC circuits (Temam Fig. 5 and paragraphs [0082-0083]).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Fang using Temam and configure the circuit to include a respective shift register for each of the column of the matrix multiplier for temporarily storing the results and for sequentially outputting the result for each column while outputting the results for each row in parallel such that all the results can be written out to the output storage in a number of clock cycles corresponding to the number of rows (Temam Fig. 5 and paragraph [0083] and Fang page 10 middle S604) so that a next calculation cycle can be performed while the write operations are being performed.
Therefore, the combination of Fang as modified in view of Temam teaches further comprising a plurality of shift registers coupled in common to the MAC units, each of the shift registers being coupled to an output of the accumulator circuit within a respective one of the MAC circuits within each of the MAC units.
Regarding claim 7, Fang as modified in view of Temam teaches all the limitations of claim 6 as stated above. Further, Fang as modified in view of Temam teaches wherein the plurality of shift registers consists of a quantity of shift registers that matches quantity of broadcast data paths constituted by the plurality of broadcast data paths (Fang Fig. 4; Temam Fig. 5; see also claim 6 analysis where four shift registers are included one for each column which matches the number of broadcast datapaths (i.e., four data lines from 410-413 to the computing units). The motivation to combine is the same as claim 6.
Regarding claims 16-17, they are directed to a method practiced by the device of claims 6-7 respectively. All steps performed by the method of claims 16-17 would be practiced by the device of claims 6-7 respectively. Claims 6-7 analysis applies equally to claims 16-17 respectively.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Fang as applied to claims 1 and 11 above, and further in view of Manzo (US 20190026077 A1).
Regarding claim 8, Fang teaches all the limitations of claim 1 as stated above. Further, Fang teaches wherein each of the MAC circuits within each of the plurality of MAC units implements an operational pipeline in which the input data value and the shared one of the weighting values are (Fang Fig. 5 and page 8 bottom paragraph):
received within a given one of the MAC circuits during a first timing cycle of the plurality of timing cycles (Fang Fig. 5 and page 8 bottom paragraph);
multiplied to generate a multiplication product within the sequence of multiplication products during a (Fang Fig. 5 and page 8 bottom paragraph); and
accumulated into the sum of constituent multiplication products during a (Fang Fig. 5 and page 8 bottom paragraph).
Fang does not explicitly teach the input data value and the shared one of the weighting values are multiplied to generate a multiplication product within the sequence of multiplication products during a second timing cycle of the plurality of timing cycles; and accumulated into the sum of constituent multiplication products during a third timing cycle of the plurality of timing cycles.
However, on the same field of endeavor, Manzo discloses a multiply accumulate pipeline circuitry wherein multiplication is performed in a first clock cycle and accumulation is performed in a second clock cycle (Manzo Figs. 2-3 and paragraphs [0034-0035]).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, to modify Fang using Manzo and configure the system to perform the multiplication operation in a next clock cycle after storing the operands in the registers and perform the accumulation operation in a following clock cycle after the multiplication operation in order to implement a system capable of implementing a fused floating point multiply accumulate operation (Manzo paragraphs [0002, 0030]).
Therefore, the combination of Fang as modified in view of Manzo teaches wherein each of the MAC circuits within each of the plurality of MAC units implements an operational pipeline in which the input data value and the shared one of the weighting values are: received within a given one of the MAC circuits during a first timing cycle of the plurality of timing cycles; multiplied to generate a multiplication product within the sequence of multiplication products during a second timing cycle of the plurality of timing cycles; and accumulated into the sum of constituent multiplication products during a third timing cycle of the plurality of timing cycles, wherein the first, second and third timing cycles transpire sequentially.
Regarding clam 18, it is directed to a method practiced by the device of claim 8. All steps performed by the method of claim 18 would be practiced by the device of claim 8. Claim 8 analysis applies equally to claim 18.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Hanrahan et al. (US 20220012058 A1), Roy et al. (US 20210150313 A1) are related to performing matrix multiplication (multiply-accumulate operations) using an array of MAC circuits by broadcasting the input and weights.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Carlo Waje whose telephone number is (571)272-5767. The examiner can normally be reached 9:00-6:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Carlo Waje/Examiner, Art Unit 2151 (571)272-5767