Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 05/02/2023, 11/05/2024, 06/05/2025, 12/19/2025, and 12/19/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements mentioned above are being considered by the examiner, except for the references that have been lined through in the IDS submitted on 06/05/2025 and 12/19/2025 because the applicant has not provided a translated copy of the cited references.
Specification
The disclosure is objected to because of the following informalities:
The specification is objected to as failing to comply with 37 CFR 1.52(b)(6) because the paragraphs in the specification “individually and consecutively numbered using Arabic numerals, so as to unambiguously identify each paragraph.”
Appropriate correction is required.
Drawings
The drawings are objected to under 37 CFR 1.83(a). The drawings must show every feature of the invention specified in the claims. Therefore, the “skipping ineffectual terms mapped outside a defined accumulator width” of claim 12, and the “storing floating-point values in groups” of claim 18, must be shown or the feature(s) canceled from the claim(s). No new matter should be entered.
Furthermore, The drawings are objected to as failing to comply with 37 CFR 1.84(o) because figure 5 lack suitable descriptive legends. In the 1st block of Fig. 5, the box labeled “MAX” is unclear if it is meant to be the comparator. In the 2nd block of Fig. 5, it is unclear what the two boxes labelled “neg” are. In the 3rd block of Fig. 5, it is unclear what the unlabeled box receiving the input “eMAX” is, and in the 3rd block, and the box “eacc”, it is unclear if it is some sort of register, register file, buffer, memory, or something else.
Furthermore, the drawings are objected to as failing to comply with 37 CFR 1.84(a)(1) because Fig. 4-19 are blurry and difficult to read and requires solid black lines.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
“input module” as in claim 11;
“exponent module” as in claim 11;
“reduction module” as in claim 11;
“accumulation module” in claim 11.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
The “exponent module” as in claim 11, “to add exponents of the first data stream A and the second data stream B in pairs to produce product exponents, and to determine a maximum exponent using a comparator” is being interpreted as the structure of the box labeled “Exponent” of figure 5, including each component inside the “Exponent” labeled box, including all of the internal connections, and including the same inputs/outputs, and connection to the block of circuitry labeled “Shift&Reduce”. Furthermore, the “reduction module” as in claim 11, “to determine a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and use an adder tree to reduce the operands in the second data stream into a single partial sum” is being interpreted as the structure of the box labeled “Shift&Reduce” of figure 5, including each component inside the “Shift&Reduce” labeled box, including all of the internal connections, and including the same inputs/outputs, including inputs from connection to the (as interpreted above) Exponent module, and including outputs to a connecting block of circuitry labeled “Accumulation”. Furthermore, the “accumulation module” as in claim 11, “to add the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values, and to output the accumulated values” is being interpreted as the structure of the box labeled “Accumulation” of figure 5, including each component inside the “Accumulation” labeled box, including all of the internal connections, and including the same inputs/outputs, including inputs from connection to the (as interpreted above) reduction module, and including the same outputs as shown in figure 5.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 11-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
The claim limitation of “input module” of claim 11 invokes 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph. However, the written description fails to provide an adequate description of the structure, material, or acts to perform the claimed functions of these limitations. See rejection under 35 USC 112(b) below for further details as to the lack of structure. Claims 12-20 inherits the same deficiency as claim 11 based on dependence.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 1, claim 1 recites the limitations of: “A method for accelerating multiply-accumulate (MAC) floating-point units during training or inference of deep learning networks”. It is unclear the training from, “training or inference of deep learning networks”, is meant to be directed towards deep learning training, or some other type of training. For purposes of examination, the Examiner interprets the limitation to mean that the training is directed towards training deep learning networks.
Furthermore, claim 1 recites the limitation of: “A method for accelerating multiply-accumulate (MAC) floating-point units”, it is unclear if “floating-point units” is meant to be understood as floating point values, floating point vectors, floating point matrices, or if it is directed towards meaning accelerating circuitry which calculates multiply-accumulate operations on floating point values. For purposes of examination, the Examiner interprets the limitation to mean that it is directed towards meaning accelerating circuitry upon which they calculate multiply-accumulate operations on floating point values.
Furthermore, claim 1 recites the limitations of: “determining a number of bits by which each significand in the second data stream has to be shifted prior to accumulation”. The limitation of “to be shifted” is a limitation describing a future event potentially taking place, therefore it is unclear whether or not the shifting of each significand in the second data stream is positively recited or not.
Furthermore, claim 1 recites the limitations of: “adding product exponent deltas to the corresponding term in the first data stream”. The Examiner interprets the “product exponent” to be referring to the earlier limitation of, “adding exponents of the first data stream A and the second data stream B in pairs to produce product exponents”. An exponent delta is the difference between two exponents. It is unclear what the delta, difference, is between. The limitation of “product exponent deltas” implies that at least one of the values which is used in finding a delta is the product exponent values, however it is unclear if the second value in finding the delta is one from the first data stream, A, or from the second data stream, B, or from the maximum exponent, or from some other exponent value.
Furthermore, claim 1 recites the limitations of: “determining a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and using an adder tree to reduce the operands in the second data stream into a single partial sum”. It is unclear if the determining of a number of bits by which each significand in the second data stream has to be shifted prior to accumulation is determined by only “by adding product exponent deltas to the corresponding term in the first data stream” or by both “by adding product exponent deltas to the corresponding term in the first data stream” and by “using an adder tree to reduce the operands in the second stream into a single partial sum”.
Furthermore, claim 1 recites the limitations of: “determining a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and using an adder tree to reduce the operands in the second data stream into a single partial sum”. Claim 1 also recites the limitation of, “adding the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”. It is unclear if the “accumulation” of the limitation, “determining a number of bits by which each significand in the second data stream has to be shifted prior to accumulation” is referring to the second data stream being reduced to a single partial sum through the adder tree, or if the accumulation is meant to be referring to the accumulated values of the limitation, “adding the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”.
Furthermore, claim 1 recites the limitations of: “adding the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”. It is unclear if the corresponding aligned value is a value from the first data stream, A, the second data stream, B, the value of the maximum exponent, or some other value.
Furthermore, claim 1 recites the limitations of: “adding the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”. The metes and bounds of the limitation is indefinite because it merely recites a use, “using the maximum exponent”, without any active, positive steps delimiting how the use is actually practiced, see MPEP 2173.05(q).
Furthermore, claim 1 recites the limitations of: “the first data stream”. It is unclear if “the first data stream” is the same data stream as the “first input data stream A”, or some other data stream. For purposes of examination, the Examiner interprets “the first data stream” to be the same as the “first input data stream A”.
Furthermore, claim 1 recites the limitations of: “the second data stream”. It is unclear if “the second data stream” is the same data stream as the “second input data stream B” or if it is some other data stream. For purposes of examination, the Examiner interprets “the second data stream” to be the same as the “second input data stream B”.
Furthermore, claim 1 recites the limitations of: “adding product exponent deltas to the corresponding term in the first data stream”. It is unclear if it is meant to be understood as a plurality of exponent deltas are added to a corresponding term in the first data stream, or if it is meant to be understood as there are a plurality of product exponent deltas, and each individual product exponent delta is added to a respective corresponding term in the first data stream. For purposes of examination, the Examiner interprets the limitation to mean that there are a plurality of product exponent deltas, and each individual product exponent delta is added to a respective corresponding term in the first data stream.
Claims 2-10 inherit the same deficiencies of claim 1 based on dependence.
With regards to claim 2, claim 2 recites the limitations of: “wherein determining the number of bits by which each significand in the second data stream has to be shifted prior to accumulation”. The limitation of “to be shifted” is a limitation describing a future event potentially taking place, therefore it is unclear whether or not the shifting of each significand in the second data stream is positively recited or not.
Furthermore, claim 2 recites the limitations of: “the second data stream”. It is unclear if “the second data stream” is the same data stream as the “second input data stream B” or if it is some other data stream. For purposes of examination, the Examiner interprets “the second data stream” to be the same as the “second input data stream B”.
With regards to claim 4, claim 4 recites the limitations of: “wherein adding the exponents and determining the maximum exponent are shared among a plurality of MAC floating-point units”. It is unclear if it is meant to be understood as the added exponents and the determined maximum exponent are shared among a plurality MAC floating-point units, or if it is meant to be understood as the act of adding the exponents and the act of determining the maximum exponent are done using the plurality of MAC floating-point units.
Furthermore, claim 4 recites the limitations of: “wherein adding the exponents”. It is unclear if the exponents of the limitation are referring to the exponents of data stream A, exponents of data stream B, or if it is referring to the addition of product exponent deltas. For purposes of examination, the Examiner interprets the limitation to be referring to the adding of data stream A and data stream B.
With regards to claim 5, claim 5 recites the limitations of: “wherein the exponents are set to a fixed value”. It is unclear if the exponents of the limitation are referring to either the exponents of data stream A, the exponents of data stream B, the maximum exponent, or the product exponents, or if it is meant to be referring to all of the exponents of data stream A, and exponents of data stream B and the maximum exponent and the product exponents. For purposes of examination, the Examiner interprets the limitation to be referring to all of the exponents of data stream A, and the exponents of data stream B and the maximum exponent and the product exponents.
Furthermore, claim 5 recites the limitations of: “wherein the exponents are set to a fixed value”. It is unclear if the limitation is meant to be understood as all of the exponents are set to a same singular fixed value, “a fixed value”, or if it is meant to be understood as each of the exponents are set to respective fixed values. For purposes of examination, the Examiner interprets the limitation to mean that the exponents are set to respective fixed values.
With regards to claim 6, claim 6 recites the limitations of: “further comprising storing floating-point values in groups”. It is unclear what values are the floating point values, if it is meant to be understood as the values of the data stream A, or the data stream B or the exponents of either data stream, or if it is the product exponents or if it is the maximum exponent or if it is the product exponent deltas, or if it is the partial sum, or if it is the accumulated values, or if it is all of the values including data stream A, and data stream B, and exponents of both streams, and product exponents and the maximum exponent and product exponent deltas and the partial sum and the accumulated values. For purposes of examination, the Examiner interprets it to mean all of the values including data stream A, and data stream B, and exponents of both streams, and product exponents and the maximum exponent and product exponent deltas and the partial sum and the accumulated values.
Claim 7 inherits the same deficiencies as claim 6 based on dependence.
With regards to claim 8, claim 8 recites the limitations of: “wherein using the comparator comprises comparing the maximum exponent to a threshold of an accumulator bit-width”. Claim 8 is dependent on claim 1, claim 1 recites a limitation of “determining a maximum exponent using a comparator”. The limitation of claim 8 of “the comparator comprises comparing the maximum exponent”, uses the maximum exponent as an already calculated value, while the limitation of claim 1, “determining a maximum exponent using a comparator” the maximum exponent is the output, or value calculated from the comparator. It is unclear if the limitation of claim 8 is meant to be understood as the comparator outputs a maximum exponent and that output is put back into the comparator as an input, and then it is compared to a threshold value, or if it is meant to be understood as the input to the comparator isn’t at the time known as the maximum exponent, but instead is a value which is compared to a threshold of an accumulator bit-width or if there is a second comparator where a first determines the maximum exponent and a second comparator compares the maximum exponent to a threshold of an accumulator bit-width, or if the comparator has multiple different comparison or determining operations which includes both determining a maximum exponent and comparing the determined exponent to a threshold of an accumulator bit-width.
Claims 9-10 inherit the same deficiencies as claim 8 based on dependence.
With regards to claim 11, claim 11 recites the limitation of: “input module” which invokes 35 USC 112(f) or pre-AIA 35 USC 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material or acts for performing the entire claimed function. The specification refers to the “input module” in terms of what it does, its function, versus what it is, its structure, see applicant’s specification [0011], [0032], and [0058].
Furthermore, claim 11 recites the limitations of: “A system for accelerating multiply-accumulate (MAC) floating-point units during training or inference of deep learning networks”. It is unclear the training from, “training or inference of deep learning networks”, is meant to be directed towards deep learning training, or some other type of training. For purposes of examination, the Examiner interprets the limitation to mean that the training is directed towards training deep learning networks.
Furthermore, claim 11 recites the limitation of: “A system for accelerating multiply-accumulate (MAC) floating-point units”, it is unclear if “floating-point units” is meant to be understood as floating point values, floating point vectors, floating point matrices, or if it is directed towards meaning accelerating circuitry which calculates multiply-accumulate operations on floating point values. For purposes of examination, the Examiner interprets the limitation to mean that it is directed towards meaning accelerating circuitry upon which they calculate multiply-accumulate operations on floating point values.
Furthermore, claim 11 recites the limitations of: “determine a number of bits by which each significand in the second data stream has to be shifted prior to accumulation”. The limitation of “to be shifted” is a limitation describing a future event potentially taking place, therefore it is unclear whether or not the shifting of each significand in the second data stream is positively recited or not.
Furthermore, claim 11 recites the limitations of: “adding product exponent deltas to the corresponding term in the first data stream”. The Examiner interprets the “product exponent” to be referring to the earlier limitation of, “add exponents of the first data stream A and the second data stream B in pairs to produce product exponents”. An exponent delta is the difference between two exponents. It is unclear what the delta, difference, is between. The limitation of “product exponent deltas” implies that at least one of the values which is used in finding a delta is the product exponent values, however it is unclear if the second value in finding the delta is one from the first data stream, A, or from the second data stream, B, or from the maximum exponent, or from some other exponent value.
Furthermore, claim 11 recites the limitations of: “determine a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and use an adder tree to reduce the operands in the second data stream into a single partial sum”. It is unclear if the determining of a number of bits by which each significand in the second data stream has to be shifted prior to accumulation is determined by only “by adding product exponent deltas to the corresponding term in the first data stream” or by both “by adding product exponent deltas to the corresponding term in the first data stream” and by “use an adder tree to reduce the operands in the second data stream into a single partial sum”.
Furthermore, claim 11 recites the limitations of: “determine a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and use an adder tree to reduce the operands in the second data stream into a single partial sum”. Claim 11 also recites the limitation of, “add the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”. It is unclear if the “accumulation” of the limitation, “determine a number of bits by which each significand in the second data stream has to be shifted prior to accumulation” is referring to the second data stream being reduced to a single partial sum through the adder tree, or if the accumulation is meant to be referring to the accumulated values of the limitation, “add the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”.
Furthermore, claim 11 recites the limitations of: “add the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”. It is unclear if the corresponding aligned value is a value from the first data stream, A, the second data stream, B, the value of the maximum exponent, or some other value.
Furthermore, claim 11 recites the limitations of: “add the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values”. The metes and bounds of the limitation is indefinite because it merely recites a use, “using the maximum exponent”, without any active, positive steps delimiting how the use is actually practiced, see MPEP 2173.05(q).
Furthermore, claim 11 recites the limitations of: “the first data stream”. It is unclear if “the first data stream” is the same data stream as the “first input data stream A”, or some other data stream. For purposes of examination, the Examiner interprets “the first data stream” to be the same as the “first input data stream A”.
Furthermore, claim 11 recites the limitations of: “the second data stream”. It is unclear if “the second data stream” is the same data stream as the “second input data stream B” or if it is some other data stream. For purposes of examination, the Examiner interprets “the second data stream” to be the same as the “second input data stream B”.
Furthermore, claim 11 recites the limitations of: “adding product exponent deltas to the corresponding term in the first data stream”. It is unclear if it is meant to be understood as a plurality of exponent deltas are added to a corresponding term in the first data stream, or if it is meant to be understood as there are a plurality of product exponent deltas, and each individual product exponent delta is added to a respective corresponding term in the first data stream. For purposes of examination, the Examiner interprets the limitation to mean that there are a plurality of product exponent deltas, and each individual product exponent delta is added to a respective corresponding term in the first data stream.
Claims 12-20 inherit the same deficiencies of claim 11 based on dependence.
With regards to claim 14, claim 14 recites the limitations of: “wherein the exponent module, the reduction module, and the accumulation module are located on a processing unit and wherein adding the exponents and determining the maximum exponent are shared among a plurality of processing units”. Claim 14 is dependent on claim 11, claim 11 recites a limitation of: “an exponent module to add exponents of the first data stream A and the second data stream B in pairs to produce product exponents, and to determine a maximum exponent using a comparator”. Claim 14 recites a limitation regarding the exponent module being located on a (singular) processing unit. Claim 11 recites a limitation regarding the exponent module being what computes adding the exponents and the determining of the maximum exponent. However, claim 14 also recites a limitation regarding the adding of the exponents and the determining of the maximum exponent are shared among a plurality of processing units. It is unclear how the limitation of sharing of adding the exponents and determining the maximum exponent coincides with the limitation of the adding of the exponents and the determining of the maximum exponent being done by the exponent module which is located on a singular processing unit.
Claims 15-16 inherit the same deficiencies as claim 14 based on dependence.
Indication of Allowable Subject Matter
Claims 1-10 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action.
Claims 11-20 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, and U.S.C. 112(a) rejection set forth in this Office action.
The following is a statement of reasons for the indication of allowable subject matter regarding claims 1-10:
With regards to claim 1, the applicant claims a method for accelerating multiply-accumulate (MAC) floating-point units during training or inference of deep learning networks, the method comprising:
receiving a first input data stream A and a second input data stream B; adding exponents of the first data stream A and the second data stream B in pairs to produce product exponents; determining a maximum exponent using a comparator; determining a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and using an adder tree to reduce the operands in the second data stream into a single partial sum; adding the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values; and outputting the accumulated values.
The primary reason for indication of allowable subject matter is the above italicized claim limitations in combination with the remaining claim limitations including intervening claims.
The following is a statement of reasons for the indication of allowable subject matter regarding claims 11-20:
With regards to claim 11, the applicant claims a system for accelerating multiply-accumulate (MAC) floating-point units during training or inference of deep learning networks, the system comprising one or more processors in communication with data memory to execute:
an input module to receive a first input data stream A and a second input data stream B; an exponent module to add exponents of the first data stream A and the second data stream B in pairs to produce product exponents, and to determine a maximum exponent using a comparator; a reduction module to determine a number of bits by which each significand in the second data stream has to be shifted prior to accumulation by adding product exponent deltas to the corresponding term in the first data stream and use an adder tree to reduce the operands in the second data stream into a single partial sum; and an accumulation module to add the partial sum to a corresponding aligned value using the maximum exponent to determine accumulated values, and to output the accumulated values.
The primary reason for indication of allowable subject matter is the above italicized claim limitations in combination with the remaining claim limitations including intervening claims.
Lee Hun Jae et al (“Design of Floating-Point MAC Unit for computing DNN Applications in PIM”, 2020 INTERNATIONAL CONFERENCE ON ELECTRONICS, INFORMATION, AND COMMUNICAITON (ICEIC), IEEE, 19 January 2020), hereinafter, “Lee” discloses calculations using floating point multiply-accumulate operations for training of deep learning networks (Abstract). Lee further discloses receiving two data streams (Fig. 3, fig. 4 regarding inputs A and B). Lee further discloses adding the exponents of the two input data streams (Fig. 4, regarding EA + EB). Lee further discloses determining a maximum exponent using a comparator (Fig. 4, regarding Exp. Comparator). Lee further discloses adding exponent products to a value to align the value (Fig. 4; Page 3 column 2, section labeled “Accumulation Unit” regarding addition is performed by aligning the mantissa according to the comparison result; Page 5 column 1 regarding shifting the number of a mantissa). However, Lee fails to teach or suggest the italicized claim limitations in combination with the remaining claim limitations as referenced above. Lee fails to teach or suggest determining number of bits by which each significand of the second data stream is shifted before accumulation by adding the product exponents (of the two inputs) to the corresponding terms in the first data stream and by using an adder tree to reduce the second data stream to a singular partial sum value. Lee further fails to teach or suggest adding the partial sum from the adder tree to a corresponding value which uses the determined maximum exponent to determine accumulated values.
Urbanski et al. (U.S. Patent application publication 2021/0263993 A1), hereinafter, “Urbanski” discloses calculations using floating point multiply-accumulate operations for training of deep learning networks ([0196] regarding use of deep learning operations; Abstract). Urbanski further discloses two input vectors ([0045] regarding vector multiply operations; [0065] regarding input vectors A, and B). Urbanski further discloses comparing exponents and determining a maximum exponent (Fig. 10 regarding MAX EXPONENT DETERMINER 1040). Urbanski further discloses using the determined maximum exponent to shift mantissa products (Fig. 10 regarding MAX EXPONENT DETERMINER 1040 output to shift circuits; [0065] regarding the max exponent used to shift mantissa products). Urbanski further discloses using an adder tree. However, Urbanski fails to teach or suggest the italicized claim limitations in combination with the remaining claim limitations as referenced above. Urbanski fails to teach or suggest determining number of bits by which each significand of the second data stream is shifted before accumulation by adding the product exponents (of the two inputs) to the corresponding terms in the first data stream and by using an adder tree to reduce the second data stream to a singular partial sum value. Urbanski further fails to teach or suggest adding the partial sum from the adder tree to a corresponding value which uses the determined maximum exponent to determine accumulated values.
Kaul et al. (U.S. Patent application publication 2019/0294415 A1) hereinafter, “Kaul” discloses calculations using floating point multiply-accumulate operations for machine learning ([0001]). Kaul further discloses two inputs (Fig. 1 inputs ‘a’ and ‘b’). Kaul further discloses adding the input exponent values together (Fig. 1 regarding ea being summed with eb), and determining a maximum exponent (Fig. 1 regarding MAXIMUM EXPONENT 30). Kaul further discloses determining exponent deltas (Fig. 7 regarding subtracting exponent values from the max exponent), and aligning floating point numbers using the exponent deltas (Fig. 6; Fig. 7). Kaul further discloses use of an adder tree to reduce a value to a singular value (Fig. 6). However, Kaul fails to teach or suggest the italicized claim limitations in combination with the remaining claim limitations as referenced above. Kaul fails to teach or suggest determining number of bits by which each significand of the second data stream is shifted before accumulation by adding the product exponents (of the two inputs) to the corresponding terms in the first data stream and by using an adder tree to reduce the second data stream to a singular partial sum value. Kaul further fails to teach or suggest adding the partial sum from the adder tree to a corresponding value which uses the determined maximum exponent to determine accumulated values.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JEROME ANTHONY KLOSTERMAN II whose telephone number is (571)272-0541. The examiner can normally be reached Monday-Friday 8:30am-3:30pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Caldwell can be reached at 571-272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.A.K./ Examiner, Art Unit 2182 /EMILY E LAROCQUE/ Primary Examiner, Art Unit 2182