Detailed Action
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is non-final and is in response to claims filed on 01/05/2026 via RCE. Claims 1-20 are pending examination. Claims 1-20 are currently amended.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 01/05/2026 has been entered.
Response to Arguments
Claim Objections
Applicant has amended the claims at issue, and, therefore, the previous objections have been withdrawn.
Rejections under 35 U.S.C. 101
Applicant’s arguments, see Remarks 9-12, filed 01/05/2026, with respect to the rejections under 35 U.S.C. 101 of claims 1-20 have been fully considered and are persuasive. The Rejection of claims 1-20 has been withdrawn.
Specifically the recitation of the neural network in the claims and the improvements of the neural network integrate the abstract ideas into a practical application.
Rejections under 35 U.S.C. 103
Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Claim 19 recites “means of a neural network” and and the related structure in the specification is at paragraph [005] “a system, including: a circuit configured to multiply a first number by a second number, the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.”
Claim 20 recites “means of a neural network” and the related structure in the specification is at paragraph [005] “a system, including: a circuit configured to multiply a first number by a second number, the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.”
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 10, and 19 are rejected as being unpatentable over Willcock et al. (US 20220156344 A1) hereinafter Willcock, in view of Darvish et al. (US 20200210840 A1) hereinafter Darvish further in view of Piskorski et al. (“Customizing CPU instructions for embedded vision systems”) hereinafter Piskorski.
With regards to claim 1, Willcock teaches A system, comprising: a circuit of a neural network including at least a first multiplier circuit and a second multiplier circuit, (Willcock [0004]: n general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock Fig. 2: shows multiple multiplication units; Willcock [0031]: the systolic array 206 is used for neural network computations)
The first multiplier circuit and the second multiplier circuit respectively including first and second inputs to receive respectively first and second bits of a first number and a second number [for a first layer of the neural network] and output a first product (Willcock [0004]: In general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock [0046]: The cell includes multiplication circuitry 306 and summation circuitry 308. The multiplication circuitry 306 can determine the product of the matrix elements stored in the input registers 302 and 304. For example, the multiplication circuitry 306 can determine a product by multiplying the element of the input matrix stored in the input register 302 by the element of the input matrix stored in the input register 304; Willcock Fig. 2: shows multiple multiplication units with 2 inputs each)
Willcock fails to teach for a first layer of the neural network, and to receive respectively third and fourth bits of a third number and a fourth number for a second layer of the neural network and output a second product, wherein precision of the second product is lower than a precision of the first product, and wherein, the neural network generates an inference based on the first product and the second product.
However, Darvish teaches the multiplication of the first and second numbers of Willcock is for a first layer of the neural network (Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format)
and to receive respectively third and fourth bits of a third number and a fourth number for a second layer of the neural network and output a second product, (Darvish [0042]: An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor... As another example, during a forward-propagation mode, the input tensor can be an output term from a layer adjacent to (e.g., preceding) the given layer or weights of the given layer; Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format, forward propagating values from the first layer of the neural network to a second layer of the neural network, determining a training performance metric for the neural network, selecting an adjusted precision parameter for the neural network, and modifying the neural network based on the adjusted precision parameter)
wherein precision of the second product is lower than a precision of the first product (Darvish [0042]: An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor... As another example, during a forward-propagation mode, the input tensor can be an output term from a layer adjacent to (e.g., preceding) the given layer or weights of the given layer; Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format, forward propagating values from the first layer of the neural network to a second layer of the neural network, determining a training performance metric for the neural network, selecting an adjusted precision parameter for the neural network, and modifying the neural network based on the adjusted precision parameter)
wherein, the neural network generates an inference based on the first product and the second product (Darvish [0041]: For example, block floating-point formats can be used to accelerate computations performed in training and inference operations using the neural network accelerator).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock with the layers of the neural network and precisions of Darvish. One of ordinary skill in the art would be motivated to make this combination because in this manner, the neural network accelerator can potentially be made smaller and more efficient than a comparable accelerator that uses only a normal-precision floating-point format. A smaller and more efficient accelerator may have increased computational performance and/or increased energy efficiency as taught by Darvish (Darvish [0042]).
Willcock in view of Darvish fails to teach the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.
However, Piskorski teaches the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
With regards to claim 10, Willcock teaches A method comprising: in a circuit of a neural network including at least a first multiplier circuit and a second multiplier circuit, the first multiplier circuit and the second multiplier circuit respectively including first and second inputs, (Willcock [0004]: In general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock [0046]: The cell includes multiplication circuitry 306 and summation circuitry 308. The multiplication circuitry 306 can determine the product of the matrix elements stored in the input registers 302 and 304. For example, the multiplication circuitry 306 can determine a product by multiplying the element of the input matrix stored in the input register 302 by the element of the input matrix stored in the input register 304; Willcock Fig. 2: shows multiple multiplication units with 2 inputs each)
receiving by the first and second inputs respectively first and second bits of a first number and a second number [for a first layer of the neural network] and outputting a first product; (Willcock [0004]: In general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock [0046]: The cell includes multiplication circuitry 306 and summation circuitry 308. The multiplication circuitry 306 can determine the product of the matrix elements stored in the input registers 302 and 304. For example, the multiplication circuitry 306 can determine a product by multiplying the element of the input matrix stored in the input register 302 by the element of the input matrix stored in the input register 304).
Willcock fails to teach for a first layer of the neural network, and receiving by the first and second inputs respectively third and fourth bits of a third number and a fourth number for a second layer of the neural network and outputting a second product, wherein precision of the second product is lower than a precision of the first product, and and generating, by the neural network, an inference based on the first product and the second product.
However, Darvish teaches the multiplication of the first and second numbers of Willcock is for a first layer of the neural network (Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format)
and receiving by the first and second inputs respectively third and fourth bits of a third number and a fourth number for a second layer of the neural network and outputting a second product, (Darvish [0042]: An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor... As another example, during a forward-propagation mode, the input tensor can be an output term from a layer adjacent to (e.g., preceding) the given layer or weights of the given layer; Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format, forward propagating values from the first layer of the neural network to a second layer of the neural network, determining a training performance metric for the neural network, selecting an adjusted precision parameter for the neural network, and modifying the neural network based on the adjusted precision parameter)
wherein precision of the second product is lower than a precision of the first product (Darvish [0042]: An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor... As another example, during a forward-propagation mode, the input tensor can be an output term from a layer adjacent to (e.g., preceding) the given layer or weights of the given layer; Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format, forward propagating values from the first layer of the neural network to a second layer of the neural network, determining a training performance metric for the neural network, selecting an adjusted precision parameter for the neural network, and modifying the neural network based on the adjusted precision parameter)
and generating, by the neural network, an inference based on the first product and the second product (Darvish [0041]: For example, block floating-point formats can be used to accelerate computations performed in training and inference operations using the neural network accelerator).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock with the layers of the neural network and precisions of Darvish. One of ordinary skill in the art would be motivated to make this combination because in this manner, the neural network accelerator can potentially be made smaller and more efficient than a comparable accelerator that uses only a normal-precision floating-point format. A smaller and more efficient accelerator may have increased computational performance and/or increased energy efficiency as taught by Darvish (Darvish [0042]).
Willcock in view of Darvish fails to teach the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.
However, Piskorski teaches the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
With regards to claim 19, Willcock teaches A system, comprising: means of a neural network including at least a first multiplier circuit and a second multiplier circuit, (Willcock [0004]: n general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock Fig. 2: shows multiple multiplication units; Willcock [0031]: the systolic array 206 is used for neural network computations)
The first multiplier circuit and the second multiplier circuit respectively including first and second inputs to receive respectively first and second bits of a first number and a second number [for a first layer of the neural network] and output a first product (Willcock [0004]: In general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock [0046]: The cell includes multiplication circuitry 306 and summation circuitry 308. The multiplication circuitry 306 can determine the product of the matrix elements stored in the input registers 302 and 304. For example, the multiplication circuitry 306 can determine a product by multiplying the element of the input matrix stored in the input register 302 by the element of the input matrix stored in the input register 304; Willcock Fig. 2: shows multiple multiplication units with 2 inputs each)
Willcock fails to teach for a first layer of the neural network, and to receive respectively third and fourth bits of a third number and a fourth number for a second layer of the neural network and output a second product, wherein precision of the second product is lower than a precision of the first product, and wherein, the neural network generates an inference based on the first product and the second product.
However, Darvish teaches the multiplication of the first and second numbers of Willcock is for a first layer of the neural network (Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format)
and to receive respectively third and fourth bits of a third number and a fourth number for a second layer of the neural network and output a second product, (Darvish [0042]: An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor... As another example, during a forward-propagation mode, the input tensor can be an output term from a layer adjacent to (e.g., preceding) the given layer or weights of the given layer; Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format, forward propagating values from the first layer of the neural network to a second layer of the neural network, determining a training performance metric for the neural network, selecting an adjusted precision parameter for the neural network, and modifying the neural network based on the adjusted precision parameter)
wherein precision of the second product is lower than a precision of the first product (Darvish [0042]: An input tensor for the given layer can be converted from a normal-precision floating-point format to a quantized-precision floating-point format. A tensor operation can be performed using the converted input tensor... As another example, during a forward-propagation mode, the input tensor can be an output term from a layer adjacent to (e.g., preceding) the given layer or weights of the given layer; Darvish [0165]: In some examples of the disclosed technology, a method of configuring a computer system to implement a neural network includes implementing a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format, forward propagating values from the first layer of the neural network to a second layer of the neural network, determining a training performance metric for the neural network, selecting an adjusted precision parameter for the neural network, and modifying the neural network based on the adjusted precision parameter)
wherein, the neural network generates an inference based on the first product and the second product (Darvish [0041]: For example, block floating-point formats can be used to accelerate computations performed in training and inference operations using the neural network accelerator).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock with the layers of the neural network and precisions of Darvish. One of ordinary skill in the art would be motivated to make this combination because in this manner, the neural network accelerator can potentially be made smaller and more efficient than a comparable accelerator that uses only a normal-precision floating-point format. A smaller and more efficient accelerator may have increased computational performance and/or increased energy efficiency as taught by Darvish (Darvish [0042]).
Willcock in view of Darvish fails to teach the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.
However, Piskorski teaches the first number being represented as: a sign bit five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
Claims 2, 3, 11, 12, and 20 are rejected as being unpatentable over Willcock, in view of Darvish, further in view of Piskorski, further in view of Rzayev et al. (“DeepRecon: Dynamically Reconfigurable Architecture for Accelerating Deep Neural Networks”) hereinafter Rzayev.
With regards to claim 2, Willcock in view of Darvish further in view of Piskorski teaches all of the limitations of claim 1 above. Willcock further teaches wherein the circuit comprises: a third multiplier circuit; and a fourth multiplier circuit, (Willcock [0004]: n general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock Fig. 2: shows multiple multiplication units).
Willcock fails to teach each of the first multiplier circuit, the second multiplier circuit, the third multiplier circuit, and the fourth multiplier circuit being a 4 bit by 8 bit multiplier circuit.
However, Rzayev does teach each of the first multiplier circuit, the second multiplier circuit, the third multiplier circuit, and the fourth multiplier circuit being a 4 bit by 8 bit multiplier circuit (Rzayev Page 121 Section 4.2: As can be seen in Figure 5, the architecture supports a 16 bit base configuration and can fracture on demand into 2x12 bit or 4x8 bit floating point multipliers).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski with the 4x8 multiplier as taught by Rzayev. One of ordinary skill in the art would be motivated to make this combination because they naturally lend themselves to efficient resource sharing. Additionally, the architecture is capable of supporting all of the intermediate bit widths by using a higher bit width configuration and tying the low-order mantissa bits low. Currently this yields the same energy-efficiency as a higher bit width configuration as taught by Rzayev (Rzayev Page 121 Section 4.2).
With regards to claim 3, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 2 above. Willcock fails to teach wherein the second number is represented as: a sign bit, five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.
However, Piskorski teaches wherein the second number is represented as: a sign bit, five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
With regards to claim 11, Willcock in view of Darvish further in view of Piskorski teaches all of the limitations of claim 10 above. Willcock further teaches wherein the circuit comprises: a third multiplier circuit; and a fourth multiplier circuit, (Willcock [0004]: n general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock Fig. 2: shows multiple multiplication units).
Willcock fails to teach each of the first multiplier circuit, the second multiplier circuit, the third multiplier circuit, and the fourth multiplier circuit being a 4 bit by 8 bit multiplier circuit.
However, Rzayev does teach each of the first multiplier circuit, the second multiplier circuit, the third multiplier circuit, and the fourth multiplier circuit being a 4 bit by 8 bit multiplier circuit (Rzayev Page 121 Section 4.2: As can be seen in Figure 5, the architecture supports a 16 bit base configuration and can fracture on demand into 2x12 bit or 4x8 bit floating point multipliers).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski with the 4x8 multiplier as taught by Rzayev. One of ordinary skill in the art would be motivated to make this combination because they naturally lend themselves to efficient resource sharing. Additionally, the architecture is capable of supporting all of the intermediate bit widths by using a higher bit width configuration and tying the low-order mantissa bits low. Currently this yields the same energy-efficiency as a higher bit width configuration as taught by Rzayev (Rzayev Page 121 Section 4.2).
With regards to claim 12, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 11 above. Willcock fails to teach wherein the second number is represented as: a sign bit, five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa.
However, Piskorski teaches wherein the second number is represented as: a sign bit, five exponent bits, and seven mantissa bits, representing an eight-bit full mantissa (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
With regards to claim 20, Willcock in view of Darvish further in view of Piskorski teaches all of the limitations of claim 19 above. Willcock further teaches wherein the means of the neural network comprises: a third multiplier circuit; and a fourth multiplier circuit, (Willcock [0004]: n general, one innovative aspect of the subject matter described in this specification can be embodied in a matrix multiplication unit that includes multiple cells arranged in a systolic array; Willcock Fig. 2: shows multiple multiplication units).
Willcock fails to teach each of the first multiplier circuit, the second multiplier circuit, the third multiplier circuit, and the fourth multiplier circuit being a 4 bit by 8 bit multiplier circuit.
However, Rzayev does teach each of the first multiplier circuit, the second multiplier circuit, the third multiplier circuit, and the fourth multiplier circuit being a 4 bit by 8 bit multiplier circuit (Rzayev Page 121 Section 4.2: As can be seen in Figure 5, the architecture supports a 16 bit base configuration and can fracture on demand into 2x12 bit or 4x8 bit floating point multipliers).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski with the 4x8 multiplier as taught by Rzayev. One of ordinary skill in the art would be motivated to make this combination because they naturally lend themselves to efficient resource sharing. Additionally, the architecture is capable of supporting all of the intermediate bit widths by using a higher bit width configuration and tying the low-order mantissa bits low. Currently this yields the same energy-efficiency as a higher bit width configuration as taught by Rzayev (Rzayev Page 121 Section 4.2).
Claims 4, 5, 13, and 14 are rejected as being unpatentable over Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul et al. (EP 3396524 A1) hereinafter Kaul.
With regards to claim 4, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 3 above. Willcock fails to teach wherein the circuit is configured, in a first configuration, to multiply the mantissa of the first number and the mantissa of the second number using the first multiplier circuit and the second multiplier circuit.
However, Kaul does teach wherein the circuit is configured, in a first configuration, to multiply the mantissa of the first number and the mantissa of the second number using the first multiplier circuit and the second multiplier circuit (Kaul [0167]: the signed multiplier 1702A-1702B ..., which are used for both integer and floating-point modes; Kaul Fig. 17A: shows the mantissas of A and B being fed into multipliers 1702A and 1702B).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with using two multipliers to multiply the mantissas as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mantissas to be multiplied in parallel, increasing efficiency.
With regards to claim 5, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul teaches all of the limitations of claim 4 above. Willcock fail to teach wherein the second product is an approximate product of the third number and the fourth number, each of the third number and the fourth number being represented as: a sign bit, five exponent bits, and ten mantissa bits, representing an 11-bit full mantissa.
However, Kaul teaches wherein the second product is an approximate product of the third number and the fourth number, (Kaul [0183]: In one embodiment the FPU encoding and configuration module 1834 can also configure the floating-point encoding methods supported by the floating point units. In addition to IEEE 754 floating-point standards for half, single, and double precision encoding for floating point values, a myriad of alternative encoding formats may be supported based on the dynamic range of the data that is currently being processed. For example, based on the dynamic range and/or distribution a given dataset, the data may be quantized more accurately from higher to lower precision by using greater than or fewer bits for exponent or mantissa data)
each of the third number and the fourth number being represented as: a sign bit, (Kaul [0153]: 1-bit sign)
five exponent bits, (Kaul [0153]: a 5-bit exponent)
and ten mantissa bits, representing an 11-bit full mantissa (Kaul [0153]: and a 10-bit fractional portion).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Kaul with calculating the approximate value as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it increases efficiency by not using the whole FP number, reducing the number of calculations needed.
With regards to claim 13, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 12 above. Willcock fails to teach wherein the circuit is configured, in a first configuration, to multiply the mantissa of the first number and the mantissa of the second number using the first multiplier circuit and the second multiplier circuit.
However, Kaul does teach wherein the circuit is configured, in a first configuration, to multiply the mantissa of the first number and the mantissa of the second number using the first multiplier circuit and the second multiplier circuit (Kaul [0167]: the signed multiplier 1702A-1702B ..., which are used for both integer and floating-point modes; Kaul Fig. 17A: shows the mantissas of A and B being fed into multipliers 1702A and 1702B).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with using two multipliers to multiply the mantissas as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mantissas to be multiplied in parallel, increasing efficiency.
With regards to claim 14, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul teaches all of the limitations of claim 13 above. Willcock fail to teach wherein the second product is an approximate product of the third number and the fourth number, each of the third number and the fourth number being represented as: a sign bit, five exponent bits, and ten mantissa bits, representing an 11-bit full mantissa.
However, Kaul teaches wherein the second product is an approximate product of the third number and the fourth number, (Kaul [0183]: In one embodiment the FPU encoding and configuration module 1834 can also configure the floating-point encoding methods supported by the floating point units. In addition to IEEE 754 floating-point standards for half, single, and double precision encoding for floating point values, a myriad of alternative encoding formats may be supported based on the dynamic range of the data that is currently being processed. For example, based on the dynamic range and/or distribution a given dataset, the data may be quantized more accurately from higher to lower precision by using greater than or fewer bits for exponent or mantissa data)
each of the third number and the fourth number being represented as: a sign bit, (Kaul [0153]: 1-bit sign)
five exponent bits, (Kaul [0153]: a 5-bit exponent)
and ten mantissa bits, representing an 11-bit full mantissa (Kaul [0153]: and a 10-bit fractional portion).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul with calculating the approximate value as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it increases efficiency by not using the whole FP number, reducing the number of calculations needed.
Claims 6 and 15 are rejected as being unpatentable over Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer et al. (US 20180081631 A1) hereinafter Langhammer further in view of Hastings et al. (“PSoC® 3 and PSoC 5LP: Getting More Resolution from 8-Bit DACs”) hereinafter Hastings.
With regards to claim 6, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul teaches all of the limitations of claim 5 above. Willcock fails to teach wherein the circuit is configured, in a second configuration, to: multiply eight bits of the full mantissa of the third number by eight bits of the full mantissa of the fourth number using the first multiplier circuit and the second multiplier circuit, multiply three bits of the full mantissa of the third number by eight bits of the full mantissa of the fourth number using the third multiplier circuit, and multiply eight bits of the full mantissa of the third number by three bits of the full mantissa of the fourth number using the fourth multiplier circuit.
However, Langhammer teaches wherein the circuit is configured, in a second configuration, to: multiply [eight bits] of the full mantissa of the third number by [eight bits] of the full mantissa of the fourth number (Langhammer [0046]: specialized processing block 216 may receive the 27 most significant bits (MSBs) of A and B (i.e., A[54:28] and B[54:28])).
multiply [three bits] of the full mantissa of the third number by [eight bits] of the full mantissa of the fourth number using the third multiplier circuit, (Langhammer [0046]: specialized processing block 212 may receive the 27 least significant bits (LSBs) of A and the 27 most significant bits (MSBs) of B (i.e., A[27:1] and B[54:28]))
and multiply [eight bits] of the full mantissa of the third number by [three bits] of the full mantissa of the fourth number using the fourth multiplier circuit (Langhammer [0046]: specialized processing block 214 may receive the 27 most significant bits (MSBs) of A and the 27 least significant bits (LSBs) of B (i.e., A[54:28] and B[27:1]))
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul with splitting the mantissa as taught by Langhammer. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to efficiently implement fixed-point operations, single-precision floating-point operations, and double-precision floating-point operations as taught by Langhammer (Langhammer [0028].
Langhammer fails to teach splitting the mantissas into 8 and 3 bits.
However, Hastings teaches splitting the mantissas into 8 and 3 bits (Hastings Page 5 paragraph 6: The 8 most significant bits… the 3 least significant bits).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer with splitting the mantissas into 8 and 3 bits as taught by Hastings. One of ordinary skill in the art would be motivated to make this combination because it allows for less bits to be needed for the calculation increasing efficiency and allowing more data formats to be used.
Langhammer in view of Hasting fails to teach using the first multiplier circuit and the second multiplier circuit.
However, Kaul does teach using the first multiplier circuit and the second multiplier circuit (Kaul [0167]: the signed multiplier 1702A-1702B ..., which are used for both integer and floating-point modes; Kaul Fig. 17A: Fig. 17A shows the mantissas of A and B being fed into multipliers 1702A and 1702B).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings with using two multipliers to multiply the mantissas as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mantissas to be multiplied in parallel, increasing efficiency.
With regards to claim 15, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul teaches all of the limitations of claim 14 above. Willcock fails to teach wherein the circuit is configured, in a second configuration, to: multiply eight bits of the full mantissa of the third number by eight bits of the full mantissa of the fourth number using the first multiplier circuit and the second multiplier circuit, multiply three bits of the full mantissa of the third number by eight bits of the full mantissa of the fourth number using the third multiplier circuit, and multiply eight bits of the full mantissa of the third number by three bits of the full mantissa of the fourth number using the fourth multiplier circuit.
However, Langhammer teaches wherein the circuit is configured, in a second configuration, to: multiply [eight bits] of the full mantissa of the third number by [eight bits] of the full mantissa of the fourth number (Langhammer [0046]: specialized processing block 216 may receive the 27 most significant bits (MSBs) of A and B (i.e., A[54:28] and B[54:28])).
multiply [three bits] of the full mantissa of the third number by [eight bits] of the full mantissa of the fourth number using the third multiplier circuit, (Langhammer [0046]: specialized processing block 212 may receive the 27 least significant bits (LSBs) of A and the 27 most significant bits (MSBs) of B (i.e., A[27:1] and B[54:28]))
and multiply [eight bits] of the full mantissa of the third number by [three bits] of the full mantissa of the fourth number using the fourth multiplier circuit (Langhammer [0046]: specialized processing block 214 may receive the 27 most significant bits (MSBs) of A and the 27 least significant bits (LSBs) of B (i.e., A[54:28] and B[27:1]))
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul with splitting the mantissa as taught by Langhammer. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to efficiently implement fixed-point operations, single-precision floating-point operations, and double-precision floating-point operations as taught by Langhammer (Langhammer [0028].
Langhammer fails to teach splitting the mantissas into 8 and 3 bits.
However, Hastings teaches splitting the mantissas into 8 and 3 bits (Hastings Page 5 paragraph 6: The 8 most significant bits… the 3 least significant bits).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer with splitting the mantissas into 8 and 3 bits as taught by Hastings. One of ordinary skill in the art would be motivated to make this combination because it allows for less bits to be needed for the calculation increasing efficiency and allowing more data formats to be used.
Langhammer in view of Hasting fails to teach using the first multiplier circuit and the second multiplier circuit.
However, Kaul does teach using the first multiplier circuit and the second multiplier circuit (Kaul [0167]: the signed multiplier 1702A-1702B ..., which are used for both integer and floating-point modes; Kaul Fig. 17A: Fig. 17A shows the mantissas of A and B being fed into multipliers 1702A and 1702B).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings with using two multipliers to multiply the mantissas as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for the mantissas to be multiplied in parallel, increasing efficiency.
Claims 7 and 16 are rejected as being unpatentable over Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings further in view of Kuang et al. (Design of Power-Efficient Configurable Booth Multiplier) hereinafter Kuang.
With regards to claim 7, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings teaches all of the limitations of claim 6 above. Willcock fails to teach wherein the circuit is configured not to calculate a product of [three least significant bits] of the full mantissa of the third number and the [three least significant bits] of the full mantissa of the fourth number.
However, Kuang does teach wherein the circuit is configured not to calculate a product of [three least significant bits] of the full mantissa of the third number and the [three least significant bits] of the full mantissa of the fourth number (Kuang Page 574 Section III B: the computation for the n-1 least significant bits of the 2n-bit output product can be disabled).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings with not multiplying the least significant bits as taught by Kuang. One of ordinary skill in the art would be motivated to make this combination to further reduce power consumption as taught by Kuang (Kuang Page 574 Section III B).
Willcock in view of Kuang fails to teach three least significant bits.
However, Hastings teaches three least significant bits (Hastings Page 5 paragraph 6: the 3 least significant bits).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings further in view of Kuang with the 3 least significant bits as taught by Hastings. One of ordinary skill in the art would be motivated to make this combination because it allows for less bits to be needed for the calculation increasing efficiency and allowing more data formats to be used.
With regards to claim 16, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings teaches all of the limitations of claim 15 above. Willcock fails to teach wherein the circuit is configured not to calculate a product of [three least significant bits] of the full mantissa of the third number and the [three least significant bits] of the full mantissa of the fourth number.
However, Kuang does teach wherein the circuit is configured not to calculate a product of [three least significant bits] of the full mantissa of the third number and the [three least significant bits] of the full mantissa of the fourth number (Kuang Page 574 Section III B: the computation for the n-1 least significant bits of the 2n-bit output product can be disabled).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Langhammer further in view of Hastings with not multiplying the least significant bits as taught by Kuang. One of ordinary skill in the art would be motivated to make this combination to further reduce power consumption as taught by Kuang (Kuang Page 574 Section III B).
Willcock in view of Kuang fails to teach three least significant bits.
However, Hastings teaches three least significant bits (Hastings Page 5 paragraph 6: the 3 least significant bits).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul, further in view of Langhammer further in view of Hastings further in view of Kuang with the 3 least significant bits as taught by Hastings. One of ordinary skill in the art would be motivated to make this combination because it allows for less bits to be needed for the calculation increasing efficiency and allowing more data formats to be used.
Claims 8, 9, 17, and 18 are rejected as being unpatentable over Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul further in view of Divakar et al. (US 20200097799 A1) hereinafter Divakar.
With regards to claim 8, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 2 above. Willcock fails to teach wherein the second number is an 8-bit integer and the circuit is configured to multiply 8 bits of the full mantissa of the first number by the second number.
However, Divakar does teach wherein the second number is an 8-bit integer (Divakar [0038]: the other operand 445 is in 8-bit integer (INT8) format)
and the circuit is configured to multiply [8 bits] of the full mantissa of the first number by the second number (Divakar [0038]: multiply two other operands 440, 445, wherein one of the operands 440 is in the FP16 format and the other operand 445 is in 8-bit integer (INT8) format).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Kaul with multiplying an integer by a floating point as taught by Divakar. One of ordinary skill in the art would be motivated to make this combination because it would allow for greater flexibility as it allows for multiple data formats to be multiplied together.
Willcock in view of Divakar fails to teach 8 bits.
However, Piskorski does teach 8 bits (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
Willcock in view of Divakar further in view of Piskorski fails to teach using the first multiplier circuit and the second multiplier circuit.
However, Kaul does teach using the first multiplier circuit and the second multiplier circuit (Kaul [0167]: the signed multiplier 1702A-1702B ..., which are used for both integer and floating-point modes; Kaul Fig. 17A: Fig. 17A shows the mantissas of A and B being fed into multipliers 1702A and 1702B).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with using the two multipliers as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for portions of the numbers to be multiplied in parallel, increasing efficiency.
With regards to claim 9, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 2 above. Willcock fails to teach wherein the second number is a 4-bit integer and the circuit is configured to multiply 8 bits of the full mantissa of the first number by the second number.
However, Divakar does teach wherein the second number is a 4-bit integer (Divakar [0053]: such as data formats of different types (e.g., floating point, integer (or other fixed point number of a corresponding Q format), etc.) and/or data formats of different precisions (e.g., 4-bit, 8-bit, 16-bit, etc.))
and the circuit is configured to multiply [8 bits] of the full mantissa of the first number by the second number (Divakar [0038]: multiply two other operands 440, 445, wherein one of the operands 440 is in the FP16 format and the other operand 445 is in 8-bit integer (INT8) format; Divakar [0053]: such as data formats of different types (e.g., floating point, integer (or other fixed point number of a corresponding Q format), etc.) and/or data formats of different precisions (e.g., 4-bit, 8-bit, 16-bit, etc.)).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with multiplying an integer by a floating point as taught by Divakar. One of ordinary skill in the art would be motivated to make this combination because it would allow for greater flexibility as it allows for multiple data formats to be multiplied together.
Willcock in view of Divakar fails to teach 8 bits.
However, Piskorski does teach 8 bits (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
Willcock in view of Divakar further in view of Piskorski fails to teach using the first multiplier circuit.
However, Kaul does teach using the first multiplier circuit and the second multiplier circuit (Kaul [0163]: that is generated by a signed 16bx16b multiplier; Kaul Fig. 15A: shows the numbers being fed into one multiplier).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with using the two multipliers as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for portions of the numbers to be multiplied in parallel, increasing efficiency.
With regards to claim 17, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 11 above. Willcock fails to teach wherein the second number is an 8-bit integer and the circuit is configured to multiply 8 bits of the full mantissa of the first number by the second number.
However, Divakar does teach wherein the second number is an 8-bit integer (Divakar [0038]: the other operand 445 is in 8-bit integer (INT8) format)
and the circuit is configured to multiply [8 bits] of the full mantissa of the first number by the second number (Divakar [0038]: multiply two other operands 440, 445, wherein one of the operands 440 is in the FP16 format and the other operand 445 is in 8-bit integer (INT8) format).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with multiplying an integer by a floating point as taught by Divakar. One of ordinary skill in the art would be motivated to make this combination because it would allow for greater flexibility as it allows for multiple data formats to be multiplied together.
Willcock in view of Divakar fails to teach 8 bits.
However, Piskorski does teach 8 bits (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
Willcock in view of Divakar further in view of Piskorski fails to teach using the first multiplier circuit and the second multiplier circuit.
However, Kaul does teach using the first multiplier circuit and the second multiplier circuit (Kaul [0167]: the signed multiplier 1702A-1702B ..., which are used for both integer and floating-point modes; Kaul Fig. 17A: Fig. 17A shows the mantissas of A and B being fed into multipliers 1702A and 1702B).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with using the two multipliers as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for portions of the numbers to be multiplied in parallel, increasing efficiency.
With regards to claim 18, Willcock in view of Darvish further in view of Piskorski further in view of Rzayev teaches all of the limitations of claim 11 above. Willcock fails to teach wherein the second number is a 4-bit integer and the circuit is configured to multiply 8 bits of the full mantissa of the first number by the second number.
However, Divakar does teach wherein the second number is a 4-bit integer (Divakar [0053]: such as data formats of different types (e.g., floating point, integer (or other fixed point number of a corresponding Q format), etc.) and/or data formats of different precisions (e.g., 4-bit, 8-bit, 16-bit, etc.))
and the circuit is configured to multiply [8 bits] of the full mantissa of the first number by the second number (Divakar [0038]: multiply two other operands 440, 445, wherein one of the operands 440 is in the FP16 format and the other operand 445 is in 8-bit integer (INT8) format; Divakar [0053]: such as data formats of different types (e.g., floating point, integer (or other fixed point number of a corresponding Q format), etc.) and/or data formats of different precisions (e.g., 4-bit, 8-bit, 16-bit, etc.)).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev with multiplying an integer by a floating point as taught by Divakar. One of ordinary skill in the art would be motivated to make this combination because it would allow for greater flexibility as it allows for multiple data formats to be multiplied together.
Willcock in view of Divakar fails to teach 8 bits.
However, Piskorski does teach 8 bits (Piskorski Page 2 Section II: It has a 5-bit exponent and a 7-bit mantissa (figure 1); Piskorski Fig. 1: shows the F13 format with a sign bit, a 5-bit exponent and a 7-bit mantissa).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teaching of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with 13 bit floating point format as taught by Piskorski. One of ordinary skill in the art would be motivated to make this combination because the 16-bit floating-point adder is the time-critical operator. Switching from F16 to F13 (7-bit mantissa) or F14 (8-bit mantissa) directly provides a speedup, since such a format saves one level of adders at each step of computation as taught by Piskorski (Piskorski Page 3 Section V B).
Willcock in view of Divakar further in view of Piskorski fails to teach using the first multiplier circuit.
However, Kaul does teach using the first multiplier circuit and the second multiplier circuit (Kaul [0163]: that is generated by a signed 16bx16b multiplier; Kaul Fig. 15A: shows the numbers being fed into one multiplier).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Willcock in view of Darvish further in view of Piskorski further in view of Rzayev further in view of Divakar with using the two multipliers as taught by Kaul. One of ordinary skill in the art would be motivated to make this combination because it would allow for portions of the numbers to be multiplied in parallel, increasing efficiency.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jakob O Gudas whose telephone number is (571)272-0695. The examiner can normally be reached Monday-Thursday: 7:30AM-5:00PM Friday: 7:30AM-4:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.O.G./Examiner, Art Unit 2151
/James Trujillo/Supervisory Patent Examiner, Art Unit 2151