Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The “activation accelerator” as recited in claim 1 is interpreted, under broadest reasonable interpretation in light of the Specification, to refer to hardware. The Specification at [0053] states “An activation accelerator is hardware that is configured to accelerate non-linear activation operations”. The units of the “one or more activation pipelines” of the “activation accelerator” in claim 1, under broadest reasonable interpretation in light of the Specification, are also hardware. The Specification at [0062] further states “Each activation pipeline 304 comprises a range conversion unit 306, an index generation unit 308, a LUT interface unit 310 and an interpolation unit 312. Each unit 306, 308, 310, 312 may be implemented in hardware”.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-6, 9, 12, and 14-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Hickman et al. (from IDS: US Pub. No. 2019/0384575, published Dec. 2019, hereinafter “Hickman”).
Regarding claim 1, Hickman teaches an activation accelerator for use in a neural network accelerator (Hickman, [0020] and [0026] – “Matrix processing unit 104 may include circuitry to perform functions to accelerate computations associated with matrices (e.g., for deep learning applications).”), the activation accelerator comprising:
a look-up table configured to store a plurality of values representing a non-linear activation function (Hickman, [0002] – “Examples of unary functions include transcendental functions (e.g., tanh, log 2, exp 2, sigmoid)” and [0017] – “Unary functions may be realized entirely or in part with lookup tables (LUTs) present in a processor.” – teaches a look-up table configured to store a plurality of values representing a non-linear activation function (non-linear activation functions such as tanh and sigmoid may be realized with lookup tables)); and
one or more activation pipelines (Hickman, [0017] – “Unary functions may be realized entirely or in part with lookup tables (LUTs) present in a processor. In some systems, the LUTs may also provide the flexibility required to implement customized functions. Some processors may provide multiple different tabulated functions (e.g., functions that may utilize a lookup based on the input) that are selectable via an instruction field.” and in [0025] - teaches one or more activation pipelines (tabulated functions selectable via instruction field, processing elements include one or more instruction pipelines)), each activation pipeline comprising:
a range conversion unit configured to:
receive an input value, an input offset, and an absolute value flag (Hickman, [0041] – “ Offset Value Register—This register stores an offset value (e.g., a pre-calculated integer value provided by a user) used to derive an input value's offset within the range.” and in [0044] – teaches receiving an input value, an input offset (stores offset value used to derived input value offset within range), and an absolute value flag (as in Hickman at [0044], teaches a symmetry mode register which, when set to y-axis, evaluates the function using the absolute value of the input, thus the symmetry mode is an absolute value flag)), and
generate a converted value from the input value, by, when the absolute value flag is not set, combining the input value and the input offset, and, when the absolute value flag is set, generating an absolute value of the combination of the input value and the input offset (Hickman, [0041] – “Offset Value Register—This register stores an offset value (e.g., a pre-calculated integer value provided by a user) used to derive an input value's offset within the range. In an embodiment, the pre-calculated integer value may be subtracted from the input value to determine the input value's offset into the range.” and in [0044] – teaches generating a converted value from the input value, by, when the absolute value flag is not set (symmetry mode set to none), combining the input value and the input offset (offset value subtracted from input value), and, when the absolute value flag is set (symmetry mode set to y-axis), generating an absolute value of the combination of the input value and the input offset (when symmetry mode is set to y-axis, function is evaluated using absolute value of input and offset to produce final output));
an index generation unit configured to generate an index and an interpolation value from the converted value (Hickman, [0017] – “For example, the index to the LUT may simply be a right-shifted input value and the output for the function may be the value present at that location in the LUT or a linear interpolation between the selected value and the following value.”, [0054] – “For example, the control module 412 may determine whether a lookup is to be performed. If a lookup is to be performed, the control module 412 calculates an address (depicted as “table index”) into the LUT 410 based on the input value x and information available in the control registers.”, and in [0064] – teaches an index generation unit configured to generate an index and an interpolation value from the converted value (control module calculates an address, or table index, based on input value, offset value, and information available in control registers, thus generating an index from the converted value. Output from LUT may be the value present at index in LUT or a linear interpolation between the selected value and the following value, thus generating an interpolation value from the converted value));
a look-up table interface unit configured to retrieve multiple values from the look-up table based on the index (Hickman, [0054] – “For example, the control module 412 may determine whether a lookup is to be performed. If a lookup is to be performed, the control module 412 calculates an address (depicted as “table index”) into the LUT 410 based on the input value x and information available in the control registers.”, and in [0064] – teaches a look-up table interface unit configured to retrieve multiple values from the look-up table based on the index (performs look-up in look-up table using address, or table index, to retrieve multiple values)); and
an interpolation unit configured to generate an estimate of a result of the non-linear activation function for the input value from an interpolation output generated by interpolating between the multiple values retrieved from the look-up table based on the interpolation value (Hickman, [0017] – “Unary functions may be realized entirely or in part with lookup tables (LUTs) present in a processor. In some systems, the LUTs may also provide the flexibility required to implement customized functions… For example, the index to the LUT may simply be a right-shifted input value and the output for the function may be the value present at that location in the LUT or a linear interpolation between the selected value and the following value.” – teaches an interpolation unit configured to generate an estimate of a result of the non-linear activation function (unary function realized on LUT) for the input value from an interpolation output generated by interpolating between the multiple values retrieved from the look-up table based on the interpolation value (output for the function may be the value present at that location in the LUT or a linear interpolation between the selected value and the following value, thus interpolating between multiple values retrieved from the LUT based on the linear interpolation)).
Claims 18-20 incorporate substantively all the limitations of claim 1 in a method and a non-transitory computer readable storage medium, and are rejected on similar grounds as above. Hickman teaches the storage medium, processors, activation accelerator, and integrated circuits of these claims at [0020] – [0026].
Regarding claim 2, Hickman teaches the activation accelerator of claim 1, wherein the range conversion unit is configured to generate the combination of the input value and the input offset by subtracting the input offset from the input value (Hickman, [0043] – “This register stores a value representing a shift amount applied to an input value's offset within a range (e.g., the value obtained after subtracting the offset value from the input value).” – teaches wherein the range conversion unit is configured to generate the combination of the input value and the input offset by subtracting the input offset from the input value (generates the combination of the input value and the offset by subtracting offset from the input)).
Regarding claim 3, Hickman teaches the activation accelerator of claim 2, wherein the range conversion unit is configured to generate the absolute value of the combination of the input value and the input offset by, when the input value is less than the input offset, subtracting the input value from the input offset, and, when the input value is greater than or equal to the input offset, subtracting the input offset from the input value (Hickman, [0041] and [0044] – teaches wherein the range conversion unit is configured to generate the absolute value of the combination of the input and the input offset by, when the input value is less than the input offset, subtracting the input value from the input offset, and, when the input value is greater than or equal to the input offset, subtracting the input offset from the input value (teaches subtracting offset from input and taking absolute value of the input, which is equivalent to subtracting input from offset and taking absolute value of the result)).
Regarding claim 4, Hickman teaches the activation accelerator of claim 1, wherein the input offset is a signed integer (Hickman, [0019] – “Any suitable FP numbers may be used in various embodiments, where a FP number may include significand (also referred to as mantissa) and exponent bits. The FP number may also include a sign bit.” and in [0041] – “This register stores an offset value (e.g., a pre-calculated integer value provided by a user) used to derive an input value's offset within the range.” – teaches wherein the input offset is a signed integer (teaches values that include a sign bit, and offset values may be an integer, thus teaching wherein the input offset is a signed integer)).
Regarding claim 5, Hickman teaches the activation accelerator of claim 1, wherein the absolute value flag is a Boolean value (Hickman, [0044] – “When the symmetry mode is none, no symmetry optimization is applied. When the symmetry mode is y-axis, the function may be evaluated using the absolute value of the input value.” – teaches wherein the absolute value flag is a Boolean value (when symmetry mode is set to none, or 0, no symmetry optimization is applied. When symmetry mode is set to y-axis, or 1, function is evaluated using absolute value for the input value)).
Regarding claim 6, Hickman teaches the activation accelerator of claim 1, wherein: the range conversion unit is further configured to generate a negative flag, wherein the negative flag is set if the combination of the input value and the input offset is negative, and the negative flag is not set if the combination of the input value and the input offset is not negative (Hickman, [0041] and [0044] – teaches wherein the range conversion unit is further configured to generate a negative flag, wherein the negative flag is set if the combination of the input value and input offset is negative, and the negative flag is not set if the combination of the input value and input offset is not negative (if input value is negative, subtracting offset from a negative input results in a negative value, and thus a negative flag is set and indicates that sign of output must be flipped when symmetry is set to origin mode. If input value is positive, then subtracting offset from positive value results in a positive value, and thus the negative flag is not set and indicates that the sign of the output does not need to be flipped when symmetry is set to origin mode)); and
the interpolation unit is further configured to receive a reverse sign flag and reverse a sign of the estimated result when the reverse sign flag is set, and the negative flag is set (Hickman, [0044] – “When the symmetry mode is origin, the function may be evaluated using the absolute value of the input value and the sign of the output is then flipped if the original input was negative to produce the final output.” – teaches wherein the interpolation unit is further configured to receive a reverse sign flag (symmetry mode set to origin, and thus a reverse sign flag is set) and reverse a sign of the estimated result when the reverse sign flag is set and the negative flag is set (if original input is negative, and thus a negative sign flag is set, then the sign of the output is flipped to produce a final output, thus reversing a sign of the estimated result when the negative flag and reverse sign flag, or origin symmetry mode, is set)).
Regarding claim 9, Hickman teaches the activation accelerator of claim 1, wherein:
the index generation unit comprises a clamp unit which is configured to receive a version of the converted value and a clamp flag (Hickman, [0040] and in [0048] – “Compression Mode—This register specifies a decompression algorithm for a lookup table entry (if the coefficients are compressed in the lookup table). In a particular embodiment, two compression modes are used.” – teaches wherein an index generation unit comprises a clamp unit which is configured to receive a version of the converted value (receives input value within a range to perform table lookup) and a clamp flag (defines two compression modes, thus receiving a clamp flag)); and,
when the clamp flag is set, clamp the version of the converted value to a clamp range to generate a clamp output, and when the clamp flag is not set, use all or a portion of the version of the converted value as the clamp output (Hickman, [0048] – “A first compression mode puts a limitation on the range of coefficients, but generates precise outputs, whereas a second compressions mode does not constrain the range of coefficients and thus allows a full range of floating point inputs, at the cost of less precise outputs.” – teaches when the clamp flag is set (first compression mode set), clamp the version of the converted value to a clamp range to generate a clamp output (puts limitation of range of coefficients and thus limits a range of inputs, thus teaching clamp the version of the converted value to a clamp range to generate a clamp output), and when clamp flag is not set (second compression mode, thus first compression mode is not set), use all or a portion of the version of the converted value as the clamp output (second compression mode do not constrain range of LUT outputs and thus allows a full range of floating point inputs, thus using all of the version of the converted value as the clamp output)); and
the index and the interpolation value are generated from the clamp output (Hickman, [0017], [0048], and in [0054] – teaches wherein the index and the interpolation value are generated from the clamp output (as in [0048] the compression mode defines if full range of floating inputs is enabled, and as in [0017] and [0054] determines the lookup table index and interpolation value from the full range of inputs or from the limited range, and thus the index and interpolation value are generated from the clamp output)).
Regarding claim 12, Hickman teaches the activation accelerator of claim 1, wherein the plurality of values are stored in the look- up table in accordance with a mode of a plurality of modes (Hickman, [0019], [0027], [0057] – “Although the FMAs 406 and 408 are illustrated as single precision FMAs, the FMAs may be configured to operate on any suitable number format. The output of FMA 408 may be supplied to post-processing module 414 for any processing to be performed before the final output value is output by the arithmetic engine 400.” – teaches wherein the plurality of values are stored in the look-up table in accordance with a mode of a plurality of modes (values in look-up table stored according to suitable number format of a plurality of suitable number formats, such as the number formats designated by Hickman at [0019])); and
the index generation unit is configured to receive information identifying the mode of the plurality of modes and generate the index and the interpolation value from the converted value based on the identified mode (Hickman, [0017], [0019] – “As described above, the input for a unary function may be a FP number. Any suitable FP numbers may be used in various embodiments, where a FP number may include significand (also referred to as mantissa) and exponent bits. The FP number may also include a sign bit. As various examples, the FP number may conform to a minifloat format (e.g., an 8 bit format), a half-precision floating-point format (FP16), a Brain Floating Point format (bfloat16), a single-precision floating-point format (FP32), a double-precision floating-point format (FP64), or other suitable FP format.”, and in [0054] – teaches wherein the index generation unit is configured to receive information identifying the mode of the plurality of modes and generate the index and the interpolation value from the converted value based on the identified mode (input received a suitable FP number format of a plurality of suitable FP number formats, and as in [0054] table index is generated from input and information available in control registers and thus the index is generated from the converted value based on the identified mode, or number format. Linear interpolation is generated from LUT values corresponding to index which is generated from the input and information available at corresponding registers, thus the linear interpolation is generated from the converted value based on the identified mode, or number format)).
Regarding claim 14, Hickman teaches the activation accelerator of claim 1, wherein the interpolation unit is configured to receive information identifying a power of two to be applied to the interpolation output and apply the power of two to the interpolation output (Hickman, [0038] – “The lookup mode specifies that a table lookup should be performed and the resulting power series should be calculated to generate the output of the function.” and in [0047] – “ Function Mode Register—specifies function-specific optimization (if any). For example, for some well known functions (e.g., sqrt(x), 1/x, 1/sqrt(x), log.sub.2(x), 2.sup.x), the exponent of the result (or a value very close to it) can be derived algorithmically with reasonably trivial extra logic (e.g. 8-bit integer addition).” – teaches wherein the interpolation unit is configured to receive information identifying a power of two to be applied to the interpolation output and apply the power of two to the interpolation output (resulting power series is calculated to generate output of the function, the exponent of the result of a well-known function such as 2.sup.x may be derived algorithmically, thus teaching receiving information identifying a power of two, 2.sup.x, to be applied to the interpolation output, output of the function, and applying the power of two to the interpolation output, exponent of result derived algorithmically)).
Regarding claim 15, Hickman teaches the activation accelerator of claim 1, wherein the input value is a 32-bit value (Hickman, [0019] – “As described above, the input for a unary function may be a FP number. Any suitable FP numbers may be used in various embodiments, where a FP number may include significand (also referred to as mantissa) and exponent bits. The FP number may also include a sign bit. As various examples, the FP number may conform to … a single-precision floating-point format (FP32)…” – teaches wherein the input value is a 32-bit value).
Regarding claim 16, Hickman teaches the activation accelerator of claim 1, wherein the input offset is a 32-bit value (Hickman, [0019], [0041], and in [0057-0058] – “Although the FMAs 406 and 408 are illustrated as single precision FMAs, the FMAs may be configured to operate on any suitable number format.” – teaches wherein the input offset is a 32-bit value (input values may be a FP number, any suitable FP number may be used in various embodiments, determines an offset value for input values, FMAs configured to operate on any suitable number format, thus the input offset may be a 32-bit value)).
Regarding claim 17, Hickman teaches the activation accelerator of claim 1, wherein the activation accelerator comprises a plurality of activation pipelines and each activation pipeline of the plurality of activation pipelines has access to a look-up table that is configured to store the plurality of values representing the non-linear activation function (Hickman, [0002], [0017] – “Unary functions may be realized entirely or in part with lookup tables (LUTs) present in a processor. In some systems, the LUTs may also provide the flexibility required to implement customized functions. Some processors may provide multiple different tabulated functions (e.g., functions that may utilize a lookup based on the input) that are selectable via an instruction field.”, and in [0025] – teaches wherein the activation accelerator comprises a plurality of activation pipelines and each activation pipeline of the plurality of activation pipelines has access to a look-up table that is configured to store the plurality of values representing the non-linear activation function (processors provide multiple different tabulated functions, thus a plurality of activation pipelines, and the multiple different tabulated functions may utilize a lookup based on the input, thus each activation pipeline of the plurality of activation pipelines has access to a lookup table. The lookup table is configured to store a plurality of values representing a unary function, which is a non-linear activation function)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hickman in view of Prebeck et al. (NPL from IDS: A Smart HW-Accelerator for Non-uniform Linear Interpolation of ML-Activation Functions, published Aug. 2022, hereinafter “Prebeck”).
Regarding claim 7, Hickman teaches the activation accelerator of claim 6.
Hickman fails to explicitly teach wherein the interpolation unit is further configured to receive an output offset and, generate the estimated result by combining the interpolation output and the output offset, wherein the interpolation unit is configured to, when the reverse sign flag is not set or the negative flag is not set, generate the estimate result by adding the output offset and the interpolation output, and, when the reverse sign flag is set and the negative flag is set, generate an estimated result with a reverse sign by subtracting the interpolation output from the output offset.
However, analogous to the field of the claimed invention, Prebeck teaches: wherein
the interpolation unit is further configured to receive an output offset and, generate the estimated result by combining the interpolation output and the output offset (Prebeck, Section 3.3 Paragraph 2 – “With an eye toward further optimizing the datapath, x- and y-offsets are introduced as reconfigurable features. Some functions exhibit symmetry around a shifted origin. For instance, the Sigmoid function is symmetric with respect to an origin that is shifted by 0.5 in the positive -direction. In this regard, offsetting in both x- and y-directions is in some functions essential to preserve the optimizations offered by symmetry. Additionally, x-offset is used to displace the x-values of the considered interpolation points and bring them closer to zero.” and in Section 3.5 Paragraph 2 – “In the same manner, the offset feature demands a subtractor from either the input side if the offset is in the x-direction, or from the output side if it’s in the y-direction, or from both sides if it’s in the x- and y-directions.” – teaches wherein the interpolation unit is further configured to receive an output offset (y-offset) and, generate the estimated result by combining the interpolation output and the output offset (combines y-offset with y-values of considered interpolation points, thus generating an estimated result by combining the interpolation output and the y-offset)),
wherein the interpolation unit is configured to, when the reverse sign flag is not set or the negative flag is not set, generate the estimate result by adding the output offset and the interpolation output, and, when the reverse sign flag is set and the negative flag is set, generate an estimated result with a reverse sign by subtracting the interpolation output from the output offset (Prebeck, Section 3.3 Paragraph 2 – “With an eye toward further optimizing the datapath, x- and y-offsets are introduced as reconfigurable features. Some functions exhibit symmetry around a shifted origin. For instance, the Sigmoid function is symmetric with respect to an origin that is shifted by 0.5 in the positive -direction. In this regard, offsetting in both x- and y-directions is in some functions essential to preserve the optimizations offered by symmetry. Additionally, x-offset is used to displace the x-values of the considered interpolation points and bring them closer to zero.” and in Section 3.5 Paragraph 2 – “In the same manner, the offset feature demands a subtractor from either the input side if the offset is in the x-direction, or from the output side if it’s in the y-direction, or from both sides if it’s in the x- and y-directions.” – teaches wherein the interpolation is configured to, when the reverse sign flag is not set or the negative flag is not set (shift in the positive-direction), generate the estimate result by adding the output offset and the interpolation output (adds y-offset to y-value in positive direction), and, when the reverse sign flag is set and the negative flag is set (thus shift in the negative direction), generate an estimated result with a reverse sign by subtracting the interpolation output from the output offset (adds y-offset to y-value that is in the negative direction, which is equivalent to subtracting the y-value from the y-offset)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the output offset and estimated results of Prebeck to the interpolation outputs, estimated results, and flags of Hickman. Doing so would utilize offsets to preserve optimizations offered by symmetry (Prebeck, Section 3.3) and cause a reduction in total area footprint thus leading to smaller comparator sizes (Prebeck, Section 4.2).
Regarding claim 8, the combination of Hickman and Prebeck teaches the activation accelerator of claim 7, wherein the output offset is a signed integer (Prebeck, Fig. 2 and in Section 3.4 Paragraph 2 – “To restrain from burdening the system with complex components like an FPU, real values are converted to integer values. The domain of the converted integer values grows by a factor of 10 for every considered decimal point. This factor is multiplied by all FP -intercept and interpolation -values that result from the search-based optimization algorithm.” – teaches wherein the output offset is a signed integer (real values are converted to integer values)).
Therefore, it would have been obvious to person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the output offset of Prebeck to the outputs, flags, and activation accelerator of Hickman. Doing so would utilize offsets to preserve optimizations offered by symmetry (Prebeck, Section 3.3) and cause a reduction in total area footprint thus leading to smaller comparator sizes (Prebeck, Section 4.2).
Claim(s) 10-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hickman in view of Li et al. (FOR from IDS: EP 3480742 A1, published May 2019, hereinafter “Li”).
Regarding claim 10, Hickman teaches the activation accelerator of claim 9.
Hickman fails to explicitly teach the index generation unit is configured to receive information identifying a first subset of bits of the converted value and a second subset of bits of the converted value, generate the index from the first subset of bits of the converted value, and generate the interpolation value from the second subset of bits of the converted value, wherein the information identifying the first subset of bits of the converted value and the second subset of bits of the converted value comprises a shift amount and a shift direction, and the index generation unit comprises an alignment unit configured to generate the version of the converted value by shifting the converted value in the identified shift direction by the identified shift amount.
However, analogous to the field of the claimed invention, Li teaches: wherein
the index generation unit is configured to receive information identifying a first subset of bits of the converted value and a second subset of bits of the converted value, generate the index from the first subset of bits of the converted value, and generate the interpolation value from the second subset of bits of the converted value (Li, [0013] – A predefined number of most significant bits of the first input value may be used as the lookup address into the lookup table and the remaining bits of the first input value are used in the interpolating between the corresponding pair of values of the activation function.” – teaches wherein the index generation unit is configured to receive information identifying a first subset of bits of the converted value and a second subset of bits of the converted value, generate the index from the first subset of bits of the converted value, and generate the interpolation value from the second subset of bits of the converted value (predefined number of most significant bits used as lookup address, or index, and remaining bits are used as the interpolation value)),
wherein the information identifying the first subset of bits of the converted value and the second subset of bits of the converted value comprises a shift amount and a shift direction, and the index generation unit comprises an alignment unit configured to generate the version of the converted value by shifting the converted value in the identified shift direction by the identified shift amount (Li, [0013] and in [0080] – “In some implementations the interval may be variable such that a greater number of data points are provided where the underlying activation function is changing more quickly. If the interval is fixed, there is the advantage that the value of x.sub.1 - x.sub.0 is represented by the value of 1 left-shifted by the number of most significant bits used as a lookup into the lookup table. This allows the above equation to be more simply implemented in hardware as follows: y=y0+y1−y0∗LSBs>>number of MSBs” – teaches wherein the information identifying the first subset of bits of the converted value and the second subset of bits of the converted value comprises a shift amount and a shift direction (the value of x.sub.1 - x.sub.0 is represented by the value of 1 left-shifted by the number of most significant bits, and as in equation, LSBs>>number of MSBs, thus teaching a shift amount and a shift direction based on the subsets of bits), and the index generation unit comprises an alignment unit configure to generate the version of the converted value by shifting the converted value in the identified direction by the identified shift amount (the value of x.sub.1 - x.sub.0 is represented by the value of 1 left-shifted by the number of most significant bits used as a lookup into the lookup table, thus generating the converted value by shifting the converted value in the identified (left) direction by the identified shift amount)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the first and second subsets of bits and alignment unit of Li to the converted value and activation accelerator of Hickman. Doing so would provide efficient implementation of deep neural networks in a power efficient manner (Li,[0004]) and provide a simpler implementation of linear interpolation in hardware (Li, [0080]).
Regarding claim 11, Hickman teaches the activation accelerator of claim 9.
Hickman fails to explicitly teach wherein the index generation unit further comprises a split unit configured to extract a first predetermined set of bits from the clamp output as the index, and a second predetermined set of bits from the clamp output as the interpolation value.
However, analogous to the field of the claimed invention, Li teaches: wherein
the index generation unit further comprises a split unit configured to extract a first predetermined set of bits from the clamp output as the index, and a second predetermined set of bits from the clamp output as the interpolation value (Li, [0013] – A predefined number of most significant bits of the first input value may be used as the lookup address into the lookup table and the remaining bits of the first input value are used in the interpolating between the corresponding pair of values of the activation function.” and in [0033] – “On the DNN requiring the activation module to implement an activation function using the lookup table, the ReLU unit may be configured to clamp input values received at the activation module which lie outside the determined range of input values at the closest extreme of the determined range of input values, the clamped input values being subsequently passed to the lookup table for implementation of the activation function.” – teaches wherein the index generation unit further comprises a split unit configured to extract a first predetermined set of bits from the clamp output as the index, and a second predetermined set of bits from the clamp output as the interpolation value (ReLU unit configured to clamp input values received at activation modules, thus generating clamp outputs, and a predefined number of most significant bits of first input, or clamp output, is used as a lookup address, or index, into lookup table and the remaining bits of first input, or clamp output, are used as an interpolation value)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the first and second predetermined set of bits of Li to the clamp output and interpolation values of Hickman. Doing so would provide efficient implementation of deep neural networks in a power efficient manner (Li,[0004]) and provide a simpler implementation of linear interpolation in hardware (Li, [0080]).
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hickman in view of Nottbeck et al. (NPL from IDS: Implementation of high-performance, sub-microsecond deep neural networks on FPGAs for trigger applications, published Aug. 2019, hereinafter “Nottbeck”).
Regarding claim 13, Hickman teaches the activation accelerator of claim 1.
Hickman fails to explicitly teach wherein the interpolation unit is configured to generate the interpolation output by interpolating between the plurality of values from the look-up table based on the interpolation value and rounding in accordance with a rounding mode, wherein the interpolation unit is further configured to receive information identifying a rounding mode of a plurality of rounding modes and perform the rounding in accordance with the identified rounding mode.
However, analogous to the field of the claimed invention, Nottbeck teaches: wherein
the interpolation unit is configured to generate the interpolation output by interpolating between the plurality of values from the look-up table based on the interpolation value and rounding in accordance with a rounding mode (Nottbeck, Section 3.5 Paragraph 3 – “The multiplication of two 8 bit values and final addition would cost approximately 70 LUTs, leaving a total of ∼ 140 LUTs per activation unit to implement a relatively precise look-up with 256 sample points and linear interpolation in between. Figure 7 shows a case study, where we linearly interpolated the comparably complicated tanh and sigmoid functions. To demonstrate the loose precision requirements, we used only sixteen sample points within the shown intervals, and rounded to a granularity of only 2−6 during all steps of the interpolation (i.e. the value and derivate samples themselves were rounded, the interpolation position was rounded and multiplication and addition results were rounded).” – teaches wherein the interpolation unit is configured to generate the interpolation output by interpolating between the plurality of values from the look-up table based on the interpolation value (generates interpolation output by interpolating between plurality of values from the look-up table, such as 256 sample points and interpolating in between, based on the interpolation value) and rounding in accordance with a rounding mode (rounded to a granularity of only 2-6 during all steps of the interpolation)),
wherein the interpolation unit is further configured to receive information identifying a rounding mode of a plurality of rounding modes and perform the rounding in accordance with the identified rounding mode (Nottbeck, Section 3.5 Paragraph 3 – “The multiplication of two 8 bit values and final addition would cost approximately 70 LUTs, leaving a total of ∼ 140 LUTs per activation unit to implement a relatively precise look-up with 256 sample points and linear interpolation in between. Figure 7 shows a case study, where we linearly interpolated the comparably complicated tanh and sigmoid functions. To demonstrate the loose precision requirements, we used only sixteen sample points within the shown intervals, and rounded to a granularity of only 2−6 during all steps of the interpolation (i.e. the value and derivate samples themselves were rounded, the interpolation position was rounded and multiplication and addition results were rounded).” – teaches wherein the interpolation unit is further configured to receive information identifying a rounding mode of a plurality of rounding modes (rounded to granularity of 2−6 of a plurality of rounding modes, loose precision requirements) and perform the rounding in accordance with the identified rounding mode (rounded to granularity 2−6 during all steps of interpolation)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the rounding modes and rounding of interpolation outputs of Nottbeck to the interpolation outputs, interpolation values, and accelerator of Hickman. Doing so would provide the benefit of reduced implementation complexity, which saves resources (Nottbeck, Section 2.3) and would provide accurate approximations of non-linear functions while utilizing low precision and low sample points (Nottbeck, Section 3.5).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Pasca et al. (US Pub. No. 2026/0267369, filed May 2023) teaches systems and methods for approximating numerical computations in deep neural networks using a look-up table based on input elements. Teaches numerical computation approximator circuitry to perform softmax-based, or other non-linear function, lookup table generation. Teaches range-reduction for degree-d approximation based on quantized input data.
Abdelsalam et al. (NPL: An Efficient FPGA-based Overlay Inference Architecture for Fully Connected DNNs, published 2018) teaches a Single hidden layer Neural Network multiplication-free overlay architecture with fully-connected DNN-level performance. Teaches quantizing inputs with stochastic rounding, and provides a hardware implementation of the tanh function using a Look-Up Table. Teaches wherein the tanh function is quantized for positive inputs.
Ortega-Zamorano et al. (NPL: High precision FPGA implementation of neural network activation functions, published 2014) teaches FPGA implementations of Sigmoid and Exponential functions using a lookup tale with a linear interpolation procedure. Teaches wherein the size of the lookup table is determined by the precision of the inputs. Teaches determining first and second subsets of bits of the input.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOUIS C NYE whose telephone number is 571-272-0636. The examiner can normally be reached Monday - Friday 9:00AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOUIS CHRISTOPHER NYE/Examiner, Art Unit 2141
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141