DETAILED ACTION
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 31-60 are rejected under 35 U.S.C. 103 as being unpatentable over Heinecke (US 2019/0,079,767) in view of Oklobdzija (US 11,429,349).
Referring to claims 31, 37, 43 and 49, Heinecke discloses an apparatus (fig. 1, system 100) comprising:
decoder circuitry (fig. 1, decode circuitry 109) to decode a single instruction (fig. 2, instruction 201), the single instruction include fields for an opcode (fig. 2, opcode 202), an indication of a first plurality of packed data source Bfloat16 BF16 data elements (fig. 2, first source location 206/212A), an indication of a second plurality of packed data source BF16 data elements (fig. 2, second source location 208/212B) and an indication of a third plurality of packed data source BF16 data elements (fig. 2, destination location 204/218).
Oklobdzija discloses an indication of a first plurality of packed data source Bfloat16 BF16 data elements (fig. 2, operand A), an indication of a second plurality of packed data source BF16 data elements (fig. 2, operand-B) and an indication of a third plurality of packed data source BF16 data elements (fig. 2, operand-C), wherein each of the BF16 packed data source data elements (fig. 1, BFloat16 110) comprises an 8-bit exponent value (4:2-9, 8-bit exponent), a 7-bit mantissa value (4:2-9, 7-bit significand), and a 1-bit sign value (4:2-9, one sign bit);
execution circuitry (fig. 2, floating point multiply-add, accumulate unit with carry-save accumulator in BF16 and FP32 format; 4:22-24) to execute the decoded single instruction according to the opcode to perform a fused multiply accumulate operation (fig. 2, operations) using the first, second and third packed data source BF16 data elements to generate a plurality of packed data result BF16 data elements (fig. 2, final result in FP32 or BF16 format 229),
wherein to perform the fused multiply accumulate operation, the execution circuitry is to:
multiply (fig. 2, multiplier 210) the first plurality of packed data source BF16 data elements (fig. 2, operand-A) with corresponding packed data source BF16 data elements of the second plurality of packed data source BF16 data elements (fig. 2, operand-B) to generate a corresponding plurality of products (fig. 2, output 221); and
add (fig. 2, adder 230 and accumulator 240) each product of the corresponding plurality of products (fig. 2, output 221) with a corresponding packed data source BF16 data element of the third plurality of packed data source BF16 data elements (fig. 2, operand-C) to generate a corresponding packed data result BF16 data element of the plurality of packed data result BF16 data elements (fig. 2, C/S-ACC output 226).
Heinecke and Oklobdzija are analogous art because they are from the same field of endeavor in floating point vector instructions. Before the time of the filing, it would have been obvious to a person of ordinary skill in the art, having the teaching of Heinecke and Oklobdzija before him or her to modify the execution circuitry of Heinecke to include the floating-point multiply-add accumulate unit of Oklobdzija, thereafter the floating-point numbers execution of multiply-add in BF16 and FP32 are implemented as a single instruction hardware. The suggestion and/or motivation for doing so would be obtaining the advantage of improved operation speed (2:6-11) as suggested by Oklobdzija. Therefore, it would have been obvious to combine Heinecke with Oklobdzija to obtain the invention as specified in the instant application claims.
As to claims 32, 38, 44 and 50, Heinecke discloses the apparatus of claim 31, comprising:
a first vector register (fig. 2, left of second source 212B) to store a portion of the second plurality of packed data source BF16 data elements (fig. 2, second source 208); and
a second vector register (fig. 2, right of second source 212B) to store a portion of the second plurality of packer source BF16 data elements (fig. 2, second source 208).
As to claims 33, 39, 45 and 51, Heinecke discloses the apparatus of claim 32, comprising:
a third vector register (fig. 2, destination 218) to store a portion of the third plurality of packed data source BF16 data elements (fig. 2, destination 204).
As to claims 34, 40, 46 and 52, Heinecke discloses the apparatus of claim 33, wherein the third vector register is to additionally store a corresponding portion of the plurality of packed data result BF16 data elements (fig. 2, destination 218 stores result of accumulator 216).
As to claims 35, 41, 47 and 53, Heinecke discloses the apparatus of claim 31, wherein the opcode is to indicate the fused multiply accumulation operation (fig. 2, execution 214; fig. 3A, FP32 FMA, Fused Multiply-Add) is a per-data element position multiplication (fig. 2, multiply 215A/215B) of the first plurality of packed data source BF16 data elements (fig. 2, first source 212A) with the corresponding plurality of packed data source BF16 data elements to generate the corresponding plurality of products with infinite precision (para.0046, infinite precision).
As to claims 36, 42, 48 and 54, Heinecke discloses the apparatus of claim 35, wherein each product of the corresponding plurality of product is added (fig. 2, adder 216) to the corresponding packed data source BF16 data element of the third plurality of packed data source BF16 data elements (fig. 2, destination 218) to generate an infinite precision addition result (para.0046, infinite precision).
Oklobdzija discloses result is to be rounded (fig. 9A, rounding 270a) to generate a corresponding packed data result BF16 data element (fig. 9A, conversion into BF16 270a) of the plurality of packed data result BF16 data elements. (See TSM analysis above.)
As to claim 55, Heinecke discloses the system of claim 49, wherein the decoder circuitry and execution circuitry are a part of a core of a multicore processor (fig. 11, processor 1100).
As to claim 56, Heinecke discloses the system of claim 49, wherein the execution is to round the corresponding packed data result BF16 data element of the plurality of packed data result BF16 data elements.
As to claim 57, Oklobdzija discloses the system of claim 56, wherein the execution circuitry is to round (fig. 9A, rounding 270a) to a nearest even. (See TSM analysis above.)
As to claim 58, Heinecke discloses the system of claim 49, wherein the single instruction is to include a field for a predication register (para.0140, predicating resulting vector).
As to claim 59, Heinecke discloses the system of claim 49, wherein the single instruction is to include a field for a write mask register (fig. 3A, writemask).
As to claim 60, Heinecke discloses the system of claim 49, wherein the execution circuitry is fused multiply-accumulate circuitry (fig. 2, execution 214; fig. 3A, FP32 FMA, Fused Multiply-Add).
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner Cheng-Yuan Tseng whose telephone number is (571)272-9772, and fax number is (571)273-9772. The examiner can normally be reached on Monday through Friday from 09:00 to 17:30 Eastern Time. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached on (571)272-2330. The fax phone number for the organization where this application or proceeding is assigned is (571)273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at (866)217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call (800)786-9199 (IN USA OR CANADA) or (571)272-1000.
/CHENG YUAN TSENG/Primary Examiner, Art Unit 2615