DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The abstract of the disclosure is objected to because of the following informalities:
Second-to-last line: An “a” is missing before “product matrix” and should be added.
A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
The disclosure is objected to because of the following informalities:
[0053]: “be8-bit signed integer” is missing a space and should be corrected to read “be 8-bit signed integer”.
[00137]: Insert “computing” after second instance of “(throughput)”.
[00147]: “a register maps” is grammatically incorrect.
[00226]: The example embodiments include language used in the claims. For similar reasoning set forth in the objections/rejections below, these paragraphs should be updated as the claims are updated, particularly where incorrect or unclear.
Appropriate correction is required.
Drawings
The drawings are objected to because of the following informalities:
Fig. 2: Change “THIRD VECTOR REGISTER 208” to “THIRD VECTOR REGISTER 210” to be consistent with the specification.
Fig. 6, 626: Remove the second instance of “FETCH SINGLE INSTANCE…”
Fig. 6, 626: Delete first instance of “HAVING” in line 3.
Fig. 7, 733: Insert “FLOATING-POINT” between “SINGLE” and “INSTRUCTION” for all instances of “SINGLE INSTRUCTION” to match with “SINGLE FLOATING-POINT INSTRUCTION” in 734.
Fig. 7, 733: Delete “INSTRUCTION FETCH INSTANCE OF” in line 2.
Fig. 7, 733: Delete second instance of “HAVING” in line 3.
Examiner makes the following recommendations to the drawings:
Figs. 6 and 7: Insert articles (a, an, the, etc.) prior to each noun to improve the readability of the drawings.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 4, 13-15, and 18 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1, 1, 1, 1, 25, and 24 of U.S. Patent No. 12,314,717 (hereinafter Pat. ‘717) in view of Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers) and Arm (Arm A64 Instruction Set for A-profile architecture).
Regarding claim 1, Pat. ‘717 teaches an apparatus comprising:
decoder circuitry to decode an instruction (See claim 1), the instruction to indicate a first register to store a first matrix having two rows by eight columns , to indicate a storage location to store a second matrix having eight rows by two columns data elements, and to indicate a vector register to store a third matrix having two rows by two columns of data elements (Claim 1: The first matrix is an MxK matrix, which would include the matrix being a 2x8 matrix. The second matrix is a KxN matrix, which would include the matrix (with K=8) being an 8x2 matrix. The combination of matrices would indicate the third matrix being a 2x2 matrix. The second matrix is held in a register, which is a type of storage location); and
execution circuitry coupled with the decoder circuitry, the execution circuitry to perform operations corresponding to the instruction (see claim 1), including to:
generate a result matrix having two rows by two columns , the result matrix representing an accumulation of the third matrix with a product matrix generated from a matrix multiplication using the first and second matrices (Claim 1: A dot product of the first matrix and the second matrix is performed, before accumulating the resultant of the dot product with the third matrix to generate a result matrix).
Pat. ‘717 does not teach that the data elements of the first and second matrices are 8-bit floating-point data elements.
Note that the claim indicates the third matrix holding data elements four times the size of the data elements in the first and second matrices (see claim 1)
Wang teaches 8-bit floating point data elements and 32-bit floating point data elements (Pages 1-2 and 5-7, Under “Abstract”, “Introduction”, and “Experimental Results”:8-bit floating-point (FP) data elements may be used as input operands in computations such as GEMM computations. Single-precision (32-bits) is commonly used for DNN training, which is more accurate but also more computationally intensive compared to FP8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Pat. ‘717 with the teachings of Wang to have the first and second matrices comprise of 8-bit floating-point data elements and the third matrix comprise of 32-bit (single-precision) floating-point data elements (resulting in a 128-bit size register for each register). One of ordinary skill may prefer using 8-bit floating-point data elements as they can be used for training and inferencing of neural networks.
Pat. ‘717, in view of Wang, still does not teach that the one or more registers are vector registers.
Arm teaches to hold matrices in vector registers (Page 1933: An instruction, such as BFMMLA, uses vector registers for a first matrix operand (indicated as Zn), a second matrix operand (indicated as Zm), and the result/accumulator (indicated as Zda)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang, with the teachings of Arm to have stored the matrices in vector registers. By storing a matrix in a vector register, it would remove the need to have multiple registers correspond to parts of a matrix, which may be preferred by one of ordinary skill.
Pat. ‘717, in view of Wang and Arm, still does not teach to store the result matrix in the vector register comprising of the third matrix.
Arm also teaches to store the result matrix in the vector register comprising of the third matrix (Page 1933: After the operations of the BFMMLA instruction has completed, the results are stored in the Zda operand (which is a vector register)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further combined the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Arm to have stored the result matrix in the third vector register. One of ordinary skill would have recognized that to use the results of the instruction executed for further, it would need to be written back in a location such as a register.
Regarding claim 4, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix, and the 8-bit floating-point data elements of the second matrix, each have five exponent bits and two explicit mantissa bits (Pat. ‘717, claim 1; Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Regarding claim 13, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1, wherein the instruction allows the storage location to be a third vector register but does not allow the storage location to be in memory (Pat. ‘717, claim 1; Arm, Page 1933: The BFMMLA instruction only uses a scalable vector register for the second matrix as the second operand Zm. Therefore, the instruction does not allow the second operand Zm to be in memory).
Regarding claim 14, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm, does not currently teach that the first vector register has a second 128-bit lane to store a fourth matrix having two rows by eight columns of 8-bit floating-point data elements, the storage location has a second 128 bits to store a fifth matrix having eight rows by two columns of 8-bit floating-point data elements, and the second vector register has a second 128-bit lane to store a sixth matrix having two rows by two columns of 32-bit single-precision floating-point data elements, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is further to:
generate a second result matrix having two rows by two columns of 32-bit single- precision floating-point result data elements, the second result matrix representing an accumulation of the sixth matrix with a product matrix generated from a matrix multiplication using the fourth and fifth matrices; and
store the second result matrix in the second 128-bit lane of the second vector register.
Arm also teaches that the first vector register has a second 128-bit lane to store a fourth matrix having two rows by four columns of 16-bit floating-point data elements (Page 1933-1934: The input vector register operand Zn can include a second 128-bit segment (i.e., a lane) to include a second 4x2 matrix of 16-bit floating-point data elements (see pseudocode loop). The second matrix in the vector register operand Zn as the fourth matrix), the storage location has a second 128 bits to store a fifth matrix having four rows by two columns of 16-bit floating-point data elements (Page 1933-1934: The input vector register operand Zm can include a second 128-bit segment to include a second 2x4 matrix of 16-bit floating-point data elements (see pseudocode loop). The second matrix in the vector register operand Zm as the fifth matrix), and the second vector register has a second 128-bit lane to store a sixth matrix having two rows by two columns of 32-bit single-precision floating-point data elements (Page 1933-1934: The output vector register operand Zda can include a second 128-bit segment (i.e., a lane) to include a second 2x2 matrix of 32-bit (i.e., single precision floating-point data elements (see pseudocode loop). The second matrix in the vector register operand Zda as the sixth matrix);
generate a second result matrix having two rows by two columns of 32-bit single- precision floating-point result data elements, the second result matrix representing an accumulation of the sixth matrix with a product matrix generated from a matrix multiplication using the fourth and fifth matrices (Page 1933: The BFMMLA instruction will multiply both second matrices from both operands Zn and Zm to generate an intermediate product value, before accumulating the intermediate product value into a resulting 2x2 single-precision matrix in the second segment of operand Zda); and
store the second result matrix in the second 128-bit lane of the second vector register (Page 1933: The resulting 2x2 32-bit floating-point matrix is to be stored in the second segment of the destination operand Zda).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Arm to have included a fourth matrix having two rows by eight columns of 8-bit floating-point data elements in the first vector register, to include a fifth matrix having eight rows by two columns of 8-bit floating-point data elements in the storage location, to include a sixth matrix having two rows by two columns of 32-bit single-precision floating-point data elements in the second register, to generate a second result matrix having two rows by two columns of 32-bit single-precision floating-point result data elements, the second result matrix representing an accumulation of the sixth matrix with a product matrix generated from a matrix multiplication using the fourth and fifth matrices, and storing the second result matrix in the second 128-bit lane of the second vector register. Given the circuitry of Pat. ‘717, in view of Wang and Arm, one of ordinary skill would have recognized that to generate a second result matrix using the product of fourth and fifth matrix and accumulation of the sixth matrix, they would need to duplicate the circuitry to accommodate the additional inputs. Hence, duplication of parts, i.e., duplicating the multiply and accumulate circuitry, and changes in size/proportion, i.e., changing the sizes of the vector register and storage location to accommodate the additional matrices, are deemed routine expedients, not patentable distinctions (MPEP 2144.04(IV)(A) and MPEP 2144.04(VI)(B)).
Regarding claim 15, the claim is rejected on the same premises of claim 1 using claim 25 of Pat. ‘717.
Regarding claim 18, the claim is rejected on the same premises of claim 1 using claim 24 of Pat. ‘717.
Claims 2, 12, and 17 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1, 1, and 25 of Pat. ‘717 in view of Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Arm (Arm A64 Instruction Set for A-profile architecture), and Boswell, et al. (US 20180321938 A1).
Regarding claim 2, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1, wherein the execution circuitry, to generate and store the result matrix, is to:
for each column n of the two columns of the second matrix, and for each row m of the two rows of the first matrix (Pat. ‘717, see claim 1):
generate eight products, including to multiply the eight data elements corresponding to the row m and the eight data elements corresponding to the column n;
generate a 32-bit single-precision floating-point result data element, including to accumulate the eight products with a data element from a corresponding row m of the two rows, and a corresponding column n of the two columns, of the third matrix (Pat. ‘717, claim 1: Given that the first matrix can be a 2x8 matrix and the second matrix can be an 8x2 matrix, eight dot products would be generated and accumulated alongside a corresponding element of the 2x2 third matrix to generate a 32-bit floating-point result element); and
store the 32-bit single-precision floating-point result data element in the 128-bit lane of the third vector register at a position corresponding to the row m and the column n of the third matrix (Pat. ‘717; claim 1, Arm, Page 1933-1934: In the current combination each 32-bit floating-point result element is to be stored in their respective spot in the 2x2 result matrix).
Pat. ‘717, in view of Wang and Arm, does not teach to convert eight data elements from the row m of the first matrix to eight corresponding converted data elements each having more than eight bits, and convert eight data elements from the column n of the second matrix to eight corresponding converted data elements each having more than eight bits.
Boswell teaches to convert data elements of the first matrix to converted data elements each having more than eight bits, and convert eight data elements of the second matrix to eight corresponding converted data elements each having more than eight bits (Figs. 7 and 13, [0137]: In pipeline stage 1301, the conversion/encoding logic 1315 receives elements of two input vectors A and B (which comprise of matrices, see Fig. 7), and convert all input values to half-precision floating-point format, which indicates that each of the elements are 16-bits).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Boswell to have converted the data elements of the first and second matrix prior to performing matrix multiplication and accumulation. A processor may not have the hardware to perform on 8-bit floating-point data types, therefore, one of ordinary skill may be inclined to increase the data type to a type that a processor has the capability to perform on, which may be preferred by one of ordinary skill.
Regarding claim 12, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm, does not teach to generate the result matrix, is to generate all products of the matrix multiplication using the first and second matrices before accumulation of any of said all products of the matrix multiplication with the third matrix.
Boswell teaches to generate all products of a first matrix and a second matrix before accumulation of any of said all products of the matrix multiplication with the third matrix (Fig. 13 and [0136-0139]: The figure shows a pipeline where at the second stage pipeline 1302, the multiplication of operands A and B (which are matrices) is to be performed prior to the accumulation of products with the accumulator (operand C, which is a matrix), which occurs at the third stage pipeline 1303).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Boswell to have generated all products of the matrix before accumulating said all products with the third matrix. By having the execution circuitry produce the product of the two source matrices prior to accumulating with the accumulator matrix, it would allow pipelining of the execution circuitry with respect to the multiply-accumulate circuitry (See Fig. 13, pipeline stage 1302 and pipeline stage 1303), which may be preferred by one of ordinary skill.
Regarding claim 17, the claim is rejected on the same premises of claim 12 using claim 25 of Pat. ‘717.
Claim 3 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of Pat. ‘717 in view of Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Arm (Arm A64 Instruction Set for A-profile architecture), and Sun et al. (Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks).
Regarding claim 3, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix, and the 8-bit floating-point data elements of the second matrix, each have five exponent bits and two explicit mantissa bits (Pat. ‘717, claim 1; Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Pat. ‘717, in view of Wang and Arm, does not teach that the floating-point data elements of the matrices each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Sun to have made the 8-bit floating-point data elements have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
Claims 5-6, 16, and 19 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 1, 25, and 24 of Pat. ‘717 in view of Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Arm (Arm A64 Instruction Set for A-profile architecture), Sun et al. (Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks), and Van Zee (Supporting mixed-datatype matrix multiplication within the BLIS framework).
Regarding claim 5, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of the second matrix each have five exponent bits and two explicit mantissa bits (Pat. ‘717, claim 1; Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Pat. ‘717, in view of Wang and Arm, does not teach that the floating-point data elements of the first matrix each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Sun to have made the 8-bit floating-point data elements of the first matrix have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Pat. ‘717, in view of Wang, Arm, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang, Arm, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Regarding claim 6, Pat. ‘717, in view of Wang and Arm, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of the second matrix each have five exponent bits and two explicit mantissa bits (Pat. ‘717, claim 1; Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Pat. ‘717, in view of Wang and Arm, does not teach that the floating-point data elements of the second matrix each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Sun to have made the 8-bit floating-point data elements of the second matrix have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Pat. ‘717, in view of Wang, Arm, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang, Arm, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Regarding claim 16, Pat. ‘717, in view of Wang and Arm, teaches the method of claim 15, wherein the 8-bit floating-point data elements of one of the first and second matrices each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of another of the first and second matrices each have five exponent bits and two explicit mantissa bits (Pat. ‘717, claim 25; Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Pat. ‘717, in view of Wang and Arm, does not teach that the one of the first and second matrices each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Pat. ‘717, in view of Wang and Arm, with the teachings of Sun to have made the 8-bit floating-point data elements of one of the first and second matrices have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Pat. ‘717, in view of Wang, Arm, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang, Arm, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Regarding claim 19, the claim is rejected on the same premises as claim 16 using claim 24 of Pat. ‘717.
Claims 7-10 and 20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 1, 1, 1, and 24 of Pat. ‘717 in view of Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Arm (herein Arm1) (Arm A64 Instruction Set for A-profile architecture), Arm (herein Arm2) (Arm Architecture Reference Manual for A-profile architecture), and Heinecke et al. (US 20210286620 A1).
Regarding claim 7, Pat. ‘717, in view of Wang and Arm1, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm1, does not teach that the apparatus further comprises a floating-point control register having one or more fields to specify a floating-point round mode to be used for floating-point operations, and wherein the execution circuitry, to generate the result matrix, is to perform floating-point rounding according to a round to nearest even (RNE) round mode regardless of whether the one or more fields specify that the floating-point round mode is the RNE round mode.
Arm2 teaches a floating-point control register having one or more fields to specify a floating-point round mode to be used for floating-point operations, wherein the one or more fields specify that the floating-point round mode is the round to nearest even (RNE) round mode (Pages A1-58 and A1-(64-65): The floating-point control register FPCR includes a rounding mode support field FPCR.Rmode. When the field indicates a Round to Nearest mode, it will round to the nearest value, and if the two nearest floating-point numbers bracketing the value before rounding are equally near, the value will be rounded to the number with an even least significant digit. Therefore, the Round to Nearest mode is a round to nearest even mode).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm1, with the teachings of Arm2 to have a floating-point control register to indicate a round to nearest even mode. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which the rounding should be modified based on certain needs.
Pat. ‘717, in view of Wang, Arm1, and Arm2, still does not teach to perform floating-point rounding according to a round to nearest even (RNE) round mode regardless of whether the one or more fields specify that the floating-point round mode is the RNE round mode.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Pat. ‘717, in view of Wang, Arm1, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 8, Pat. ‘717, in view of Wang and Arm1, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm1, does not teach that the apparatus further comprises a floating-point control register having one or more fields to specify whether input denormal values are to be treated as zero, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is not to treat the input denormal values as zero regardless of whether the one or more fields specify that the input denormal values are to be treated as zero.
Arm2 teaches a floating-point control register having one or more fields to specify whether input denormal values are to be treated as zero, wherein one or more fields specify that the input denormal values are to be treated as zero (Pages A1-58, A1-60, C5-767, and C5-772: The floating-point control register FPCR includes a flushing of denormalized inputs to zero field FPCR.FIZ. When the field indicates flushing denormalized inputs to zero, the denormalized inputs will be flushed to zero and will be treated as such).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm1, with the teachings of Arm2 to have a floating-point control register to indicate input denormal values to be treated as zero. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which denormalized inputs are to be treated differently based on certain needs.
Pat. ‘717, in view of Wang, Arm1, and Arm2, still does not teach to not treat the input denormal values as zero regardless of whether the one or more fields specify that the input denormal values are to be treated as zero.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Pat. ‘717, in view of Wang, Arm1, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 9, Pat. ‘717, in view of Wang and Arm1, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm1, does not teach that the apparatus further comprises one or more fields to specify whether denormal results are to be made zero, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is to make the denormal results zero regardless of whether the one or more fields specify that the denormal results are to be made zero.
Arm2 teaches a floating-point control register having one or more fields to specify whether denormal results are to be made zero, wherein one or more fields specify that the denormal results are to be made zero (Pages A1-58, A1-(60-61), and C5-(767-768): The floating-point control register FPCR includes a flushing denormalized values to zero field FPCR.FZ. When the field indicates flushing denormalized values to zero, the denormalized input and output values are flushed to zero (i.e., made zero)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm1, with the teachings of Arm2 to have a floating-point control register to indicate denormal results to be made zero. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which denormalized outputs are to be treated differently based on certain needs.
Pat. ‘717, in view of Wang, Arm1, and Arm2, still does not teach to make the denormal results zero regardless of whether the one or more fields specify that the denormal results are to be made zero.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Pat. ‘717, in view of Wang, Arm1, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 10, Pat. ‘717, in view of Wang and Arm1, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm1, does not teach that the apparatus further comprises a floating-point control register having one or more fields to specify whether floating-point exceptions are to be reported, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is not to report the floating-point exceptions regardless of whether the one or more fields specify that the floating-point exceptions are to be reported.
Note that the BFMMLA instruction may reference a control register (Page 1933: The instruction may reference to bits of FPCR.EBF, which is a control register)
Arm2 teaches a floating-point control register having one or more fields to specify whether floating-point exceptions are to be reported, wherein one or more fields specify that the floating-point exceptions are to be reported (Pages A1-58, A1-66, C5-767, and C5-(769-771): The floating-point control register FPCR includes a plurality of fields to set exception traps, which include FPCR.IDE, FPCR.IXE, FPCR.UFE, FPCR.OFE, FPCR.DZE, and FPCR.IOE. When the one or more fields are set to indicated untrapped exception handling, the one or more exceptions are to be reported to a floating-point status register FPSR).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm1, with the teachings of Arm2 to have a floating-point control register to report floating-point exceptions. One of ordinary skill may appreciate having better control over the exceptions during floating-point operations as there may be instances in which the exceptions should or should not be reported based on certain needs.
Pat. ‘717, in view of Wang, Arm1, and Arm2, still does not teach to not to report the floating-point exceptions regardless of whether the one or more fields specify that the floating-point exceptions are to be reported.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Pat. ‘717, in view of Wang, Arm1, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 20, Pat. ‘717, in view of Wang and Arm1, teaches the system of claim 18.
Pat. ‘717, in view of Wang and Arm1, does not teach that the system further comprises a floating-point control register having one or more fields to specify whether denormal values in inputs to floating-point operations are to be treated as zero, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is not to treat denormal values in inputs to floating-point operations as zero regardless of whether the one or more fields specify that denormal values in inputs to floating-point operations are to be treated as zero.
Arm2 teaches a floating-point control register having one or more fields to specify whether denormal values in inputs to floating-point operations are to be treated as zero, wherein one or more fields specify denormal values in inputs to floating-point operations are to be treated as zero (Pages A1-58, A1-60, C5-767, and C5-772: The floating-point control register FPCR includes a flushing of denormalized inputs for floating-point operations to zero field FPCR.FIZ. When the field indicates flushing denormalized inputs for floating-point operations to zero, the denormalized inputs will be flushed to zero and will be treated as such during floating-point operations).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Pat. ‘717, in view of Wang and Arm1, with the teachings of Arm2 to have a floating-point control register to indicate input denormal values of floating-point operations to be treated as zero. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which denormalized inputs are to be treated differently based on certain needs.
Pat. ‘717, in view of Wang, Arm1, and Arm2, still does not teach to not treat the denormal values in inputs to floating-point operations as zero regardless of whether the one or more fields specify that denormal values in inputs to floating-point operations are to be treated as zero.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Pat. ‘717, in view of Wang, Arm1, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill
Claim 11 is rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of Pat. ‘717 in view of Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Arm (herein Arm1) (Arm A64 Instruction Set for A-profile architecture), Arm (herein Arm2) (Arm Architecture Reference Manual for A-profile architecture), and Heinecke et al. (US 20210286620 A1).
Regarding claim 11, Pat. ‘717, in view of Wang and Arm1, teaches the apparatus of claim 1.
Pat. ‘717, in view of Wang and Arm1, does not teach that the apparatus comprises a floating-point control register, and wherein the execution circuitry is to complete the performance of the operations corresponding to the instruction without accessing the floating-point control register.
Arm2 teaches a floating-point control register that does not need to be accessed during the performance of an operation corresponding to an instruction (Pages A1-58, A1-66, C5-767, C5-(769-771), and C5-(776-778): The floating-point status register FPSR controls the exception bits. Therefore, it’s a floating-point control register. When the FPCR.IDE, FPCR.IXE, FPCR.UFE, FPCR.OFE, FPCR.DZE, and FPCR.IOE fields of the FPCR indicate the exceptions to be trapped, the respective exception bits in the FPSR are not set. Therefore, the register would not accessed during the performance of operations corresponding to instructions).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Pat. ‘717, in view of Wang and Arm1, with the teachings of Arm2, to not access the exception bits of a floating-point control register during the performance of an operation corresponding to an instruction. One of ordinary skill may prefer trapping exceptions instead of raising exception flags as to immediately address the exceptions through means such as handling the exceptions through a handler.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1, 15, and 18 are an apparatus claim, a method claim, and a system claim, respectively. Therefore, the claims are directed to a process, machine, manufacture or composition of matter.
Under Prong One of Step 2A of the 2019 Revised Patent Subject Matter Eligibility Guidance (“2019 PEG”), claim 1 recites “a first matrix having two rows by eight columns of 8-bit floating-point data elements”, “a second matrix having eight rows by two columns of 8-bit floating-point data elements”, “a third matrix having two rows by two columns of 32-bit single-precision floating-point data elements”, “generate a result having two rows by two columns of 32-bit single-precision floating-point result data elements, the result matrix representing an accumulation of the third matrix with a product matrix generated from a matrix multiplication using the first and second matrices”. Such limitations cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). Accordingly, the claim recites an abstract idea.
Under Prong Two of Step 2A, this judicial exception is not integrated into a practical application. The elements “decoder circuitry to decode an instruction, the instruction to indicate a first vector… to indicate a storage location… and to indicate a second vector register”, and “execution circuitry coupled with the decoder circuitry, the execution circuitry to perform operations corresponding to the instruction” are recited at a high level of generality, i.e., generic computer components performing generic functions, which amount(s) to no more than mere instructions to apply the exception using generic computer elements (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). Alternatively, the elements amount to no more than generally linking the abstract idea to a technological environment/field of use (e.g., SIMD/vector/parallel computing areas) (MPEP 2106.05(h)), which does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The elements “a first vector register having a 128-bit lane to store a first matrix”, “a storage location having 128 bits to store a second matrix”, “a second vector register having a 128-bit lane to store a third matrix”, and “store the result matrix in the 128-bit lane of the second register” are considered to be an insignificant step of storing data in memory (See MPEP 2106.05(d)(II)(iv), storing and retrieving information in memory)), which do not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). Thus, the elements fail to integrate the judicial exception into a practical application.
Under Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed previously, with respect to Step 2A Prong Two, the elements “decoder circuitry to decode an instruction, the instruction to indicates a first vector register… to indicate a storage location… and to indicate a second vector register”, and “execution circuitry coupled with the decoder circuitry, the execution circuitry to perform operations corresponding to the instruction” amount to no more than mere instructions to apply the exception using generic computer elements (See MPEP 2106.05(f)), or alternatively, the elements amount to no more than generally linking the abstract idea to a technological environment/field of use (e.g., SIMD/vector/parallel computing areas) (MPEP 2106.05(h)). The elements “a first vector register having a 128-bit lane to store a first matrix”, “a storage location having 128 bits to store a second matrix”, “a second vector register having a 128-bit lane to store a third matrix”, and “store the result matrix in the 128-bit lane of the second register” are considered to be an insignificant step of storing data in memory (See MPEP 2106.05(d)(II)(iv), storing and retrieving information in memory), and is deemed to be considered well-understood, routine, and conventional by the courts (MPEP 2106.05(d); See Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93). Accordingly, this claim is not patent-eligible under 35 U.S.C. 101.
Regarding claim 2, the claim recites “for each column n of the two columns of the second matrix, and for each row m of the two rows of the first matrix: convert eight data elements from the row m of the first matrix to eight corresponding converted data elements each having more than eight bits, and convert eight data elements from the column n of the second matrix to eight corresponding converted data elements each having more than eight bits; generate eight products, including to multiply the eight converted data elements corresponding to the row m and the eight converted data elements corresponding to the column n; generate a 32-bit single-precision floating-point result data element, including to accumulate the eight products with a data element from a corresponding row m of the two rows, and a corresponding column n of the two columns, of the third matrix”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 3, the claim recites “the 8-bit floating-point data elements of the first matrix, and the 8-bit floating-point data elements of the second matrix, each have four exponent bits and three explicit mantissa bits”. Such limitations further cover mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 4, the claim recites “the 8-bit floating-point data elements of the first matrix, and the 8-bit floating-point data elements of the second matrix, each have five exponent bits and two explicit mantissa bits”. Such limitations further cover mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 5, the claim recites “the 8-bit floating-point data elements of the first matrix each have four exponent bits and three explicit mantissa bits, and wherein the 8-bit floating-point data elements of the second matrix each have five exponent bits and two explicit mantissa bits”. Such limitations further cover mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 6, the claim recites “the 8-bit floating-point data elements of the first matrix each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of the second matrix each have four exponent bits and three explicit mantissa bits”. Such limitations further cover mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 7, the claim recites “one or more fields to specify a floating-point round mode to be used for floating-point operations” and “to generate the result matrix, is to perform floating-point rounding according to a round to nearest even (RNE) round mode regardless of whether the one or more fields specify that the floating-point round mode is the RNE round mode”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim also recites “a floating-point control register”, which is recited at a high level of generality, which amounts to no more than mere instructions to apply the exception using elements recited at a high level (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 8, the claim recites “one or more fields to specify whether input denormal values are to be treated as zero” and “to perform the operations corresponding to the instruction, is not to treat the input denormal values as zero regardless of whether the one or more fields specify that the input denormal values are to be treated as zero”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim also recites “a floating-point control register”, which is recited at a high level of generality, which amounts to no more than mere instructions to apply the exception using elements recited at a high level (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 9, the claim recites “one or more fields to specify whether denormal results are to be made zero” and “to perform the operations corresponding to the instruction, is to make the denormal results zero regardless of whether the one or more fields specify that the denormal results are to be made zero”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim also recites “a floating-point control register”, which is recited at a high level of generality, which amounts to no more than mere instructions to apply the exception using elements recited at a high level (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 10, the claim recites “one or more fields to specify whether floating-point exceptions are to be reported” and “to perform the operations corresponding to the instruction, is not to report the floating-point exceptions regardless of whether the one or more fields specify that the floating-point exceptions are to be reported”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim also recites “a floating-point control register”, which is recited at a high level of generality, which amounts to no more than mere instructions to apply the exception using elements recited at a high level (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 11, the claim recites “complete the performance of the operations corresponding to the instruction without accessing the floating-point control register”. Such limitation further covers mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim also recites “a floating-point control register”, which is recited at a high level of generality, which amounts to no more than mere instructions to apply the exception using elements recited at a high level (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 12, the claim recites “generate all products of the matrix multiplication using the first and second matrices before accumulation of any of said all products of the matrix multiplication with the third matrix”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 13, the claim recites “the instruction allows the storage location to be a third vector register but does not allow the storage location to be in memory”, which is recited at a high level of generality, i.e., generic computer elements, which amounts to no more than mere instructions to apply the exception using generic computer elements (See MPEP 2106.05(f)) and does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). Alternatively, the element amounts to no more than generally linking the abstract idea to a technological environment/field of use (e.g., SIMD/vector/parallel computing areas) (MPEP 2106.05(h)), which does not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claim 14, the claim recites “a fourth matrix having two rows by eight columns of 8-bit floating-point data elements”, “a fifth matrix having eight rows by two columns of 8-bit floating-point data elements”, “a sixth matrix having two rows by two columns of 32-bit single-precision floating-point data elements”, and “generate a second result matrix having two rows by two columns of 32-bit single- precision floating-point result data elements, the second result matrix representing an accumulation of the sixth matrix with a product matrix generated from a matrix multiplication using the fourth and fifth matrices”. Such limitations further cover mathematical concepts such as mathematical relationships, mathematical formulas/equations, or mathematical calculations and/or mental processes that are concepts performed in the human mind or with pen and paper (including an observation, evaluation, judgement, or opinion). The claim also recites “the first vector register has a second 128-bit lane to store a fourth matrix”, “the storage location has a second 128 bits to store a fifth matrix”, “the second vector register has a second 128-bit lane to store a sixth matrix”, and “store the second result matrix in the second 128-bit lane of the second vector register”, which are elements that are considered to be an insignificant step of storing data in memory (See MPEP 2106.05(d)(II)(iv), storing and retrieving information in memory), which do not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)) and are deemed to be considered well-understood, routine, and conventional by the courts (MPEP 2106.05(d); See Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claims 15-17, the claims recite a method similar to the apparatus of claims 1, 5, and 12, respectively. Therefore, the claims are rejected on the same premises.
Regarding claim 18, the claim is mostly rejected for the same reasons as claim 1. The claim also recites “processor” and “dynamic random access memory (DRAM) coupled with the processor”, which are elements recited at a high level of generality, i.e., generic computer elements, which amount to no more than mere instructions to appl the exception using generic computer elements (MPEP 2106.05(f)) and do not integrate the judicial exception into a practical application (See MPEP 2106.04(d)(I)). The claim fails to provide an element that would integrate the judicial exception into a practical application under Step 2A Prong Two and does not amount to anything significantly more under Step 2B. Accordingly, the claim is not patent-eligible.
Regarding claims 19-20, the claims recite a method similar to the apparatus of claims 5, and 8, respectively. Therefore, the claims are rejected on the same premises.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4, 13, and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (Arm A64 Instruction Set for A-profile architecture) and Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers).
Regarding claim 1, Symes teaches an apparatus (Fig. 1: CPU 4) comprising:
decoder circuitry to decode an instruction (Fig. 1 and [0041]: CPU 4 includes processing pipeline 12, which includes a decoding stage (hence a decoder must be included in the processing pipeline) to decode an instruction); and
execution circuitry coupled with the decoder circuitry (Fig. 1 and [0041]: The processing pipeline 12 also includes an execute stage (hence executing circuitry must be included in the processing pipeline, Since the execute stage receives decoded instruction from the decode stage, the decoder must be coupled to the execution circuitry), the execution circuitry to perform operations corresponding to the instruction (Fig. 1 and [0041]: The execute stage executes decoded instructions, which would perform operations corresponding to the decoded instruction).
Symes does not teach that the instruction to indicate a first vector register having a 128-bit lane to store a first matrix having two rows by eight columns of 8-bit floating-point data elements, to indicate a storage location having 128 bits to store a second matrix having eight rows by two columns of 8-bit floating-point data elements, and to indicate a second vector register having a 128-bit lane to store a third matrix having two rows by two columns of 32-bit single-precision floating-point data elements and for the execution circuitry to generate a result matrix having two rows by two columns of 32-bit single-precision floating-point result data elements, the result matrix representing an accumulation of the third matrix with a product matrix generated from a matrix multiplication using the first and second matrices and store the result matrix in the 128-bit lane of the second vector register in response to the instruction.
Note that Symes indicated that instructions of an instruction set can be performed on the processor (see [0041]).
Arm teaches an instruction of the Arm A64 ISA which indicates a first vector register having a 128-bit lane to store a first matrix having two rows by four columns of 16-bit floating-point data elements (Page 1933: The first operand of the BFMMLA instruction (indicated as Zn) is a vector register comprising of a 2 column by 4 row matrix of 16-bit floating point data elements, which makes up a 128-bit lane vector. The vector register comprising of the first matrix as the first register), to indicate a storage location having 128 bits to store a second matrix having four rows by two columns of 16-bit floating-point data elements (Page 1933: The second operand of the BFMMLA instruction (indicated as Zm) is a vector register comprising of a 4 column by 2 row matrix of 16-bit floating point data elements, which makes up a 128-bit lane vector. The vector register comprising of the second matrix as the storage location), and to indicate a second vector register having a 128-bit lane to store a third matrix having two rows by two columns of 32-bit single-precision floating-point data elements (Page 1933: The third operand of the BFMMLA instruction (indicated as Zda) is a vector register comprising of a 2 column by 2 row matrix of single-precision (i.e., 32-bit) floating point data elements, which makes up a 128-bit lane vector. The vector register comprising of the third matrix as the second register);
generate a result matrix having two rows by two columns of 32-bit single-precision floating-point result data elements, the result matrix representing an accumulation of the third matrix with a product matrix generated from a matrix multiplication using the first and second matrices (Page 1933: The BFMMLA instruction will multiply both matrices from both operands Zn and Zm to generate an intermediate product value, before accumulating the intermediate product value into a resulting 2x2 single-precision matrix in operand Zda); and
store the result matrix in the 128-bit lane of the second vector register (Page 1933: The resulting 2x2 32-bit floating-point matrix is to be stored in the destination operand Zda).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm, to have the processor executed instructions of the Arm A64 ISA, which would have included the BFMMLA instruction. The A64 ISA includes benefits of having a fixed-size instruction set, which in turn, reduces the complexity of decoders, which may be preferred by one of ordinary skill compared to other ISAs like x86.
Symes, in view of Arm, still does not teach that the first matrix is to have two rows by eight columns of 8-bit floating-point data elements and the second matrix is to have eight rows by two columns of 8-bit floating-point data elements.
Wang teaches to use 8-bit floating point data elements (Pages 1 and 5-6, Under “Abstract” and “Experimental Results”: 8-bit floating-point data elements may be used as input operands in computations such as GEMM computations)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Symes, in view of Arm, with the teachings of Wang to have the first and second matrices comprise of 8-bit floating-point data elements. One of ordinary skill may prefer using 8-bit data elements as it allows for quicker computations and conserves space compared to 16-bit floating-point data elements.
Symes, in view of Wang, still does not teach that the first and second matrices have two rows by eight columns and eight rows by two columns, respectively.
Arm also teaches an instruction which uses 2x8 and 8x2 matrices (Page 1737: The UMMLA instruction uses a 2x8 8-bit integer matrix and an 8x2 8-bit integer matrix as source operands of the instruction).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified the teachings of Symes, in view of Arm and Wang, with the teachings of Arm to have the first and second matrices have two rows by eight columns and eight rows by two columns, respectively. By increasing the size of the matrices, more data could be processed at the same time, which may be preferred by one of ordinary skill. Furthermore, changes in size/proportion, i.e., changing the size of the matrices, is considered to be a routine expedient, not a patentable distinction (MPEP 2144.04(IV)(A)).
Regarding claim 4, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix, and the 8-bit floating-point data elements of the second matrix, each have five exponent bits and two explicit mantissa bits (Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Regarding claim 13, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the instruction allows the storage location to be a third vector register but does not allow the storage location to be in memory (Arm, Page 1933: The BFMMLA instruction only uses a scalable vector register for the second matrix as the second operand Zm. Therefore, the instruction does not allow the second operand Zm to be in memory).
Regarding claim 14, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the first vector register has a second 128-bit lane to store a fourth matrix having two rows by eight columns of 8-bit floating-point data elements (Arm, Page 1933-1934: In the current combination, the input vector register operand Zn would include a second 128-bit segment (i.e., a lane) to include a second 8x2 matrix (see pseudocode loop in Arm) of 8-bit floating-point data elements. The second matrix in the vector register operand Zn as the fourth matrix), the storage location has a second 128 bits to store a fifth matrix having eight rows by two columns of 8-bit floating-point data elements (Arm, Page 1933-1934: In the current combination, the input vector register operand Zm would include a second 128-bit segment to include a second 2x8 matrix (see pseudocode loop in Arm) of 8-bit floating-point data elements. The second matrix in the vector register operand Zm as the fifth matrix), and the second vector register has a second 128-bit lane to store a sixth matrix having two rows by two columns of 32-bit single-precision floating-point data elements (Arm, Page 1933-1934: In the current combination, the output vector register operand Zda would include a second 128-bit segment (i.e., a lane) to include a second 2x2 matrix of 32-bit (i.e., single precision floating-point) data elements (see pseudocode loop in Arm). The second matrix in the vector register operand Zda as the sixth matrix), and wherein the execution circuitry, to perform the operations corresponding to the instruction, is further to:
generate a second result matrix having two rows by two columns of 32-bit single- precision floating-point result data elements, the second result matrix representing an accumulation of the sixth matrix with a product matrix generated from a matrix multiplication using the fourth and fifth matrices (Arm, Page 1933: In the current combination, the BFMMLA instruction will multiply both second matrices from both operands Zn and Zm to generate an intermediate product value, before accumulating the intermediate product value into a resulting 2x2 single-precision matrix in the second segment of operand Zda); and
store the second result matrix in the second 128-bit lane of the second vector register (Arm, Page 1933: The resulting 2x2 32-bit floating-point matrix is to be stored in the second segment of the destination operand Zda).
Regarding claim 15, the claim recites a method similar to the apparatus of claim 1. Therefore, the claim is rejected on the same premises.
Claims 2, 12, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), and Boswell et al. (US 20180321938 A1).
Regarding claim 2, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the execution circuitry, to generate and store the result matrix, is to:
for each column n of the two columns of the second matrix, and for each row m of the two rows of the first matrix:
generate eight products, including to multiply the eight data elements corresponding to the row m and the eight data elements corresponding to the column n (Arm, Page 1933-1934 and 4994: In the current combination and with respect to the pseudocode of BFMatMulAdd, the execution circuitry would generate eight products by multiplying each of the data elements of row m (interpreted as “i” in the pseudocode) of the first matrix of the first operand Zn and the data elements of column n (interpreted as “j” in the pseudocode) of the second matrix of the second operand Zm. In the pseudocode of BFMatMulAdd, the inner k-loop would load two pairs of elements from each matrix and performs two dot products at a time);
generate a 32-bit single-precision floating-point result data element, including to accumulate the eight products with a data element from a corresponding row m of the two rows, and a corresponding column n of the two columns, of the third matrix; and
store the 32-bit single-precision floating-point result data element in the 128-bit lane of the third vector register at a position corresponding to the row m and the column n of the third matrix (Arm, Page 1933-1934 and 4994: In the current combination and with respect to the BFMatMulAdd pseudocode, after each k-loop, the summation of the eight products and the accumulator would be a single-precision data element, which is then stored in a position of the result matrix to be stored in operand Zda in a corresponding row m of the two rows and column n of the two columns).
Symes, in view of Arm and Wang, does not teach to convert eight data elements from the row m of the first matrix to eight corresponding converted data elements each having more than eight bits, and convert eight data elements from the column n of the second matrix to eight corresponding converted data elements each having more than eight bits.
Boswell teaches to convert data elements of the first matrix to converted data elements each having more than eight bits, and convert eight data elements of the second matrix to eight corresponding converted data elements each having more than eight bits (Figs. 7 and 13, [0137]: In pipeline stage 1301, the conversion/encoding logic 1315 receives elements of two input vectors A and B (which comprise of matrices, see Fig. 7), and convert all input values to half-precision floating-point format, which indicates that each of the elements are 16-bits).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm and Wang, with the teachings of Boswell to have converted the data elements of the first and second matrix prior to performing matrix multiplication and accumulation. A processor may not have the hardware to perform on 8-bit floating-point data types, therefore, one of ordinary skill may be inclined to increase the data type to a type that a processor has the capability to perform on, which may be preferred by one of ordinary skill.
Regarding claim 12, Symes, in view of Arm and Wang, teaches the apparatus of claim 1.
Symes, in view of Arm and Wang, does not teach to generate the result matrix, is to generate all products of the matrix multiplication using the first and second matrices before accumulation of any of said all products of the matrix multiplication with the third matrix.
Boswell teaches to generate all products of a first matrix and a second matrix before accumulation of any of said all products of the matrix multiplication with the third matrix (Fig. 13 and [0136-0139]: The figure shows a pipeline where at the second stage pipeline 1302, the multiplication of operands A and B (which are matrices) is to be performed prior to the accumulation of products with the accumulator (operand C, which is a matrix), which occurs at the third stage pipeline 1303).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm and Wang, with the teachings of Boswell to have generated all products of the matrix before accumulating said all products with the third matrix. By having the execution circuitry produce the product of the two source matrices prior to accumulating with the accumulator matrix, it would allow pipelining of the execution circuitry with respect to the multiply-accumulate circuitry (See Fig. 13, pipeline stage 1302 and pipeline stage 1303), which may be preferred by one of ordinary skill.
Regarding claim 17, the claim recites a method similar to the apparatus of claim 12. Therefore, the claim is rejected on the same premises.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), and Sun et al. (Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks).
Regarding claim 3, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix, and the 8-bit floating-point data elements of the second matrix, each have five exponent bits and two explicit mantissa bits (Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Symes, in view of Arm and Wang, does not teach that the floating-point data elements of the matrices each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Symes, in view of Arm and Wang, with the teachings of Sun to have made the 8-bit floating-point data elements have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
Claims 5-6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Sun et al. (Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks), and Van Zee (Supporting mixed-datatype matrix multiplication within the BLIS framework)
Regarding claim 5, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of the second matrix each have five exponent bits and two explicit mantissa bits (Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Symes, in view of Arm and Wang, does not teach that the floating-point data elements of the first matrix each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Symes, in view of Arm and Wang, with the teachings of Sun to have made the 8-bit floating-point data elements of the first matrix have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Symes, in view of Arm, Wang, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm, Wang, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Regarding claim 6, Symes, in view of Arm and Wang, teaches the apparatus of claim 1, wherein the 8-bit floating-point data elements of the first matrix each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of the second matrix each have five exponent bits and two explicit mantissa bits (Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Symes, in view of Arm and Wang, does not teach that the floating-point data elements of the second matrix each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Symes, in view of Arm and Wang, with the teachings of Sun to have made the 8-bit floating-point data elements of the second matrix have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Symes, in view of Arm, Wang, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm, Wang, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Regarding claim 16, Symes, in view of Arm and Wang, teaches the method of claim 15, wherein the 8-bit floating-point data elements of one of the first and second matrices each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of another of the first and second matrices each have five exponent bits and two explicit mantissa bits (Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Symes, in view of Arm and Wang, does not teach that the one of the first and second matrices each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Symes, in view of Arm and Wang, with the teachings of Sun to have made the 8-bit floating-point data elements of one of the first and second matrices have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Symes, in view of Arm, Wang, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm, Wang, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Claims 7-10 are rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (herein Arm1) (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Arm (herein Arm2) (Arm Architecture Reference Manual for A-profile architecture), and Heinecke et al. (US 20210286620 A1).
Regarding claim 7, Symes, in view of Arm1 and Wang, teaches the apparatus of claim 1.
Symes, in view of Arm1 and Wang, does not teach that the apparatus further comprises a floating-point control register having one or more fields to specify a floating-point round mode to be used for floating-point operations, and wherein the execution circuitry, to generate the result matrix, is to perform floating-point rounding according to a round to nearest even (RNE) round mode regardless of whether the one or more fields specify that the floating-point round mode is the RNE round mode.
Arm2 teaches a floating-point control register having one or more fields to specify a floating-point round mode to be used for floating-point operations, wherein the one or more fields specify that the floating-point round mode is the round to nearest even (RNE) round mode (Pages A1-58 and A1-(64-65): The floating-point control register FPCR includes a rounding mode support field FPCR.Rmode. When the field indicates a Round to Nearest mode, it will round to the nearest value, and if the two nearest floating-point numbers bracketing the value before rounding are equally near, the value will be rounded to the number with an even least significant digit. Therefore, the Round to Nearest mode is a round to nearest even mode).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm1 and Wang, with the teachings of Arm2 to have a floating-point control register to indicate a round to nearest even mode. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which the rounding should be modified based on certain needs.
Symes, in view of Arm1, Wang, and Arm2, still does not teach to perform floating-point rounding according to a round to nearest even (RNE) round mode regardless of whether the one or more fields specify that the floating-point round mode is the RNE round mode.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Symes, in view of Arm1, Wang, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 8, Symes, in view of Arm1 and Wang, teaches the apparatus of claim 1.
Symes, in view of Arm1 and Wang, does not teach that the apparatus further comprises a floating-point control register having one or more fields to specify whether input denormal values are to be treated as zero, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is not to treat the input denormal values as zero regardless of whether the one or more fields specify that the input denormal values are to be treated as zero.
Arm2 teaches a floating-point control register having one or more fields to specify whether input denormal values are to be treated as zero, wherein one or more fields specify that the input denormal values are to be treated as zero (Pages A1-58, A1-60, C5-767, and C5-772: The floating-point control register FPCR includes a flushing of denormalized inputs to zero field FPCR.FIZ. When the field indicates flushing denormalized inputs to zero, the denormalized inputs will be flushed to zero and will be treated as such).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm1 and Wang, with the teachings of Arm2 to have a floating-point control register to indicate input denormal values to be treated as zero. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which denormalized inputs are to be treated differently based on certain needs.
Symes, in view of Arm1, Wang, and Arm2, still does not teach to not treat the input denormal values as zero regardless of whether the one or more fields specify that the input denormal values are to be treated as zero.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Symes, in view of Arm1, Wang, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 9, Symes, in view of Arm1 and Wang, teaches the apparatus of claim 1.
Symes, in view of Arm1 and Wang, does not teach that the apparatus further comprises one or more fields to specify whether denormal results are to be made zero, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is to make the denormal results zero regardless of whether the one or more fields specify that the denormal results are to be made zero.
Arm2 teaches a floating-point control register having one or more fields to specify whether denormal results are to be made zero, wherein one or more fields specify that the denormal results are to be made zero (Pages A1-58, A1-(60-61), and C5-(767-768): The floating-point control register FPCR includes a flushing denormalized values to zero field FPCR.FZ. When the field indicates flushing denormalized values to zero, the denormalized input and output values are flushed to zero (i.e., made zero)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm1 and Wang, with the teachings of Arm2 to have a floating-point control register to indicate denormal results to be made zero. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which denormalized outputs are to be treated differently based on certain needs.
Symes, in view of Arm1, Wang, and Arm2, still does not teach to make the denormal results zero regardless of whether the one or more fields specify that the denormal results are to be made zero.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Symes, in view of Arm1, Wang, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Regarding claim 10, Symes, in view of Arm1 and Wang, teaches the apparatus of claim 1.
Symes, in view of Arm1 and Wang, does not teach that the apparatus further comprises a floating-point control register having one or more fields to specify whether floating-point exceptions are to be reported, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is not to report the floating-point exceptions regardless of whether the one or more fields specify that the floating-point exceptions are to be reported.
Note that the BFMMLA instruction may reference a control register (Page 1933: The instruction may reference to bits of FPCR.EBF, which is a control register)
Arm2 teaches a floating-point control register having one or more fields to specify whether floating-point exceptions are to be reported, wherein one or more fields specify that the floating-point exceptions are to be reported (Pages A1-58, A1-66, C5-767, and C5-(769-771): The floating-point control register FPCR includes a plurality of fields to set exception traps, which include FPCR.IDE, FPCR.IXE, FPCR.UFE, FPCR.OFE, FPCR.DZE, and FPCR.IOE. When the one or more fields are set to indicated untrapped exception handling, the one or more exceptions are to be reported to a floating-point status register FPSR).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm1 and Wang, with the teachings of Arm2 to have a floating-point control register to report floating-point exceptions. One of ordinary skill may appreciate having better control over the exceptions during floating-point operations as there may be instances in which the exceptions should or should not be reported based on certain needs.
Symes, in view of Arm1, Wang, and Arm2, still does not teach to not to report the floating-point exceptions regardless of whether the one or more fields specify that the floating-point exceptions are to be reported.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Symes, in view of Arm1, Wang, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (herein Arm1) (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), and Arm (herein Arm2) (Arm Architecture Reference Manual for A-profile architecture).
Regarding claim 11, Symes, in view of Arm1 and Wang, teaches the apparatus of claim 1.
Symes, in view of Arm1 and Wang, does not teach that the apparatus comprises a floating-point control register, and wherein the execution circuitry is to complete the performance of the operations corresponding to the instruction without accessing the floating-point control register.
Arm2 teaches a floating-point control register that does not need to be accessed during the performance of an operation corresponding to an instruction (Pages A1-58, A1-66, C5-767, C5-(769-771), and C5-(776-778): The floating-point status register FPSR controls the exception bits. Therefore, it’s a floating-point control register. When the FPCR.IDE, FPCR.IXE, FPCR.UFE, FPCR.OFE, FPCR.DZE, and FPCR.IOE fields of the FPCR indicate the exceptions to be trapped, the respective exception bits in the FPSR are not set. Therefore, the register would not accessed during the performance of operations corresponding to instructions).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Symes, in view of Arm1 and Wang, with the teachings of Arm2, to not access the exception bits of a floating-point control register during the performance of an operation corresponding to an instruction. One of ordinary skill may prefer trapping exceptions instead of raising exception flags as to immediately address the exceptions through means such as handling the exceptions through a handler.
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), and Dhoble et al. (US 20220201009 A1).
Regarding claim 18, the claim is mostly rejected for the same reasons as claim 1. Symes, in view of Arm and Wang, also teaches a system (Fig. 1 and [0040]: Processing system 2) comprising:
a processor (Fig. 1 and [0040]: CPU 4) and memory coupled with the processor (Fig. 1 and [0040]: Shared memory 10).
Symes, in view of Arm and Wang, does not teach that the memory is DRAM
Dhoble teaches a processor coupled to DRAM (Fig. 1 and [0024]: Processor(s) 101 mis coupled to system memory 105, which may be DRAM).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Symes, in view of Arm and Wang, with the teachings of Dhoble to have made the shared memory of Symes be DRAM. One of ordinary skill would recognize that DRAM provides high storage density and is cost-effective, compared to other types of memories such as SRAM.
Claims 19 is rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Dhoble et al. (US 20220201009 A1), Sun et al. (Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks), and Van Zee (Supporting mixed-datatype matrix multiplication within the BLIS framework)
Regarding claim 19, Symes, in view of Arm, Wang, and Dhoble, teaches the system of claim 18, wherein the 8-bit floating-point data elements of one of the first and second matrices each have five exponent bits and two explicit mantissa bits, and wherein the 8-bit floating-point data elements of another of the first and second matrices each have five exponent bits and two explicit mantissa bits (Wang, page 3, Under “New Reduced Precision Floating Point Formats: FP8 and FP16: In the current combination, the 8-bit floating-point data elements of the source matrices use a floating-point format of 1 sign bit, 5 exponent bits, and 2 mantissa bits).
Symes, in view of Arm, Wang, and Dhoble, does not teach that the one of the first and second matrices each have four exponent bits and three explicit mantissa bits.
Sun teaches 8-bit floating point data elements having four exponent bits and three explicit mantissa bits (Page 1, Section 1, Paragraphs 1 and Page 3, Section 1.2, Paragraph 2: FP8 format with 1 bit for sign, 4 bits for exponent, and 3 bits for mantissa).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Symes, in view of Arm, Wang, and Dhoble, with the teachings of Sun to have made the 8-bit floating-point data elements of one of the first and second matrices have four bits for an exponent and three bits for an explicit mantissa. By changing the mantissa to have three bits instead of two bits and the exponent to have four bits instead of five, the floating-point data elements cold be more precise, in exchange of decreasing the number range, which may be preferred by one of ordinary skill.
However, Symes, in view of Arm, Wang, Dhoble, and Sun, does not teach to perform matrix multiplication between two matrices comprising of different floating-point types.
Van Zee teaches GEMM with matrices which may comprise of different floating-point types (Pages 3 and 10-11, Fig. 1, Section 5: The matrices A, B, and C, can be stored in different precision types (in this case, single or double precision) and are converted to match before multiplication occurs).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm, Wang, Dhoble, and Sun, with the teachings of Van Zee to have performed matrix multiplication with the first matrix comprising of a first floating-point format type different from the second matrix of a second floating-point format type. By allowing multiplication of matrices comprising of different floating-point types, there would be no need to first convert one of the matrices to match the floating-point data format of the other matrix, which may be appreciated by one of ordinary skill.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Symes et al. (US 20230080578 A1) in view of Arm (herein Arm1) (Arm A64 Instruction Set for A-profile architecture), Wang et al. (Training Deep Neural Networks with 8-bit Floating Point Numbers), Dhoble et al. (US 20220201009 A1), Arm (herein Arm2) (Arm Architecture Reference Manual for A-profile architecture), and Heinecke et al. (US 20210286620 A1).
Regarding claim 20, Symes, in view of Arm1, Wang, and Dhoble, teaches the system of claim 18.
Symes, in view of Arm1, Wang, and Dhoble, does not teach that the system further comprises a floating-point control register having one or more fields to specify whether denormal values in inputs to floating-point operations are to be treated as zero, and wherein the execution circuitry, to perform the operations corresponding to the instruction, is not to treat denormal values in inputs to floating-point operations as zero regardless of whether the one or more fields specify that denormal values in inputs to floating-point operations are to be treated as zero.
Arm2 teaches a floating-point control register having one or more fields to specify whether denormal values in inputs to floating-point operations are to be treated as zero, wherein one or more fields specify denormal values in inputs to floating-point operations are to be treated as zero (Pages A1-58, A1-60, C5-767, and C5-772: The floating-point control register FPCR includes a flushing of denormalized inputs for floating-point operations to zero field FPCR.FIZ. When the field indicates flushing denormalized inputs for floating-point operations to zero, the denormalized inputs will be flushed to zero and will be treated as such during floating-point operations).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Symes, in view of Arm1, Wang, and Dhoble, with the teachings of Arm2 to have a floating-point control register to indicate input denormal values of floating-point operations to be treated as zero. One of ordinary skill may appreciate having better control over the behaviors of the floating-point operations as there may be instances in which denormalized inputs are to be treated differently based on certain needs.
Symes, in view of Arm1, Wang, Dhoble, and Arm2, still does not teach to not treat the denormal values in inputs to floating-point operations as zero regardless of whether the one or more fields specify that denormal values in inputs to floating-point operations are to be treated as zero.
Heinecke teaches an instruction with a field that overrides a register value of a control register (Fig. 25A-B and [0222]: The instruction includes a round operation control field 2558, which allows it to override the configuration as indicated in a control register).
It would have been obvious to one of ordinary skill in the art before the effective filing date to have combined the teachings of Symes, in view of Arm1, Wang, Dhoble, and Arm2, with the teachings of Heinecke to have the BFMMLA instruction include a field to indicate overriding a configuration indicated by the floating-point control register. One of ordinary skill may appreciate having the instruction indicate a configuration to override the control register as there may be instances in which rather than changing the values of the control register, it would be simpler to have the instruction override the configuration, providing flexibility to one of ordinary skill.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Low-precision Floating-point Arithmetic for High-performance FPGA-based CNN Acceleration: Wu et al. teaches to perform arithmetic operations using 8-bit floating-point data.
8-bit precision for training deep learning systems: Wang discusses 8-bit floating-point matrix multiplication and 16-bit floating-point matrix accumulation (See under Chunk-based accumulation).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMILIO ALCANTARA-RAMOS whose telephone number is (571)272-4211. The examiner can normally be reached Mon-Fri 8:30-5:00 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571)270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/E.A./Examiner, Art Unit 2183
/David J. Huisman/Primary Examiner, Art Unit 2183