DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement filed 10/13/2023 has been considered except where lined through because the listed foreign document was not provided.
Claim Interpretation
Herein, a “computer readable storage medium” is interpreted, as defined in paragraph [0121] of Applicant’s specification such that “A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as electrical signals transmitted through a wire, radio waves or other freely propagating electromagnetic waves, or electromagnetic waves propagating through a wave transmission medium (e.g., a wave guide or fiber-optic cable).”
Claim Objections
Claims 5, 12, 14, 16 are objected to because of the following informalities:
Claim 5 on line 1, change “first number greater than” to “first number is greater than”.
Claim 12, change “wherein the third processor configured to compute the dot product comprises the third processor further configured to” to “wherein the third processor is further configured to”.
Claim 14, change “wherein the SD Splitter configured to generate the second column-split matrix comprises the SD splitter further configured to” to “wherein the SD Splitter is further configured to”.
Claim 14, change “wherein the SD Splitter configured to generate the second row-split matrix comprises the SD splitter further configured to” to “wherein the SD Splitter is further configured to”.
Claim 14, change “wherein the second processor configured to compute the second partial dot product comprises the second processor further configured to add” to “wherein the second processor is further configured to add”.
Claim 17, change “the first processor configured to compute the first partial dot product comprises the first processor further configured to compute” to “the first processor is further configured to compute”.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 11-20 will be addressed first.
Regarding claim 11, at Step 1, the claim is directed to a computing system, which is a statutory category of invention (Machine).
At Step 2A Prong 1, Examiner notes that the claims are directed towards an abstract idea. The claim language has been reproduced below:
A computing system, the computing system comprising:
a first processor; a second processor; a third processor;
and, a Shared Dimension (SD) Splitter, the SD Splitter configured to:
determine that a left side matrix and a right side matrix have a shared dimension number, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows (mathematical relationship and/or mental process);
generate, based on the determining that the left side matrix and the right side matrix have the shared dimension number, a first column-split matrix and a first row-split matrix (mathematical calculation and/or mental process), the first column-split matrix comprising a first number of columns among the shared dimension number of columns of the left side matrix (mathematical relationship), the first row-split matrix comprising the first number of rows among the shared dimension number of rows of the right side matrix (mathematical relationship);
and, generate, based on the determining that the left side matrix and the right side matrix have the shared dimension number, a second column-split matrix and a second row-split matrix (mathematical calculation and/or mental process), the second column-split matrix comprising a second number of columns among the shared dimension number of columns of the left side matrix (mathematical relationship), the second row-split matrix comprising the second number of rows among the shared dimension number of rows of the right side matrix (mathematical relationship);
wherein the first processor is configured to compute a first partial dot product comprising a sum of products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix (mathematical calculation);
wherein the second processor is configured to compute, concurrent with the first processor computing the first partial dot product, a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix (mathematical calculation);
and, wherein the third processor is configured to compute a dot product comprising a sum of the first partial dot product and the second partial dot product (mathematical calculation).
Determining whether two matrices have a shared dimension number can be reasonably performed in the human mind based on the process described in paragraph [0041] by observing the dimension lengths. Generating the first row-split matrix, first column-split matrix, second row-split matrix, and second column-split matrix based on the determining can be reasonably performed in the human mind with pencil and paper based on the process discussed in Fig. 1 and paragraph [0044], wherein the generating is merely splitting the parent matrices.
At Step 2A Prong 2, the additional elements are bolded above. The additional elements do not integrate the abstract ideas into a practical application because the computer elements, which are recited at a high level of generality, provide conventional computer functions that do not impose any meaningful limits on practicing the abstract ideas. See MPEP 2106.05(f). The limitations first processor, second processor, third processor, and Shared Dimension (SD) Splitter are merely generic computer components performing the functions such that it is the equivalent of reciting “apply it” to the judicial exception. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
At Step 2B, the additional elements do not, alone or in combination, amount to significantly more than the recited judicial exception. As set forth in step 2A prong 2 analysis, the first processor, second processor, third processor, and Shared Dimension (SD) Splitter are the equivalent of adding the words “apply it” to the judicial exception and are mere instructions to implement the abstract idea on a computer. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, which do not provide an inventive concept. The claim is not eligible.
Regarding claim 12, it is directed to the mathematical concept of “to input the first partial dot product from the memory to compute the sum of the first partial dot product and the second partial dot product”.
Under Step 2A Prong 2, the claim recites additional elements a memory; wherein the first processor is further configured to output the first partial dot product to the memory. The additional elements do not integrate the abstract ideas into a practical application because the memory is recited at a high level of generality and do not impose any meaningful limits on practicing the abstract idea. Furthermore, the output the first partial dot product to the memory is an insignificant extra-solution activity of data outputting and storing data. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
Under Step 2B, the additional elements do not, alone or in combination, amount to significantly more than the recited judicial exception. As set forth in step 2A prong 2 analysis, the function of storing information in memory is recognized by the courts as well-understood routine and conventional. See MPEP 2106.05(d)(II). Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, which do not provide an inventive concept. The claim is not eligible.
Regarding claim 13, under Step 2A Prong 2, the claim recites additional element the first processor and second processor comprise different processors. The additional element does not integrate the abstract ideas into a practical application because the relationship of the first and second processors are recited at a high level of generality and do not impose any meaningful limits on practicing the abstract idea. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
Under Step 2B, the additional elements do not, alone or in combination, amount to significantly more than the recited judicial exception. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, which do not provide an inventive concept. The claim is not eligible.
Regarding claims 14-15, the claims merely recite functions for generating the second column-split matrix and second row-split matrix, and computing the second partial dot product that further mathematically limit the mathematical concepts, or provide additional mathematical functions, of claim 11. They do not include additional elements that would require further analysis under steps 2A prong 2 and step 2B.
Regarding claim 16, it is directed to the mathematical concept and/or mental process of “add a first product, among products included in the first partial dot product”.
Under Step 2A Prong 2, the claim recites additional element an accumulator. The additional element does not integrate the abstract ideas into a practical application because the accumulator is recited at a high level of generality and do not impose any meaningful limits on practicing the abstract idea. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
Under Step 2B, the additional elements do not, alone or in combination, amount to significantly more than the recited judicial exception. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, which do not provide an inventive concept. The claim is not eligible.
Regarding claims 17-18, the claims merely recite functions for multiply-accumulate (MACC) computation, and sum of products computation that further mathematically limit the mathematical concepts, or provide additional mathematical functions, of claim 11. They do not include additional elements that would require further analysis under steps 2A prong 2 and step 2B.
Regarding claim 19, it is directed to the mathematical concept and/or mental process of “compute the first partial dot product as the MACC computation”.
Under Step 2A Prong 2, the claim recites additional element a MACC arithmetic logic unit (ALU). The additional element does not integrate the abstract ideas into a practical application because the ALU is recited at a high level of generality and do not impose any meaningful limits on practicing the abstract idea. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
Under Step 2B, the additional elements do not, alone or in combination, amount to significantly more than the recited judicial exception. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, which do not provide an inventive concept. The claim is not eligible.
Regarding claim 20, under Step 2A Prong 2, the claim recites additional element at least one of the first processor, the second processor, and the third processor comprises a processor of a reconfigurable dataflow unit. The additional element does not integrate the abstract ideas into a practical application because the relationship of the first and second processors are recited at a high level of generality and do not impose any meaningful limits on practicing the abstract idea. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
Under Step 2B, the additional elements do not, alone or in combination, amount to significantly more than the recited judicial exception. Even when considered in combination, these additional elements represent mere instructions to apply an exception and insignificant extra-solution activity, which do not provide an inventive concept. The claim is not eligible.
Regarding claims 1-8, the claims are directed to a method that would be practiced by the computing system of claims 11-12, 18, 14-17, and 20, respectively. All steps performed by the method of claims 1-8 are executed by the computing system in claims 11-12, 18, 14-17, and 20 as configured. The analysis of claims 11-12, 18, 14-17, and 20 applies equally to claims 1-8.
Regarding claims 9-10, the claims are directed to a computer readable storage medium having program instructions that would be practiced by the computing system of claims 11 and 14, respectively. All steps performed by the program instructions of claims 9-10 are executed by the computing system in claims 11 and 14 as configured. The analysis of claims 11 and 14 applies equally to claims 9-10.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Double Patenting Rejection #1
Claim 1-20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 8, 11, 8, 12, 14, 16, 20, 15, 8, 12, 8, 11, 10, 12, 14, 16, 20, 8, 19, 15, respectively of copending Application No. 18378278 in view of Liu et al. (US 20230169144 A1, hereinafter “Liu”).
Claim 8 of copending Application No. 18378278 teaches the limitations of claim 11 as underlined in the table below
18105695
18378278
A computing system, the computing system comprising: a first processor; a second processor; a third processor;
A computing system, the computing system comprising: a first matrix processing unit (MPU); a second MPU; and a third MPU,
and,
based on the determining that the left side matrix and the right side matrix have the shared dimension number, a first column-split matrix and a first row-split matrix, the first column-split matrix comprising a first number of columns among the shared dimension number of columns of the left side matrix, the first row-split matrix comprising the first number of rows among the shared dimension number of rows of the right side matrix;
wherein the first MPU is configured to: receive, based on a left side matrix and a right side matrix having a shared dimension number,column elements of a row of a first column-split matrix and row elements of a column of a first row-split matrix, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows, the first column-split matrix comprising a first number of columns among columns of a left side matrix, the first row-split matrix comprising the first number of rows among rows of a right side matrix;
based on the determining that the left side matrix and the right side matrix have the shared dimension number, a second column-split matrix and a second row-split matrix, the second column-split matrix comprising a second number of columns among the shared dimension number of columns of the left side matrix, the second row-split matrix comprising the second number of rows among the shared dimension number of rows of the right side matrix;
wherein the second MPU is configured to: receive, based on the left side matrix and the right side matrix having the shared dimension,column elements of a row of a second column-split matrix and row elements of a column of a second row-split matrix, the second column-split matrix comprising a second number of columns among the shared dimension number of columns of the left side matrix, the second row-split matrix comprising the second number of rows among the shared dimension number of rows of the right side matrix;
wherein the first processor is configured to compute a first partial dot product comprising a sum of products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix;
and, compute a first partial dot product comprising a sum of products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix;
wherein the second processor is configured to compute, concurrent with the first processor computing the first partial dot product, a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix;
and, compute, concurrent with the first MPU computing the first partial dot product, a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix;
and, wherein the third processor is configured to compute a dot product comprising a sum of the first partial dot product and the second partial dot product.
and, wherein the third MPU is configured to compute a dot product comprising a sum of the first partial dot product and the second partial dot product.
As to claim 11, claim 8 of copending application 18378278 does not explicitly teach a Shared Dimension (SD) Splitter, the SD Splitter configured to: determine that a left side matrix and a right side matrix have a shared dimension number, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows;or generating the column-split and row-split matrices.
However, in the same field of endeavor, Liu discloses a Shared Dimension (SD) Splitter, the SD Splitter configured to: determine that a left side matrix and a right side matrix have a shared dimension number, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows (Liu: Fig. 1-2b, wherein matrix A corresponds with the left side matrix and matrix B corresponds with the right side matrix; [0086]); and generating the column=split and row=split matrices (Liu: Fig. 1-2b, [0099]).
Accordingly, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify Claim 8 of copending application 18378278 using Liu and configure the computing system to generate split-matrices based on the shared dimension number because Liu improves operation efficiency during matrix multiplication (Liu: abstract).
Therefore, the combination of copending application 18378278 in view of Liu teaches a computing system that generates column-split and row-split matrices based on the shared dimension of left and right matrices.
As to claim 12, claim 11 of copending application 18378278 teaches The computing system of claim 11, wherein the computing system further comprises a memory (“The computing system of claim 8, wherein the computing system further comprises a memory”); wherein the first processor is further configured to output the first partial dot product to the memory (“wherein the first MPU is further configured to output the first partial dot product to the memory”); and; wherein the third processor configured to compute the dot product comprises the third processor further configured to input the first partial dot product from the memory to compute the sum of the first partial dot product and the second partial dot product (“and, wherein the third MPU configured to compute the dot product comprises the third MPU further configured to: input the first partial dot product from the memory; and, add the first partial dot product input from the memory to the second partial dot product to compute the sum of the first partial dot product and the second partial dot product.”).
As to claim 13, claim 10 of copending application 18378278 teaches The computing system of claim 11, wherein the first processor and the second processor comprise different processors (“The computing system of claim 8, wherein the first MPU and the second MPU comprise different MPUs.”).
As to claim 14, claim 12 of copending application 18378278 teaches The computing system of claim 11, wherein the first number is greater than the second number (“The computing system of claim 8, wherein the first number is greater than the second number)”;
wherein the SD Splitter configured to generate the second column-split matrix comprises the SD splitter further configured to generate, based on the first number greater than the second number, the second column-split matrix to further comprise an all-zeros column, each element of the all-zeros column having value zero; wherein the SD splitter configured to generate the second row-split matrix comprises the SD splitter further configured to generate, based on the first number greater than the second number, the second row-split matrix to further comprise an all-zeros row, each element of the all-zeros row having value zero (“wherein, based on the first number greater than the second number, the second column-split matrix further comprises an all-zeros column, each element of the all-zeros column having value zero; wherein, based on the first number greater than the second number, the second row-split matrix further comprises an all-zeros row, each element of the all-zeros row having the value zero”);
and, wherein the second processor configured to compute the second partial dot product comprises the second processor further configured to add, to the second partial dot product, a product of a row element of the all-zeros column of the row of the second column-split matrix multiplied by a corresponding column element of the all-zeros row of the column of second row-split matrix (“and, wherein the second MPU configured to compute the second partial dot product comprises the second MPU further configured to add, to the second partial dot product, a product of a row element of the all-zeros column of the second column-split matrix multiplied by a column element of the all-zeros row of the second row-split matrix.”).
As to claim 15, claim 14 of copending application 18378278 teaches The computing system of claim 11, wherein the first number is greater than the second number (“wherein the first number is greater than the second number”);
and, wherein the second processor configured to compute the second partial dot product comprises the second processor adding, based on the first number greater than the second number, a value of zero to the second partial dot product (“and, wherein the second MPU configured to compute the second partial dot product comprises the second MPU further configured to add, based on the first number greater than the second number, a value of zero to the second partial dot product.”).
As to claim 16, claim 16 of copending application 18378278 teaches The computing system of claim 11, wherein the computing system comprises an accumulator (“wherein the first MPU comprises a first accumulator”);
and, wherein the first processor configured to compute the first partial dot product comprises the first processor further configured to add a first product, among products included in the first partial dot product, to the accumulator (“and, wherein the first MPU configured to compute the first partial dot product comprises the first MPU further configured to add, to the first accumulator, products among the products of column elements of the row of the second column-split matrix multiplied by corresponding row elements of the column of the second row-split matrix.”.
As to claim 17, claim 20 of copending application 18378278 teaches The computing system of claim 16, wherein the first processor configured to compute the first partial dot product comprises the first processor further configured to compute the first partial dot product as a multiply-accumulate (MACC) computation (claim 19: “wherein the first MPU configured to compute the first partial dot product comprises the first MPU further configured to perform a multiply- accumulate (MACC) computation to compute the first partial dot product.”), the MACC computation comprising: computing, by the first processor, the first product; adding, by the first processor, the first product to the accumulator; computing, by the first processor, a second product, among the products of the column elements of the row of the first column-split matrix multiplied by the corresponding row elements of the column of the first row-split matrix; and, adding, by the first processor, the second product to the accumulator (“wherein the first MPU configured to perform the MACC computation comprises the MACC ALU configured to perform the MACC computation.”).
As to claim 18, claim 8 of copending application 18378278 teaches The computing system of claim 11, wherein the dot product comprises a sum of products of all elements of the row of the first column-split matrix multiplied by all corresponding row elements of the column of the first row-split matrix (claim 8: “a first partial dot product comprising a sum of products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix”; “a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix”).
As to claim 19, claim 20 of copending application 18378278 teaches The computing system of claim 17,wherein the first processor comprises a MACC arithmetic logic unit (ALU); and, wherein the first processor configured to compute the first partial dot product as the MACC computation comprises the MACC ALU configured to compute the first partial dot product as the MACC computation (claim 19: “wherein the first MPU configured to compute the first partial dot product comprises the first MPU further configured to perform a multiply- accumulate (MACC) computation to compute the first partial dot product.”).
As to claim 20, claim 15 of copending application 18378278 teaches The computing system of claim 11, wherein at least one of the first processor, the second processor, and the third processor comprises a processor of a reconfigurable dataflow unit (“The computing system of claim 8, wherein at least one of the first MPU, the second MPU,and the third MPU comprises a reconfigurable dataflow unit.”).
As to claims 1-8, the claims are directed to a method that would be practiced by the computing system of claims 11-12, 18, 14-17, and 20, respectively. All steps performed by the method of claims 1-8 are executed by the computing system in claims 11-12, 18, 14-17, and 20 as configured. The analysis of claims 11-12, 18, 14-17, and 20 applies equally to claims 1-8.
As to claims 9-10, the claims are directed to a computer readable storage medium having program instructions that would be practiced by the computing system of claims 11 and 14, respectively. All steps performed by the program instructions of claims 9-10 are executed by the computing system in claims 11 and 14 as configured. The analysis of claims 11 and 14 applies equally to claims 9-10.
Double Patenting Rejection #2
Claim 1-20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 9, 15, 16, 12, 14, 11, 16, 20, 9, 12, 9, 15, 9, 12, 14, 11, 16, 16, 9, 20, respectively of copending Application No. 18378293 in view of Liu et al. (US 20230169144 A1, hereinafter “Liu”).
Claim 9 of copending Application No. 18378293 teaches the limitations of claim 11 as underlined in the table below
18105695
18378293
A computing system, the computing system comprising: a first processor; a second processor; a third processor;
A Matrix Processing Unit (MPU) included in a computing system, the MPU comprising: a first Multiply-Accumulate (MACC) Arithmetic Logic Unit (ALU); a second MACC ALU; and, a first adder ALU
and,
based on the determining that the left side matrix and the right side matrix have the shared dimension number, a first column-split matrix and a first row-split matrix, the first column-split matrix comprising a first number of columns among the shared dimension number of columns of the left side matrix, the first row-split matrix comprising the first number of rows among the shared dimension number of rows of the right side matrix;
wherein the first MACC ALU is configured to: receive, based on a left side matrix and a right side matrix having a shared dimension, number column elements of a row of a first column-split matrix and row elements of a column of a first row-split matrix, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows, the first column-split matrix comprising a first number of columns among columns of a left side matrix, the first row-split matrix comprising the first number of rows among rows of a right side matrix;
based on the determining that the left side matrix and the right side matrix have the shared dimension number, a second column-split matrix and a second row-split matrix, the second column-split matrix comprising a second number of columns among the shared dimension number of columns of the left side matrix, the second row-split matrix comprising the second number of rows among the shared dimension number of rows of the right side matrix;
wherein the second MACC ALU is configured to: receive, based on the left side matrix and the right side matrix having the shared dimension,column elements of a row of a second column-split matrix and row elements of a column of a second row-split matrix, the second column-split matrix comprising a second number of columns among the shared dimension number of columns of the left side matrix, the second row-split matrix comprising the second number of rows among the shared dimension number of rows of the right side matrix;
wherein the first processor is configured to compute a first partial dot product comprising a sum of products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix;
and, compute a first partial dot product comprising a sum of first row-column products, the first row- column products comprising products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix;
wherein the second processor is configured to compute, concurrent with the first processor computing the first partial dot product, a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix;
and, compute, concurrent with the first MACC ALU computing the first partial dot product, a second partial dot product comprising products of a sum of second row-column products, the second row-column products comprising column elements of a row of the second column- split matrix multiplied by corresponding row elements of a column of the second row-split matrix;
and, wherein the third processor is configured to compute a dot product comprising a sum of the first partial dot product and the second partial dot product.
and, wherein the first adder ALU is configured to: input the first partial dot product and the second partial dot product; and, compute a dot product comprising a sum of the first partial dot product and the second partial dot product.
As to claim 11, claim 9 of copending application 18378293 does not explicitly teach a Shared Dimension (SD) Splitter, the SD Splitter configured to: determine that a left side matrix and a right side matrix have a shared dimension number, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows;or generating the column-split and row-split matrices.
However, in the same field of endeavor, Liu discloses a Shared Dimension (SD) Splitter, the SD Splitter configured to: determine that a left side matrix and a right side matrix have a shared dimension number, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows (Liu: Fig. 1-2b, wherein matrix A corresponds with the left side matrix and matrix B corresponds with the right side matrix; [0086]); and generating the column-split and row-split matrices (Liu: Fig. 1-2b, [0099]).
Accordingly, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify Claim 9 of copending application 18378293 using Liu and configure the computing system to generate split-matrices based on the shared dimension number because Liu improves operation efficiency during matrix multiplication (Liu: abstract).
Therefore, the combination of copending application 18378293 in view of Liu teaches a computing system that generates column-split and row-split matrices based on the shared dimension of left and right matrices.
As to claim 12, claim 15 of copending application 18378293 teaches The computing system of claim 11, wherein the computing system further comprises a memory; wherein the first processor is further configured to output the first partial dot product to the memory (“the first MACC ALU is configured to output, to a first memory,”); and; wherein the third processor configured to compute the dot product comprises the third processor further configured to input the first partial dot product from the memory to compute the sum of the first partial dot product and the second partial dot product (“wherein the first adder ALU configured to add the first partial dot product to the second partial dot product comprises the first adder ALU further configured to: input the first partial dot product, from the first memory; and, input the second partial dot product from the second memory.”).
As to claim 13, claim 9 of copending application 18378293 teaches The computing system of claim 11, wherein the first processor and the second processor comprise different processors (the first MACC ALU and second MACC ALU as claimed are interpreted as two distinct MACC ALUs).
As to claim 14, claim 12 of copending application 18378293 teaches The computing system of claim 11, wherein the first number is greater than the second number (“The computing system of claim 9, wherein the first number is greater than the second number)”;
wherein the SD Splitter configured to generate the second column-split matrix comprises the SD splitter further configured to generate, based on the first number greater than the second number, the second column-split matrix to further comprise an all-zeros column, each element of the all-zeros column having value zero; wherein the SD splitter configured to generate the second row-split matrix comprises the SD splitter further configured to generate, based on the first number greater than the second number, the second row-split matrix to further comprise an all-zeros row, each element of the all-zeros row having value zero (“wherein, based on the first number greater than the second number, the second column-split matrix further comprises an all-zeros column, each element of the all-zeros column having value zero; wherein, based on the first number greater than the second number, the second row-split matrix further comprises an all-zeros row, each element of the all-zeros row having the value zero”);
and, wherein the second processor configured to compute the second partial dot product comprises the second processor further configured to add, to the second partial dot product, a product of a row element of the all-zeros column of the row of the second column-split matrix multiplied by a corresponding column element of the all-zeros row of the column of second row-split matrix (“and, wherein the second MACC ALU configured to compute the second partial dot product comprises the second MACC ALU further configured to add to the second partial dot product, based on the first number greater than the second number, a product of a row element of the all- zeros column of the second column-split matrix multiplied by a column element of the all- zeros row of the second row-split matrix.”).
As to claim 15, claim 14 of copending application 18378293 teaches The computing system of claim 11, wherein the first number is greater than the second number (“wherein the first number is greater than the second number”);
and, wherein the second processor configured to compute the second partial dot product comprises the second processor adding, based on the first number greater than the second number, a value of zero to the second partial dot product (“and, wherein the second MACC ALU configured to compute the second partial dot product comprises the second MACC ALU further configured to add, based on the first number greater than the second number, a value of zero to the second partial dot product.”).
As to claim 16, claim 11 of copending application 18378293 teaches The computing system of claim 11, wherein the computing system comprises an accumulator (“wherein the third MACC ALU comprises an accumulator”);
and, wherein the first processor configured to compute the first partial dot product comprises the first processor further configured to add a first product, among products included in the first partial dot product, to the accumulator and wherein the first adder ALU configured to add the first partial dot product, output from the first MACC ALU, to the second partial dot product comprises the first adder ALU further configured to add the first partial dot product, output from the first MACC ALU to the accumulator.”).
As to claim 17, claim 16 of copending application 18378293 teaches The computing system of claim 16, wherein the first processor configured to compute the first partial dot product comprises the first processor further configured to compute the first partial dot product as a multiply-accumulate (MACC) computation (“wherein the first MACC ALU comprises a multiplier ALU and a second adder ALU”), the MACC computation comprising: computing, by the first processor, the first product; adding, by the first processor, the first product to the accumulator; computing, by the first processor, a second product, among the products of the column elements of the row of the first column-split matrix multiplied by the corresponding row elements of the column of the first row-split matrix; and, adding, by the first processor, the second product to the accumulator (“wherein the first MACC ALU configured to compute the sum of the first row-column products comprises the second adder ALU configured to compute a sum of the first product and the second product.”).
As to claim 18, claim 16 of copending application 18378293 teaches The computing system of claim 11, wherein the dot product comprises a sum of products of all elements of the row of the first column-split matrix multiplied by all corresponding row elements of the column of the first row-split matrix ( “input, to the multiplier ALU, a first column element, a second column element, a first row element, and a second row element,the first column element and the second column element among the column elements of the row of the first column-split matrix, the first row element and the second row element among the corresponding row elements of the column of the first column-split matrix”; “the second adder ALU configured to compute a sum of the first product and the second product”).
As to claim 19, claim 16 of copending application 18378293 teaches The computing system of claim 17,wherein the first processor comprises a MACC arithmetic logic unit (ALU); and, wherein the first processor configured to compute the first partial dot product as the MACC computation comprises the MACC ALU configured to compute the first partial dot product as the MACC computation (claim 9: “wherein the first MACC ALU is configured to: … compute a first partial dot product comprising a sum of first row-column products, the first row- column products comprising products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix;”).
As to claim 20, claim 20 of copending application 18378293 teaches The computing system of claim 11, wherein at least one of the first processor, the second processor, and the third processor comprises a processor of a reconfigurable dataflow unit (“wherein the MPU comprises a reconfigurable dataflow unit.”).
As to claims 1-8, the claims are directed to a method that would be practiced by the computing system of claims 11-12, 18, 14-17, and 20, respectively. All steps performed by the method of claims 1-8 are executed by the computing system in claims 11-12, 18, 14-17, and 20 as configured. The analysis of claims 11-12, 18, 14-17, and 20 applies equally to claims 1-8.
As to claims 9-10, the claims are directed to a computer readable storage medium having program instructions that would be practiced by the computing system of claims 11 and 14, respectively. All steps performed by the program instructions of claims 9-10 are executed by the computing system in claims 11 and 14 as configured. The analysis of claims 11 and 14 applies equally to claims 9-10.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 6-9, 11-14, 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 20230169144 A1, hereinafter “Liu”) in view of Lu et al. (US 10073816 B1, hereinafter “Lu”).
As per claim 1, Liu teaches A method, the method comprising: determining, by a computing system, that a left side matrix and a right side matrix have a shared dimension number, the left side matrix comprising the shared dimension number of columns and the right side matrix comprising the shared dimension number of rows (Liu: Fig. 1-2b, wherein matrix A corresponds with the left side matrix and matrix B corresponds with the right side matrix; [0086]);
generating, by the computing system, based on the determining that the left side matrix and the right side matrix have the shared dimension number, a first column-split matrix and a first row-split matrix, the first column-split matrix comprising a first number of columns among the shared dimension number of columns of the left side matrix, the first row-split matrix comprising the first number of rows among the shared dimension number of rows of the right side matrix (Liu: Fig. 1-2b, wherein the first column-split matrix corresponds with the left portion of matrix A of six elements, and the first row-split matrix corresponds with the top portion of matrix B of nine elements; [0099]);
generating, by the computing system, based on the determining that the left side matrix and the right side matrix have the shared dimension number, a second column-split matrix and a second row-split matrix, the second column-split matrix comprising a second number of columns among the shared dimension number of columns of the left side matrix, the second row-split matrix comprising the second number of rows among the shared dimension number of rows of the right side matrix (Liu: Fig. 1-2b, wherein the second column-split matrix corresponds with the right portion of matrix A of two elements, and the second row-split matrix corresponds with the bottom portion of matrix B of three elements; [0099]);
computing, by the computing system, a first partial dot product comprising a sum of products of column elements of a row of the first column-split matrix multiplied by corresponding row elements of a column of the first row-split matrix (Liu: Fig. 1-3 element S1-12; [0110]);
However, while Liu discloses partitioning the input matrices for computation on processing elements (Figs. 1-2b, 1-3), Liu does not explicitly disclose circuitry processing the partitions nor that the processing is performed concurrently. Thus, Liu does not teach computing, by the computing system, concurrent with the computing system computing the first partial dot product, a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix; and, computing, by the computing system, a dot product comprising a sum of the first partial dot product and the second partial dot product.
Lu teaches computing, by the computing system, concurrent with the computing system computing the first partial dot product, a second partial dot product comprising a sum of products of column elements of a row of the second column-split matrix multiplied by corresponding row elements of a column of the second row-split matrix (Lu: Fig. 6 element 620; col 8 lines 40-45);
and, computing, by the computing system, a dot product comprising a sum of the first partial dot product and the second partial dot product (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35).
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the system of Liu with the with the contraction engine of Lu. One would have been motivated to combine these references because both references disclose matrix multiplication on partitioned matrices, and the hardware parallelism of Lu allows for scalable architecture (Lu: col 14 lines 15-18).
As per claim 2, Liu/Lu further teaches The method of claim 1, wherein the computing system computing the first partial dot product comprises the computing system storing the first partial dot product in a memory of the computing system (Lu: Fig. 6 the MAC registers storing the result);
and, wherein the computing system computing the dot product comprises: inputting the first partial dot product from the memory to an adder component of the computing system (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35);
inputting, to the adder component, the second partial dot product (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35);
and, computing, by the adder component, the sum of the first partial dot product and the second partial dot product (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35).
As per claim 3, Liu/Lu further teaches The method of claim 1, wherein the dot product comprises a sum of products of all elements of the row of the first column-split matrix multiplied by all corresponding row elements of the column of the first row-split matrix (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35).
As per claim 6, Liu/Lu further teaches The method of claim 1, wherein the computing system comprises an accumulator (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35);
and, wherein the computing system computing the first partial dot product comprises the computing system adding a first product, among the products of the column elements of the row of the first column-split matrix multiplied by the corresponding row elements of the column of the first row-split matrix, to the accumulator (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35).
As per claim 7, Liu/Lu further teaches The method of claim 6, wherein the computing system computing the first partial dot product further comprises the computing system computing the first partial dot product as a multiply-accumulate computation (Lu: Fig. 6 element 640; col 8 lines 44-46), the multiply-accumulate computation comprising: computing, by the computing system, the first product (Lu: Fig. 6 element 640; col 8 lines 44-46);
adding, by the computing system, the first product to the accumulator (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35);
computing, by the computing system, a second product, among the products of the column elements of the row of the first column-split matrix multiplied by the corresponding row elements of the column of the first row-split matrix (Lu: Fig. 6 element 640; col 8 lines 44-46);
and, adding, by the computing system, the second product to the accumulator (Lu: Fig. 6 element 616; Fig. 7 element 716; col 9 lines 33-35).
As per claim 8, Liu/Lu further teaches The method of claim 1, wherein the computing system computing the first partial dot product further comprises a first reconfigurable dataflow unit, included in the computing system, computing the first partial dot product (Lu: Fig. 6 elements 620; wherein one of the outer product units correspond with a first reconfigurable dataflow unit);
and, wherein the computing system computing the dot product further comprises a second reconfigurable dataflow unit, included in the computing system, computing the sum of the first partial dot product and the second partial dot product (Lu: Fig. 6 elements 620; wherein another outer product unit corresponds with a second reconfigurable dataflow unit).
As per claim 9, the claim is directed to a computer readable storage medium that implements the same or similar features as the method of claim 1, and is therefore rejected for at least the same reasons therein.
As per claims 11-12, the claims are directed to a computing system that implements the same or similar features as the method of claims 1-2, respectively, and are therefore rejected for at least the same reasons therein. Furthermore, Liu/Lu further teaches a first processor; a second processor; and a third processor (Lu: Fig. 6 element 610, wherein the processors correspond to elements within the contraction engine); and a Shared Dimension (SD) Splitter (Lu: Fig. 6 element 612; col 8 lines 55-58).
As per claim 13, Liu/Lu further teaches The computing system of claim 11, wherein the first processor and the second processor comprise different processors (Lu: Fig. 6 elements 620; wherein the outer product units correspond with processors).
As per claims 16-18, the claims are directed to a computing system that implements the same or similar features as the method of claims 6-7 and 3, respectively, and are therefore rejected for at least the same reasons therein. Furthermore, Liu/Lu further teaches an accumulator (Lu: Fig. 6 element 616).
As per claim 19, Liu/Lu further teaches The computing system of claim 17, wherein the first processor comprises a MACC arithmetic logic unit (ALU) (Lu: Fig. 6 element 640; col 8 lines 44-46);
and, wherein the first processor configured to compute the first partial dot product as the MACC computation comprises the MACC ALU configured to compute the first partial dot product as the MACC computation (Lu: Fig. 6 element 620; col 8 lines 53-55).
As per claim 20, the claim is directed to a computing system that implements the same or similar features as the method of claim 8, and is therefore rejected for at least the same reasons therein.
Claims 4-5, 10, 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Liu/Lu in further view of Scott et al. (US 20190266218 A1, hereinafter “Scott”).
As per claim 4, Liu/Lu further teaches The method of claim 1, wherein the first number is greater than the second number (Liu: Fig. 1-2b, wherein the first and second number corresponds to the portioning discussed in claim 1);
and, wherein the computing system computing the second partial dot product comprises the computing system adding, to the second partial dot product a product of a row element of the all-zeros column of the row of the second column-split matrix multiplied by a corresponding column element of the all-zeros row of the column of second row-split matrix. (Liu: Fig. 1-3 element S1-12; [0110]),
However, while Liu discloses partitioning matrices that are not equal sizes (Fig. 1-2b) and processing intermediate results (Fig. 1-3), Liu does not disclose how uneven partitions are handled such that intermediate results can be combined. Thus, Liu does not teach wherein the computing system generating the second column-split matrix comprises the computing system generating, based on the first number greater than the second number, the second column-split matrix to further comprise an all-zeros column, each element of the all- zeros column having value zero; wherein the computing system generating the second row-split matrix comprises the computing system generating, based on the first number greater than the second number, the second row-split matrix to further comprise an all-zeros row, each element of the all-zeros row having value zero;
Scott teaches wherein the computing system generating the second column-split matrix comprises the computing system generating, based on the first number greater than the second number, the second column-split matrix to further comprise an all-zeros column, each element of the all- zeros column having value zero (Scott: Fig. 2 element 210; [0035]);
wherein the computing system generating the second row-split matrix comprises the computing system generating, based on the first number greater than the second number, the second row-split matrix to further comprise an all-zeros row, each element of the all-zeros row having value zero (Scott: Fig. 2 element 210; [0035]);
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the system of Liu with the with the zero-padding teaching of Scott, such that the input partitions have the same dimensions. One would have been motivated to combine these references because both references disclose matrix multiplication, and combining prior art elements according to known methods to yield predictable results (ensuring partial dot product outputs have the same dimensions in order to sum them).
As per claim 5, Liu/Lu further teaches The method of claim 1, wherein the first number greater than the second number (Liu: Fig. 1-2b, wherein the first and second number corresponds to the portioning discussed in claim 1);
However, while Liu discloses partitioning matrices that are not equal sizes (Fig. 1-2b) and processing intermediate results (Fig. 1-3), Liu does not disclose how uneven partitions are handled such that intermediate results can be combined. Thus, Liu does not teach and, wherein the computing system computing the second partial dot product comprises the computing system adding to the second partial dot product, based on the first number greater than the second number, a value of zero.
Scott teaches and, wherein the computing system computing the second partial dot product comprises the computing system adding to the second partial dot product, based on the first number greater than the second number, a value of zero (Scott: Fig. 2 element 210; [0035]).
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the system of Liu with the with the zero-padding teaching of Scott, such that the input partitions have the same dimensions for at least the same reasons as discussed for claim 4.
As per claim 10, the claim is directed to a computer readable storage medium that implements the same or similar features as the method of claim 4, and is therefore rejected for at least the same reasons therein.
As per claims 14-15, the claims are directed to a computing system that implements the same or similar features as the method of claims 4-5, respectively, and are therefore rejected for at least the same reasons therein.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHAT N LE whose telephone number is (571)272-0546. The examiner can normally be reached Monday-Friday 8:30AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew T Caldwell can be reached at (571) 272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/P.N.L./
Phat LeExaminer, Art Unit 2182 (571) 272-0546
/Carlo Waje/Examiner, Art Unit 2151