DETAILED ACTION
This action is in response to the filing on 09/09/2024. Claims 1-20 are pending and have been considered below.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement filed 09/09/2024 fails to comply with 37 CFR 1.98(a)(2), which requires a legible copy of each cited foreign patent document; each non-patent literature publication or that portion which caused it to be listed; and all other information or that portion which caused it to be listed. It has been placed in the application file, but the information referred to therein has not been considered.
The following non-patent literature are missing a copy: cite no. 129, 130, 139, 149, 158, 159, and 163.
Specification
The disclosure is objected to because of the following informalities:
Para. 7, line 2, recites "4 or 8- -bits in kernels", contains extra whitespace.
Para. 19, line 3, recites "LLaMa large language model", should recite -- LLM large language model --.
Para. 29, 2nd to last line, recites "multiplication to", contains extra whitespace.
Appropriate correction is required.
Claim Objections
Claims 1, 7, 9, 15, 17, and 18 objected to because of the following informalities:
Claim 1, line 7, recites "produce a product", extra whitespace between 'a' and 'product'.
Claim 1, line 8, recites "product vector", extra whitespace between 'product' and 'vector'.
Claims 7, 9, 15, and 18 have inconsistent formatting and are missing idents on line 1.
Claim 15 has inconsistent formatting and is missing an indent on lines 2 and 3.
Claim 17 has consistent formatting and is missing an indent on line 4.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent Claims 1, 9, and 17
Step 1:
Claims 1, 9, and 17 recite a method, system and method, respectively; therefore, they are directed to one of the four categories of statutory subject matter (process/method, machine/product/apparatus, manufacture, or composition of matter).
Step 2A Prong 1:
Claims 1 and 9 recite a method and system comprising:
for a matrix A, for each row in A, for each unique value z appearing in one or more locations in the row in A: summing the set of rows in a matrix B where the set of rows in matrix B correspond to the indices of z in the row in A, the summing producing a vector; multiplying the vector by the unique value z to produce a product vector; and adding the product vector to a row in an output matrix C which corresponds to the row in A. — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mental process, or a concept that can be performed in the human mind with the use of a physical aid (e.g. pen and paper), including observation, evaluation, judgement or opinion (see MPEP § 2106.04(a)(2)(III)). Or a mathematical concept, specifically mathematical calculations (see MPEP § 2106.04(a)(2)(I)(C).
Claim 17 recites a method comprising:
A method of executing a neural network (NN), the method comprising: summing activation values in an activation tensor which correspond to one unique value in a weight tensor; and multiplying the resulting sum by the unique value. — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mental process, or a concept that can be performed in the human mind with the use of a physical aid (e.g. pen and paper), including observation, evaluation, judgement or opinion (see MPEP § 2106.04(a)(2)(III)). Or a mathematical concept, specifically mathematical calculations (see MPEP § 2106.04(a)(2)(I)(C).
Step 2A Prong 2:
This judicial exception is not integrated into a practical application.
Claim 1 recites the additional elements of:
A method of executing a neural network (NN), the method comprising, using a computer processor: — This element amounts to no more than mere instructions to apply a judicial exception on a computer (see MPEP § 2106.05(f)).
Claim 9 recites the additional elements of:
A system for executing a neural network (NN), the system comprising: a memory; a computer processor to: — This element amounts to no more than mere instructions to apply a judicial exception on a computer (see MPEP § 2106.05(f)).
Step 2B:
The claims do not contain significantly more than the judicial exception.
Claim 1 recites the additional elements of:
A method of executing a neural network (NN), the method comprising, using a computer processor: — This element amounts to no more than mere instructions to apply a judicial exception on a computer (see MPEP § 2106.05(f)).
Claim 9 recites the additional elements of:
A system for executing a neural network (NN), the system comprising: a memory; a computer processor to: — This element amounts to no more than mere instructions to apply a judicial exception on a computer (see MPEP § 2106.05(f)).
As such claims 1, 9, and 17 are not patent eligible.
Dependent Claims 2-8, 10-16, and 18-20
Step 1:
Claims 2-8, 10-16, and 18-20, recite a method, system, and method, respectively; therefore, they are directed to one of the four categories of statutory subject matter (process/method, machine/product/apparatus, manufacture, or composition of matter).
Step 2A Prong 1:
Claims 2-8, 10-16, and 18-20 merely narrow the previously cited abstract idea limitations. For the reasons described above with respect to independent claims 1, 9, and 17 this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claim(s) disclose similar limitations described for the independent claim(s) above and do not provide anything more than the abstract idea.
Claims 2 and 10 recite a method and system comprising:
wherein each value in A is quantized such that each value in A is represented by three or fewer bits; and wherein each value in B is not quantized, such that each value in B is represented by eight or more bits — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mathematical concept specifically a mathematical relationship (see MPEP § 2106.04(a)(2)(I)(A)).
Claims 3, 11, and 19 recite a method, system, and method comprising:
comprising performing inference on the NN by accepting an input to the NN and producing an output from the NN — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mental process, or a concept that can be performed in the human mind with the use of a physical aid (e.g. pen and paper), including observation, evaluation, judgement or opinion (see MPEP § 2106.04(a)(2)(III)).
Claims 4, 12, and 20 recite a method, system, and method comprising:
wherein the summing uses add CPU instructions and the multiplication uses multiply CPU instructions — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mathematical concept (see MPEP § 2106.04(a)(2)(I)), specifically a mathematical algorithm being applied on a general purpose computer (see MPEP § 2106.05(f)(2)(i)).
Claims 5, 13, and 18 recite a method, system, and method comprising:
wherein the summing uses vector add instructions and the multiplication uses vector multiply instructions — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mathematical concept (see MPEP § 2106.04(a)(2)(I)), specifically a mathematical algorithm being applied on a general purpose computer (see MPEP § 2106.05(f)(2)(i)).
Claims 6 and 14 recite a method and system comprising:
comprising producing code for the summing and multiplying based on an input of the matrix A — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mathematical concept (see MPEP § 2106.04(a)(2)(I)), specifically a mathematical algorithm being applied on a general purpose computer (see MPEP § 2106.05(f)(2)(i)).
Claims 8 and 16 recite a method and system comprising:
wherein summing the set of rows in the matrix B comprises partitioning each of the set of rows in the matrix B such that the summing occurs for each partition — Under its broadest reasonable interpretation, this limitation encompasses the abstract idea of a mental process, or a concept that can be performed in the human mind with the use of a physical aid (e.g. pen and paper), including observation, evaluation, judgement or opinion (see MPEP § 2106.04(a)(2)(III)). Or a mathematical concept, specifically mathematical relationships (see MPEP § 2106.04(a)(2)(I)(A)).
Step 2A Prong 2:
This judicial exception is not integrated into a practical application.
Claims 7 and 15 recite the additional element of:
wherein matrix A stores weights or a kernel — This element amounts to no more than insignificant extra-solution activity in the form of mere data gathering and output (see MPEP § 2106.05(g)), and is well-known, understood, routine, conventional, activity (see MPEP § 2106.05(d)(II), performing repetitive calculations).
Step 2B:
The claims do not contain significantly more than the judicial exception.
Claims 7 and 15 recite the additional element of:
wherein matrix A stores weights or a kernel — This element amounts to no more than insignificant extra-solution activity in the form of mere data gathering and output (see MPEP § 2106.05(g)), and is well-known, understood, routine, conventional, activity (see MPEP § 2106.05(d)(II), performing repetitive calculations).
As such claims 2-8, 10-16, and 18-20 are not patent eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 17 and 19-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Moshovos et al. (US 2022/0092382 A1), hereinafter Moshovos.
Regarding claim 17, Moshovos teaches A method of executing a neural network (NN), the method comprising: summing activation values in an activation tensor which correspond to one unique value in a weight tensor; and multiplying the resulting sum by the unique value (Moshovos discloses using a computer system with a processor and computer-readable instructions for the neural network computations and storage [see Moshovos, para. 20, 21, and 26]. Moshovos further discloses performing a multiplication between a weight tensor and activation tensor such that the corresponding activations can be accumulated into centroid sums and then multiplying the sums by the weights [see Moshovos, para. 124-125 and Eq. 2-5]).
Regarding claim 19, Moshovos teaches all the limitations of claim 17 and further teaches:
comprising performing inference on the NN by accepting an input to the NN and producing an output from the NN (In some embodiments, the method further includes generating an output from performing one or more multiply-accumulate operations A.sub.1B.sub.1+ . . . +A.sub.nB.sub.n on input vectors A and input vectors B, wherein n is the n-th input vector and wherein one or more of input vectors B are each one of the representative values, by accumulating input vectors A to an accumulated sum of input vectors A per input vector B having the same representative value and subsequently multiplying each of the accumulated sums of input vectors A by the representative value of the input vector B. [see Moshovos, para. 12]).
Regarding claim 20, Moshovos teaches all the limitations of claim 17 and further teaches:
wherein the summing uses add CPU instructions and the multiply uses multiply CPU instructions (Moshovos discloses using a computer system with a processor and computer-readable instructions for the neural network computations and storage [see Moshovos, para. 20, 21, and 26], as well as performing accumulation and multiply-accumulate MAC operations for the neural network [see Moshovos, para. 124-125 and Eq. 2-5]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Moshovos et al. (US 2022/0092382 A1), hereinafter Moshovos, in view of Boswell et al. (US 2018/0321938 A1), hereinafter Boswell.
Regarding claim 1, Moshovos teaches A method of executing a neural network (NN), the method comprising, using a computer processor: (Moshovos discloses using a computer system with a processor and computer-readable instructions for the neural network computations and storage [see Moshovos, para. 20, 21, and 26]):
summing the vector B where the indices in vector B correspond to the index of z in matrix A, the summing producing a vector; multiplying the vector by the unique value z to produce a product vector; and adding the product vector to an output vector C which corresponds to the vector A (Moshovos discloses performing a multiplication between a weight matrix and activation vector such that the corresponding activations can be accumulated into centroid sums and then multiplying the sums by the weights [see Moshovos, para. 124-125 and Eq. 2-5]).
However, Moshovos fails to teach for a matrix A, for each row in A, for each unique value z appearing in one or more locations in the row in A: summing the set of rows in a matrix B where the set of rows in matrix B correspond to the indices of z in the row in A, the summing producing a vector; multiplying the vector by the unique value z to produce a product vector; and adding the product vector to a row in an output matrix C which corresponds to the row in A.
In the same field of endeavor, Boswell teaches:
for a matrix A, for each row in A, for each unique value z appearing in one or more locations in the row in A (Boswell discloses a method of matrix multiply and accumulate (MMA) operation between a multiplicand input matrix A and multiplier input matrix B with a third collector matrix C to accumulate the results [see Boswell, para. 25], and that "each element of the result matrix is generated by calculating at least one dot product of corresponding pairs of vectors stored in the plurality of operand collectors" [see Boswell, para. 27]).
It would have been obvious to one of ordinary skill, in the art at the time before the effective filing date of the invention to incorporate for a matrix A, for each row in A, for each unique value z appearing in one or more locations in the row in A as suggested in Boswell into Moshovos because both methods perform multiply-accumulate operations (see Moshovos, Abstract; see Boswell, Abstract) and to further teach for a matrix A, for each row in A, for each unique value z appearing in one or more locations in the row in A: summing the set of rows in a matrix B where the set of rows in matrix B correspond to the indices of z in the row in A, the summing producing a vector; multiplying the vector by the unique value z to produce a product vector; and adding the product vector to a row in an output matrix C which corresponds to the row in A because it would have been obvious to one of ordinary skill in the art before the effective filing date that the multiply-accumulate operations of Moshovos between a weight matrix and input vector [see Moshovos, para. 124-125 and Eq. 2-5] could be extended to matrix multiply-accumulate operations such that input and output matrices could be used in addition to the weight matrix similar to Boswell [see Boswell, para. 25]. Thus, the summing would process per Eq. 5 of Moshovos, summing the rows of the input matrix such that values corresponding to the same weight representative values are added together before being multiplied by the representative value and resulting in the product vector to be added to the output matrix. Incorporating the teaching of Boswell into Moshovos would accelerate matrix operations as executed by a processor (see Boswell, para. 24).
Regarding claim 2, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein each value in A is quantized such that each value in A is represented by three or fewer bits; and wherein each value in B is not quantized, such that each value in B is represented by eight or more bits (Moshovos discloses that the weights are quantized from 32-bit to 3-bits using a dictionary of 8 representative values (e.g., centroids) and storing as 3-bit indexes [see Moshovos, para. 47 and 53], and that the computation of the activation vector uses 32-bit activation sums without quantization [see Moshovos, para. 124]).
Regarding claim 3, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
comprising performing inference on the NN by accepting an input to the NN and producing an output from the NN (In some embodiments, the method further includes generating an output from performing one or more multiply-accumulate operations A.sub.1B.sub.1+ . . . +A.sub.nB.sub.n on input vectors A and input vectors B, wherein n is the n-th input vector and wherein one or more of input vectors B are each one of the representative values, by accumulating input vectors A to an accumulated sum of input vectors A per input vector B having the same representative value and subsequently multiplying each of the accumulated sums of input vectors A by the representative value of the input vector B. [see Moshovos, para. 12]).
Regarding claim 4, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the summing uses add CPU instructions and the multiplication uses multiply CPU instructions (Moshovos discloses using a computer system with a processor and computer-readable instructions for the neural network computations and storage [see Moshovos, para. 20, 21, and 26], as well as performing accumulation and multiply-accumulate MAC operations for the neural network [see Moshovos, para. 124-125 and Eq. 2-5]).
Regarding claim 5, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein the summing uses vector add instructions and the multiplication uses vector multiply instructions (Boswell discloses using single-instruction multiple thread architecture (SIMT) using the combination of different vectors to compute the matrix multiply-accumulate (MMA) operations [see Boswell, para. 108]).
Regarding claim 6, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
comprising producing code for the summing and multiplying based on an input of the matrix A (Moshovos discloses using a computer system with a processor and computer-readable instructions for the neural network computations and storage [see Moshovos, para. 20, 21, and 26], as well as performing accumulation and multiply-accumulate MAC operations [see Moshovos, para. 124-125 and Eq. 2-5]. It would have been obvious to one of ordinary skill in the art before the effective filing date to produce code for the summing and multiplying based on the input of matrix A, because the values in A correspond to different values V in Eq. 5, thus, the code has to be produced to add and multiply different registers at each operation, it won’t always use the same values, or always group the same indices of A.).
Regarding claim 7, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein matrix A stores weights or a kernel (Moshovos discloses that the weights are quantized from 32-bit to 3-bits using a dictionary of 8 representative values (e.g., centroids) and storing as 3-bit indexes [see Moshovos, para. 47 and 53], and that the 3-bit values are stored in a weight matrix [see Moshovos, para. 124]).
Regarding claim 8, the combination of Moshovos and Boswell as applied in claim 1 above teaches all the limitations of claim 1 and further teaches:
wherein summing the set of rows in the matrix B comprises partitioning each of the set of rows in the matrix B such that the summing occurs for each partition (It would have been obvious to one of ordinary skill in the art before the effective filing date that the multiply-accumulate operations of Moshovos between a weight matrix and input vector [see Moshovos, para. 124-125 and Eq. 2-5] could be extended to matrix multiply-accumulate operations such that input and output matrices could be used in addition to the weight matrix similar to Boswell [see Boswell, para. 25]. Thus, the summing would process per Eq. 5 of Moshovos, summing the rows of the input matrix such that values corresponding to the same weight representative values are added together before being multiplied by the representative value and resulting in the product vector to be added to the output matrix).
Regarding claim 9, claim 9 contains substantially similar limitations to those found in claim 1. Therefore it is rejected for the same reason as claim 1 above. Additionally, the combination of Moshovos and Boswell further teaches:
A system for executing a neural network (NN), the system comprising: a memory; a computer processor to: (Moshovos discloses using a computer system with a processor and computer-readable instructions for the neural network computations and storage [see Moshovos, para. 20, 21, and 26]).
Regarding claim 10, claim 19 contains substantially similar limitations to those found in claim 2 above. Consequently, claim 10 is rejected for the same reasons.
Regarding claim 11, claim 11 contains substantially similar limitations to those found in claim 3 above. Consequently, claim 11 is rejected for the same reasons.
Regarding claim 12, claim 12 contains substantially similar limitations to those found in claim 4 above. Consequently, claim 12 is rejected for the same reasons.
Regarding claim 13, claim 13 contains substantially similar limitations to those found in claim 5 above. Consequently, claim 13 is rejected for the same reasons.
Regarding claim 14, claim 14 contains substantially similar limitations to those found in claim 6 above. Consequently, claim 14 is rejected for the same reasons.
Regarding claim 15, claim 15 contains substantially similar limitations to those found in claim 7 above. Consequently, claim 15 is rejected for the same reasons.
Regarding claim 16, claim 16 contains substantially similar limitations to those found in claim 8 above. Consequently, claim 16 is rejected for the same reasons.
Regarding claim 18, Moshovos as applied in claim 17 above teaches all the limitations of claim 17 and further teaches:
However, Moshovos fails to teach wherein the summing uses vector add instructions and the multiplying uses vector multiply instructions.
In the same field of endeavor, Boswell teaches:
wherein the summing uses vector add instructions and the multiplying uses vector multiply instructions (Boswell discloses using single-instruction multiple thread architecture (SIMT) using the combination of different vectors to compute the matrix multiply-accumulate (MMA) operations [see Boswell, para. 108]).
It would have been obvious to one of ordinary skill, in the art at the time before the effective filing date of the invention to incorporate wherein the summing uses vector add instructions and the multiplying uses vector multiply instructions as suggested in Boswell into Moshovos because both methods perform multiply-accumulate operations (see Moshovos, Abstract; see Boswell, Abstract). Incorporating the teaching of Boswell into Moshovos would accelerate matrix operations as executed by a processor (see Boswell, para. 24).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Liu et al. (US 2020/0326938 A1) teaches a processor for sparse matrix computations with instruction sets requiring fewer multiply/accumulate operations as well as single-instruction multiple data (SIMD) processing.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAKE BREEN whose telephone number is (571)272-0456. The examiner can normally be reached Monday - Friday, 7:00 AM - 3:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.T.B./Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143