Prosecution Insights
Last updated: August 18, 2026
Application No. 18/619,392

WAVE LEVEL MATRIX MULTIPLY INSTRUCTIONS

Non-Final OA §102§103§112
Filed
Mar 28, 2024
Priority
Apr 03, 2023 — provisional 63/493,972
Examiner
VICARY, KEITH E
Art Unit
2183
Tech Center
2100 — Computer Architecture & Software
Assignee
Advanced Micro Devices Inc.
OA Round
3 (Non-Final)
58%
Grant Probability
Moderate
3-4
OA Rounds
1y 6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 58% of resolved cases
58%
Career Allowance Rate
399 granted / 692 resolved
+2.7% vs TC avg
Strong +41% interview lift
Without
With
+41.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
40 currently pending
Career history
740
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
34.6%
-5.4% vs TC avg
§102
11.4%
-28.6% vs TC avg
§112
37.3%
-2.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 692 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending in this office action and presented for examination. Claims 1-3, 5-10, 12-17, and 19-20 are newly amended by the response received March 31, 2026. In claim 1, line 3, a comma appears to be removed without appropriate strikethrough. In claim 7, line 2, a comma appears to be added without appropriate underlining. In claim 10, line 3, a comma appears to be added without appropriate underlining. In claim 14, line 3, “wherein” appears to be added without appropriate underlining. In claim 15, line 8, a comma appears to be removed without appropriate strikethrough. In claim 17, line 3, a comma appears to be added without appropriate underlining. Examiner requests that future amendments be made in the appropriate manner conveyed in MPEP 714 to avoid confusion and potential notices of non-compliant amendment. See paragraph 7 of the office action dated October 20, 2025. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 1 recites the limitation “a register file comprising circuitry configured to store operand data for vector operations … fetch, from the register file only once, a first plurality of values” in lines 2-8 (with further recited limitations regarding the first plurality of values). However, a claim may lack written description support when a broad genus claim is presented but the disclosure only describes a narrow species with no evidence that the genus is contemplated. In the instant case, Examiner submits that the claim is a broad genus claim (by using the language "register file", which encompasses both “scalar register file” 332 and “vector register file 330”, in the context of the remaining language of the limitation), but the original disclosure (e.g., paragraph [0014], “The circuitry of the parallel data processing circuit performs a matrix multiplication operation using source operands accessed only once from a vector register file”; paragraph [0018], “Therefore, the data of the rows and columns of matrices 110 and 120 are retrieved only once from the vector register file”; original claim 1, “fetch, from the vector register file only once, a first plurality of values”) only describes a narrow species (fetching, from the “vector” register file only once, a first plurality of values) with no evidence that the genus is contemplated. Examiner notes that while the claim recites a register file “comprising circuitry configured to store operand data for vector operations”, this further language does not appear limit the recited register file to be a vector register file in particular. For example, paragraphs [0033]-[0035] and [0039] appear to convey that a scalar register file can store scalar operand data for a vector operation. Note that claim 2 recites the similar limitation “fetch, from the register file only once, a second plurality of values” in line 4. Claim 1 recites the limitation “perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 9-11. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., paragraph [0018]) does not appear to provide support for performing, using the arithmetic logic circuit, a first operation by reusing just one of the first plurality of values for at least two iterations of computations used to perform the first operation, which is a scenario encompassed by the claim language in view of the “one or more” language. For example, the original disclosure (e.g., paragraph [0014]) does not appear to provide support for performing, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for “at least two” iterations of computations used to perform the first operation. Claims 2-7 are rejected for failing to alleviate the rejections of claim 1 above. Claim 7 recites the limitation “use, in a pipeline stage of the plurality of arithmetic logic circuits, each of the first plurality of values of the first matrix or each of the second plurality of values of the third matrix a number of times equal to a number of columns of the second matrix or the fourth matrix, respectively” in lines 10-13. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., [0018], [0028], [0046]) does not appear to provide support for a pipeline stage of a plurality of arithmetic logic circuits using each of the first plurality of values of the first matrix or each of the second plurality of values of the third matrix a number of times equal to a number of columns of the second matrix or the fourth matrix, respectively. Claim 8 recites the limitation “storing, by circuitry of a register file, operand data for vector operations; … fetching, by the processing circuit from the register file only once, a first plurality of values” in lines 1-6 (with further recited limitations regarding the first plurality of values). However, a claim may lack written description support when a broad genus claim is presented but the disclosure only describes a narrow species with no evidence that the genus is contemplated. In the instant case, Examiner submits that the claim is a broad genus claim (by using the language "register file", which encompasses both “scalar register file” 332 and “vector register file 330”, in the context of the remaining language of the limitation), but the original disclosure (e.g., paragraph [0014], “The circuitry of the parallel data processing circuit performs a matrix multiplication operation using source operands accessed only once from a vector register file”; paragraph [0018], “Therefore, the data of the rows and columns of matrices 110 and 120 are retrieved only once from the vector register file”; original claim 1, “fetch, from the vector register file only once, a first plurality of values”) only describes a narrow species (fetching, from the “vector” register file only once, a first plurality of values) with no evidence that the genus is contemplated. Examiner notes that while the claim recites “storing, by circuitry of a register file, operand data for vector operations”, this further language does not appear limit the recited register file to be a vector register file in particular. For example, paragraphs [0033]-[0035] and [0039] appear to convey that a scalar register file can store scalar operand data for a vector operation. Note that claim 9 recites the similar limitation “fetching, by the processing circuit from the register file only once, a second plurality of values” in lines 4-5. Claim 8 recites the limitation “performing, using an execution pipeline of the processing circuit comprising an arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 7-12. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., paragraph [0018]) does not appear to provide support for performing, using an execution pipeline of the processing circuit comprising an arithmetic logic circuit, a first operation by reusing just one of the first plurality of values for at least two iterations of computations used to perform the first operation, which is a scenario encompassed by the claim language in view of the “one or more” language. For example, the original disclosure (e.g., paragraph [0014]) does not appear to provide support for performing, using an execution pipeline of the processing circuit comprising an arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for “at least two” iterations of computations used to perform the first operation. Claims 9-14 are rejected for failing to alleviate the rejections of claim 8 above. Claim 14 recites the limitation “using, in a pipeline stage of the plurality of arithmetic logic circuits, each of the first plurality of values of the first matrix or each of the second plurality of values of the third matrix a number of times equal to a number of columns of the second matrix or the fourth matrix, respectively” in lines 10-13. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., [0018], [0028], [0046]) does not appear to provide support for a pipeline stage of a plurality of arithmetic logic circuits using each of the first plurality of values of the first matrix or each of the second plurality of values of the third matrix a number of times equal to a number of columns of the second matrix or the fourth matrix, respectively. Claim 15 recites the limitation “a register file comprising circuitry configured to store operand data for vector operations … fetch, from the register file only once, a first plurality of values” in lines 7-14 (with further recited limitations regarding the first plurality of values). However, a claim may lack written description support when a broad genus claim is presented but the disclosure only describes a narrow species with no evidence that the genus is contemplated. In the instant case, Examiner submits that the claim is a broad genus claim (by using the language "register file", which encompasses both “scalar register file” 332 and “vector register file 330”, in the context of the remaining language of the limitation), but the original disclosure (e.g., paragraph [0014], “The circuitry of the parallel data processing circuit performs a matrix multiplication operation using source operands accessed only once from a vector register file”; paragraph [0018], “Therefore, the data of the rows and columns of matrices 110 and 120 are retrieved only once from the vector register file”; original claim 1, “fetch, from the vector register file only once, a first plurality of values”) only describes a narrow species (fetching, from the “vector” register file only once, a first plurality of values) with no evidence that the genus is contemplated. Examiner notes that while the claim recites a register file “comprising circuitry configured to store operand data for vector operations”, this further language does not appear limit the recited register file to be a vector register file in particular. For example, paragraphs [0033]-[0035] and [0039] appear to convey that a scalar register file can store scalar operand data for a vector operation. Note that claim 16 recites the similar limitation “fetch, from the register file only once, a second plurality of values” in line 4. Claim 15 recites the limitation “perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 15-17. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., paragraph [0018]) does not appear to provide support for performing, using the arithmetic logic circuit, a first operation by reusing just one of the first plurality of values for at least two iterations of computations used to perform the first operation, which is a scenario encompassed by the claim language in view of the “one or more” language. For example, the original disclosure (e.g., paragraph [0014]) does not appear to provide support for performing, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for “at least two” iterations of computations used to perform the first operation. Claims 16-20 are rejected for failing to alleviate the rejections of claim 15 above. Claim 19 recites the limitation “The computing system as recited in claim 17, wherein the computing system further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits, wherein responsive to the first instruction, each of the plurality of arithmetic logic circuits is configured to: receive the first plurality of values of the first matrix and the second matrix used as source operands; and perform a matrix multiplication operation of a fused multiply add (FMA) operation using at least a first multiplier circuit and a second multiplier circuit, each having a size less than a size of the first plurality of values of the first matrix and the second matrix” in lines 1-10. Claim 15, upon which claim 19 is indirectly dependent, recites the limitation “a second processor comprising: a register file comprising circuitry configured to store operand data for vector operations; an execution pipeline comprising an arithmetic logic circuit; and circuitry configured to: responsive to a first instruction of the one or more kernels with a first type; fetch, from the register file only once, the first plurality of values; and perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 6-17. The metes and bounds of claim 19 encompasses the possibility that the overall computing system, but not necessarily the second processor, comprises the recited plurality of execution pipelines having the plurality of arithmetic logic circuits configured to perform the recited functionality; however, the original disclosure (e.g., FIG. 5) does not appear to provide support for this possibility. Claim 20 is rejected for failing to alleviate the rejection of claim 19 above. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites the limitation “A processor comprising: a register file comprising circuitry configured to store operand data for vector operations; an execution pipeline comprising an arithmetic logic circuit; and circuitry, wherein responsive to a first instruction, the circuitry is configured to: fetch, from the register file only once, a first plurality of values; and perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 1-11. However, the specification appears to conflict with the aforementioned claimed subject matter. For example, while the claim recites that an execution pipeline comprises an arithmetic logic circuit, the specification (e.g., paragraph [0047]) appears to convey that an ALU comprises an execution pipeline. For example, while the claim appears to recite the circuitry of line 6 as a separate element from the register file in line 2 and the execution pipeline in line 3, Figure 3 shows the vector processing circuit 310A comprising vector register file 330 and vector ALU 350. A claim, although clear on its face, may also be indefinite when a conflict or inconsistency between the claimed subject matter and the specification disclosure renders the scope of the claim uncertain as inconsistency with the specification disclosure or prior art teachings may make an otherwise definite claim take on an unreasonable degree of uncertainty. Therefore, because the aforementioned claimed subject matter appears to conflict with the specification in the manner explained above, the claim is indefinite. Note that claim 5 recites the similar limitation “the processor further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits” in lines 1-2. Note that claim 7 recites the similar limitation “the processor further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits” in lines 1-2. Claim 1 recites the limitation “the circuitry” in line 7. However, it is indefinite as to whether the antecedent basis for this limitation is “circuitry” in claim 1, line 2, or “circuitry” in claim 1, line 6. Note that this limitation is also recited in claim 2, line 3; claim 3, line 1; and claim 7, line 3. Claims 2-7 are rejected for failing to alleviate the rejections of claim 1 above. Claim 7 recites the limitation “the first plurality of values of the first matrix” in lines 10-11. However, there is insufficient antecedent basis for this limitation in the claims. Claim 7 recites the limitation “the second plurality of values of the third matrix” in lines 11-12. However, there is insufficient antecedent basis for this limitation in the claims. Claim 8 recites the limitation “A method, comprising: storing, by circuitry of a register file, operand data for vector operations; responsive to receiving, by a processing circuit, a first instruction of a first type: fetching, by the processing circuit from the register file only once, a first plurality of values; and performing, using an execution pipeline of the processing circuit comprising an arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 1-12. However, the specification appears to conflict with the aforementioned claimed subject matter. For example, while the claim recites that an execution pipeline comprises an arithmetic logic circuit, the specification (e.g., paragraph [0047]) appears to convey that an ALU comprises an execution pipeline. For example, while the claim appears to recite the processing circuit of line 5 as a separate element from the vector register file in line 5, Figure 3 shows the vector processing circuit 310A comprising vector register file 330. A claim, although clear on its face, may also be indefinite when a conflict or inconsistency between the claimed subject matter and the specification disclosure renders the scope of the claim uncertain as inconsistency with the specification disclosure or prior art teachings may make an otherwise definite claim take on an unreasonable degree of uncertainty. Therefore, because the aforementioned claimed subject matter appears to conflict with the specification in the manner explained above, the claim is indefinite. Note that claim 12 recites the similar limitation “the processing circuit further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits” in lines 1-2. Note that claim 14 recites the similar limitation “the processing circuit further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits” in lines 1-3. Claims 9-14 are rejected for failing to alleviate the rejection of claim 8 above. Claim 9 recites the limitation “The method as recited in claim 8, responsive to receiving, by the processing circuit, a second instruction of a second type different from the first type of the first instruction: fetching … performing …” in lines 1-9. However, the metes and bounds of this limitation are grammatically indefinite. For example, it is indefinite as to whether the method is being recited to comprise the fetching and performing steps. Claims 10-14 are rejected for failing to alleviate the rejection of claim 9 above. Claim 10 recites the limitation “The method as recited in claim 9, further comprising fetching, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operation, by the processing circuit, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands” in lines 1-5. Claim 9, upon which claim 10 is dependent, recites the limitation “fetching, by the processing circuit from the register file only once, a second plurality of values” in lines 4-5. Claim 8, upon which claim 9 is dependent, recites the limitation “fetching, by the processing circuit from the register file only once, a first plurality of values” in lines 5-6. Therefore, it is indefinite as to whether claim 10 (in the context of claim 8) entails fetching the first plurality of values once or twice, in view of the “further” language in claim 10, line 1. Similarly, it is indefinite as to whether claim 10 (in the context of claim 9) entails fetching the second plurality of values once or twice, in view of the “further” language in claim 10, line 1. Claims 11-14 are rejected for failing to alleviate the rejection of claim 10 above. Claim 15 recites the limitation “a second processor comprising: a register file comprising circuitry configured to store operand data for vector operations; an execution pipeline comprising an arithmetic logic circuit; and circuitry configured to: responsive to a first instruction of the one or more kernels with a first type; fetch, from the register file only once, the first plurality of values; and perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation” in lines 6-17. However, the specification appears to conflict with the aforementioned claimed subject matter. For example, while the claim recites that an execution pipeline comprises an arithmetic logic circuit, the specification (e.g., paragraph [0047]) appears to convey that an ALU comprises an execution pipeline. For example, while the claim appears to recite the circuitry of line 11 as a separate element from the register file in line 7 and the execution pipeline in line 8, Figure 3 shows the vector processing circuit 310A comprising vector register file 330 and vector ALU 350. A claim, although clear on its face, may also be indefinite when a conflict or inconsistency between the claimed subject matter and the specification disclosure renders the scope of the claim uncertain as inconsistency with the specification disclosure or prior art teachings may make an otherwise definite claim take on an unreasonable degree of uncertainty. Therefore, because the aforementioned claimed subject matter appears to conflict with the specification in the manner explained above, the claim is indefinite. Note that claim 19 recites the similar limitation “the computing system further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits” in lines 1-3. Claim 15 recites the limitation “first instruction of the one or more kernels with a first type” in line 12. However, it is indefinite as to whether it is the first instruction, or the one or more kernels, with the first type. Claims 16-20 are rejected for failing to alleviate the rejections of claim 15 above. Claim 16 recites the limitation “the circuitry” in line 3. However, it is indefinite as to whether the antecedent basis for this limitation is “circuitry” in claim 15, line 2, “circuitry” in claim 15, line 4; “circuitry” in claim 15, line 7; or “circuitry” in claim 15, line 11. Note that this limitation is also recited in claim 17, line 1. Claims 17-20 are rejected for failing to alleviate the rejection of claim 16 above. The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claims 3-7, 10-14, and 17-20 are rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim 3 recites the limitation “The processor as recited in claim 2, wherein the circuitry is configured to fetch, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operation, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands” in lines 1-5. Claim 2, upon which claim 3 is dependent, recites the limitation “fetch, from the register file only once, a second plurality of values” in line 4. Claim 1, upon which claim 2 is dependent, recites the limitation “fetch, from the register file only once, a first plurality of values” in line 8. Therefore, claim 3 appears to fail to include all the limitations of the claim upon which it depends, because claim 1 recites that the first plurality of values is fetched from the register file only once and claim 2 recites that the second plurality of values is fetched from the register file only once, whereas claim 3 appears to encompass the possibility that, rather than the first plurality of values or the second plurality of values being fetched from the register file only once, the first plurality of values or the second plurality of values may be fetched from the register file an additional time(s) following each element of a resulting matrix being updated by the first operation or the second operation. Claims 4-7 are rejected for failing to alleviate the rejection of claim 3 above. Claim 10 recites the limitation “The method as recited in claim 9, further comprising fetching, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operation, by the processing circuit, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands” in lines 1-5. Claim 9, upon which claim 10 is dependent, recites the limitation “fetching, by the processing circuit from the register file only once, a second plurality of values” in lines 4-5. Claim 8, upon which claim 9 is dependent, recites the limitation “fetching, by the processing circuit from the register file only once, a first plurality of values” in lines 5-6. Therefore, claim 10 appears to fail to include all the limitations of the claim upon which it depends, because claim 8 recites that the first plurality of values is fetched from the register file only once and claim 9 recites that the second plurality of values is fetched from the register file only once, whereas claim 10 appears to encompass the possibility that, rather than the first plurality of values or the second plurality of values being fetched from the register file only once, the first plurality of values or the second plurality of values may be fetched from the register file an additional time(s) following each element of a resulting matrix being updated by the first operation or the second operation. Claims 11-14 are rejected for failing to alleviate the rejection of claim 10 above. Claim 17 recites the limitation “The computing system as recited in claim 16, wherein the circuitry is configured to fetch, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operation, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands” in lines 1-5. Claim 16, upon which claim 17 is dependent, recites the limitation “fetch, from the register file only once, a second plurality of values” in line 4. Claim 15, upon which claim 16 is dependent, recites the limitation “fetch, from the register file only once, the first plurality of values” in line 14. Therefore, claim 17 appears to fail to include all the limitations of the claim upon which it depends, because claim 15 recites that the first plurality of values is fetched from the register file only once and claim 16 recites that the second plurality of values is fetched from the register file only once, whereas claim 17 appears to encompass the possibility that, rather than the first plurality of values or the second plurality of values being fetched from the register file only once, the first plurality of values or the second plurality of values may be fetched from the register file an additional time(s) following each element of a resulting matrix being updated by the first operation or the second operation. Claims 18-20 are rejected for failing to alleviate the rejection of claim 17 above. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1 and 8 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhang et al. (Zhang) (US 20220206749 A1). Consider claim 1, Zhang discloses a processor ([0029], line 6, processor) comprising: a register file comprising circuitry configured to store operand data for vector operations ([0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel); an execution pipeline comprising an arithmetic logic circuit ([0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400); and circuitry, wherein responsive to a first instruction ([0053], line 12, single instruction), the circuitry is configured to: fetch, from the register file only once, a first plurality of values ([0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations); and perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation ([0048], lines 1-6, in Block 504, the data reuse unit 321 determines the multiple data subsets from the data set, so as to respectively input the multiple data subsets into the multiple dot product data units 322-1 to 322-n. The two data subsets inputted into the two adjacent dot product data units include a portion of the same data; [0049], lines 1-4, in Block 506, each dot product data unit of the multiple dot product data units 322-1 to 322-n performs the dot product operation on the inputted data subset, so as to generate the dot product operation result; [0050], lines 1-5, in Block 508, each dot product data unit of the multiple dot product data units 322-1 to 322-n generates the current cumulative result of the dot product data unit based on the previous cumulative result of the dot product data unit and the dot product operation result; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 6-8, otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation). Consider claim 8, Zhang discloses a method, comprising: storing, by circuitry of a register file, operand data for vector operations ([0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel); responsive to receiving, by a processing circuit ([0029], line 6, processor), a first instruction of a first type ([0053], line 12, single instruction): fetching, by the processing circuit from the register file only once, a first plurality of values ([0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel); and performing, using an execution pipeline of the processing circuit comprising an arithmetic logic circuit ([0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400), a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation ([0048], lines 1-6, in Block 504, the data reuse unit 321 determines the multiple data subsets from the data set, so as to respectively input the multiple data subsets into the multiple dot product data units 322-1 to 322-n. The two data subsets inputted into the two adjacent dot product data units include a portion of the same data; [0049], lines 1-4, in Block 506, each dot product data unit of the multiple dot product data units 322-1 to 322-n performs the dot product operation on the inputted data subset, so as to generate the dot product operation result; [0050], lines 1-5, in Block 508, each dot product data unit of the multiple dot product data units 322-1 to 322-n generates the current cumulative result of the dot product data unit based on the previous cumulative result of the dot product data unit and the dot product operation result; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 6-8, otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 2-7 and 9-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (in the case of claims 2-7, as applied to claim 1 above; in the case of claims 9-14, as applied to claim 8 above), and further in view of Chen et al. (Chen) (US 20200272687 A1). Consider claim 2, Zhang discloses the processor as recited in claim 1 (see above), wherein responsive to an instruction, the circuitry is further configured to: fetch, from the register file only once, a second plurality of values ([0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations); and perform, using the arithmetic logic circuit, an operation by reusing one or more of the second plurality of values for at least two iterations of computations used to provide the operation ([0048], lines 1-6, in Block 504, the data reuse unit 321 determines the multiple data subsets from the data set, so as to respectively input the multiple data subsets into the multiple dot product data units 322-1 to 322-n. The two data subsets inputted into the two adjacent dot product data units include a portion of the same data; [0049], lines 1-4, in Block 506, each dot product data unit of the multiple dot product data units 322-1 to 322-n performs the dot product operation on the inputted data subset, so as to generate the dot product operation result; [0050], lines 1-5, in Block 508, each dot product data unit of the multiple dot product data units 322-1 to 322-n generates the current cumulative result of the dot product data unit based on the previous cumulative result of the dot product data unit and the dot product operation result; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 6-8, otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation). However, Zhang does not disclose that the aforementioned instruction is a second instruction of a second type different from a first type of the first instruction, and responsive to the second instruction, a second operation different from the first operation is performed. On the other hand, Chen discloses an instruction is a second instruction of a second type different from a first type of a first instruction, and responsive to the second instruction, a second operation different from a first operation is performed ([0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to combine the teaching of Chen with the invention of Zhang in order to increase processing capability via supporting different types of instructions. (Alternatively, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Zhang with the invention of Chen in order to improve efficiency; see paragraph [0014] of Zhang.) Consider claim 3, the overall combination entails the processor as recited in claim 2 (see above), wherein the circuitry is configured to fetch, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operand, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 1-8, in Block 512, each dot product data unit of the dot product data units 322-1 to 322-n writes the current cumulative result of the dot product data unit to the general register 310 to serve as the convolution operation result when it is determined in the Block 510 that the convolution operation is over. Otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). Consider claim 4, the overall combination entails the processor as recited in claim 3 (see above), wherein the first operation is a fused multiply add (FMA) operation (Zhang, Figure 4, which shows a fused multiply add operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) and the second operation is a dot product operation (Zhang, [0049], line 3, dot product operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). Consider claim 5, the overall combination entails the processor as recited in claim 3 (see above), wherein the processor further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits (Zhang, [0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400), wherein responsive to the first instruction ([0053], line 12, single instruction), each of the plurality of arithmetic logic circuits is configured to: receive the first plurality of values of the first matrix and the second matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and perform a matrix (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) multiplication operation of a fused multiply add (FMA) operation (Zhang, Figure 4, which shows a fused multiply add operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) using at least a first multiplier circuit and a second multiplier circuit (Zhang, Figure 4, which shows the multiplier circuits; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations), each having a size less than a size of the first plurality of values of the first matrix and the second matrix (Zhang, [0059], lines 5-6, the dot product data unit 322-1 may perform a dot product operation of A1*B1+A2*B2+A3*B3). Consider claim 6, the overall combination entails the processor as recited in claim 5 (see above), wherein responsive to the second instruction (Zhang, [0053], line 12, single instruction; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations), each of the plurality of arithmetic logic circuits is configured to: receive the second plurality of values of the third matrix and the fourth matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and perform a matrix (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) multiplication operation of a dot product operation (Zhang, Figure 4, which shows a dot product operation, which entails multiplication; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) using the first multiplier circuit and the second multiplier circuit (Zhang, Figure 4, which shows the multiplier circuits). Consider claim 7, the overall combination entails the processor as recited in claim 3 (see above), wherein the processor further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits (Zhang, [0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400), wherein the circuitry is further configured to: store the first plurality of values or the second plurality of values in a plurality of storage elements for reuse by the plurality of arithmetic logic circuits (Zhang, [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and use, in a pipeline stage of the plurality of arithmetic logic circuits, each of the first plurality of values of the first matrix or each of the second plurality of values of the third matrix a number of times equal to a number of columns of the second matrix or the fourth matrix, respectively (Chen, [0025], line 9, for example, matrix multiplication operations; Zhang, [0030], lines 2-5, the data and the operation result include, but are not limited to, numerical values, such as a pixel matrix, a convolution operation result matrix, a convolution kernel, and so on). Consider claim 9, Zhang discloses the method as recited in claim 8 (see above), responsive to receiving, by the processing circuit, an instruction: fetching, by the processing circuit from the register file only once, a second plurality of values ([0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations); and performing, using the arithmetic logic circuit, an operation by reusing one or more of the second plurality of values for at least two iterations of computations used to provide the operation ([0048], lines 1-6, in Block 504, the data reuse unit 321 determines the multiple data subsets from the data set, so as to respectively input the multiple data subsets into the multiple dot product data units 322-1 to 322-n. The two data subsets inputted into the two adjacent dot product data units include a portion of the same data; [0049], lines 1-4, in Block 506, each dot product data unit of the multiple dot product data units 322-1 to 322-n performs the dot product operation on the inputted data subset, so as to generate the dot product operation result; [0050], lines 1-5, in Block 508, each dot product data unit of the multiple dot product data units 322-1 to 322-n generates the current cumulative result of the dot product data unit based on the previous cumulative result of the dot product data unit and the dot product operation result; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 6-8, otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation). However, Zhang does not disclose that the aforementioned instruction is a second instruction of a second type different from a first type of the first instruction, and responsive to receiving the second instruction, a second operation different from the first operation is performed. On the other hand, Chen discloses an instruction is a second instruction of a second type different from a first type of a first instruction, and responsive to receiving the second instruction, a second operation different from a first operation is performed ([0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to combine the teaching of Chen with the invention of Zhang in order to increase processing capability via supporting different types of instructions. (Alternatively, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Zhang with the invention of Chen in order to improve efficiency; see paragraph [0014] of Zhang.) Consider claim 10, the overall combination entails the method as recited in claim 9 (see above), further comprising fetching, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operand, by the processing circuit, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 1-8, in Block 512, each dot product data unit of the dot product data units 322-1 to 322-n writes the current cumulative result of the dot product data unit to the general register 310 to serve as the convolution operation result when it is determined in the Block 510 that the convolution operation is over. Otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). Consider claim 11, the overall combination entails the method as recited in claim 10 (see above), wherein the first operation is a fused multiply add (FMA) operation (Zhang, Figure 4, which shows a fused multiply add operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) and the second operation is a dot product operation (Zhang, [0049], line 3, dot product operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). Consider claim 12, the overall combination entails the method as recited in claim 10 (see above), wherein the processing circuit further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits (Zhang, [0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400), wherein responsive to the first instruction ([0053], line 12, single instruction), wherein responsive to the first instruction, the method further comprises, by each of the plurality of arithmetic logic circuits: receiving the first plurality of values of the first matrix and the second matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and performing the first operation as a matrix (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) multiplication operation of a fused multiply add (FMA) operation (Zhang, Figure 4, which shows a fused multiply add operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) using at least a first multiplier circuit and a second multiplier circuit (Zhang, Figure 4, which shows the multiplier circuits; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations), each having a size less than a size of the first plurality of values of the first matrix and the second matrix (Zhang, [0059], lines 5-6, the dot product data unit 322-1 may perform a dot product operation of A1*B1+A2*B2+A3*B3). Consider claim 13, the overall combination entails the method as recited in claim 12, wherein responsive to the second instruction (Zhang, [0053], line 12, single instruction; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations), the method further comprises, by each of the plurality of arithmetic logic circuits: receiving the second plurality of values of the third matrix and the fourth matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and performing the second operation as a matrix (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) multiplication operation of a dot product operation (Zhang, Figure 4, which shows a dot product operation, which entails multiplication; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) using the first multiplier circuit and the second multiplier circuit (Zhang, Figure 4, which shows the multiplier circuits). Consider claim 14, the overall combination entails the method as recited in claim 10 (see above), wherein the processing circuit further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits (Zhang, [0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400), wherein the method further comprises: storing, by the processing circuit, the first plurality of values or the second plurality of values in a plurality of storage elements for reuse by the plurality of arithmetic logic circuits (Zhang, [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and using, in a pipeline stage of the plurality of arithmetic logic circuits, each of the first plurality of values of the first matrix or each of the second plurality of values of the third matrix a number of times equal to a number of columns of the second matrix or the fourth matrix, respectively (Chen, [0025], line 9, for example, matrix multiplication operations; Zhang, [0030], lines 2-5, the data and the operation result include, but are not limited to, numerical values, such as a pixel matrix, a convolution operation result matrix, a convolution kernel, and so on). Consider claim 15, Zhang discloses a computing system comprising: a memory comprising circuitry configured to store at least a first plurality of values ([0013], line 9, memory); and a second processor ([0029], line 6, processor) comprising: a register file comprising circuitry configured to store operand data for vector operations ([0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel); an execution pipeline comprising an arithmetic logic circuit ([0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400); and circuitry configured to: responsive to a first instruction ([0053], line 12, single instruction): fetch, from the register file only once, the first plurality of values ([0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations); and perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation ([0048], lines 1-6, in Block 504, the data reuse unit 321 determines the multiple data subsets from the data set, so as to respectively input the multiple data subsets into the multiple dot product data units 322-1 to 322-n. The two data subsets inputted into the two adjacent dot product data units include a portion of the same data; [0049], lines 1-4, in Block 506, each dot product data unit of the multiple dot product data units 322-1 to 322-n performs the dot product operation on the inputted data subset, so as to generate the dot product operation result; [0050], lines 1-5, in Block 508, each dot product data unit of the multiple dot product data units 322-1 to 322-n generates the current cumulative result of the dot product data unit based on the previous cumulative result of the dot product data unit and the dot product operation result; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 6-8, otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation). However, Zhang does not disclose the memory is configured to store a plurality of instructions, and a first processor comprising circuitry configured to launch one or more kernels comprising the plurality of instructions, the first instruction being of the one or more kernels. On the other hand, Chen discloses memory configured to store a plurality of instructions and data ([0049], lines 14-19, in various implementations, the program instructions are stored on any of a variety of non-transitory computer readable storage mediums. The storage medium is accessible by a computing system during use to provide the program instructions to the computing system for program execution; FIG. 1, memory device(s) 140), and, alongside a second processor ([0019], lines 3-8, in one implementation, processor 105N is a data parallel processor with a highly parallel architecture. Data parallel processors include graphics processing units (GPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and so forth), and a first processor comprising circuitry configured to launch one or more kernels comprising the plurality of instructions, a first instruction being of the one or more kernels ([0025], lines 1-7, CPU (not shown) of computing system 200 launches kernels to be performed on GPU 205. Command processor 235 receives kernels from the host CPU and uses dispatch unit 250 to issue corresponding wavefronts to compute units 255A-N. In one implementation, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit; [0019], lines 1-2, processor 105A is a general purpose processor, such as a central processing unit (CPU)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to combine the teaching of Chen with the invention of Zhang in order to increase processing capability via supporting multiple processors. (Alternatively, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Zhang with the invention of Chen in order to improve efficiency; see paragraph [0014] of Zhang.) Consider claim 16, the overall combination entails the computing system as recited in claim 15 (see above), wherein responsive to an instruction, the circuitry is further configured to: fetch, from the register file only once, a second plurality of values (Zhang, [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations); and perform, using the arithmetic logic circuit, an operation by reusing one or more of the second plurality of values for at least two iterations of computations used to provide the operation ([0048], lines 1-6, in Block 504, the data reuse unit 321 determines the multiple data subsets from the data set, so as to respectively input the multiple data subsets into the multiple dot product data units 322-1 to 322-n. The two data subsets inputted into the two adjacent dot product data units include a portion of the same data; [0049], lines 1-4, in Block 506, each dot product data unit of the multiple dot product data units 322-1 to 322-n performs the dot product operation on the inputted data subset, so as to generate the dot product operation result; [0050], lines 1-5, in Block 508, each dot product data unit of the multiple dot product data units 322-1 to 322-n generates the current cumulative result of the dot product data unit based on the previous cumulative result of the dot product data unit and the dot product operation result; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 6-8, otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation). However, the overall combination does not entail that the aforementioned instruction is a second instruction of a second type different from a first type of the first instruction, and responsive to the second instruction, a second operation different from the first operation is performed. On the other hand, Chen further discloses an instruction is a second instruction of a second type different from a first type of a first instruction, and responsive to the second instruction, a second operation different from a first operation is performed ([0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, to combine the further teaching of Chen with the previously explained combination of Zhang and Chen in order to increase processing capability via supporting different types of instructions. (Alternatively, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Zhang with the invention of Chen in order to improve efficiency; see paragraph [0014] of Zhang.) Consider claim 17, the overall combination entails the computing system as recited in claim 16 (see above), wherein the circuitry is configured to fetch, from the register file only once, until each element of a resulting matrix is updated by the first operation or the second operand, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; [0052], lines 1-2, in Block 510, it is determined whether the convolution operation has ended; [0053], lines 1-8, in Block 512, each dot product data unit of the dot product data units 322-1 to 322-n writes the current cumulative result of the dot product data unit to the general register 310 to serve as the convolution operation result when it is determined in the Block 510 that the convolution operation is over. Otherwise, return to the Block 504 to continue performing the cycle on the data to be calculated in the convolution operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). Consider claim 18, the overall combination entails the computing system as recited in claim 17 (see above), wherein the first operation is a fused multiply add (FMA) operation (Zhang, Figure 4, which shows a fused multiply add operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) and the second operation is a dot product operation (Zhang, [0049], line 3, dot product operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations). Consider claim 19, the overall combination entails the computing system as recited in claim 17 (see above), wherein the computing system further comprises a plurality of execution pipelines having a plurality of arithmetic logic circuits (Zhang, [0032], line 8, multiple dot product data units 322-1 to 322-n; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400), wherein responsive to the first instruction ([0053], line 12, single instruction), each of the plurality of arithmetic logic circuits is configured to: receive the first plurality of values of the first matrix and the second matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and perform a matrix (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) multiplication operation of a fused multiply add (FMA) operation (Zhang, Figure 4, which shows a fused multiply add operation; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) using at least a first multiplier circuit and a second multiplier circuit (Zhang, Figure 4, which shows the multiplier circuits; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations), each having a size less than a size of the first plurality of values of the first matrix and the second matrix (Zhang, [0059], lines 5-6, the dot product data unit 322-1 may perform a dot product operation of A1*B1+A2*B2+A3*B3). Consider claim 20, the overall combination entails the computing system as recited in claim 19 (see above), wherein responsive to the second instruction (Zhang, [0053], line 12, single instruction; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations), each of the plurality of arithmetic logic circuits is configured to: receive the second plurality of values of the third matrix and the fourth matrix used as source operands (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations); and perform a matrix (Zhang, [0029], line 4, general register 310; [0047], lines 1-3, in Block 502, the data reuse unit 321 reads from the general register 310 and temporarily stores the data set used for the multiple convolution operations; [0056], lines 1-3, the data reuse unit 321 may read from the general register 310 and temporarily store the above-mentioned 5×5 pixel matrix and the 3×3 convolution kernel; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) multiplication operation of a dot product operation (Zhang, Figure 4, which shows a dot product operation, which entails multiplication; [0036], lines 1-2, FIG. 4 is a schematic block diagram of a dot product data unit 400; Chen, [0025], lines 5-13, each compute unit 255A-N includes an adaptive multi-instruction type matrix operations unit. For example, the adaptive multi-instruction type matrix operations unit performs matrix multiplication operations, dot product operations, and fused multiply add (FMA) operations. Additionally, in various implementations, the adaptive, multi-instruction type matrix operations unit performs other types of matrix, arithmetic, or bitwise operations) using the first multiplier circuit and the second multiplier circuit (Zhang, Figure 4, which shows the multiplier circuits). Response to Arguments Applicant on page 9 argues: ‘In the present Final Office Action, it is suggested that the term "register file" in the claims can be either the scalar register file 332 or the vector register file 330 of Figure 3 of the Specification. Claims 1, 8 and 15 have been amended to include "a register file comprising circuitry configured to store operand data for vector operations." Support for the claim amendments may be found in at least paragraphs 12 and 39. Therefore, it is believed that claims 1-3, 8-10 and 15-17 have been amended in a manner to overcome the rejections.’ However, Examiner notes that while the claim recites a register file “comprising circuitry configured to store operand data for vector operations”, this further language does not appear to limit the recited register file to be a vector register file in particular. For example, paragraphs [0033]-[0035] and [0039] appear to convey that a scalar register file can store scalar operand data for a vector operation. Applicant on page 9 argues: ‘In the present Final Office Action, it is suggested that the Specification describes the ALU 350 includes an execution pipeline, whereas claim 1 recites "a plurality of execution pipelines, having a plurality of arithmetic logic circuits." However, as shown in Figure 3 of the Specification, each of the lanes 320A-320C includes a corresponding vector ALU such as vector ALU 350 and paragraph 36 describes "the hardware of lane 320C is an instantiation of the hardware of lane 320A. The components in lanes 320A-320C operate in lockstep."’ However, the reproduced portion of paragraph 36 does not appear to disclose the terminology “execution pipeline”, and Figure 3 likewise does not use the terminology “execution pipeline” to refer to elements 320A-320C. Examiner notes that the terminology “execution pipeline” is used elsewhere in the specification (e.g., paragraph [0047]) to refer to an element that is different from, for example, a “lane” as conveyed in the reproduced portion of paragraph 36, and therefore Examiner submits that the terminology “execution pipeline” in the claims would not be understood to be equivalent to the “lane” in the reproduced portion of paragraph 36. Examiner notes that the meaning of every term used in any of the claims should be apparent from the descriptive portion of the specification with clear disclosure as to its import, and further notes that the terms and phrases used in claims must find clear support or antecedent basis in the description so that the meaning of the terms in the claims may be ascertainable by reference to the description. Applicant on page 9 argues: ‘Paragraph 37 of the Specification further describes "the parallel computational lanes 320A-320C operate in lockstep. In various implementations, the data flow within each of the lanes 320A-320C is pipelined. Pipeline registers are used for storing intermediate results. Within a given row across lanes 320A-320C, vector arithmetic logic unit (ALU) 350 includes the same circuitry and functionality, and operates on the same instruction, but different data associated with a different thread."’ However, the reproduced portion of paragraph 37 does not appear to disclose the terminology “execution pipeline”. Examiner notes that the terminology “execution pipeline” is used elsewhere in the specification (e.g., paragraph [0047]) to refer to an element that is different from, for example, a “lane” as conveyed in the reproduced portion of paragraph 37, and therefore Examiner submits that the terminology “execution pipeline” in the claims would not be understood to be equivalent to the “lane” in the reproduced portion of paragraph 37. Examiner notes that the meaning of every term used in any of the claims should be apparent from the descriptive portion of the specification with clear disclosure as to its import, and further notes that the terms and phrases used in claims must find clear support or antecedent basis in the description so that the meaning of the terms in the claims may be ascertainable by reference to the description. Applicant on page 9 argues: ‘Further, paragraph 40 of the Specification describes "Selection circuit 342 also includes multiplexers and possible crossbar circuitry to route source operands to particular inputs of operations being performed by vector ALU 350. In various implementations, lane 320A is organized as a multi-stage pipeline. Intermediate sequential elements, such as staging flip-flop circuits, registers, or latches, are not shown for ease of illustration."’ However, the reproduced portion of paragraph 40 does not appear to disclose the terminology “execution pipeline”. Examiner notes that the terminology “execution pipeline” is used elsewhere in the specification (e.g., paragraph [0047]) to refer to an element that is different from, for example, a “lane” as conveyed in the reproduced portion of paragraph 40, and therefore Examiner submits that the terminology “execution pipeline” in the claims would not be understood to be equivalent to the “lane” in the reproduced portion of paragraph 40. Examiner notes that the meaning of every term used in any of the claims should be apparent from the descriptive portion of the specification with clear disclosure as to its import, and further notes that the terms and phrases used in claims must find clear support or antecedent basis in the description so that the meaning of the terms in the claims may be ascertainable by reference to the description. Applicant on page 9 argues: “Accordingly, with each of the parallel computational lanes 320A-320C operating as a multi-stage pipeline, Applicant submits the Specification supports the features "a plurality of execution pipelines, having a plurality of arithmetic logic circuits." However, Examiner notes that the terminology “execution pipeline” is used elsewhere in the specification (e.g., paragraph [0047]) to refer to an element that is different from, for example, a parallel computational lane operating as a multi-stage pipeline, and therefore Examiner submits that the terminology “execution pipeline” in the claims would not be understood to be equivalent to a parallel computational lane operating as a multi-stage pipeline. Examiner notes that the meaning of every term used in any of the claims should be apparent from the descriptive portion of the specification with clear disclosure as to its import, and further notes that the terms and phrases used in claims must find clear support or antecedent basis in the description so that the meaning of the terms in the claims may be ascertainable by reference to the description. Examiner submits that the reproduced portions addressed above do not convey that the arithmetic logic circuit is in the “execution pipeline”, when “execution pipeline” is interpreted in a manner consistent with how that particular terminology is used in the specification. Applicant across pages 9-10 argues: ‘Regardless, claim 1 has been amended to include its original features such as "a plurality of execution pipelines, each comprising a corresponding arithmetic logic circuit."’ While the aforementioned language may be present in the original claims, Examiner submits that the language is indefinite for the reasons provided in the associated indefinite rejection. Applicant on page 10 argues: ‘In the present Final Office Action, it is suggested that claim 1 has indefinite features such as "a plurality of execution pipelines, having a plurality of arithmetic logic circuits, each comprising circuitry configured to execute at least two different types of instructions." It is suggested that it is indefinite as to whether the plurality of execution pipelines or the plurality of arithmetic logic circuits execute at least two different types of instructions. The Specification describes the plurality of arithmetic logic circuits executes at least two different types of instructions (e.g., the fused multiply add (FMA) operation and the dot product (inner product) operation). However, claim 1 has been amended to include "a plurality of execution pipelines, each comprising a corresponding arithmetic logic circuit." Claims 8 and 15 have been amended in a similar manner. Therefore, it is believed that claims 1, 8 and 15 have been amended in a manner to overcome the rejections.’ In view of the aforementioned amendments, the associated previously presented indefinite rejections are withdrawn. Applicant across pages 10-11 argues: “In the present Final Office Action, it is suggested that the Specification does not support the claim 1 features "perform, using the plurality of arithmetic logic circuits, a first operation by reusing the first plurality of values for at least two iterations of computations used to perform the first operation." On page 6 of the present Final Office Action, it is suggested that "Figure 3 does not show a plurality of arithmetic logic circuits operating on a reused first plurality of values fetched from a register file." However, Figure 1 and paragraph 18 of the Specification describe "Data items 142 are data items that have been retrieved from the vector register file. In some implementations, data items 142 are retrieved from the vector register file and stored in temporary storage elements such as flip-flop circuits, registers, random-access memory, or other. Therefore, the data of the rows and columns of matrices 110 and 120 are retrieved only once from the vector register file . . .the data of the rows and columns of matrices 110 and 120 are retrieved only once from the vector register file. The data are shared and reused between iterations of computations performed to provide the matrix multiplication operation." Paragraph 19 of the Specification describes "Data items 140 are data items that have been retrieved from the vector register file and sent to vector ALU 150 that performs one of a fused multiply add (FMA) operation and a dot product (inner product) operation." Therefore, the vector ALU 150 reuses multiple data items of matrices 110 and 120 that have been retrieved only once from the vector register file. Paragraph 41 of the Specification describes "vector ALU 350 has the same functionality as vector ALU 150 (of FIG. 1) and each of the multiple vector ALUs 250 (of FIG. 2)." Therefore, vector ALU 350 of Figure 3 of the Specification performs reuse of the data, since the "data are shared and reused between iterations of computations performed to provide the matrix multiplication operation." Paragraph 28 of the Specification provides further support and an example of the number of clock cycles required to generate the data items of the output matrix C (matrix 130). Accordingly, the Specification supports the claim 1 features "perform, using the plurality of arithmetic logic circuits, a first operation by reusing the first plurality of values for at least two iterations of computations used to perform the first operation." Regardless, claim 1 has been amended to include the original features such as "perform, using the corresponding arithmetic logic circuit of each of the plurality of execution pipelines, a first operation." Examiner notes that claim 1 does not appear to have been amended in the manner set forth above. Nevertheless, in view of the manner in which claim 1 has been amended, the associated previously presented indefinite rejection is withdrawn. Applicant on page 11 argues: ‘On page 7 of the present Final Office Action, it is suggested that the claim 5 features "the values of the first matrix" and "the values of the second matrix" are indefinite. Claim 5 has been amended to include "receive the first plurality of values of the first matrix and the second matrix used as source operands" and "each having a size less than a size of the first plurality of values of the first matrix and the second matrix." Claim 6, which is dependent on claim 5, has been amended to include "receive the second plurality of values of the third matrix and the fourth matrix used as source operands." Claim 5 is dependent on claim 3, which has been amended to include features previously recited in claim 7 such as "wherein the circuitry is configured to fetch, from the vector register file only once, until each element of a resulting matrix is updated by the first operation or the second operation, the first plurality of values as data of a first matrix and a second matrix used as source operands or the second plurality of values as data of a third matrix and a fourth matrix used as source operands." The two matrixes used as source operands are shown as matrices 110 and 120 of Figure 1 of the Specification. Claims 10, 12-13, 17 and 19-20 have been amended in a similar manner. Therefore, it is believed that claims 5-6, 12-13 and 19-20 have been amended in a manner to overcome the rejections.’ In view of the aforementioned amendments, the associated previously pending indefinite rejections are withdrawn. Applicant on page 12 argues: ‘In the present Final Office Action, regarding claim 9, it is suggested that "it is indefinite as to whether the method is being recited to comprise the fetching and performing steps." Claim 9 has been amended in a similar manner as claim 8.’ However, the amendment to claim 9 does not appear to address the rationale for indefiniteness. Examiner notes that claim 9 does not appear to recite “further comprising” language or “wherein” language or other language commonly used to further narrow a claim. Applicant on page 12 argues: ‘It is also suggested, regarding claim 10, "it is indefinite as to whether claim 10 (in the context of claim 8) entails fetching the first plurality of values once or twice, in view of the "further" language in claim 10." As described earlier, claim 10 has been amended to include features previously recited in claim 7 such as fetching from the vector register file only once "until each element of a resulting matrix is updated by the first operation or the second operation."’ However, the previously presented rejection appears to remain applicable, in view of, for example, the “further” language followed by a further step. Examiner submits that it is unclear by Applicant’s remarks whether the method is intended to entail one step of fetching or two steps of fetching. Applicant on page 12 argues: ‘Regarding claims 12-13, it is suggested that these claims are indefinite due to whether the first operation and the matrix multiplication operation of a fused multiply add (FMA) operation are separate operations in claim 12 and whether the second operation and the matrix multiplication operation of a dot product operation are separate operations. Claims 12 and 13 have been amended to include "performing the first operation as a matrix multiplication operation of a fused multiply add (FMA) operation" and "performing the second operation as a matrix multiplication operation of a dot product operation", respectively. It is believed the claims have been amended in a manner to further clarify the features and overcome the rejections.’ In view of the aforementioned amendments, the previously presented indefinite rejections are withdrawn. Applicant on page 13 argues: “Each of the two adjacent dot product data units fail to reuse data from the previous iteration of the convolution operation. Therefore, Zhang fails to disclose or suggest at least the above highlighted features.” Examiner notes that the claims do not appear to recite “previous” language. In addition, see the response to argument in the following numbered paragraph. Applicant on page 14 argues: ‘The above disclosures describe the dot product data unit 322-1 uses the three pairs of data (A1, B1), (A2, B2), and (A3, B3) in the first cycle of the first convolution operation. The dot product data unit 322-1 uses the three pairs of data (A6, B4), (A7, B5), and (A8, B6) in the second cycle of the first convolution operation. Therefore, the dot product data unit 322-1 does not reuse data for at least two iterations of computations used to perform the first convolution operation. Additionally, the dot product data unit 322-1 uses the three pairs of data (Al1 B7), (A12, B8),and (A13, B9) in the third cycle of the first convolution operation. (See Zhang, para. 61). Again, the dot product data unit 322-1 does not reuse data for at least two iterations of computations used to perform the first convolution operation. Similarly, the dot product data unit 322-2 of Zhang does not reuse data for at least two iterations of computations used to perform the second convolution operation. Therefore, Zhang fails to disclose or suggest the claim 1 features "the circuitry is configured to . . . perform, using the arithmetic logic circuit, a first operation by reusing one or more of the first plurality of values for at least two iterations of computations used to perform the first operation."’ However, Examiner first notes that Zhang discloses a 3x3 convolution operation on a 5x5 pixel matrix (see paragraph [0055]), and further notes that Zhang explicitly steps through some, but not all, of the 3x3 convolution operation on a 5x5 pixel matrix. (Examiner submits that such a convolution operation was well-known to a person of ordinary skill in the art and that a person of ordinary skill in the art would recognize a 3x3 convolution operation on a 5x5 pixel matrix to, for example, entail analogous calculations involving A5, A10, A15, and A16-A25. Examiner notes that Applicant likewise notes the implicit calculation involving A5 in page 17 of the remarks.) Examiner submits that performing the entirety of the 3x3 convolution operation on a 5x5 pixel matrix, in the context of the disclosure of Zhang, implicitly entails, for example, dot product data unit 322-1 performing additional calculations akin to paragraphs [0059]-[0061], except with the sliding window shifted downward (such that a further 3x3 window of the 5x5 pixel matrix is A6, A7, A8, A11, A12, A13, A16, A17, A18, and a still further 3x3 window of the 5x5 pixel matrix is A11, A12, A13, A16, A17, A18, A21, A22, and A23). As such, Examiner submits that A6, A7, A8, A11, A12, A13, A16, A17, and A18 are reused. Examiner further notes that A3, for example, is used in an iteration explicitly described in paragraph [0059], an iteration explicitly described in paragraph [0063], and an implicit iteration that would involve A3, A4, and A5. Therefore, A3 is also reused. While each of the aforementioned iterations may occur in the same cycle, Examiner submits that the claims do not preclude such a scenario. Applicant on page 14 argues: “The arithmetic unit 320 of Zhang can utilize data reusing, but only for the portion of the same data among two data subsets inputted into two adjacent dot product data units.” However, as explained above, such data reusing may still teach the recited claim language under the broadest reasonable interpretation. In addition, Examiner submits that Zhang implicitly teaches a same dot product data unit reusing data, as explained above. Applicant on page 15 argues: “As independent claims 8 and 15 include features similar to claim 1, claims 8 and 15 are patentably distinguished from the cited art for similar reasons. As each of the dependent claims includes the features of the independent claims on which they depend, each of the dependent claims is patentably distinct for at least the above reasons.” Examiner’s responses to arguments above with respect to claim 1 are likewise applicable to the arguments directed towards claims 8 and 15. Applicant on page 15 argues: “As described earlier, Zhang discloses reusing a portion of the two data subsets inputted into two adjacent dot product data units and this portion includes the same data among the two data subsets, rather than the entirety of the two data subsets. Therefore, Zhang fails to disclose or suggest the further features of claim 2.” However, the claims do not appear to mandate functionality that corresponds to “the entirety of the two data subsets” rather than what Zhang discloses. Applicant across pages 16-17 argues: ‘The above disclosure merely describes the dot product data unit 322-1 performs the dot product operation with a subset of a matrix corresponding to the 5x5 pixel matrix shown in Table 1 and a subset of a matrix corresponding to the 3 x3 convolution kernel shown in Table 2. However, Zhang nowhere discloses the multiplier circuits of the dot product data unit 322-1 has "a size less than a size of the values of the first matrix and less than a size of the values of the second matrix." Regarding the use of the dot product data unit 322-1, Zhang discloses the dot product data unit 400 includes three multipliers 410-1 to 410-3 and "Each of the multipliers 410 is configured to multiply one pair of inputted data, so as to generate a product, and input the product to a corresponding adder 420. A pair of data includes, for example, a pixel and a weight in the convolution kernel." (See Zhang, para. 37). The three multipliers 410-1 to 410-3 of Zhang operate on the entirety of the pair of inputted data, not "a size less than a size of the first plurality of values of the first matrix and the second matrix." Therefore, for at least these further reasons, claim 5 is patentably distinguishable from the cited art. Claims 12 and 19 include similar features and are similarly patentably distinguishable.’ However, Examiner submits that the size of one pixel or one weight is less than an overall size of the multiple pixels in the pixel matrix and the overall size of the multiple weights in the convolution kernel. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEITH E VICARY whose telephone number is (571)270-1314. The examiner can normally be reached Monday to Friday, 9:00 AM to 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571)270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KEITH E VICARY/ Primary Examiner, Art Unit 2183
Read full office action

Prosecution Timeline

Show 1 earlier event
Jul 02, 2025
Non-Final Rejection mailed — §102, §103, §112
Sep 30, 2025
Response Filed
Oct 20, 2025
Final Rejection mailed — §102, §103, §112
Feb 18, 2026
Applicant Interview (Telephonic)
Feb 18, 2026
Examiner Interview Summary
Mar 31, 2026
Request for Continued Examination
Apr 03, 2026
Response after Non-Final Action
Jul 29, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699569
SELF-PROVISIONING AND FLEXIBLE HARDWARE ACCELERATOR ARCHITECTURE
2y 2m to grant Granted Aug 04, 2026
Patent 12688397
PARTITIONABLE DIGITAL HARDWARE SYSTEM FOR IMPLEMENTING RECURRENT NEURAL NETWORKS
5y 2m to grant Granted Jul 21, 2026
Patent 12663994
Supporting Multiple Vector Lengths with Configurable Vector Register File
2y 11m to grant Granted Jun 23, 2026
Patent 12663990
Apparatus and Method for Remote Atomic Floating Point Operations
2y 2m to grant Granted Jun 23, 2026
Patent 12657031
Coprocessor Prefetcher
1y 10m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
58%
Grant Probability
99%
With Interview (+41.0%)
3y 11m (~1y 6m remaining)
Median Time to Grant
High
PTA Risk
Based on 692 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month