DETAILED ACTION
Claims 1-20 are pending.
The office acknowledges the following papers:
Claims and remarks filed on 3/24/2026.
Allowable Subject Matter
Claims 15-20 are allowed.
Maintained Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-2 and 11 are rejected under 35 U.S.C. 102(a)(1 & 2) as being anticipated by Diard (U.S. 9,536,275).
As per claim 1:
Claim 1 essentially recites the same limitations of claim 11. Claim 1 additionally recites the following limitations:
a system interconnect (Diard: Figure 1 element 105, column 3 lines 12-18); and
a general-purpose parallel processing engine coupled with the system interconnect (Diard: Figures 1-2 elements 105 and 202(0)-202(1), column 3 lines 12-18 and column 4 lines 30-43).
As per claim 2:
Diard disclosed the accelerator device as in claim 1, wherein the first pipeline is to operate concurrently with the second pipeline (Diard: Figure 3 elements 302(0)-302(P-1), column 6 lines 3-12, column 6 lines 21-34, and column 7 lines 1-14)(The processing engines are configured to execute in parallel with each other. Separate instructions can be dispatched to different processing engines for parallel execution. Additionally, a single instruction can be dispatched to the processing engines for concurrent processing.).
As per claim 11:
Diard disclosed a method comprising:
providing a matrix engine having multiple pipelines including a first pipeline and a second pipeline (Diard: Figure 3 elements 302(0)-302(P-1), column 6 lines 3-12)(The processing core (i.e. matrix engine) includes processing engines that include pipelined functional units. A first processing engine reads upon the first pipeline and a second processing engine reads upon the second pipeline.);
sharing a common input between the first pipeline and the second pipeline of the multiple pipelines (Diard: Figure 3 element 312, column 7 lines 1-14)(The instruction unit provides a common input to the processing engines.);
associating a first output memory with the first pipeline and a second output memory with the second pipeline (Diard: Figures 2-3 elements 204 and 306, column 4 lines 37-43 and column 6 lines 35-44)(The processing engines can write to the shared memory (i.e. output memory). The PP memory can receive result data to write back to system memory. Both memories are associated with the processing engines.); and
configuring common output circuitry to output from one of the first output memory and the second output (Diard: Figures 2-3 element 214, column 4 lines 21-29)(The memory interface is a common output circuitry that outputs data from the shared memory to the PP memory, as well as outputs data from the PP memory to the shared memory.).
Maintained Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 3-10 and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Diard (U.S. 9,536,275), in view of Beu et al. (U.S. 2021/0389948).
As per claim 3:
Diard disclosed the accelerator device as in claim 2, wherein the matrix engine includes the common input that is shared between the first pipeline and the second pipeline (Diard: Figure 3 element 312, column 7 lines 1-14)(The instruction unit provides a common input to the processing engines.).
Diard failed to teach wherein the matrix engine includes a first set of multiple inputs associated with an accumulator value, a second set of multiple inputs associated with row data for a matrix multiply operation.
However, Beu combined with Diard disclosed wherein the matrix engine includes a first set of multiple inputs associated with an accumulator value, a second set of multiple inputs associated with row data for a matrix multiply operation (Beu: Figures 10-11, paragraphs 100-107)(Diard: Figure 3 elements 302(0)-302(P-1), column 6 lines 3-12)(Beu disclosed a systolic array for processing matrix operations. The combination implements the systolic array within each processing engine of Diard for performing matrix operations. The adders receive accumulator inputs and the multipliers receive row inputs (e.g. kernel/input data).).
It would have been obvious to one of ordinary skill in the art to combine the teachings of Diard and Beu. Both references were directed to the problems of performing operations on array data in parallel in a data processor. Thus, one of ordinary skill in the art at the time of the effective filing date would have been motivated to incorporate the Beu teachings of multiple inputs associated with an accumulator value and multiple inputs associated with row data at least to input the data for array data processing in parallel which would speed processing and increase throughput.
As per claim 4:
Diard and Beu disclosed the accelerator device as in claim 3, wherein the common input is associated with column data for the matrix multiply operation (Beu: Figures 10-11, paragraphs 100-107)(Diard: Figure 3 element 312, column 7 lines 1-14)(The instruction unit provides a common input to the processing engines. The combination allows for the processing engines to perform matrix multiplication operations. This allows for the matrix multiply operations sent to the processing engines on the common input to be associated with the multipliers of the systolic array receiving column inputs (e.g. input/kernel data). The combination allows for SIMD operations to perform the matrix multiplication operations over a plurality of threads within the processing engines.).
As per claim 5:
Diard disclosed the accelerator device as in claim 1.
Diard failed to teach wherein the general-purpose parallel processing engine is to fetch an instruction to perform operations associated with a matrix instruction and decode the instruction into multiple sub-instructions.
However, Beu combined with Diard disclosed wherein the general-purpose parallel processing engine is to fetch an instruction to perform operations associated with a matrix instruction and decode the instruction into multiple sub-instructions (Beu: Figures 1-2 and 9-11 elements 30 and 46, paragraphs 49-50, 62-63, 70, and 99-107)(Diard: Figure 3 elements 302(0)-302(P-1), column 6 lines 3-12 and column 7 lines 1-14)(Beu disclosed a decoder to decode processor instructions, including matrix instructions, for execution on matrix circuitry. Diard disclosed an instruction unit that outputs different/single instructions to processing engines for execution. The combination allows for Diard to output different/single matrix instructions, as in Beu, to the processing engines for execution. The combination allows for decoding the different/single matrix instruction into multiple sub-instructions that are each processed by a thread within a processing engine.).
It would have been obvious to one of ordinary skill in the art to combine the teachings of Diard and Beu. Both references were directed to the problems of performing operations on array data in parallel in a data processor. Thus, one of ordinary skill in the art at the time of the effective filing date would have been motivated to incorporate the Beu teachings of multiple inputs associated with an accumulator value and multiple inputs associated with row data at least to input the data for array data processing in parallel which would speed processing and increase throughput.
As per claim 6:
Diard and Beu disclosed the accelerator device as in claim 5, wherein to decode the instruction into the multiple sub-instructions includes to generate a first set of sub-instructions for execution by the first pipeline and to generate a second set of sub-instructions for execution by the second pipeline (Beu: Figures 1-2 and 9-11 elements 30 and 46, paragraphs 49-50, 62-63, 70, and 99-107)(Diard: Figure 3 elements 302(0)-302(P-1), column 6 lines 3-12 and column 7 lines 1-14)(The combination allows for Diard to output different/single matrix instructions, as in Beu, to the processing engines for execution. The combination allows for decoding the different/single matrix instruction into multiple sub-instructions that are each processed by a thread within a processing engine. First and second processing engines receives different sets of threads to execute, which are different portions of outer product operations.).
As per claim 7:
Diard and Beu disclosed the accelerator device as in claim 6, wherein the first set of sub-instructions and the second set of sub-instructions reference a common set of registers to store data associated with the common input shared between the first pipeline and the second pipeline (Beu: Figures 1-2 element 34, paragraphs 42 and 62)(Diard: Figure 3 elements 302(0)-302(P-1), column 6 lines 3-12 and column 7 lines 1-14)(Beu disclosed matrix instructions referencing register data as inputs. The combination allows for Diard to process matrix multiplication instructions using common register data.).
As per claim 8:
Diard and Beu disclosed the accelerator device as in claim 7, wherein the matrix engine includes circuitry to:
read, by the first pipeline, a first set of matrix elements specified by operands of the first set of sub-instructions and a second set of matrix elements specified by operands of the second set of sub- instructions; store, by the first pipeline, a first sub-set of matrix elements to memory within the matrix engine, the memory accessible by the second pipeline; relay, by the first pipeline, a second sub-set of matrix elements to the second pipeline; perform, by the first pipeline, processing operations specified by the first set of sub-instructions; and perform, by the second pipeline, processing operations specified by the second set of sub-instructions (Beu: Paragraphs 58-62)(The mapping of the architectural register to physical registers for storage and retrieval and performing sub-operations of instruction(s) using subdivided data corresponds to this limitation.).
As per claim 9:
Diard and Beu disclosed the accelerator device as in claim 7, wherein to generate the first set of sub-instructions for execution by the first pipeline includes to determine a first set of registers that store operands for the first set of sub-instructions and determine a second set of registers that store operands for the second set of sub-instructions (Beu: Paragraphs 0058-0062)| the mapping of the architectural register to physical registers for storage and retrieval and performing sub- operations of instruction(s) using subdivided data corresponds to this limitation].
As per claim 10:
Diard and Beu disclosed the accelerator device as in claim 9, wherein the common output circuitry is to write output from one of the first output memory and the second output memory to a register file associated with the matrix engine (Beu: Figure 1, elements 34 and 48-50)(Diard: Figures 2-3 element 214, column 4 lines 21-29)(Note the register file (34) is depicted as coupled to the memory via load store unit (LDST) which is coupled to the memory. Therefore one of ordinary skill in the art would have been motivated to store/retrieve data to/from memory and register file to ensure that the data needed for processing was available with the fastest access such as in registers and ensured that when the register file was full storing data not currently being used in memory so it was not lost. The memory interface is a common output circuitry that outputs data from the shared memory to the PP memory, as well as outputs data from the PP memory to the shared memory. This allows for indirect writing to the register file via the shared memory.).
As per claim 12:
Diard disclosed the method of claim 11.
Diard failed to teach reading operand data for an instruction to be executed by the matrix engine from a register file associated with the matrix engine, the operand data including matrix elements associated with the instruction; executing a first portion of the instruction via the first pipeline and concurrently executing a second portion of the instruction via the second pipeline; and writing output of the first portion of the instruction and the second portion of the instruction to the register file.
However, Beu combined with Diard disclosed reading operand data for an instruction to be executed by the matrix engine from a register file (34) associated with the matrix engine, the operand data including matrix elements associated with the instruction (Beu: Paragraphs 62 and 67);
executing a first portion of the instruction via the first pipeline (Beu: Figures 9-10, paragraph 35) and concurrently executing a second portion of the instruction via the second pipeline (Beu: Paragraphs 35 and 64); and
writing output of the first portion of the instruction and the second portion of the instruction to the register file (Diard: Column 6 line 61-column 7 line 6).
It would have been obvious to one of ordinary skill in the art to combine the teachings of Diard and Beu. Both references were directed to the problems of performing operations on array data in parallel in a data processor. Thus, one of ordinary skill in the art at the time of the effective filing date would have been motivated to incorporate the Beu teachings of multiple inputs associated with an accumulator value and multiple inputs associated with row data at least to input the data for array data processing in parallel which would speed processing and increase throughput.
As per claim 13:
Diard and Beu disclosed the method of claim 12, wherein reading operand data for the instruction includes:
reading, by the first pipeline, matrix elements associated with the first portion of the instruction and the second portion of the instruction (Beu: Paragraph 70);
storing, by the first pipeline, a first sub-set of matrix elements to memory within the matrix engine, the memory accessible by the second pipeline; and relaying, by the first pipeline, a second sub-set of matrix elements to the second pipeline (Diard: Column 6 line 61-column 7 line 28)[note in the embodiment where data is needed by multiple threads one of ordinary skill would have been motivated to relay the results from one pipeline to a second pipeline at least to reduce the time for the secondary pipe to access the required data and therefore reduce processing time and increase throughput].
As per claim 14:
Diard and Beu disclosed the method of claim 12, wherein writing output to the register file includes:
writing output from the first pipeline to a first output buffer associated with the first pipeline writing output from the second pipeline to an output buffer associated with the second pipeline (Diard: Column 6 line 61-column 7 line 28)(Beu and Diard did not expressly detail writing output of first and second pipeline to first and second output buffers and writing output from the first output buffer and the second output buffer to the register file via a common output. However one of ordinary skill would have been motivated to store the output of the first and second pipelines in separate buffers and then storing the data to the register file at least to relax the timing necessary for storing data to the register file and retrieving data from the register file. This would have simplified the control of storing and retrieving data and a reduced system cost.).
Response to Arguments
The arguments presented by Applicant in the response, received on 3/24/2026 are not considered persuasive.
Applicant argues regarding claims 1 and 11:
“Claim 1 recites "a general-purpose parallel processing engine coupled with the system interconnect, the general-purpose parallel processing engine comprising a matrix engine." Diard, by contrast, discloses a general-purpose parallel processing subsystem for graphics rendering and general-purpose computing tasks. Diard, Column 3, Lines 45-55. Diard's processing subsystem is not a matrix accelerator and does not include a matrix engine as required by the claims. Diard makes no mention of matrix operations, matrix multiply operations, or any specialized matrix processing hardware. The absence of any matrix-specific functionality in Diard is a fundamental deficiency that precludes anticipation of the claimed invention.”
This argument is not found to be persuasive for the following reason. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., matrix operations, matrix multiply operations, and matrix processing hardware) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
Applicant argues regarding claims 1 and 11:
“Furthermore, claims 1 and 11 require: (1) a first output memory specifically associated with a first pipeline; (2) a second output memory specifically associated with a second pipeline; and (3) common output circuitry that is configurable to selectively output from one of these dedicated output memories. Diard does not disclose this claimed structure. Diard's memory interface 214 is a general-purpose memory access interface that enables cores to read from or write to various external memory devices. Diard, Column 4, Lines 21-28. Diard expressly states that memory interface 214 "can be of generally conventional design." Diard, Column 4, Line 29.
Diard does not disclose dedicated output memories associated with respective pipelines as required by the claims. Diard's local register file 304 stores "local input data, intermediate results, and the like," not dedicated output memories associated with pipelines. Diard, Column 6, Lines 28-34. Furthermore, Diard's shared memory 306 "is shared among all of the processing engines 302 in core 208" rather than being a dedicated output memory associated with a specific pipeline. Diard, Column 6, Lines 42-44.
The Examiner's assertion that memory interface 214 provides a common output for memory 204 and memory 306 conflates general memory access functionality with the specific common output circuitry recited by the claims. Claims 1 and 11 require common output circuitry that is "configurable to output from one of the first output memory and the second output memory," where each output memory is specifically associated with a respective pipeline. Diard's memory interface 214 merely provides general memory access to external memory devices and does not selectively output from dedicated pipeline-associated output memories as claimed.”
This argument is not found to be persuasive for the following reason. Both of the PP Memory and the shared memory are associated with the processing engines that include functional unit pipelines. Both memories can be read as output memories as they both are configured to store result data after result data in registers is written to memory. The claims include no limitations that state the first and second output memories are dedicated memories to the corresponding first and second pipelines. The memory interface is capable of receiving output data from both the shared memory and the PP memory for storage. Thus, reading upon the claimed limitations.
Applicant argues regarding claims 3-4:
“Additionally, the combination of Diard and Beu fails to disclose or suggest the specific dual pipeline parallel systolic array architecture recited by the claims. Claim 3 recites "the matrix engine includes a first set of multiple inputs associated with an accumulator value, a second set of multiple inputs associated with row data for a matrix multiply operation, and the common input that is shared between the first pipeline and the second pipeline." Claim 4 further recites "the common input is associated with column data for the matrix multiply operation." The dual pipeline parallel systolic array includes an input for a Srcl operand that is a common input that is shared between the two systolic array pipelines, where the Srcl input inputs column data that is used by the two systolic array pipelines to perform matrix multiply operations in which two sets of matrix row data are multiplied by a single set of column data, with separate inputs provided for SrcO (accumulator value) inputs.
The Examiner has alleged that Beu's FIGS. 10-11 teach the claimed matrix engine inputs. Applicant respectfully disagrees. Beu's disclosure relates to mixed-element-size instructions where operands have different data element sizes. See Beu, paragraph [0070]. Beu's systolic array architecture in FIGS. 10-11 does not disclose a dual pipeline parallel systolic array having a first pipeline and a second pipeline that share a common input for column data while having separate accumulator inputs and separate row data inputs. Beu's approach repurposes existing registers to accommodate different element sizes within a single processing path, rather than providing two parallel pipelines that share a common Srcl input while having separate SrcO and Src2 inputs as recited by claims 3 and 4.
The claimed architecture provides increased throughput for matrix operations without incurring the power and area costs associated with two separate and fully independent systolic arrays." In contrast, Beu's systolic array "consists of an array of MAC processing elements (PEs), which communicate operands and results using local register-to-register communication only." Beu, paragraph [0104]. Beu's approach involves reusing "part of the matrix multiplication hardware... to do twice as many multiplies of narrower width" within a single processing path. Beu, paragraph [0088]. This is fundamentally different from providing two parallel pipelines with a shared common input as claimed.”
This argument is not found to be persuasive for the following reason. The systolic array of Beu clearly shows data element inputs to the array that are part of the source matrices and source accumulation data. The combination allows for the systolic arrays to be within the processing engines and the source inputs for rows/columns to be sent to the processing engines. Additionally, the combination allows for the matrix instruction to be send to the processing engines on the common input. Thus, reading upon the claimed limitations.
Applicant argues regarding claim 8:
“Claim 8 recites that the matrix engine includes circuitry to "read, by the first pipeline, a first set of matrix elements specified by operands of the first set of sub-instructions and a second set of matrix elements specified by operands of the second set of sub-instructions" and to "store, by the first pipeline, a first sub-set of matrix elements to memory within the matrix engine, the memory accessible by the second pipeline" and to "relay, by the first pipeline, a second sub-set of matrix elements to the second pipeline." In other words, the first pipeline can perform register reads for both pipelines and can store a first matrix element to memory within the dual pipeline parallel systolic array and relay a second matrix element to the second pipeline. Neither Diard nor Beu discloses this specific architecture where the first pipeline reads matrix elements for both pipelines and relays data to the second pipeline. The Examiner's citation to Beu paragraphs 0058-0062 relates to general register mapping and does not teach the specific data flow architecture where the first pipeline reads operand data for both pipelines and relays a subset of matrix elements to the second pipeline.”
This argument is not found to be persuasive for the following reason. The combination allows for the systolic arrays to be implemented within each of the processing engines of Diard. Upon executing matrix operations, matrix data is loaded from the local register file to the systolic arrays. Writebacks of matrix results are sent back to the local register file. Intermediate results are reloaded back into the systolic array from the local register file. Loading intermediate results to any other processing engine allows for relaying the results. Thus, reading upon the claimed limitation.
Applicant argues regarding claim 14:
“Claim 14 recites "writing output from the first pipeline to a first output buffer associated with the first pipeline; writing output from the second pipeline to a second output buffer associated with the second pipeline; and writing output from the first output buffer and the second output buffer to the register file via a common output." When a consolidated output is used, output operations from the left output buffer are immediately output, with output operations from the right output buffer being delayed and output after the output from the left systolic array pipeline. The Examiner acknowledges that "Beu and Diard did not expressly detail writing output of first and second pipeline to first and second output buffers and writing output from the first output buffer and the second output buffer to the register file via a common output." The Examiner's assertion that one of ordinary skill would have been motivated to implement this feature "to relax the timing necessary for storing data to the register file" is unsupported by any teaching in the cited references. The Federal Circuit has stated that "rejections on obviousness cannot be sustained with mere conclusory statements; instead, there must be some articulated reasoning with some rational underpinning to support the legal conclusion of obviousness." In re Kahn, 441 F.3d 977, 988, 78 USPQ2d 1329, 1336 (Fed. Cir. 2006). The Examiner's conclusory statement regarding motivation constitutes impermissible hindsight reconstruction.”
This argument is not found to be persuasive for the following reason. The use of buffers is well-known to one of ordinary skill in the art to temporarily store data prior to it being written to memory. For example, pipeline registers are very common data buffers that store data just immediately prior to writeback in register files. Thus, the rejection is maintained.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The following is text cited from 37 CFR 1.111(c): In amending in reply to a rejection of claims in an application or patent under reexamination, the applicant or patent owner must clearly point out the patentable novelty which he or she thinks the claims present in view of the state of the art disclosed by the references cited or the objections made. The applicant or patent owner must also show how the amendments avoid such references or objections.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB A. PETRANEK whose telephone number is (571)272-5988. The examiner can normally be reached on M-F 8:00-4:30.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached on (571) 270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB PETRANEK/Primary Examiner, Art Unit 2183