DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Action is non-final and is in response to the claims filed 02/27/2026. Claims 1-20 are currently pending, of which claims 1-20 are currently rejected.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 02/27/2026 has been entered.
Response to Arguments
Applicant’s arguments filed on 02/27/2026 have been fully considered.
35 U.S.C. 103: Applicant’s arguments regarding the 35 U.S.C. 103 rejection have been fully considered.
Applicant argues in page 9 that Hosseinzadeh alone does not teach separate input and weight scratchpads that feed data to a buffer. Applicant specifically argues “Hosseinzadeh's memory unit 120 stores both input data and weights but does not teach separate input and weight scratchpads that feed into a buffer from which the analog accelerator performs matrix multiplication, let alone the now-recited language. See Hosseinzadeh, [00438].
Applicant further argues in page 9 that Raghavan alone does not teach input data and weight scratchpads as separate scratchpads that feed data to a buffer. Applicant specifically argues “Raghavan's scratchpad memory stores computed results, not input data and weights in separate scratchpads that feed into a buffer, let alone the now-recited language. Raghavan, 7:50-59.”
Applicant further argues in page 9 that neither Gao nor Linu teach the newly added limitations. Applicant specifically argues “Gao's SRAM buffers are local to each tile/engine and do not correspond to the claimed architecture of separate weight and input scratchpads feeding into a buffer, let alone the now-recited language. Gao, Section 2.2, page 810. Linu does not address memory architecture for neural network processing.”
Applicant also argues in pages 9 and 10 no combination of Hosseinzadeh, Raghavan, Linu and Gao would have satisfied the language now recited in claims 1, 11, and 14.
Examiner respectfully disagrees. Hosseinzadeh teaches the DAC unit 130 acting as a buffer to feed input and weight data to the optical processor 140. See Hosseinzadeh: ¶00413. Raghavan also teaches using two registers to respectively store two matrices for matrix multiplication. See Raghavan: Abstract. Combination would cause for Hosseinzadeh to store input and weight data in separate registers (scratchpads) before transmitting data to DAC unit 130 (buffer). See new grounds of rejection necessitated by amendments for motivations to combine Hosseinzadeh in view of Linu, Raghavan, and Gao to teach amended claims 1 and 11, and Hosseinzadeh in view of Raghavan, and Gao to teach amended claim 14.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-10 and 14-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 does not positively recite structural elements for the apparatus claim. It is unclear how the listed elements “an input scratchpad”, “a weight scratchpad”, and “a buffer” are positively recited as listed elements reciting the structural elements of the apparatus claim 1. Appropriate correction is required. See MPEP 2115.
Claims 2-10 recite the same deficiency as claim 1 by reason of dependence.
Claim 14 does not positively recite structural elements for the apparatus claim. It is unclear how the listed elements “an input scratchpad”, “a weight scratchpad”, and “a buffer” are positively recited as listed elements reciting the structural elements of the apparatus claim 14. Appropriate correction is required. See MPEP 2115.
Claims 15-20 recite the same deficiency as claim 14 by reason of dependence.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-13 are rejected under 35 U.S.C. 103 as being unpatentable over Hosseinzadeh, Arash (WIPO International Publication No.: WO 2020149953 A1), hereinafter “Hosseinzadeh”, in view of Linu M Jiji in NPL: Architecture implementation of optimized DSP accelerator with modified booth recoder in FPGA (https://ieeexplore.ieee.org/document/7830138) hereinafter “Linu”, further in view of Raghavan et al. (U.S. Patent No.: US 10521225 B2), hereinafter “Raghavan”, further in view of Mingyu Gao in NPL: Tangram: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators (https://dl.acm.org/doi/pdf/10.1145/3297858.3304014), hereinafter “Gao”.
With regards to Claim 1, Hosseinzadeh teaches:
Hosseinzadeh teaches:
A processing system, comprising an analog accelerator arranged to perform matrix multiplication (Fig. 1A, e.g., Optical processor 140 (analog accelerator) performs matrix multiplication in the OMM unit);
… a controller coupled to … the analog accelerator … (Fig. 1A, e.g., Controller 110 is coupled to optical processor 140 (analog accelerator)), wherein the controller is configured to process a first layer of a multi-layer neural network (¶00453, e.g., Process 200 processes data through the first hidden layer (first layer); Fig. 2A), wherein processing the first layer comprises:
storing an input data set … and a weight matrix … (¶00401, e.g., Controller 110 receives ANN computation request, and stores the input dataset and the neural network weights (weight matrix) in the memory unit 120);
storing a first portion of the weight matrix … in a buffer and storing at least a first portion of the input data set … in the buffer (¶00413, e.g., DAC unit 130 (buffer) is configured to buffer analog signals. DAC unit 130 receives multiple weight control signals and digital input vectors);
controlling the analog accelerator to perform a first matrix multiplication to produce a first output data block (¶00399, e.g., Controller 110 controls ANN computations; Fig. 2A, e.g., Flowchart of ANN computation. Step 240 obtains optical output vector (first output data block) from OMM unit, which performs first matrix multiplication) … ;
storing a second portion of the weight matrix … in the buffer and storing at least a second portion of the input data set … in the buffer (¶00413, e.g., DAC unit 130 (buffer) is configured to buffer analog signals. DAC unit 130 receives multiple weight control signals and digital input vectors);
controlling the analog accelerator to perform a second matrix multiplication to produce a second output data block (Fig. 2B, e.g., Multiple vectors are processed through the OMM unit, hence a second matrix multiplication is performed producing a second output data block; ¶00456, e.g., Multiple vectors are processed, hence multiple matrix multiplications) … ; and
… controlling the [controller] to process the first output data block using a non-linear operation (¶00447, e.g., nonlinear transformations are performed on the first digital output by the controller 110).
Hosseinzadeh does not teach:
a digital processor; and
a controller coupled to both the analog accelerator and the digital processor wherein the controller is configured to:
storing an input data set in an input scratchpad and a weight matrix in a weight scratchpad;
storing a first portion of the weight matrix stored in the weight scratchpad in a buffer and storing at least a first portion of the input data set stored in the input scratchpad in the buffer;
controlling the analog accelerator to perform a first matrix multiplication to produce a first output data block using the first portion of the weight matrix stored in the buffer and at least the first portion of the input data set stored in the buffer;
storing a second portion of the weight matrix stored in the weight scratchpad in the buffer and storing at least a second portion of the input data set stored in the input scratchpad in the buffer;
controlling the analog accelerator to perform a second matrix multiplication to produce a second output data block using the second portion of the weight matrix stored in the buffer and at least the second portion of the input data set stored in the buffer; and
subsequent to completion of the first matrix multiplication and prior to completion of the second matrix multiplication, controlling the digital processor to process the first output data block using a non-linear operation.
However, in the same field of endeavor, Linu teaches how a Digital Signal Processor (DSP) may be used as an accelerator to execute certain functions instead of executing them in the CPU. Linu explains "Digital Signal Processors are special type of microprocessors with its architecture optimized for signal processing applications. Hardware acceleration has been proved as an extremely promising implementation strategy for the digital signal processing (DSP) and multimedia application domain. An accelerator module can be attached to processor core for enhancing performance. It enhances the performance or functionality by executing certain function in the accelerator instead of executing in the processor core." (Abstract)
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the DSP as an accelerator as taught by Linu with the system 100 as taught by Hosseinzadeh. One would have been motivated to combine these references because both references disclose accelerating signal processing, and Linu enhances the model of Hosseinzadeh by adding a DSP to the system 100 to execute digital operations, (nonlinear operations) instead of having them executed in the controller 110, and the motivation to add the DSP is because "It enhances the performance or functionality by executing certain function in the accelerator instead of executing in the processor core" (Linu: Abstract). The combination of Hosseinzadeh as modified by Linu teaches the controller coupled to a DSP (digital processor), and would control the digital processor to process the first output data block using a non-linear operation.
Hosseinzadeh in view of Linu do not teach:
storing an input data set in an input scratchpad and a weight matrix in a weight scratchpad;
storing a first portion of the weight matrix stored in the weight scratchpad in a buffer and storing at least a first portion of the input data set stored in the input scratchpad in the buffer;
controlling the analog accelerator to perform a first matrix multiplication to produce a first output data block using the first portion of the weight matrix stored in the buffer and at least the first portion of the input data set stored in the buffer;
storing a second portion of the weight matrix stored in the weight scratchpad in the buffer and storing at least a second portion of the input data set stored in the input scratchpad in the buffer;
controlling the analog accelerator to perform a second matrix multiplication to produce a second output data block using the second portion of the weight matrix stored in the buffer and at least the second portion of the input data set stored in the buffer; and
subsequent to completion of the first matrix multiplication and prior to completion of the second matrix multiplication, controlling the digital processor to process the first output data block using a non-linear operation.
However, Raghavan teaches:
perform a first matrix multiplication … using a first portion of the weight matrix and at least a first portion of the input data set (Fig. 2, e.g., shows portion of elements 200 and 208 (first row) of first matrix (first portion of the weight matrix) and portions of elements 202, 204, 210, and 212 (first and second column) of second matrix (first portion of input data set); Column 4 Lines 36 – 61, e.g., portions are used for matrix multiplication)
perform a second matrix multiplication … using a second portion of the weight matrix and at least a second portion of the input data set (Fig. 2, e.g., Elements in the second row of the weight matrix (second portion of weight matrix), and Elements on the third and fourth column (second portion of input data set));
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the VMA instruction as taught by Raghavan with the Optical Matrix Multiplication (OMM) unit as taught by Hosseinzadeh in view of Linu. One would have been motivated to combine these references because both references disclose accelerating signal processing for neural network matrix multiplication, and Raghavan enhances the model of Hosseinzadeh in view of Linu because “the vma instruction enables product matrix elements to be computed in fewer iterations than the typical approach.” (Raghavan: Column 5 Lines 2-4)
Raghavan further teaches:
storing an input data set in an input scratchpad and a weight matrix in a weight scratchpad (Abstract, e.g., A first register (input scratchpad) stores element values of the first matrix, and a second register (weight scratchpad) stores element values of the second matrix);
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the first and second registers to store first and second matrices as taught by Raghavan with the memory unit 120 storing input dataset and weight matrix as taught by Hosseinzadeh in view of Linu. One would have been motivated to combine these references because both references disclose accelerating signal processing for neural network matrix multiplication, and Raghavan enhances the model of Hosseinzadeh in view of Linu because “register file 308 may be accessed using multiple ports that enable concurrent read and/or write operations” (Raghavan: Column 6 Lines 23-25). Combination would cause for input and weight data stored in respective registers to be fed to DAC unit 130 (buffer) as taught by Hosseinzadeh, teaching the limitations:
storing a first portion of the weight matrix stored in the weight scratchpad in a buffer and storing at least a first portion of the input data set stored in the input scratchpad in the buffer;
… storing a second portion of the weight matrix stored in the weight scratchpad in the buffer and storing at least a second portion of the input data set stored in the input scratchpad in the buffer;
Hosseinzadeh in view of Linu in view of Raghavan do not teach:
subsequent to completion of the first matrix multiplication and prior to completion of the second matrix multiplication, control the digital processor to process the first output data block using a non-linear operation.
However, in the same field of endeavor, Gao teaches how the next layer in a neural network can start processing as soon as input data is available instead of waiting for the entire layer to finish processing. Gao explains “ ALLO modifies the intermediate fmap data access patterns, in order to allow the next layer to start processing as soon as a subset of the fmaps within a single data sample are ready. For example, in Figure 8, L-1 computes each ofmap sequentially. If the next layer (L-2) also sequentially accepts these data as its ifmaps, it can start processing after waiting for a single fmap rather than all fmaps" (Page 813, Column 1, Alternate layer loop ordering (ALLO) dataflow, First paragraph). In addition, Hosseinzadeh teaches how steps of steps of the ANN computation, including the nonlinear transformation, may be run in any order. See Hosseinzadeh ¶00436.
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the Alternate layer loop ordering (ALLO) dataflow method as taught by Gao with the Optical Matrix Multiplication (OMM) unit as taught by Hosseinzadeh in view of Linu in view of Raghavan. One would have been motivated to combine these references because both references disclose processing neural networks, and Gao enhances the model of Hosseinzadeh in view of Linu in view of Raghavan by allowing for multiple layers to be computed concurrently for a faster run time. The combination of Hosseinzadeh in view of Linu in view of Raghavan as modified by Gao would cause for the nonlinear activation function of the first matrix multiplication to be completed in order to be an input of the next layer to start processing without waiting for the current layer to finish generating ofmaps, e.g., for a second matrix multiplication to be completed. See Hosseinzadeh ¶0004 and ¶00447. The combination of Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach Claim 1 in its entirety.
With regards to Claim 2, Hosseinzadeh in view of Linu, Raghavan, and Gao teach the processing system of claim 1. Hosseinzadeh in view of Linu, Raghavan, and Gao do not teach:
wherein the analog accelerator has a first accelerator core and a second accelerator core, wherein the controller is further configured to: control the analog accelerator to perform the first matrix multiplication using the first accelerator core;
and control the analog accelerator to perform the second matrix multiplication using the second accelerator core.
However, Raghavan teaches:
wherein the [computing device] has a first accelerator core and a second accelerator core (Fig. 3, e.g., Processor Core 302 (first accelerator core) and 304 (second accelerator core)), … perform the first matrix multiplication using the first accelerator core (Column 8 Lines 16 – 24, e.g., ALU 306 Performs first matrix multiplication; Fig. 3, e.g., ALU 306 is included in the Processor Core 302);
… perform the second matrix multiplication using the second accelerator core (Column 7 Lines 54 – 57, e.g., Processor core 304 contains its own ALU for computing, for example, a second matrix multiplication; Column 7 Lines 66-67 and Column 8 Lines 1-6, e.g., Each core performs multiplications of different rows in the matrix).
Therefore, combining the computing device with the two processor cores as taught by Raghavan with the OMM unit as taught by Hosseinzadeh would yield the controller controlling the Optical processor (analog accelerator), including the OMM unit. This combination would cover the limitations in Claim 2 in its entirety. It would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the computing device with the two processor cores as taught by Raghavan with the OMM unit as taught by Hosseinzadeh. One would have been motivated to combine these references because both references disclose matrix multiplication using cores, and Raghavan enhances the model of Hosseinzadeh by allowing for matrix multiplications to be performed in parallel in different cores for faster processing.
With regards to Claim 3, Hosseinzadeh in view of Raghavan teach:
The processing system of claim 1, wherein the digital processor is configured to complete the processing of the first output data block prior to completion of the second matrix multiplication (Hosseinzadeh: ¶00570, e.g., Second matrix multiplication is done using digital output vector after nonlinear transformation (non-linear operation); ¶00447, e.g., Controller may include one or more modules (digital processors) to perform nonlinear transformation on first digital output vector).
With regards to Claim 4, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 1, wherein the first portion of the weight matrix comprises at least a first row of the weight matrix (Raghavan: Fig. 2, e.g., First Matrix uses first portion elements 200 and 208 for matrix multiplication, which are in a first row),
and wherein the controller is configured to control the analog accelerator to perform the first matrix multiplication to produce the first output data block (Hosseinzadeh: ¶00399, e.g., Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., step 240 obtains a first plurality of digitalized optical outputs corresponding to the optical output vector (first output data block) from the OMM unit (analog accelerator)) using the first row of the weight matrix (Raghavan: Column 5 Lines 5 – 25, e.g., Pseudocode A shows iteration through weight matrix rows to implement VMA (vector multiply-add) instruction (iterating through the second row as well); Implementing VMA instruction from Raghavan into the OMM unit from Hosseinzadeh would cause for first matrix multiplication to be done using the first row of the weight matrix).
With regards to Claim 5, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 4, wherein the second portion of the weight matrix comprises at least a second row of the weight matrix (Raghavan: Fig. 2, e.g., Elements in the second row of the weight matrix (second portion of weight matrix)), and wherein the controller is configured to control the analog accelerator to perform the second matrix multiplication to produce the second output data block (Hosseinzadeh: Fig. 2B, e.g., Multiple vectors are processed through the OMM unit, hence a second matrix multiplication is performed.) using the second row of the weight matrix (Column 5 Lines 5 – 25, e.g., Pseudocode A shows iteration through weight matrix rows to implement VMA (vector multiply-add) instruction (iterating through the second row as well); Implementing VMA instruction from Raghavan into the OMM unit from Hosseinzadeh would cause for second matrix multiplication to be done using the second row of the weight matrix).
With regards to Claim 6, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 1, wherein the controller is configured to control the analog accelerator to perform the first matrix multiplication (Hosseinzadeh: ¶00399, e.g., Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., step 240 obtains a first plurality of digitalized optical outputs corresponding to the optical output vector from the OMM unit (analog accelerator) by performing a first matrix multiplication) using tile parallelism (Raghavan: Column 4 Lines 23 – 29, e.g., Each element group corresponds to a tile; Column 4 Lines 42 – 52, e.g., Tile parallelism is used to perform matrix multiplication).
With regards to Claim 7, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 1, wherein the controller is configured to control the analog accelerator to perform the first matrix multiplication (Hosseinzadeh: ¶00399, e.g., Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., step 240 obtains a first plurality of digitalized optical outputs corresponding to the optical output vector from the OMM unit (analog accelerator) by performing a first matrix multiplication) using data parallelism (Raghavan: Column 4 Lines 42 – 52, e.g., Data parallelism is used to perform matrix multiplication).
With regards to Claim 8, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 1, wherein the analog accelerator comprises a photonic accelerator (Hosseinzadeh: ¶00432, e.g., Laser unit 142, modulator array 144, OMM unit 150, and Photodetectors may be a photonic integrated circuit, making it a photonic accelerator), and wherein the analog accelerator is configured to perform the first matrix multiplication at least partially in an optical domain (Hosseinzadeh: ¶00432, e.g., Optical processor 140 contains optical components; ¶00445, e.g., OMM unit 150 performs matrix multiplication using optical input vector generated from modulator array 144).
With regards to Claim 9, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 8, wherein the photonic accelerator comprises an optical multiplier configured to perform scalar multiplication in the optical domain (Hosseinzadeh: Fig. 4B, e.g., scalar multiplication is performed; ¶00485).
With regards to Claim 10, Hosseinzadeh in view of Linu in view of Raghavan in view of Gao teach:
The processing system of claim 8, wherein the photonic accelerator comprises an optical adder configured to perform scalar addition in the optical domain (Hosseinzadeh: Fig. 4B, e.g., scalar addition is performed; ¶00485).
With regards to Claims 11-13, they are directed to a method practiced by the apparatus of claims 1-3, respectively. They are rejected for the same reasons therein.
Claims 14 – 19 are rejected under 35 U.S.C. 103 as being unpatentable over Hosseinzadeh, in view of Raghavan, further in view of Gao.
With regards to Claim 14, Hosseinzadeh teaches:
A processing system configured to process a multi-layer neural network comprising first and second layers, the processing system comprising (¶0004):
an … analog accelerator … (Fig. 1A, e.g., Optical processor 140 (analog accelerator)) ; and
a controller coupled to the … analog accelerator and configured to:
store an input data set … (¶00401, e.g., Controller 110 receives ANN computation request, which includes neural network weights and input dataset; Fig.1, e.g., Controller 110 is coupled to Optical Processor 140 (analog accelerator)) and a first weight matrix associated with the first layer of the multi-layer neural network … (¶00401, e.g., Controller 110 receives ANN computation request, which includes neural network weights (first weight matrix); Fig 1A; ¶00453, e.g., Process 200 processes data, including neural network weights (first weight matrix), through the first hidden layer (first layer); Fig. 2A e.g., Process involves using first plurality of neural network weights (first weight matrix));
process the first layer of the multi-layer neural network, wherein processing the first layer comprises: (¶00453, e.g., Process 200 processes data through the first hidden layer (first layer); Fig. 2A)
storing a first portion of the first weight matrix … in a buffer and storing at least a first portion of the input data set … in the buffer (¶00413, e.g., DAC unit 130 (buffer) is configured to buffer analog signals. DAC unit 130 receives multiple weight control signals and digital input vectors);
… produce a first output data block (Fig. 2A, e.g., Step 240 outputs digitalized optical outputs) … ;
storing a second weight matrix associated with the second layer of the multi- layer neural network (¶00454, e.g., Second hidden layer (second layer) is processed by the OMM unit 150; ¶00413, e.g., DAC unit 130 (buffer) is configured to buffer analog signals. DAC unit 130 receives multiple weight control signals and digital input vectors);
storing a second portion of the first weight matrix … in the buffer and storing at least a second portion of the input data set … in the buffer (¶00413, e.g., DAC unit 130 (buffer) is configured to buffer analog signals. DAC unit 130 receives multiple weight control signals and digital input vectors);
… produce a second output data block (Fig. 2B, e.g., Multiple vectors are processed through the OMM unit, hence a second matrix multiplication is performed) … ; and
process the second layer of the multi-layer neural network (¶00454, e.g., Second hidden layer (second layer) is processed by the OMM unit 150), wherein processing the second layer comprises:
… perform a third matrix multiplication using the second weight matrix and the first output data block (¶00570, e.g., Second matrix multiplication (third matrix multiplication) is performed using digital output vector (first output data block), which is the result of the first matrix multiplication; ¶00453, e.g., Controller reconfigures OMM unit 150 to perform matrix multiplication corresponding to second plurality of neural network weights (second weight matrix) associated with the second hidden layer).
Hosseinzadeh does not teach:
a multi-core analog accelerator comprising first and second accelerator cores; and
a controller coupled to the multi-core analog accelerator and configured to:
store an input data set in an input scratchpad and a first weight matrix associated with the first layer of the multi-layer neural network in a weight scratchpad;
process the first layer of the multi-layer neural network, wherein processing the first layer comprises:
storing a first portion of the first weight matrix stored in the weight scratchpad in a buffer and storing at least a first portion of the input data set stored in the input scratchpad in the buffer;
controlling the first accelerator core to perform a first matrix multiplication to produce a first output data block using the first portion of the first weight matrix stored in the buffer and at least the first portion of the input data set stored in the buffer;
storing a second weight matrix associated with the second layer of the multi- layer neural network;
storing a second portion of the first weight matrix stored in the weight scratchpad in the buffer and storing at least a second portion of the input data set stored in the input scratchpad in the buffer;
controlling the first accelerator core to perform a second matrix multiplication to produce a second output data block using the second portion of the first weight matrix stored in the buffer and at least the second portion of the input data set stored in the buffer; and
process the second layer of the multi-layer neural network, wherein processing the second layer comprises:
subsequent to completion of the first matrix multiplication and prior to completion of the second matrix multiplication, controlling the second accelerator core to perform a third matrix multiplication using the second weight matrix and the first output data block.
However, in the same field of endeavor, Raghavan teaches how a computing device can include multiple cores for performing matrix multiplication. Raghavan explains “computing device 300 comprises processor cores 302-304. Processor core 304 may have its own ALU, register file, and/or scratchpad memory. In some embodiments, processor core 304 is organized similarly to processor core 302. Although FIG. 3 depicts two cores, embodiments disclosed herein may be implemented using any number of cores” (Column 7, Lines 53 - 59). See figure 3. In addition, Hosseinzadeh explains how submatrix multiplications can be performed in different devices (cores). See ¶00194.
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the processor cores as taught by Raghavan with the Optical Matrix Multiplication (OMM) unit as taught by Hosseinzadeh. One would have been motivated to combine these references because both references disclose processing matrix multiplications using cores, and Raghavan enhances the model of Hosseinzadeh because “each core can separately execute a machine code instruction within the same clock cycle(s) in which another core executes an instruction, thereby achieving parallelization” (Raghavan: Column 7 Lines 61 - 64), allowing for faster processing. The combination of Hosseinzadeh as modified by Raghavan teaches a multi-core analog accelerator, and would cause for the processing cores to receive input and weight data from DAC unit 130 (buffer) as taught by Hosseinzadeh.
Raghavan also teaches:
controlling the first accelerator core to perform a first matrix multiplication (Column 8 Lines 16 – 24, e.g., ALU 306 Performs matrix multiplications; Fig. 3, e.g., ALU 306 is included in the Processor Core 302; Column 4 Lines 48 – 61, e.g., First matrix multiplication) … using the first portion of the first weight matrix … and at least the first portion of the input data set … (Fig. 2, e.g., shows portion of elements 200 and 208 of first matrix (weight matrix), and portions of elements 202 and 204 of second matrix (input data set); Column 4 Lines 48 – 52, e.g., Portions are used for matrix multiplication);
and controlling the [second] accelerator core to perform a second matrix multiplication (Fig. 3, e.g., Processor cores include ALU, which is used for matrix multiplication; Column 7 Lines 66-67 and Column 8 Lines 1-6, e.g., Each core performs multiplications of different rows in the first matrix 100) … using a second portion of the first weight matrix … (Fig. 2, e.g., Elements in the second row (second portion of weight matrix)) … and at least a second portion of the input data set … (Fig. 2, e.g., Elements from the third and fourth column (second portion of the input data set) would be used for computing the rest of the elements; Column 1 Lines 47 – 65, e.g., Pseudocode shows iterations over each column of the second matrix);
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine VMA instruction as taught by Raghavan with the Optical Matrix Multiplication (OMM) unit as taught by Hosseinzadeh. One would have been motivated to combine these references because both references disclose processing matrix multiplications using cores, and Raghavan enhances the model of Hosseinzadeh because “the vma instruction enables product matrix elements to be computed in fewer iterations than the typical approach.” (Raghavan: Column 5 Lines 2-4).
Raghavan further teaches:
store an input data set in an input scratchpad and a first weight matrix … in a weight scratchpad (Abstract, e.g., A first register (input scratchpad) stores element values of the first matrix, and a second register (weight scratchpad) stores element values of the second matrix);
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the first and second registers to store first and second matrices as taught by Raghavan with the memory unit 120 storing input dataset and weight matrix as taught by Hosseinzadeh in view of Linu. One would have been motivated to combine these references because both references disclose accelerating signal processing for neural network matrix multiplication, and Raghavan enhances the model of Hosseinzadeh in view of Linu because “register file 308 may be accessed using multiple ports that enable concurrent read and/or write operations” (Raghavan: Column 6 Lines 23-25).
Hosseinzadeh in view of Raghavan does not teach:
controlling the first accelerator core to perform a second matrix multiplication to produce a second output data block using the second portion of the first weight matrix stored in the buffer and at least the second portion of the input data set stored in the buffer; and
… subsequent to completion of the first matrix multiplication and prior to completion of the second matrix multiplication, controlling the second accelerator core to perform a third matrix multiplication using the second weight matrix and the first output data block.
However, in the same field of endeavor, Gao teaches how multiple engines can be implemented to process multiple small layers instead of using multiple engines in one layer. Gao explains “Alternatively, we can use multiple engines to process multiple layers in a pipelined manner by spatially mapping the NN DAG structures [2, 25, 36, 38, 40, 44]. Such inter-layer pipelining is effective in increasing the hardware utilization when layers are small and hardware resources are abundant.” (Gao: Page 810, Column 1 Second Paragraph)
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the implementation of multiple engines to process multiple layers as taught by Gao with the Optical Matrix Multiplication (OMM) unit as taught by Hosseinzadeh in view of Raghavan. One would have been motivated to combine these references because both references disclose parallel matrix multiplication in neural networks, and Gao enhances the model of Hosseinzadeh in view of Raghavan by "increasing the hardware utilization when layers are small and hardware resources are abundant" (Gao: Page 810, Column 1 Second Paragraph). The combination of Hosseinzadeh in view of Raghavan as modified by Gao teaches the limitation “and controlling the first accelerator core to perform a second matrix multiplication to produce a second output data block using a second portion of the first weight matrix and at least a second portion of the input data set;” in its entirety.
Additionally, Gao teaches how layers can start processing data (i.e., matrix multiplication) with a subset of input data (fmap) instead of waiting for the entire input to be available (all fmaps). Gao explains “If the next layer (L-2) also sequentially accepts these data as its ifmaps, it can start processing after waiting for a single fmap rather than all fmaps." (Page 813, Alternate layer loop ordering (ALLO) dataflow, First paragraph)
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the Alternate layer loop ordering (ALLO) dataflow method as taught by Gao with the Optical Matrix Multiplication (OMM) unit as taught by Hosseinzadeh in view of Raghavan. One would have been motivated to combine these references because both references disclose processing neural networks, and Gao enhances the model of Hosseinzadeh in view of Raghavan by allowing for multiple layers to be computed concurrently for a faster run time. The combination of Hosseinzadeh in view of Raghavan as modified by Gao teach “wherein processing the second layer comprises: subsequent to completion of the first matrix multiplication and prior to completion of the second matrix multiplication, controlling the second accelerator core to perform a third matrix multiplication using the second weight matrix and the first output data block.”
With regards to Claim 15, Hosseinzadeh in view of Raghavan in view of Gao teach:
The processing system of claim 14, wherein the controller is further configured to control the second accelerator core (Hosseinzadeh: ¶00399, e.g., Controller 110 controls ANN computations; Fig. 2A, e.g., flowchart of ANN computation, where OMM unit from the Optical processor (analog accelerator) is used; Raghavan: Fig. 3, e.g., Processor cores include ALU, which is used for matrix multiplication; Implementing processor cores from Raghavan in the OMM unit from Hosseinzadeh would cause for the Controller 110 to control a core to perform matrix multiplication (second accelerator core)) to complete the third matrix multiplication subsequent to completion of the second matrix multiplication by the first accelerator core (Gao: Page 810, Column 1 Second Paragraph, e.g., Multiple engines (cores) can be implemented to process multiple small layers; Page 813, Alternate layer loop ordering (ALLO) dataflow, First paragraph, e.g., Layer L-2 would need to receive all fmaps (i.e., from first and second matrix multiplication) to finish computation of fmaps (i.e., third matrix multiplication)).
With regards to Claim 16, Hosseinzadeh in view of Raghavan in view of Gao teach:
The processing system of claim 14, wherein the first portion of the first weight matrix comprises at least a first row of the first weight matrix (Raghavan: Fig. 2, e.g., Elements 200 and 208 (first portion) are in the first row on the weight matrix), and wherein the controller is configured to control the first accelerator core (Hosseinzadeh: ¶00399, e.g., Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., step 240 obtains a first plurality of digitalized optical outputs corresponding to the optical output vector (first output data block) from the OMM unit (analog accelerator)) to perform the first matrix multiplication to produce the first output data block using the first row of the first weight matrix (Raghavan: Fig. 2, e.g., VMA instruction uses first portion elements 200 and 208 of the first matrix for matrix multiplication, which are in a first row of the first weight matrix).
With regards to Claim 17, Hosseinzadeh in view of Raghavan in view of Gao teach:
The processing system of claim 16, wherein the second portion of the first weight matrix comprises at least a second row of the first weight matrix (Raghavan: Fig. 2, e.g., Elements in the second row (second portion of weight matrix)), and wherein the controller is configured to control the first accelerator core to perform the second matrix multiplication to produce the second output data block using the second row of the first weight matrix (Hosseinzadeh: ¶00399, e.g., Controller 110 controls ANN computations; Fig. 2A, e.g., flowchart of ANN computation, where OMM unit from the Optical processor performs matrix multiplication; Raghavan: Fig. 2, e.g., second row of the product matrix (second output data block) is obtained from a matrix multiplication (second matrix multiplication) using the second row of the first matrix 100 (weight matrix)).
With regards to Claim 18, Hosseinzadeh in view of Raghavan in view of Gao teach:
The processing system of claim 17, wherein the controller is configured to control the first accelerator core to perform the first matrix multiplication (Hosseinzadeh: ¶00399 Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., step 240 obtains a first plurality of digitalized optical outputs corresponding to the optical output vector from the OMM unit (analog accelerator) by performing a first matrix multiplication) using tile parallelism (Raghavan: Column 4 Lines 23 – 29, e.g., Each element group corresponds to a tile; Column 4 Lines 48 – 52, e.g., Tile parallelism is used to perform matrix multiplication; Fig. 2).
With regards to Claim 19, Hosseinzadeh in view of Raghavan in view of Gao teach:
The processing system of claim 14, wherein the controller is configured to control the first accelerator core to perform the first matrix multiplication (Hosseinzadeh: ¶00399 Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., step 240 obtains a first plurality of digitalized optical outputs corresponding to the optical output vector from the OMM unit (analog accelerator) by performing a first matrix multiplication) using data parallelism (Raghavan: Column 4 Lines 42 – 52, e.g., Data parallelism is used to perform matrix multiplication).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Hosseinzadeh, in view of Raghavan, in view of Gao, further in view of Bunandar et al. (U.S. Patent Application Publication No.: US 20190356394 A1), hereinafter “Bunandar”.
With regards to Claim 20, Hosseinzadeh in view of Raghavan in view of Gao teach:
and wherein: controlling the first accelerator core to perform the first matrix multiplication comprises controlling the first … core to perform the first matrix multiplication in an optical domain (Hosseinzadeh: ¶00432, e.g., Optical processor 140 contains optical components; ¶00399, e.g., Controller 110 controls ANN computations; ¶00349, e.g., Fig 2A shows the process of the ANN computation; Fig. 2A, e.g., Matrix multiplication is performed in step 240; Raghavan: Fig. 3, e.g., Processor cores include ALU, which is used for matrix multiplication; Implementing processor cores from Raghavan in the OMM unit from Hosseinzadeh would cause for the Controller 110 to control a core to perform matrix multiplication (second accelerator core));
and controlling the second accelerator core to perform the third matrix multiplication comprises controlling the second … core to perform the second third multiplication in the optical domain (Hosseinzadeh: ¶00432, e.g., Optical processor 140 contains optical components; Fig. 2A, e.g., Process of the ANN computation. Matrix multiplication is performed in step 240; ¶00453, e.g., Second matrix multiplication is performed in the second layer; Gao: Page 810, Column 1 Second Paragraph, e.g., Multiple engines (cores) can be implemented to process multiple small layers. Second core would compute second layer (including third matrix multiplication)).
Hosseinzadeh in view of Raghavan in view of Gao does not teach:
wherein the first accelerator core comprises a first photonic core and the second accelerator core comprises a second photonic core,
controlling the first photonic core to perform the first matrix multiplication in an optical domain;
controlling the second photonic core to perform the second matrix multiplication in the optical domain.
However, in the same field of endeavor, Bunandar teaches how photonic cores can be used for matrix multiplication. Bunandar explains “Referring to FIG. 1-3, the photonic processor 1-103 implements matrix multiplication on an input vector represented by the n input optical pulse”, See ¶0119, and “According to some embodiments, in the photocore of a photonic processor representing only real matrices, the process 3-700 of FIG. 3-7 may be carried out”, See ¶0399. Additionally, Hosseinzadeh explains how submatrix multiplication can be performed in different device (core). See ¶00194.
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the photonic cores as taught by Bunandar with the Multi-core Optical Matrix Multiplication unit as taught by Hosseinzadeh in view of Raghavan in view of Gao. One would have been motivated to combine these references because both references disclose performing matrix multiplication using cores, and Bunandar enhances the model of Hosseinzadeh in view of Raghavan in view of Gao by providing photonic cores in order to make computations in the optical processor.
Prior Art made of Record
US 12518146 B1 – teaches processing a neural network inference circuit including computational nodes at multiple layers using processing cores.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARLOS H DE LA GARZA whose telephone number is (571)272-0474. The examiner can normally be reached Monday-Friday 9:30AM-6PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Caldwell can be reached at (571) 272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/C.H.D./
Carlos H. De La GarzaExaminer, Art Unit 2182 (571)272-0474
/ANDREW CALDWELL/Supervisory Patent Examiner, Art Unit 2182