Detailed Action
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is non-final and is in response to claims filed 01/18/2023. Claims 1-20 are pending for examination.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
With regards to claim 1, it recites “a computation circuit configured to perform a pooling operation on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal”. It is unclear how the system can perform a pooling operation on the second matrix from the first memory if the first memory outputs the first matrix and not the second matrix. For examination purposes, examiner has interpreted “a computation circuit configured to perform a pooling operation on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal” as “a computation circuit configured to perform a pooling operation on the first matrix from the first memory or the second matrix from the second memory to obtain an operation result in response to a fourth control signal”.
With regards to claim 1, it recites “serially output each element of a second matrix in the first matrix stored by the second memory is”. It is unclear if the elements are being inserted into the first matrix or if the second matrix is a portion, partition, or sub-matrix of the first matrix.
With regards to claim 12, it recites “The chip for pooling operation according to claim 1, further comprising: a data collator configured to collate a plurality of the operation results into a preset format and output in response to a sixth control signal”. It is unclear what the “output” is. For examination purposes, examiner has interpreted “The chip for pooling operation according to claim 1, further comprising: a data collator configured to collate a plurality of the operation results into a preset format and output in response to a sixth control signal” as “The chip for pooling operation according to claim 1, further comprising: a data collator configured to collate a plurality of the operation results into a preset format and output the collated data in response to a sixth control signal”.
With regards to claim 14, it recites “performing a pooling operation by a computation circuit on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal”. It is unclear how the system can perform a pooling operation on the second matrix from the first memory if the first memory outputs the first matrix and not the second matrix. For examination purposes, examiner has interpreted “performing a pooling operation by a computation circuit on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal” performing a pooling operation by a computation circuit on the first matrix from the first memory or the second matrix from the second memory to obtain an operation result in response to a fourth control signal”.
With regards to claim 14, it recites “outputting serially by a second memory each element of a second matrix in the first matrix stored by the second memory”. It is unclear if the elements are being inserted into the first matrix or if the second matrix is a portion, partition, or sub-matrix of the first matrix.
Claims 2-11, 13, and 15-20 are rejected for being dependent on an above rejected claim.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 7, and 10-19 are rejected under 35 U.S.C. 103 as being unpatentable over Zejda et al. (US 10354733 B1) hereinafter Zejda in view of Bae et al. (US 6421695 B1) hereinafter Bae further in view of Wu et al. (US 12159212 B1) hereinafter Wu.
With regards to claim 1, Zejda teaches A chip for pooling operation, comprising: [a demultiplexer comprising a first input terminal, a first output terminal, and a second output terminal and configured to output a first matrix from the first input terminal via the first output terminal or the second output terminal in response to a first control signal;] (Zejda Column 5 Lines 42-43: The hardware accelerator 116; Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler)
a first memory [connected to the first output terminal] and configured to perform a plurality of outputs to output each element of the first matrix stored by the first memory in response to a second control signal, (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 2: a read control circuit (“read control 346”); Zejda Fig. 9: shows that the matrix is output from the memory)
outputting a column of elements of the first matrix in parallel for each output; (Zejda Column 12 Lines 10-11: the matrix B buffer 904 may be in multiples of the column size of the compute array; Zejda Fig. 9 shows that the matrix is output in parallel in columns)
a second memory [connected to the second output terminal] and configured to serially output each element of a second matrix [in the first matrix] stored by the second memory in response to a third control signal; (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 8: a read control circuit (“read control 350”); Zejda Fig. 9: shows that the matrix is output from the memory)
and a computation circuit configured to perform a pooling operation on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal (Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler; Zejda Column 1 Lines 35-38: Example activation functions include the sigmoid function, the hyperbolic tangent (tan h) function, the Rectified Linear Unit (ReLU) function, and the identity function; Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit; (This shows that the system can perform an identity function and then pass the matrix to the pooling circuit)).
Zejda fails to teach [A chip for pooling operation, comprising:] a demultiplexer comprising a first input terminal, a first output terminal, and a second output terminal and configured to output a first matrix from the first input terminal via the first output terminal or the second output terminal in response to a first control signal; [a first memory] connected to the first output terminal, [a second memory] connected to the second output terminal.
However, Bae teaches [A chip for pooling operation, comprising:] a demultiplexer comprising a first input terminal, a first output terminal, and a second output terminal and configured to output a first matrix from the first input terminal via the first output terminal or the second output terminal in response to a first control signal; (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600)
[a first memory] connected to the first output terminal (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600)
[a second memory] connected to the second output terminal (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda with the demultiplexer as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
Zejda in view of Bae fails to teach [a second matrix] in the first matrix.
However, Wu teaches [a second matrix] in the first matrix (Wu Column 4 Lines 49-51: The memory 105 stores one or more tensor buffers 110 that can be used to store all, or a portion of, input data).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae with the second matrix being in the first matrix as taught by Wu. One of ordinary skill in the art would be motivated to make this combination because in various examples, employing the shared convolution of the convolution model 200 improves throughput of the corresponding neural net by reducing corresponding computation as taught by Wu (Wu Column 9 Lines 36-38). Also, this would increase the flexibility of the system as it could store different portions of the matrix depending on the needs of the system.
With regards to claim 2, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches wherein the first memory [comprises] at least one row buffer [connected to the first output terminal,] (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer storing rows of the matrix)
each row buffer being configured to store a row of elements in the first matrix (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer storing rows of the matrix).
Zejda fails to each [wherein the first memory] comprises [at least one row buffer] connected to the first output terminal.
However, Bae teaches [wherein the first memory comprises at least one row buffer] connected to the first output terminal, (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer connection as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
Zejda in view of Bae fails to teach [wherein the first memory] comprises [at least one row buffer connected to the first output terminal,].
However, Wu teaches [wherein the first memory] comprises [at least one row buffer connected to the first output terminal,] (Wu Column 4 Lines 49-51: The memory 105 stores one or more tensor buffers 110 that can be used to store all, or a portion of, input data).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the second matrix being in the first matrix as taught by Wu. One of ordinary skill in the art would be motivated to make this combination because in various examples, employing the shared convolution of the convolution model 200 improves throughput of the corresponding neural net by reducing corresponding computation as taught by Wu (Wu Column 9 Lines 36-38). Also, this would increase the flexibility of the system as it could store different portions of the matrix depending on the needs of the system.
With regards to claim 3, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 2 above. Zejda further teaches wherein each row buffer comprises one or more first-in-first-out memories that are connected in series; (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer storing rows of the matrix and the elements of the buffer being connected in series in a first in first out arrangement)
wherein the at least one row buffer comprises N row buffers, (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers)
different row buffers being configured to store different rows of elements in the first matrix, N being an integer greater than or equal to 2 (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers).
With regards to claim 4, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 3 above. Zejda further teaches wherein the N row buffers are sequentially connected in series, (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series)
a 1st row buffer of the N row buffers is connected [to the first output terminal,] (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series)
an i-th row buffer is configured to output a row of elements flowing to the i-th row buffer to an (i+1)th row buffer so that different row buffers store different rows of elements in the first matrix, (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series that flow to a next one, for example the one storing element 1 flows into the one that stores element 2)
and i is an integer greater than or equal to 1 and less than or equal to N-1 (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series that flow to a next one, for example the one storing element 1 flows into the one that stores element 2).
Zejda fails to each [a 1st row buffer of the N row buffers is connected] to the first output terminal.
However, Bae teaches [a 1st row buffer of the N row buffers is connected] to the first output terminal, (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer connection as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
With regards to claim 5, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 4 above. Zejda further teaches wherein the i-th row buffer is configured to output one element to the (i+1)th row buffer each time the element is output to the computation circuit (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series that flow to the next one while processing is occurring).
With regards to claim 7, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches wherein the computation circuit is configured to [determine the second matrix on the basis of the first matrix] from the first memory in response to the fourth control signal (Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler; Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit)
and to perform the pooling operation on the second matrix to obtain the operation result (Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler; Zejda Column 1 Lines 35-38: Example activation functions include the sigmoid function, the hyperbolic tangent (tan h) function, the Rectified Linear Unit (ReLU) function, and the identity function; Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit; (This shows that the system can perform an identity function and then pass the matrix to the pooling circuit)).
Zejda fails to teach [wherein the computation circuit is configured to] determine the second matrix on the basis of the first matrix [from the first memory in response to the fourth control signal].
However, Wu teaches [wherein the computation circuit is configured to] determine the second matrix on the basis of the first matrix [from the first memory in response to the fourth control signal] (Wu Column 4 Lines 49-51: The memory 105 stores one or more tensor buffers 110 that can be used to store all, or a portion of, input data).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with determining the second matrix as taught by Wu. One of ordinary skill in the art would be motivated to make this combination because in various examples, employing the shared convolution of the convolution model 200 improves throughput of the corresponding neural net by reducing corresponding computation as taught by Wu (Wu Column 9 Lines 36-38). Also, this would increase the flexibility of the system as it could store different portions of the matrix depending on the needs of the system.
With regards to claim 10, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches wherein the second memory is configured to [determine the second matrix from the first matrix] stored by the second memory (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 8: a read control circuit (“read control 350”); Zejda Fig. 9: shows that the matrix is output from the memory)
and serially output each element of the second matrix to the computation circuit in response to the third control signal (Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler; Zejda Column 1 Lines 35-38: Example activation functions include the sigmoid function, the hyperbolic tangent (tan h) function, the Rectified Linear Unit (ReLU) function, and the identity function; Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit; (This shows that the system can perform an identity function and then pass the matrix to the pooling circuit)).
Zejda fails to teach[wherein the second memory is configured to] determine the second matrix on the basis of the first matrix [stored by the second memory].
However, Wu teaches [wherein the second memory is configured to] determine the second matrix on the basis of the first matrix [stored by the second memory] (Wu Column 4 Lines 49-51: The memory 105 stores one or more tensor buffers 110 that can be used to store all, or a portion of, input data).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with determining the second matrix as taught by Wu. One of ordinary skill in the art would be motivated to make this combination because in various examples, employing the shared convolution of the convolution model 200 improves throughput of the corresponding neural net by reducing corresponding computation as taught by Wu (Wu Column 9 Lines 36-38). Also, this would increase the flexibility of the system as it could store different portions of the matrix depending on the needs of the system.
With regards to claim 11, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches further comprising: a multiplexer comprising a second input terminal, a third input terminal, and a third output terminal (Zejda Column 7 Lines 26-29: Outputs of the IM2COL circuit 344 and the read control circuit 346 are coupled to inputs of the multiplexer 356. An output of the multiplexer 356 is coupled to an input of the FIFOs; Zejda Fig. 3: shows the multiplexer with 2 inputs and an output)
wherein the second input terminal is connected to the first memory, the third input terminal is connected to the second memory, and the third output terminal is connected to the computation circuit (Zejda Column 7 Lines 26-29: Outputs of the IM2COL circuit 344 and the read control circuit 346 are coupled to inputs of the multiplexer 356. An output of the multiplexer 356 is coupled to an input of the FIFOs; Zejda Column 7 Lines 48-53: the input activation matrices may be read directly from the RAM 226 using the block order unit, caches, and the read control circuit 346. Alternatively, to perform convolution, for example, the input activations may be read from the RAM 226 and processed by the IM2COL circuit 344 for input to the compute array; Zejda Fig. 3: shows the multiplexer with 2 inputs and an output, the output being connected to the pooling circuit).
With regards to claim 12, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches further comprising: a data collator configured to collate a plurality of the operation results into a preset format (Zejda Column 7 Lines 38-42: An output of the max pool circuit 366 is coupled to another input of the multiplexer 368. An output of the multiplexer 368 is coupled to an input of the FIFOs 354, and an output of the FIFOs 354 is coupled to an input of the write control circuit 352; Zejda Column 11 Lines 39-41: the C sub-block accumulation into matrix C may be moved out of the compute array and performed by write control 352)
and output in response to a sixth control signal (Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit).
With regards to claim 13, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches further comprising: a first controller configured to transmit a plurality of control signals, (Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit)
the plurality of control signals comprising [the first control signal], the second control signal, the third control signal, and the fourth control signal (Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit).
Zejda fails to teach the first control signal.
However, Bae teaches the first control signal (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer connection as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
With regards to claim 14, Zejda teaches A method for pooling operation, comprising: [outputting by a demultiplexer a first matrix from a first input terminal of the demultiplexer via a first output terminal of the demultiplexer or via a second output terminal of the demultiplexer in response to a first control signal;] (Zejda Column 5 Lines 42-43: The hardware accelerator 116; Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler)
performing multiple outputs by a first memory to output each element of the first matrix stored by the first memory (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 2: a read control circuit (“read control 346”); Zejda Fig. 9: shows that the matrix is output from the memory)
and output a column of elements of the first matrix in parallel each time in response to a second control signal in the case where the first matrix is output [via the first output terminal,] (Zejda Column 12 Lines 10-11: the matrix B buffer 904 may be in multiples of the column size of the compute array; Zejda Fig. 9 shows that the matrix is output in parallel in columns)
the first memory being connected to [the first output terminal;] (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 2: a read control circuit (“read control 346”); Zejda Fig. 9: shows that the matrix is output from the memory)
outputting serially by a second memory each element of a second matrix [in the first matrix] stored by the second memory in response to a third control signal in the case where the first matrix is output [via the second output terminal,] (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 8: a read control circuit (“read control 350”); Zejda Fig. 9: shows that the matrix is output from the memory)
the second memory being connected [to the second output terminal;] (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 8: a read control circuit (“read control 350”); Zejda Fig. 9: shows that the matrix is output from the memory)
and performing a pooling operation by a computation circuit on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal (Zejda Column 7 Lines 58-60: The max pool circuit 366 may implement a max pooling function on the scaled output of the scaler; Zejda Column 1 Lines 35-38: Example activation functions include the sigmoid function, the hyperbolic tangent (tan h) function, the Rectified Linear Unit (ReLU) function, and the identity function; Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit; (This shows that the system can perform an identity function and then pass the matrix to the pooling circuit)).
Zejda fails to teach outputting by a demultiplexer a first matrix from a first input terminal of the demultiplexer via a first output terminal of the demultiplexer or via a second output terminal of the demultiplexer in response to a first control signal; [in the case where the first matrix is output] via the first output terminal, [the first memory being connected] to the first output terminal; [outputting serially by a second memory each element of a second matrix] in the first matrix [stored by the second memory in response to a third control signal in the case where the first matrix is output] via the second output terminal, [the second memory being connected to] the second output terminal.
However, Bae teaches outputting by a demultiplexer a first matrix from a first input terminal of the demultiplexer via a first output terminal of the demultiplexer or via a second output terminal of the demultiplexer in response to a first control signal; (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600)
[in the case where the first matrix is output] via the first output terminal, (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600)
[the first memory being connected] to the first output terminal; (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600)
[outputting serially by a second memory each element of a second matrix in the first matrix stored by the second memory in response to a third control signal in the case where the first matrix is output] via the second output terminal, (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600)
[the second memory being connected to] the second output terminal; (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda with the demultiplexer as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
Zejda in view of Bae fails to teach [outputting serially by a second memory each element of a second matrix] in the first matrix [stored by the second memory].
However, Wu teaches [outputting serially by a second memory each element of a second matrix] in the first matrix [stored by the second memory] (Wu Column 4 Lines 49-51: The memory 105 stores one or more tensor buffers 110 that can be used to store all, or a portion of, input data).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae with the second matrix being in the first matrix as taught by Wu. One of ordinary skill in the art would be motivated to make this combination because in various examples, employing the shared convolution of the convolution model 200 improves throughput of the corresponding neural net by reducing corresponding computation as taught by Wu (Wu Column 9 Lines 36-38). Also, this would increase the flexibility of the system as it could store different portions of the matrix depending on the needs of the system.
With regards to claim 15, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 14 above. Zejda further teaches wherein the first memory [comprises] at least one row buffer [connected to the first output terminal,] (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer storing rows of the matrix)
each row buffer being configured to store a row of elements in the first matrix (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer storing rows of the matrix).
Zejda fails to each [wherein the first memory] comprises [at least one row buffer] connected to the first output terminal.
However, Bae teaches [wherein the first memory comprises at least one row buffer] connected to the first output terminal, (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer connection as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
Zejda in view of Bae fails to teach [wherein the first memory] comprises [at least one row buffer connected to the first output terminal,].
However, Wu teaches [wherein the first memory] comprises [at least one row buffer connected to the first output terminal,] (Wu Column 4 Lines 49-51: The memory 105 stores one or more tensor buffers 110 that can be used to store all, or a portion of, input data).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the second matrix being in the first matrix as taught by Wu. One of ordinary skill in the art would be motivated to make this combination because in various examples, employing the shared convolution of the convolution model 200 improves throughput of the corresponding neural net by reducing corresponding computation as taught by Wu (Wu Column 9 Lines 36-38). Also, this would increase the flexibility of the system as it could store different portions of the matrix depending on the needs of the system.
With regards to claim 16, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 15 above. Zejda further teaches wherein each of the row buffer comprises one or more first-in-first-out memories that are connected in series; (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer storing rows of the matrix and the elements of the buffer being connected in series in a first in first out arrangement)
wherein the at least one row buffer comprises N row buffers, (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers)
different row buffers being configured to store different rows of elements in the first matrix, N being an integer greater than or equal to 2 (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers).
With regards to claim 17, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 16 above. Zejda further teaches wherein the N row buffers are sequentially connected in series, (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series)
and a 1st row buffer of the N row buffers is connected [to the first output terminal,] (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series)
the method further comprises: outputting by an i-th row buffer a row of elements flowing to the i-th row buffer to an (i+1)th row buffer so that different row buffers store different rows of elements in the first matrix, (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series that flow to a next one, for example the one storing element 1 flows into the one that stores element 2)
where i is an integer greater than or equal to 1 and less than or equal to N-1 (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series that flow to a next one, for example the one storing element 1 flows into the one that stores element 2).
Zejda fails to each [and a 1st row buffer of the N row buffers is connected] to the first output terminal.
However, Bae teaches [and a 1st row buffer of the N row buffers is connected] to the first output terminal, (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer connection as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
With regards to claim 18, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 17 above. Zejda further teaches wherein the i-th row buffer outputs one element to the (i+1)th row buffer each time the element is output to the computation circuit (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A buffer contains multiple row buffers connected in series that flow to the next one while processing is occurring).
With regards to claim 19, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 1 above. Zejda further teaches an accelerator for pooling operation, comprising: the chip for pooling operation according to claim 1 (Zejda Column 4 Lines 7-17: FIG. 1 is a block diagram depicting a system 100 for implementing neural networks, in accordance with an example of the present disclosure. The system 100 includes a computer system… and/or one or more hardware accelerators 116).
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Zejda in view of Bae further in view of Wu further in view of Zaidy et al. (US 20220223201 A1) hereinafter Zaidy.
With regards to claim 6, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 2 above. Zejda further teaches wherein the first memory further comprises a data path connected [to the first output terminal] (Zejda Column 5 Lines 42-43: The hardware accelerator 116 includes a programmable IC 228, a non-volatile memory (NVM) 224, and RAM 226; Zejda Column 7 Line 8: a read control circuit (“read control 350”); Zejda Fig. 9: shows that the matrix is output from the memory)
[and in parallel with the at least one row buffer] and configured such that a first row of elements or a last row of elements in the first matrix [flow out via the data path] (Zejda Column 12 Lines 9-10: the height (the row size, designated as MH) of the matrix A buffer 902 may be in multiples of the row size of the compute array; Zejda Fig. 9: shows the matrix A is output in rows).
Zejda fails to teach [wherein the first memory further comprises a data path connected] to the first output terminal.
However, Bae teaches [wherein the first memory further comprises a data path connected] to the first output terminal (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer connection as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
Zejda in view of Bae fails to teach and in parallel with the at least one row buffer [and configured such that a first row of elements or a last row of elements in the first matrix] flow out via the data path.
However, Zaidy teaches and in parallel with the at least one row buffer [and configured such that a first row of elements or a last row of elements in the first matrix] flow out via the data path (Zaidy [0120]: loading 317 of the weights 309 is performed from the random access memory 105 into the local memory 115, by passing the system buffer 305; Zaidy Fig. 8: shows that the datapath is in parallel to the buffer and allows the data to flow from the RAM).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the datapath as taught by Zaidy. One of ordinary skill in the art would be motivated to make this combination because the capacity of the system buffer can be used more efficiently for other types of data that may be re-fetched more frequently as taught by Zaidy (Zaidy [0096]).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Zejda in view of Bae further in view of Wu further in view of Zaidy further in view of Singh et al. (US 20210264250 A1) hereinafter Singh.
With regards to claim 20, Zejda in view of Bae further in view of Wu teaches all of the limitations of claim 19 above. Zejda further teaches A system for pooling operation, comprising: the accelerator for pooling operation according to claim 19; (Zejda Column 4 Lines 7-17: FIG. 1 is a block diagram depicting a system 100 for implementing neural networks, in accordance with an example of the present disclosure. The system 100 includes a computer system… and/or one or more hardware accelerators 116)
a direct memory access module configured to retrieve the first matrix [from the third matrix] and transmit to [the demultiplexer] in response to a seventh control signal; (Zejda Column 6 Lines 41-42: The PCIe DMA controller 304 facilitates DMA operations to the RAM 226 and the kernel 238)
and a [second] controller configured to transmit the seventh control signal and an eighth control signal, (Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit)
the eighth control signal causing the first control signal, the second control signal, the third control signal, and the fourth control signal to be transmitted (Zejda Column 7-8 Lines 65-2: The control logic 342 controls the various circuits in the processing circuits 341, such as the IM2COL circuit 344, the 3-D partitioning block order unit, the read control circuit 346, the multiplexers 356 and 368, the read control circuit 350, the scaler 364, the max pool circuit).
Zejda fails to teach the demultiplexer.
However, Bae teaches the demultiplexer (Bae Column 8 Lines 59-62: The demultiplexer DMUX outputs, in accordance with the select signal SS, the matrix data value from the 1-D IDCT core 700 to the first memory unit 300 or to the second memory unit 600).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the demultiplexer as taught by Bae. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility of the system as it would be able to choose where to store the matrices.
Zejda in view of Bae fails to teach a second controller.
However, Zaidy teaches a second controller (Zaidy [0033]: the control unit 113 can use the processing units 111 to perform vector and matrix operations in accordance with instructions. Further, the control unit 113 can load instructions and operands from the random access memory 105 through a memory interface 117 and a high speed/bandwidth connection 119).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu with the second controller as taught by Zaidy. One of ordinary skill in the art would be motivated to make this combination because it would increase the flexibility and efficiency of the system as it could choose which components to activate to perform the required operations.
Zaidy in view of Bae further in view of Zaidy fails to teach a third memory configured to store a third matrix; [a direct memory access module configured to retrieve the first matrix] from the third matrix [and transmit to the demultiplexer in response to a seventh control signal;].
However, Singh teaches a third memory configured to store a third matrix; (Singh [0022]: For example, the input data may 110 may include a data structure stored in a memory)
[a direct memory access module configured to retrieve the first matrix] from the third matrix [and transmit to the demultiplexer in response to a seventh control signal;] (Singh [0036]: In one embodiment, during the pooling operation, the feature tensor 128 is divided into batches. The feature tensor 120 is patched by depth. Pooling operations are performed on the batches from the feature tensor. The pooling operations are performed on the sub-tensors from each batch. Accordingly, each batch is divided into a plurality of sub-tensors).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings Zejda in view of Bae further in view of Wu further in view of Zaidy with the third memory and the third matrix as taught by Singh. One of ordinary skill in the art would be motivated to make this combination because the pooling unit includes hardware blocks that promote computational and area efficiency in the convolutional neural network as taught by Singh (Singh Abstract).
Allowable Subject Matter
Claims 8-9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
With regards to claim 8, the limitation of “wherein at least one column of the plurality of second elements is the same as at least one column of the plurality of first elements, and the second pooling operation comprises: performing the first operation on each column of second elements of the plurality of second elements except the at least one column of second elements to obtain a second intermediate result” differentiates the claim from the prior art.
While prior art teaches of performing pooling operations on sub-matrices with overlapping regions. For example, Du et al. (US 20180232629 A1) teaches of performing pooling operations on sub-matrices with overlapping regions. However, they perform the pooling operation on the overlapping region as well.
Prior art also teaches of using a sliding window to perform the pooling operation. However, they perform the pooling operation on the overlapping region as well.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jakob O Gudas whose telephone number is (571)272-0695. The examiner can normally be reached Monday-Thursday: 7:30AM-5:00PM Friday: 7:30AM-4:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.O.G./Examiner, Art Unit 2151
/James Trujillo/Supervisory Patent Examiner, Art Unit 2151