DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after allowance or after an Office action under Ex Parte Quayle, 25 USPQ 74, 453 O.G. 213 (Comm'r Pat. 1935). Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, prosecution in this application has been reopened pursuant to 37 CFR 1.114. Applicant's submission filed on 5/15/2026 has been entered.
The indicated allowability of claims 1-3, 5-15, 17-25, 27-30 are withdrawn in view of the newly discovered references to Kuo et al. (US 20190220742 A1) and Ovsiannikov et al. (US 20200349420 A1). Rejections based on the newly cited reference(s) follow.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 5/15/2026 has been considered except where lined through because the listed foreign document was not provided.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-7, 11-18, 22-28, 30 are rejected under 35 U.S.C. 103 as being unpatentable over Jia et al. (U.S. Pub. No. US 20230074229 A1, hereinafter “Jia”) in view of Ovsiannikov et al. (US 20200349420 A1, hereinafter “Ovsiannikov”) in further view of Kuo et al. (US 20190220742 A1, hereinafter “Kuo”).
As per claim 1, Jia teaches A method, comprising: performing a convolution with a compute-in-memory (CIM) array to generate CIM output (Jia: Fig. 17 element CIMA, [0138]);
writing at least a portion of the CIM output corresponding to a first output data channel, of a plurality of output data channels in the CIM output (Jia: [0118]).
However, while Jia discloses near-memory computation (Fig. 17; [0141]), Jia does not explicitly disclose they contain buffers nor a partitioning of input data based on buffer configuration. Thus, Jia does not teach to a first digital multiply-accumulate (DMAC) activation buffer comprises: partitioning input data to the CIM array based on a configuration of the first DMAC activation buffer; and selectively writing the CIM output to the first DMAC activation buffer based on the partitioning; reading a patch of the CIM output from the first DMAC activation buffer; reading weight data from a DMAC weight buffer; and performing multiply-accumulate (MAC) operations with the patch of CIM output and the weight data to generate a DMAC output.
Ovsiannikov teaches to a first digital multiply-accumulate (DMAC) activation buffer (Ovsiannikov: Fig. 1B element 139; [0159]) comprises: and selectively writing the CIM output to the first DMAC activation buffer based on the partitioning (Ovsiannikov: Figs. 2G-2J; [0203]; [0189] wherein the IFM cache of each MR tile receives different slices of the IFM 3D tensor based on the sub-divided sets);
reading a patch of the CIM output from the first DMAC activation buffer (Ovsiannikov: Fig. 1B element 141; [0160], [0162], wherein the ABU reads in lanes of activation data from the IFM cache);
reading weight data from a DMAC weight buffer (Ovsiannikov: Fig. 1C elements 127; [0161]);
and performing multiply-accumulate (MAC) operations with the patch of CIM output and the weight data to generate a DMAC output (Ovsiannikov: Fig. 1B element 133; [0158]).
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the near-memory computing of Jia with the Multiply-and-Reduce (MR) tiles of Ovsiannikov. One would have been motivated to combine these references because both references disclose neural network processors with near-memory computation, and MR tiles of Ovsiannikov reduces the number of memory (SRAM) reads without reducing computation throughput (Ovsiannikov: [0159]).
However, while Jia discloses near-memory computation (Fig. 17; [0141]), Jia does not explicitly disclose a partitioning of input data based on buffer configuration. Thus, Jia does not teach partitioning input data to the CIM array based on a configuration of the first DMAC activation buffer;
Kuo teaches partitioning input data to the CIM array based on a configuration of the first DMAC activation buffer (Kuo: [0018]);
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the input data movement of Jia with the data movement teaching of Kuo, wherein Jia will partition input data into multiple tiles based on limited buffer size. One would have been motivated to combine these references because both references disclose data movement in neural network processors, and the tiling of data taught in Kuo allows for cross-tile reuse of output data (Kuo: [0022]).
As per claim 2, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, further comprising generating a depthwise convolution output based on DMAC output for each respective output data channel of the plurality of output data channels (Ovsiannikov: [0186]).
As per claim 3, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, wherein performing the MAC operations with the patch of CIM output and the weight data is pipelined behind the CIM array (Ovsiannikov: [0159]) by initiating the MAC operations upon determining that a specified data element has been written to the first DMAC activation buffer (Ovsiannikov: [0162]).
As per claim 5, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, wherein selectively writing the CIM output to the first DMAC activation buffer based on the partitioning comprises: writing a first data element for a first partition to the first DMAC activation buffer; and retaining the first data element in the first DMAC activation buffer when writing data elements for a second partition (Ovsiannikov: [0159]).
As per claim 6, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, wherein selectively writing the CIM output to the first DMAC activation buffer comprises: writing a first set of data elements from the CIM output to the first DMAC activation buffer in a linear manner (Ovsiannikov: Figs. 1C-1F; [0162]);
upon reaching a boundary of a partition in the CIM output, bypassing a second set of data elements in the CIM output; and writing a third set of data elements from the CIM output to the first DMAC activation buffer in a linear manner (Ovsiannikov: [0188]).
As per claim 7, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, wherein reading the patch of CIM output from the first DMAC activation buffer is performed non-linearly and comprises: reading a first data element from the first DMAC activation buffer; and reading a second data element from the first DMAC activation buffer, wherein the second data element is not adjacent to the first data element in the first DMAC activation buffer (Ovsiannikov: [0173], wherein zero-valued elements are skipped).
As per claim 11, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, further comprising, for a second output data channel of the plurality of output data channels: determining that a patch of data in the second output data channel satisfies defined sparsity criteria; and refraining from processing the patch of data using a DMAC (Jia: [0224], [0250]).
As per claim 12, Jia/Ovsiannikov/Kuo further teaches the method of Claim 1, wherein the first DMAC activation buffer is tightly coupled to the CIM array by a dedicated path from the first output data channel to the first DMAC activation buffer (Jia: Fig. 17, [0138]; Ovsiannikov: Fig. 1 element 104; [0156]).
As per claims 13-15, 17-18, 22, they are directed to a non-transitory computer-readable medium corresponding to the method of claims 1-3, 6-7, 11, respectively, and are rejected for at least the same reasons.
As per claims 23-25, 27-28, 30, they are directed to a system corresponding to the method of claims 1-3, 6-7, 11, respectively, and are rejected for at least the same reasons.
Claims 8, 10, 19, 21, 29 are rejected under 35 U.S.C. 103 as being unpatentable over Jia/Ovsiannikov/Kuo in further view of Tullberg (U.S. Pub. No. US 20130339677 A1, hereinafter “Tullberg”).
As per claim 8, Jia/Ovsiannikov/Kuo teaches the method of Claim 1.
However, while Jia discloses SIMD controllers at the IMC output (Jia: [0126]), Jia does not explicitly disclose that they perform the reading of data from the IMC. Thus, Jia does not teach wherein writing CIM output to the first DMAC activation buffer and reading the patch of CIM output from the first DMAC activation buffer are controlled by a programmable controller.
Tullberg teaches wherein writing CIM output to the first DMAC activation buffer and reading the patch of CIM output from the first DMAC activation buffer are controlled by a programmable controller (Tullberg: [0028]).
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the near-memory computation controllers of Jia with the controller of Tullberg. One would have been motivated to combine these references because both references teach controlling dataflow, and combining prior art elements according to known methods to yield predictable results (controlling dataflow).
As per claim 10, Jia/Ovsiannikov/Kuo further teaches the method of Claim 8, wherein the programmable controller is programmed using one or more configurable variables comprising: a first variable indicating a size of the patch of CIM output (Ovsiannikov: [0203]);
a second variable indicating a size of each data element in the patch of CIM output (Ovsiannikov: [0176];
and a third variable indicating a spacing of data elements in the activation buffer (Tullberg: [0037]).
As per claim 19, it is directed to a non-transitory computer-readable medium corresponding to the method of claim 8 and is rejected for at least the same reasons.
As per claim 21, it is directed to a non-transitory computer-readable medium corresponding to the method of claim 10 and is rejected for at least the same reasons.
As per claim 29, Jia/Ovsiannikov/Kuo teaches the system of Claim 23.
However, while Jia discloses SIMD controllers at the IMC output (Jia: [0126]), Jia does not explicitly disclose that they perform the reading of data from the IMC. Thus, Jia does not teach wherein writing CIM output to the first DMAC activation buffer and reading the patch of CIM output from the first DMAC activation buffer are controlled by a programmable finite state machine (FSM).
Tullberg teaches wherein writing CIM output to the first DMAC activation buffer and reading the patch of CIM output from the first DMAC activation buffer are controlled by a programmable finite state machine (FSM) (Tullberg: [0028]).
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the near-memory computation controllers of Jia with the controller of Tullberg. One would have been motivated to combine these references because both references teach controlling dataflow, and combining prior art elements according to known methods to yield predictable results (controlling dataflow).
Claims 9, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jia/Ovsiannikov/Kuo/Tullberg in further view of Pal et al. (U.S. Pub. No. US 20070121499 A1, hereinafter “Pal”).
As per claim 9, Jia/Ovsiannikov/Kuo/Tullberg teaches the method of Claim 8.
However, while Jia discloses SIMD controllers controlling multiple blocks (Jia: [0126]), Jia does not explicitly disclose how they perform the reading of data from the IMC. Thus, Jia does not teach wherein: the programmable controller implements a read pointer and a write pointer; and the read pointer and write pointer are shared across a plurality of DMAC activation buffers, each associated with a respective output data channel of the plurality of output data channels.
Pal teaches wherein: the programmable controller implements a read pointer and a write pointer; and the read pointer and write pointer are shared across a plurality of DMAC activation buffers, each associated with a respective output data channel of the plurality of output data channels (Pal: Fig. 16, [0143] – [0145]).
Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify, with a reasonable expectation of success, the near-memory computation controller of Jia with the unified queue logic of Pal. One would have been motivated to combine these references because both references teach controlling dataflow of parallel units, and combining prior art elements according to known methods to yield predictable results (controlling dataflow of parallel units).
As per claim 20, it is directed to a non-transitory computer-readable medium corresponding to the method of claim 9 and is rejected for at least the same reasons.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHAT N LE whose telephone number is (571)272-0546. The examiner can normally be reached Monday-Friday 8:30AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew T Caldwell can be reached at (571) 272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/P.N.L./
Phat LeExaminer, Art Unit 2182 (571) 272-0546
/ANDREW CALDWELL/Supervisory Patent Examiner, Art Unit 2182