Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-10, 15, 19, 20, 25, 27 and 28 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Huang et al. (Pub. No.: US 2020/0117597).
Regarding independent claims 1, 15 and 19, Huang discloses a memory processing unit (MPU) (Fig.3) comprising: a first memory (Fig.3: memory 411) including a plurality of regions (Fig.3: Memory region); and a plurality of processing regions (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 and artificial intelligence core 430) interleaved between the plurality of regions of the first memory (Fig.3: Memory region), wherein respective processing regions (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 or artificial intelligence core 430) are coupled to adjacent ones of the plurality of regions of the first memory (Fig.3: Memory region), and wherein one or more of the plurality of processing regions (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 and artificial intelligence core 430) include a plurality of compute cores (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 or artificial intelligence core 430) ([0041]: Referring to FIG. 3, a memory 400 of the present embodiment includes memory region 411, 413, 415 and 417, row buffer blocks 412, 414, 416 and 418, a mode register 420, an artificial intelligence core 430 and a memory interface 440. In the present embodiment, the mode register 420 is coupled to the artificial intelligence core 430 and the memory interface 440 to provide a plurality of memory mode settings to the artificial intelligence core 430 and the memory interface 440, respectively. The memory interface 440 is coupled to the central processing unit core 351, the graphics processing unit core 352 and the digital signal processor core 353 via, for example, a shared bus. In the present embodiment, the artificial intelligence core 430 and the memory interface 440 each operate independently to access a memory array. The memory array includes the memory regions 411, 413, 415 and 417 as well as the row buffer blocks 412, 414, 416 and 418, and the memory regions 411, 413, 415 and 417 each include a plurality of memory banks) comprising;
one or more input/output (I/O) cores (Fig.3: ) configured to access input and output ports of the MPU ([0037]: In the present embodiment, the special function processing core may be, for example, a central processing unit (CPU) core, an image signal processor (ISP) core, a digital signal processor (DSP) core, a graphics processing unit (GPU) core or other similar special function processing core. In the present embodiment, the special function processing core is coupled to the memory interface 140 via a shared bus (or standard bus) to access the memory array 110 via the memory interface 140); and
a plurality of near memory (M) compute cores configured to compute neural network functions ([0036]: the artificial intelligence core 130 may be pre-designed to have functions and characteristics for performing a specific neural network operation. In other words, the memory 100 of the present embodiment has a function of performing an artificial intelligence operation, and the artificial intelligence core 130 and the external special function processing core can simultaneously access the memory array 110, so as to provide highly efficient data access and operation effects).
Regarding claim 2, Huang teaches wherein the plurality of regions of first memory are columnal interleaved between the plurality of processing regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 3, Huang teaches comprising: a second memory (Fig.3: memory region 413) coupled to the plurality of processing regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 4, Huang teaches wherein the second memory is configurably couplable to one or more near memory (M) compute cores in one or more of the plurality of processing regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 5, Huang teaches wherein the compute cores of one or more of the plurality of processing regions further comprises: one or more arithmetic (A) compute cores configured to compute arithmetic operations, wherein the one or more arithmetic (A) compute cores of each of one or more of the plurality of processing regions are communicatively coupled to adjacent ones of the first plurality of memory regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 6, Huang teaches wherein the one or more input/output (I/O) cores comprises: a first input/output (I/O) core configured to stream data into one of the plurality of regions of the first memory; and a second input/output (I/O) core configured to stream data out of another of the plurality of region of the first memory ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 7, Huang teaches wherein the near memory (M) compute cores include a plurality of physical channels configurable to perform computations simultaneously (see Fig.3).
Regarding claim 8, Huang teaches wherein the near memory (M) compute cores of respective ones of the plurality of processing regions are associated with one or more blocks of the second memory ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 9, Huang teaches wherein the near memory (M) compute cores include a plurality of physical channels configurable to perform computations simultaneously, and wherein the physical channels of the near memory (M) compute cores are associated with respective slices of the second memory (Fig.3 and [0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 10, Huang teaches wherein the near memory (M) compute cores include a plurality of configurable virtual channels (Fig.3 and [0037]: In the present embodiment, the special function processing core may be, for example, a central processing unit (CPU) core, an image signal processor (ISP) core, a digital signal processor (DSP) core, a graphics processing unit (GPU) core or other similar special function processing core. In the present embodiment, the special function processing core is coupled to the memory interface 140 via a shared bus (or standard bus) to access the memory array 110 via the memory interface 140).
Regarding claim 20, Huang teaches comprising: a second memory configurably couplable to one or more near memory (M) compute cores in one or more of the plurality of processing regions (Fig.3 and [0037]: In the present embodiment, the special function processing core may be, for example, a central processing unit (CPU) core, an image signal processor (ISP) core, a digital signal processor (DSP) core, a graphics processing unit (GPU) core or other similar special function processing core. In the present embodiment, the special function processing core is coupled to the memory interface 140 via a shared bus (or standard bus) to access the memory array 110 via the memory interface 140).
Regarding claim 25, Huang teaches comprising configuring operations of one or more sets of compute cores in the plurality of processing regions based on one or more neural network models ([0036]: the artificial intelligence core 130 may be pre-designed to have functions and characteristics for performing a specific neural network operation. In other words, the memory 100 of the present embodiment has a function of performing an artificial intelligence operation, and the artificial intelligence core 130 and the external special function processing core can simultaneously access the memory array 110, so as to provide highly efficient data access and operation effects).
Regarding claim 27, Huang teaches wherein the PU comprises a memory processing unit (MPU) (see Fig.3).
Regarding claim 28, Huang teaches wherein the PU comprises a neural processing unit (NPU) (see Fig.3).
Allowable Subject Matter
Claims 11-14, 16-18, 21-24 and 26 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is an examiner’s statement of reasons for allowance:
Claim 11 identifies the distinct features “wherein the near memory (M) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective near memory (M) compute core, to fetch data from the second memory or an adjacent one of a sequence of the plurality of compute cores in a respective processing region, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective near memory (M) compute core; a multiply-and-accumulate array unit configurable to perform computations and pre-channel and bias scaling; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or the other adjacent one of the sequence of the plurality of compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, and chain directions and interfaces of the fetch unit and writeback units to ports of the respective near memory (M) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 12 identifies the distinct features “wherein the arithmetic (A) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective arithmetic (A) compute core, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective arithmetic (A) compute core; arithmetic unit configurable to perform computations; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or an adjacent one of the sequence of the plurality of compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, chain directions and interfaces of the fetch unit and writeback units to ports of the respective arithmetic (A) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 13 identifies the distinct features “wherein the one or more input/output (I/O) cores include an input (I) core comprising: an input port configured to fetch data into the memory processing unit and triggers a writeback unit; the writeback unit configured to write data to an adjacent one of the plurality of regions of the first memory and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to connect the writeback unit to the adjacent one of the plurality of regions of the first memory based on configuration information", which are not taught or suggested by the prior art of records.
Claim 14 identifies the distinct features “wherein the one or more input/output (I/O) cores include an output (O) core comprising: a fetch unit configured to fetch data from an adjacent one of the plurality of region of the first memory and trigger an inter-layer-communication (ILC) unit; an output unit configured to output data out of the memory processing unit; and a switch unit configured to connect the fetch unit to the adjacent one of the plurality of regions of the first memory and the inter-layer-communication (ILC) unit based on configuration information", which are not taught or suggested by the prior art of records.
Claim 16 identifies the distinct features “a second memory coupled to the plurality of processing regions, wherein the second memory comprises a plurality of memory macros and wherein organization and storage of a weight array in a given one of the plurality of memory macros comprises: quantizing the weight array; unrolling each filter of the quantized weight array and append bias and exponent entries; reshaping the unrolled and appended filters to fit into corresponding physical channels; rotating the reshaped filters; and loading virtual channels of the rotated filters into physical channels of the given one of the memory macros; and wherein the first memory comprises an activation memory or feature memory", which are not taught or suggested by the prior art of records.
Claims 17 and 18, which respectively depend on objected-to claim 16, are allowable for at least the same reasons as claim 16.
Claim 21 identifies the distinct features “wherein the near memory (M) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective near memory (M) compute core, to fetch data from the second memory or an adjacent one of a sequence of the compute cores in a respective processing region, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective near memory (M) compute core; a multiply-and-accumulate array unit configurable to perform computations and pre-channel and bias scaling; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or the other adjacent one of the sequence of the compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, and chain directions and interfaces of the fetch unit and writeback units to ports of the respective near memory (M) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 22 identifies the distinct features “wherein the arithmetic (A) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective arithmetic (A) compute core, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective arithmetic (A) compute core; arithmetic unit configurable to perform computations; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or an adjacent one of the sequence of the compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, chain directions and interfaces of the fetch unit and writeback units to ports of the respective arithmetic (A) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 23 identifies the distinct features “wherein the one or more input/output (I/O) cores include an input (I) core comprising: an input port configured to fetch data into the memory processing unit and triggers a writeback unit; the writeback unit configured to write data to an adjacent one of the plurality of regions of the first memory and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to connect the writeback unit to the adjacent one of the plurality of regions of the first memory based on configuration information", which are not taught or suggested by the prior art of records.
Claim 24 identifies the distinct features “wherein the one or more input/output (I/O) cores include an output (O) core comprising: a fetch unit configured to fetch data from an adjacent one of the plurality of region of the first memory and trigger an inter-layer-communication (ILC) unit; an output unit configured to output data out of the memory processing unit; and a switch unit configured to connect the fetch unit to the adjacent one of the plurality of regions of the first memory and the inter-layer-communication (ILC) unit based on configuration information", which are not taught or suggested by the prior art of records.
Claim 26 identifies the distinct features “configuring dataflows including: core-to-core dataflow between adjacent compute cores in respective ones of the plurality of processing regions; memory-to-core dataflow from respective ones of the plurality of regions of the first memory to one or more cores within an adjacent one of the plurality of processing regions; core-to-memory dataflow from one or more cores within ones of the plurality of processing regions to an adjacent one of the plurality of regions of the first memory; and memory-to-core dataflow from the second memory to one or more cores of corresponding ones of the plurality of processing regions", which are not taught or suggested by the prior art of records.
Claims 11-14, 16-18, 21-24 and 26 would be allowable over the prior art of record because the claimed features as mentioned above in combination with other claimed features are not recited or suggested by the prior art of records.
Response to Arguments
Applicant’s arguments filed on 12/29/2025 have been fully considered but they are not persuasive.
1st Point of Argument
Regarding Applicant’s remarks on page 12, the applicants argue Huang nowhere describes an MPU architecture having "a plurality of processing regions interleaved between the plurality of regions of the first memory" nor does Huang describe "respective processing regions ... coupled to adjacent ones of the plurality of regions of the first memory," as described by Applicant's independent claim 1.
In response, Huang discloses that the plurality of memory regions are selectively assigned and accessed by different processing cores based on memory mode settings ([0035]). The claim does not require a specific physical arrangement of processing regions between memory regions. Under a broadest reasonable interpretation, Huang’s selective assignment and access of memory regions by different processing cores reasonably satisfies the claimed interleaving.
However, if Applicant amends the claims to explicitly recite that the plurality of processing regions are physically interleaved between the plurality of regions of the first memory, the rejection would be overcome.
2nd Point of Argument
Regarding Applicant’s remarks on page 12, the applicants argue Huang does not disclose a plurality of "processing regions" arranged between memory regions in an interleaved configuration, nor does Huang disclose a structural arrangement in which each processing region is coupled to adjacent memory regions of a first memory.
In response, Huang discloses that such processing cores are coupled to and access the memory regions via the memory interface based on mode settings ([0035] and [0041]). The claimed invention does not require a specific structural arrangement beyond the recited functional relationships.
Again, however, if Applicant amends the claims to explicitly recite that the plurality of processing regions are physically interleaved between the plurality of regions of the first memory, the rejection would be overcome.
3rd Point of Argument
Regarding Applicant’s remarks on page 13, the applicants argue that Huang does not disclose a processing region (as interleaved between memory regions) that includes a plurality of compute cores including I/O cores and a plurality of near memory compute cores as claimed, nor does Huang disclose such elements arranged in the claimed interleaved architecture.
In response, Huang disclose that such cores access the memory array via the memory interface and shared bus ([0041]), which corresponds to the claimed input/output (I/O) cores configured to access input and output ports of the MPU.
Again, if Applicant amends the claims to explicitly recite that the plurality of processing regions are physically interleaved between the plurality of regions of the first memory, the rejection would be overcome.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Troia (Patent No.: US 11,573,705) “Artificial Intelligence Accelerator”
Considered for teachings related generally to memory devices, and more particularly, to apparatuses and methods for memory with an artificial intelligence (AI) accelerator.
Does not disclose or suggest a first memory including a plurality of regions; and a plurality of processing regions interleaved between the plurality of regions of the first memory, wherein respective processing regions are coupled to adjacent ones of the plurality of regions of the first memory, and wherein one or more of the plurality of processing regions include a plurality of compute cores comprising; one or more input/output (1/O) cores configured to access input and output ports of the MPU.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication should be directed to Yong Choe at telephone number 571-270-1053 or email to yong.choe@uspto.gov. The examiner can normally be reached on M-F 8:00am to 5:00pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Rutz, Jared Ian can be reached on (571) 272-5535. Any inquiry of a general nature or relating to the status of this application should be directed to the TC 2100 whose telephone number is (571) 272-2100.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PMR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-irect.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/YONG CHOE/
Primary Examiner, Art Unit 2135