Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 07/17/2026 has been entered.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-10, 15, 19, 20, 25, 27 and 28 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-10, 15, 19, 20, 25, 27 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Huang et al. (Pub. No.: US 2020/0117597) in view of Isomura et al. (US Patent 5,898,636).
Regarding independent claims 1, 15 and 19, Huang discloses a memory processing unit (MPU) (Fig.3) comprising: a first memory (Fig.3: memory 411) including a plurality of regions (Fig.3: Memory region); and a plurality of processing regions (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 and artificial intelligence core 430) interleaved between the plurality of regions of the first memory (Fig.3: Memory region), wherein respective processing regions (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 or artificial intelligence core 430) are coupled to adjacent ones of the plurality of regions of the first memory (Fig.3: Memory region), and wherein one or more of the plurality of processing regions (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 and artificial intelligence core 430) include a plurality of compute cores (Fig.3: central processing unit core 351, graphics processing unit core 352, digital signal processor core 353 or artificial intelligence core 430) ([0041]: Referring to FIG. 3, a memory 400 of the present embodiment includes memory region 411, 413, 415 and 417, row buffer blocks 412, 414, 416 and 418, a mode register 420, an artificial intelligence core 430 and a memory interface 440. In the present embodiment, the mode register 420 is coupled to the artificial intelligence core 430 and the memory interface 440 to provide a plurality of memory mode settings to the artificial intelligence core 430 and the memory interface 440, respectively. The memory interface 440 is coupled to the central processing unit core 351, the graphics processing unit core 352 and the digital signal processor core 353 via, for example, a shared bus. In the present embodiment, the artificial intelligence core 430 and the memory interface 440 each operate independently to access a memory array. The memory array includes the memory regions 411, 413, 415 and 417 as well as the row buffer blocks 412, 414, 416 and 418, and the memory regions 411, 413, 415 and 417 each include a plurality of memory banks) comprising:
one or more input/output (I/O) cores (Fig.3: ) configured to access input and output ports of the MPU ([0037]: In the present embodiment, the special function processing core may be, for example, a central processing unit (CPU) core, an image signal processor (ISP) core, a digital signal processor (DSP) core, a graphics processing unit (GPU) core or other similar special function processing core. In the present embodiment, the special function processing core is coupled to the memory interface 140 via a shared bus (or standard bus) to access the memory array 110 via the memory interface 140); and
a plurality of near memory (M) compute cores configured to compute neural network functions ([0036]: the artificial intelligence core 130 may be pre-designed to have functions and characteristics for performing a specific neural network operation. In other words, the memory 100 of the present embodiment has a function of performing an artificial intelligence operation, and the artificial intelligence core 130 and the external special function processing core can simultaneously access the memory array 110, so as to provide highly efficient data access and operation effects).
However, Huang does not specifically teach a plurality of processing regions physically interleaved between the plurality of regions of the first memory.
Isomura teaches a plurality of processing regions (Fig.1: logic circuit blocks 8 and 10) physically interleaved between the plurality of regions (Fig.1: memory blocks 7 and 9) of the first memory (Col.2, lines 22-29: the memory block can be set at any location for the logic circuit block to allow an efficient layout design. In addition, the constitution in which the logic circuit block is located between a pair of memory blocks provides a shortest possible data path between them, thereby increasing a processing speed of the semiconductor chip and Col.4, lines 25-44: The semiconductor integrated circuit device containing memory blocks according to the present invention comprises a first logic circuit block 10 composed of gate arrays, a pair of memory blocks 7 and 9, and a second logic circuit block 8 formed in an area between the above-mentioned memory blocks).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to physically arrange Huang’s processing regions between adjacent memory regions in the manner taught by Isomura in order to shorten writing and signal propagation distances, thereby improving processing speed.
Regarding claim 2, Huang teaches wherein the plurality of regions of first memory are columnal interleaved between the plurality of processing regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 3, Huang teaches comprising: a second memory (Fig.3: memory region 413) coupled to the plurality of processing regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 4, Huang teaches wherein the second memory is configurably couplable to one or more near memory (M) compute cores in one or more of the plurality of processing regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 5, Huang teaches wherein the compute cores of one or more of the plurality of processing regions further comprises: one or more arithmetic (A) compute cores configured to compute arithmetic operations, wherein the one or more arithmetic (A) compute cores of each of one or more of the plurality of processing regions are communicatively coupled to adjacent ones of the first plurality of memory regions ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 6, Huang teaches wherein the one or more input/output (I/O) cores comprises: a first input/output (I/O) core configured to stream data into one of the plurality of regions of the first memory; and a second input/output (I/O) core configured to stream data out of another of the plurality of region of the first memory ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 7, Huang teaches wherein the near memory (M) compute cores include a plurality of physical channels configurable to perform computations simultaneously (see Fig.3).
Regarding claim 8, Huang teaches wherein the near memory (M) compute cores of respective ones of the plurality of processing regions are associated with one or more blocks of the second memory ([0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 9, Huang teaches wherein the near memory (M) compute cores include a plurality of physical channels configurable to perform computations simultaneously, and wherein the physical channels of the near memory (M) compute cores are associated with respective slices of the second memory (Fig.3 and [0035]: the plurality of memory regions are respectively selectively assigned to the special function processing core or the artificial intelligence core 130 according to a plurality of memory mode settings stored in the mode register 120).
Regarding claim 10, Huang teaches wherein the near memory (M) compute cores include a plurality of configurable virtual channels (Fig.3 and [0037]: In the present embodiment, the special function processing core may be, for example, a central processing unit (CPU) core, an image signal processor (ISP) core, a digital signal processor (DSP) core, a graphics processing unit (GPU) core or other similar special function processing core. In the present embodiment, the special function processing core is coupled to the memory interface 140 via a shared bus (or standard bus) to access the memory array 110 via the memory interface 140).
Regarding claim 20, Huang teaches comprising: a second memory configurably couplable to one or more near memory (M) compute cores in one or more of the plurality of processing regions (Fig.3 and [0037]: In the present embodiment, the special function processing core may be, for example, a central processing unit (CPU) core, an image signal processor (ISP) core, a digital signal processor (DSP) core, a graphics processing unit (GPU) core or other similar special function processing core. In the present embodiment, the special function processing core is coupled to the memory interface 140 via a shared bus (or standard bus) to access the memory array 110 via the memory interface 140).
Regarding claim 25, Huang teaches comprising configuring operations of one or more sets of compute cores in the plurality of processing regions based on one or more neural network models ([0036]: the artificial intelligence core 130 may be pre-designed to have functions and characteristics for performing a specific neural network operation. In other words, the memory 100 of the present embodiment has a function of performing an artificial intelligence operation, and the artificial intelligence core 130 and the external special function processing core can simultaneously access the memory array 110, so as to provide highly efficient data access and operation effects).
Regarding claim 27, Huang teaches wherein the PU comprises a memory processing unit (MPU) (see Fig.3).
Regarding claim 28, Huang teaches wherein the PU comprises a neural processing unit (NPU) (see Fig.3).
Allowable Subject Matter
Claims 11-14, 16-18, 21-24 and 26 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is an examiner’s statement of reasons for allowance:
Claim 11 identifies the distinct features “wherein the near memory (M) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective near memory (M) compute core, to fetch data from the second memory or an adjacent one of a sequence of the plurality of compute cores in a respective processing region, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective near memory (M) compute core; a multiply-and-accumulate array unit configurable to perform computations and pre-channel and bias scaling; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or the other adjacent one of the sequence of the plurality of compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, and chain directions and interfaces of the fetch unit and writeback units to ports of the respective near memory (M) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 12 identifies the distinct features “wherein the arithmetic (A) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective arithmetic (A) compute core, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective arithmetic (A) compute core; arithmetic unit configurable to perform computations; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or an adjacent one of the sequence of the plurality of compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, chain directions and interfaces of the fetch unit and writeback units to ports of the respective arithmetic (A) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 13 identifies the distinct features “wherein the one or more input/output (I/O) cores include an input (I) core comprising: an input port configured to fetch data into the memory processing unit and triggers a writeback unit; the writeback unit configured to write data to an adjacent one of the plurality of regions of the first memory and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to connect the writeback unit to the adjacent one of the plurality of regions of the first memory based on configuration information", which are not taught or suggested by the prior art of records.
Claim 14 identifies the distinct features “wherein the one or more input/output (I/O) cores include an output (O) core comprising: a fetch unit configured to fetch data from an adjacent one of the plurality of region of the first memory and trigger an inter-layer-communication (ILC) unit; an output unit configured to output data out of the memory processing unit; and a switch unit configured to connect the fetch unit to the adjacent one of the plurality of regions of the first memory and the inter-layer-communication (ILC) unit based on configuration information", which are not taught or suggested by the prior art of records.
Claim 16 identifies the distinct features “a second memory coupled to the plurality of processing regions, wherein the second memory comprises a plurality of memory macros and wherein organization and storage of a weight array in a given one of the plurality of memory macros comprises: quantizing the weight array; unrolling each filter of the quantized weight array and append bias and exponent entries; reshaping the unrolled and appended filters to fit into corresponding physical channels; rotating the reshaped filters; and loading virtual channels of the rotated filters into physical channels of the given one of the memory macros; and wherein the first memory comprises an activation memory or feature memory", which are not taught or suggested by the prior art of records.
Claims 17 and 18, which respectively depend on objected-to claim 16, are allowable for at least the same reasons as claim 16.
Claim 21 identifies the distinct features “wherein the near memory (M) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective near memory (M) compute core, to fetch data from the second memory or an adjacent one of a sequence of the compute cores in a respective processing region, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective near memory (M) compute core; a multiply-and-accumulate array unit configurable to perform computations and pre-channel and bias scaling; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or the other adjacent one of the sequence of the compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, and chain directions and interfaces of the fetch unit and writeback units to ports of the respective near memory (M) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 22 identifies the distinct features “wherein the arithmetic (A) compute cores comprise: a fetch unit configurable to control an operation sequence of the respective arithmetic (A) compute core, to fetch data from an adjacent one of the plurality of regions of the first memory, decrement an inter-layer-communication (ILC) counter, and trigger other units of the respective arithmetic (A) compute core; arithmetic unit configurable to perform computations; a writeback unit configurable to perform a fuse operation, send data to an other adjacent one of the plurality of regions of the first memory or an adjacent one of the sequence of the compute cores in the respective processing region, and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to configure memory accesses, chain directions and interfaces of the fetch unit and writeback units to ports of the respective arithmetic (A) compute core based on configuration information", which are not taught or suggested by the prior art of records.
Claim 23 identifies the distinct features “wherein the one or more input/output (I/O) cores include an input (I) core comprising: an input port configured to fetch data into the memory processing unit and triggers a writeback unit; the writeback unit configured to write data to an adjacent one of the plurality of regions of the first memory and to increment an inter-layer-communication (ILC) counter; and a switch unit configured to connect the writeback unit to the adjacent one of the plurality of regions of the first memory based on configuration information", which are not taught or suggested by the prior art of records.
Claim 24 identifies the distinct features “wherein the one or more input/output (I/O) cores include an output (O) core comprising: a fetch unit configured to fetch data from an adjacent one of the plurality of region of the first memory and trigger an inter-layer-communication (ILC) unit; an output unit configured to output data out of the memory processing unit; and a switch unit configured to connect the fetch unit to the adjacent one of the plurality of regions of the first memory and the inter-layer-communication (ILC) unit based on configuration information", which are not taught or suggested by the prior art of records.
Claim 26 identifies the distinct features “configuring dataflows including: core-to-core dataflow between adjacent compute cores in respective ones of the plurality of processing regions; memory-to-core dataflow from respective ones of the plurality of regions of the first memory to one or more cores within an adjacent one of the plurality of processing regions; core-to-memory dataflow from one or more cores within ones of the plurality of processing regions to an adjacent one of the plurality of regions of the first memory; and memory-to-core dataflow from the second memory to one or more cores of corresponding ones of the plurality of processing regions", which are not taught or suggested by the prior art of records.
Claims 11-14, 16-18, 21-24 and 26 would be allowable over the prior art of record because the claimed features as mentioned above in combination with other claimed features are not recited or suggested by the prior art of records.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Troia (Patent No.: US 11,573,705) “Artificial Intelligence Accelerator”
Considered for teachings related generally to memory devices, and more particularly, to apparatuses and methods for memory with an artificial intelligence (AI) accelerator.
Does not disclose or suggest a first memory including a plurality of regions; and a plurality of processing regions interleaved between the plurality of regions of the first memory, wherein respective processing regions are coupled to adjacent ones of the plurality of regions of the first memory, and wherein one or more of the plurality of processing regions include a plurality of compute cores comprising; one or more input/output (1/O) cores configured to access input and output ports of the MPU.
Any inquiry concerning this comm1unication should be directed to Yong Choe at telephone number 571-270-1053 or email to yong.choe@uspto.gov. The examiner can normally be reached on M-F 9:30am to 6:00pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Rutz, Jared Ian can be reached on (571) 272-5535. Any inquiry of a general nature or relating to the status of this application should be directed to the TC 2100 whose telephone number is (571) 272-2100.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PMR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-irect.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/YONG CHOE/
Primary Examiner, Art Unit 2135