Prosecution Insights
Last updated: August 17, 2026
Application No. 18/619,750

Host Accesses to Processing-in-Memory Oriented Data Structures

Non-Final OA §103§112
Filed
Mar 28, 2024
Examiner
NGO, THANH
Art Unit
Tech Center
Assignee
Advanced Micro Devices Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
6 currently pending
Career history
1
Total Applications
across all art units

Statute-Specific Performance

§101
21.1%
-18.9% vs TC avg
§103
52.6%
+12.6% vs TC avg
§112
26.3%
-13.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§103 §112
DETAILED ACTION This communication is in response to the application filed on March 28, 2024 in which claims 1-20 are pending in the application. Claims 1, 13, and 19 are in independent form. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings are objected to because FIG. 3 inaccurately depicts the layout 310. In the second bank grouping of the layout 310, the second row shows the matrix element designation "3,2" in both the column grouping 302c position and the column grouping 302d position. Consistent with the matrix 304 and the pattern of the remaining groupings (compare "3,0" and "3,1" in the first grouping and "3,4" and "3,5" in the third grouping), the entry in the column grouping 302d position should read "3,3," and element 3,3 of the matrix 304 is otherwise not shown in the layout 310. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification The abstract of the disclosure is objected to because it refers to the same element inconsistently. The abstract introduces "a host processing unit" but thereafter refers to "the host processor." The specification and claims consistently recite "host processor". A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). Claim Objections Claim 16 is objected to because of the following informality: As per claim 16, the claim recites "the one or more banks of the multiple processing-in-memory units," whereas claim 13, from which claim 16 depends, introduces "one or more banks of the memory." For consistent terminology, the recitation should read "the one or more banks of the memory." Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 13-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 13 recites the limitation "respective lanes of t" in the line of "store elements of a matrix . There is insufficient antecedent basis for this limitation in the claim. The processing-in-memory units of claim 13 are not recited as having lanes. Reciting the units as each having multiple lanes, as claim 11 does, would resolve the ambiguity. For examination purposes, “respective lanes” is interpreted as if the multiple processing-in-memory units were recited as each having multiple lanes, consistent with the specification. Claims 14-18, which are dependent on claim 13, are similarly rejected. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop et al. (US 2022/0206839 A1) (hereinafter Alsop) in view of Song (US 2021/0208801 A1) (hereinafter Song). As per claim 1, Alsop primarily teaches the invention as claimed including: a computing device, comprising: a memory ([0022] the memory modules 120 include N number of PIM-enabled memory elements, where each PIM-enabled memory element includes one or more banks); multiple processing-in-memory units, each processing-in-memory unit configured to access one or more banks of the memory ([0022] each PIM-enabled memory element includes one or more banks and a corresponding PIM execution unit configured to execute PIM commands and [0025] each of the N number of partitions contains compute task data for a particular PIM-enabled memory module and a corresponding PIM execution unit); a host processor configured to: receive an access request to access an element of a data structure stored in the memory ([0023] the microprocessor 110 includes threads 130 and a tasking mechanism 140; [0019] in operation, threads push compute task data to the AMAT mechanism before the compute tasks are executed; [0030] the thread determines one or more next operations to be performed on data stored at a particular memory address and generates compute task data that specifies the operation(s), a memory address, and a value); the access request including input parameters indicating a processing-in-memory unit of the multiple processing-in-memory units by which the element is accessible ([0017] the term "compute task data" refers to data that specifies one or more input parameters of such a compute task and may include, for example, one or more operations to be executed, one or more addresses or index values; [0018] the AMAT partition corresponding to the memory module containing data that will be processed by the task according to an input index parameter; [0026] the partition ID for a given compute task may be determined by some function of a subset of the address bits in the compute task data, specified by the address mapping data); and an offset of the element relative to other elements of the data structure ([0034] a base address and index offset, rather than a full address, may be supplied in the address information); Alsop does not explicitly teach: generate a memory address for the access request based on the processing-in-memory unit and the offset, the element of the data structure being accessed based on the memory address. However, Song teaches: generate a memory address for the access request based on the processing-in-memory unit and the offset ([0232] the address generator 3220 may receive a base address ADDR_B and an offset signal OFFSET from the host 3300. The address generator 3220 may generate a restored address ADDR_RE having a restored address map state with change of a column address in the base address ADDR_B in response to the offset signal OFFSET and [0236] the base address ADDR_B may include a rank address, a row address, a bank address, a column address, a channel address, a burst length, and so on), the element of the data structure being accessed based on the memory address ([0240] the weight data transmitted from the memory banks BK0~BK15 to the MAC operators MAC0~MAC15 may be selected by the second column address CA2 included in the restored address ADDR_RE). Alsop and Song are both concerned with processing-in-memory devices that perform arithmetic on bank-resident data and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song because it would provide for an address generator that builds each memory address from a base address and an offset supplied with the request. Motivation would improve system performance and memory coherency by minimizing an overlap phenomenon of memory regions of the memory operations in the PIM device as taught by Song ([0274]). As per claim 2, Song further teaches wherein the processing-in-memory unit ([0236] a bank address, a column address, a channel address; [0066] the MAC operators MAC0, . . . , and MAC7 may be disposed such that one of the even-numbered memory banks BK0, BK2, . . . , and BK14 and one of the odd-numbered memory banks BK1, BK3, . . . , and BK15 share any one of the MAC operators MAC0, . . . , and MAC7 with each other) and the offset ([0232] an offset signal OFFSET) are specified directly via the input parameters ([0232] the address generator 3220 may receive a base address ADDR_B and an offset signal OFFSET from the host 3300). As per claim 3, Song further teaches wherein the offset further indicates a particular bank of the one or more banks that the processing-in-memory unit is configured to access ([0067] the first memory bank BK0, the second memory bank BK1, and the first MAC operator MAC0 between the first memory bank BK0 and the second memory bank BK1 may constitute a first MAC unit; [0077] the address latch 260 may convert the input address I_ADDR output from the receiving driver 230 into a bank selection signal BK_S and a row/column address ADDR_R/ADDR_C; [0078] the second MAC command signal may be a first MAC read signal MAC_RD_BK0, the third MAC command signal may be a second MAC read signal MAC_RD_BK1). As per claim 10, Alsop further teaches wherein the host processor is configured to store elements of the data structure in the memory in a layout ([0038] it may be necessary to issue multiple store commands to non-contiguous strides of memory in order to ensure they map to the same memory element), the layout including interacting elements of the data structure stored at locations in the memory that are local to respective processing-in-memory units of the multiple processing-in-memory units ([0020] the AMAT mechanism ensures that all of the data that will be accessed during the processing of a PIM task are located in the same memory element and [0003] in order to operate on multiple elements of memory data in a single bank-local PIM execution unit, for example to perform reduction of elements in an array, all the memory addresses of the target operands must map to the same physical memory bank). Claim(s) 4 and 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Islam et al. (US 2022/0066662 A1) (hereinafter Islam). As per claim 4, Alsop in view of Song disclose the claimed invention as detailed above for claim 1. Song further teaches wherein the host processor is configured to generate the memory address using a physical address map ([0236] it may be assumed that the base address ADDR_B includes a row address RA, a bank address BA, a column address CA, and a channel address CHA and the base address ADDR_B is mapped in order of the row address RA, the bank address BA, the column address CA, and the channel address CHA and [0239] the restored address ADDR_RE having the same address map state as the base address ADDR_B). Alsop in view of Song do not explicitly teach: the physical address map including one or more mappings that assign bit positions of the memory address to corresponding components of the memory. However, Islam teaches: the physical address map including one or more mappings that assign bit positions of the memory address to corresponding components of the memory ([0034] FIG. 3 depicts how physical address bits are mapped for indexing inside the PIM-enabled memory using address interleaving memory mapping. For example, bit 0 represents a channel number. Bits 1 and 2 represent a bank number. Bits 3-5 represent a column number, Bits 6-7 represent a row number; [0042] the same bits of physical address as depicted in FIG. 3 are used for row and column addressing, but to decode the bits, different equations are utilized by a memory controller; [0028] memory controller 102 includes mapping logic 110 that is configured to manage the storage and access of data elements in memory structure 108). Alsop, Song, and Islam are all concerned with physical address mapping for processing-in-memory operands and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Islam because it would provide for a physical address map that assigns designated bits of the address to the channel, the bank, the column, and the row of the memory, so that elements accessed together land in the same row of one bank. Motivation would achieve superior PIM performance and energy efficiency since the mapping reduces the number of row activations as taught by Islam ([0025]). As per claim 5, Islam further teaches wherein the corresponding components of the memory include memory channels ([0034] bit 0 represents a channel number), the multiple processing-in-memory units ([0032] each bank is coupled to a separate PIM execution unit), banks of the memory ([0034] bits 1 and 2 represent a bank number), rows of the banks ([0034] bits 6-7 represent a row number), and columns of the banks ([0034] bits 3-5 represent a column number). Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Islam in view of Gupta et al. (US 6,108,745) (hereinafter Gupta). As per claim 6, Alsop in view of Song in view of Islam disclose the claimed invention as detailed above for claims 1 and 4, and further teach wherein the processing-in-memory unit and the offset are indicated by one or more numerical identifiers (Alsop [0026] the partition ID for a given compute task may be determined by some function of a subset of the address bits in the compute task data and [0034] a base address and index offset, rather than a full address, may be supplied in the address information; Song [0236] a row address RA, a bank address BA, a column address CA, and a channel address CHA). Alsop in view of Song in view of Islam do not explicitly teach: to generate the memory address, the host processor is configured to route source bits of the one or more numerical identifiers to the bit positions of the memory address in accordance with a routing protocol corresponding to a mapping of the physical address map. However, Gupta teaches: to generate the memory address, the host processor is configured to route source bits of the one or more numerical identifiers to the bit positions of the memory address in accordance with a routing protocol corresponding to a mapping of the physical address map (abstract any address bit provided by the processor can be routed to any bank, row, or column bit, and can be used to generate any rank bit; col. 9:1-2 FIG. 6 shows address bit routing circuitry 48, which can route any address bit to any bank, row, or column bit; col. 9:24-26 if rank bit 0 is "1", then address bit 0 is routed to the output of multiplexer 52 indicated by the routing stored in register 50). Alsop, Song, Islam, and Gupta are all concerned with mapping processor addresses onto the bank, row, and column organization of memory and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Islam in view of Gupta because it would provide for routing circuitry that can route any address bit from the processor to any bank, row, or column bit. Motivation would improve system adaptability because the routing supports a variety of memory sizes and interleaving schemes as taught by Gupta (Abstract). Claim(s) 7-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Islam in view of Gupta in view of Yamazaki et al. (US 5,940,342) (hereinafter Yamazaki). As per claim 7, Alsop in view of Song in view of Islam in view of Gupta disclose the claimed invention as detailed above for claims 1, 4, and 6 but do not explicitly teach: wherein the routing protocol is hardwired into the host processor. However, Yamazaki teaches: wherein the routing protocol is hardwired into the host processor (col. 31:33-35 this address conversion can be implemented not only by wiring but also by employing a shifter array, as described later in detail and col. 32:38-60 dissimilarly to a structure of fixedly deciding the address conversion mode by interconnections, flexible address conversion can be implemented through the shifter array 116), the wiring implementation fixing the conversion in the interconnections. Alsop, Song, Islam, Gupta, and Yamazaki are all concerned with converting processor addresses into bank, row, and column addresses of a multi-bank memory and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Islam in view of Gupta in view of Yamazaki because it would provide for an address conversion that is fixed in the wiring of the host. Motivation would improve efficiency of the address conversion because a mapping fixed in the wiring converts the address with no additional circuitry, delay, or control state. As per claim 8, Alsop in view of Song in view of Islam in view of Gupta disclose the claimed invention as detailed above for claims 1, 4, and 6 but do not explicitly teach: wherein the routing protocol is implemented by barrel shifters of the host processor that are reconfigurable to account for different mappings of the physical address map. However, Yamazaki teaches: wherein the routing protocol is implemented by barrel shifters of the host processor (col. 30:36-43 an address conversion part includes a barrel shifter 100a for receiving the CPU address bits A21 to A9 and shifting the same toward the least significant bit by a prescribed number of bits (four bits), and a buffer register 100b for receiving and storing the CPU address bits A8 to A0. The barrel shifter 100a outputs the bank address bits BA3 to BA0 and the row address bits RA8 to RA0, and the buffer register 100b outputs the column address bits CA8 to CA0; col. 32:38-60 the address conversion part includes a shift data register 110 for storing shift bit number information ... a decoder 114 for decoding the shift bit number information stored in the shift data register 110 ... and generating a shift control signal, and a shifter array 116 for shifting the CPU address bits A21 to A0 in accordance with the shift control signal from the decoder 114 and generating the memory controller addresses (DRAM addresses) BA3 to BA0, RA8 to RA0 and CA8 to CA0) that are reconfigurable to account for different mappings of the physical address map (col. 30:52-54 it is possible to readily cope with an arbitrary number of banks by utilizing the barrel shifter 100a and adjusting the shift bit number and col. 32:38-60 dissimilarly to a structure of fixedly deciding the address conversion mode by interconnections, flexible address conversion can be implemented through the shifter array 116 depending on application). Alsop, Song, Islam, Gupta, and Yamazaki are all concerned with converting processor addresses into bank, row, and column addresses of a multi-bank memory and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Islam in view of Gupta in view of Yamazaki because it would provide for a barrel shifter that converts the processor address bits into the bank and row address bits, with the shift amount held in a shift data register rather than fixed in the wiring. Motivation would improve adaptability of the conversion hardware because it can readily cope with an arbitrary number of banks by adjusting the shift bit number as taught by Yamazaki (col. 30:52-54). As per claim 9, the combination of references above teaches wherein the access request is received as part of a workload (Islam [0080] the memory controller recognizes the NoS information that is included with each such memory request and decodes the physical memory address per IBFS or ICFS), and to generate the memory address, the host processor is configured to: receive an indication of the mapping of the physical address map associated with the workload (Islam [0075] NoS is provided statically as per user preference at system startup time (e.g. via the basic input/output system) or dynamically to achieve flexibility and [0080] the NoS information for each data structure/page must be tracked and communicated along with a physical memory address for any memory access request to the memory controller; Song [0244] the memory mode address ADDR_MEM may be mapped in order of "row/bank/column/channel" and [0246] the MAC mode address ADDR_MAC may be mapped in order of "channel/bank/row/column"); update machine status registers of the host processor to specify the routing protocol corresponding to the mapping (Islam [0081] four possible approaches are utilized: (i) Instruction-based approach and (ii) Page Table Entry (PTE)-based approach, (iii) configuration register approach, and (iv) a mode register based approach; Gupta col. 9:12-13 register 50 is configured when the computer system is initialized, as described above and col. 9:24-26 the routing stored in register 50); and reconfigure the barrel shifters to implement the routing protocol as specified by the machine status registers (Yamazaki col. 33:29-38 the control signals φ0 to φ3 are outputted from the decoder 114 in accordance with the information stored in the shift data register 110 and the mask data register 112 shown in FIG. 32). Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Seo et al. (US 2021/0110876 A1) (hereinafter Seo). As per claim 11, Alsop in view of Song disclose the claimed invention as detailed above for claims 1 and 10 but do not explicitly teach: wherein the multiple processing-in-memory units correspond to single instruction, multiple data processing-in-memory units each having multiple lanes, the layout further including the interacting elements of the data structure stored at the locations in the memory that map to respective lanes of the multiple processing-in-memory units. However, Seo teaches: wherein the multiple processing-in-memory units correspond to single instruction, multiple data processing-in-memory units each having multiple lanes ([0058] the 16 bits of data corresponding to the 16 DRAM cells may correspond to each element of the matrix, and may be transferred to a corresponding MAC 42 and used as an operand and the PIM operator 40 may include a command generator unit 47 that receives a DRAM command and an address transmitted from a control circuit ... and then converts the command into more detailed subcommands), the parallel MACs 42 operating on respective elements under subcommands of a single received command; the layout further including the interacting elements of the data structure stored at the locations in the memory that map to respective lanes of the multiple processing-in-memory units ([0058] a DRAM array may be configured to store 256 bits of data corresponding to at least one row of a matrix, and a GIO SA 41 may receive 256 bits of read data and may be transferred to a corresponding MAC 42 and used as an operand). Alsop, Song, and Seo are all concerned with arithmetic performed by processing-in-memory operators at memory banks and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Seo because it would provide for a PIM operator in which one row read supplies every MAC with its own element in the matrix. Motivation would improve throughput of the matrix vector multiplication because each element of the read row is used as an operand by its own MAC as taught by Seo ([0058]). Claim(s) 12-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Seo in view of Gómez-Luna et al., Benchmarking a New Paradigm: An Experimental Analysis of a Real Processing-in-Memory Architecture, arXiv:2105.03814v7 (May 4, 2022) (hereinafter Gómez-Luna). As per claim 12, Alsop in view of Song in view of Seo disclose the claimed invention as detailed above for claims 1, 10, and 11 but do not explicitly teach: wherein the input parameters include element parameters indicating the element of the data structure and layout parameters indicating the layout, and the host processor is further configured to compute the processing-in-memory unit and the offset based on the element parameters and the layout parameters. However, Gómez-Luna teaches: wherein the input parameters include element parameters indicating the element of the data structure (section 4.2, p. 16 the host CPU assigns each set of consecutive rows to a DPU using linear assignment (i.e., set of rows i assigned to DPU i)) and layout parameters indicating the layout (section 4.2, p. 16 our PIM implementation of GEMV partitions the matrix across the DPUs available in the system, assigning a fixed number of consecutive rows to each DPU, while the input vector is replicated across all DPUs), and the host processor is further configured to compute the processing-in-memory unit and the offset based on the element parameters and the layout parameters (section 4.2, p. 16, the linear assignment computing the DPU from the row index and the fixed number of consecutive rows per DPU). Alsop, Song, Seo, and Gómez-Luna are all concerned with distributing data structures across processing-in-memory units for parallel computation and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Seo in view of Gómez-Luna because it would provide for a host that determines which DPU holds an element directly from the element's row and the number of rows per DPU. Motivation would scale performance linearly with the number of DPUs since the row assignment spreads the matrix evenly across the available DPUs as taught by Gómez-Luna (Section 5.1.1, p. 20). As per claim 13, Alsop primarily teaches the invention as claimed including: a system, comprising: a memory module including a memory and multiple processing-in-memory units each configured to access one or more banks of the memory ([0022] the memory modules 120 include N number of PIM-enabled memory elements, where each PIM-enabled memory element includes one or more banks and a corresponding PIM execution unit configured to execute PIM commands); and a host processor communicatively coupled to the memory module, the host processor configured to: ([0023] the microprocessor 110 includes threads 130 and a tasking mechanism 140); receive an access request to access an element of the matrix, the access request including element parameters indicating the element ([0030] the thread determines one or more next operations to be performed on data stored at a particular memory address and generates compute task data that specifies the operation(s), a memory address, and a value and [0017] one or more addresses or index values); the element of the matrix being accessed based on the processing-in-memory unit and the offset ([0045] the tasking manager 148 constructs a fully valid PIM command using the compute task data, e.g., using the address, operation and values specified by the compute task data and the ID of the PIM execution unit that corresponds to the source partition). Alsop does not explicitly teach: store elements of a matrix in the memory in a layout. However, Song teaches: store elements of a matrix in the memory in a layout ([0162] the weight data W1.1~W1.128 arrayed in the first row R1 of the weight matrix may be stored into the first row ROW0 of the first memory bank BK0 in the first channel CH0 and [0159] all of the elements W1.1~W128.128 constituting the weight matrix may be allocated to the channels CH0~CH3 and the memory banks BK0~BK15 in a way that parallelism is applicable to the channels CH0~CH3 and the memory banks BK0~BK15). Alsop and Song are both concerned with processing-in-memory devices that perform arithmetic on bank-resident data and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song because it would provide for weight data stored row by row into the memory banks, with each matrix row placed in a single bank row and the elements allocated across the channels and banks so that parallelism is applicable. Motivation would improve performance of the MAC arithmetic operation by applying parallelism across the channels and the memory banks as taught by Song ([0159]). Alsop in view of Song do not explicitly teach: the layout including interacting elements of the matrix stored at locations in the memory that map to respective lanes of the multiple processing-in-memory units. However, Seo teaches: the layout including interacting elements of the matrix stored at locations in the memory that map to respective lanes of the multiple processing-in-memory units ([0058] the specific circuit structure of the PIM operator 40 in an example in which a semiconductor memory device performs a matrix vector multiplication operation is shown; a DRAM array may be configured to store 256 bits of data corresponding to at least one row of a matrix, and a GIO SA 41 may receive 256 bits of read data; the 16 bits of data corresponding to the 16 DRAM cells may correspond to each element of the matrix, and may be transferred to a corresponding MAC 42 and used as an operand). Alsop, Song, and Seo are all concerned with arithmetic performed by processing-in-memory operators at memory banks and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Seo because it would provide for a PIM operator in which one row read supplies every MAC with its own element in the matrix. Motivation would improve throughput of the matrix vector multiplication because each element of the read row is used as an operand by its own MAC as taught by Seo ([0058]). To identify the processing-in-memory unit for a request, Alsop teaches determining the partition from bits of the supplied memory address ([0026] the partition ID for a given compute task may be determined by some function of a subset of the address bits in the compute task data, specified by the address mapping data). The inputs of that determination are the address bits themselves. The combination does not compute the unit or the offset from parameters describing the element and the layout. Determining the partition from address bits is not computing the processing-in-memory unit and the offset based on the element parameters and the layout parameters. Alsop in view of Song in view of Seo therefore do not explicitly teach: and layout parameters indicating the layout; compute, based on the element parameters and the layout parameters, a processing-in-memory unit of the multiple processing-in-memory units by which the element is accessible and an offset of the element relative to other elements of the matrix. However, Gómez-Luna teaches: and layout parameters indicating the layout (section 4.2, p. 16 our PIM implementation of GEMV partitions the matrix across the DPUs available in the system, assigning a fixed number of consecutive rows to each DPU, while the input vector is replicated across all DPUs); compute, based on the element parameters and the layout parameters, a processing-in-memory unit of the multiple processing-in-memory units by which the element is accessible and an offset of the element relative to other elements of the matrix (section 4.2, p. 16 the host CPU assigns each set of consecutive rows to a DPU using linear assignment (i.e., set of rows i assigned to DPU i)). Alsop, Song, Seo, and Gómez-Luna are all concerned with distributing data structures across processing-in-memory units for parallel computation and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Seo in view of Gómez-Luna because it would provide for a host that determines which DPU holds an element directly from the element's row and the number of rows per DPU. Motivation would scale performance linearly with the number of DPUs since the row assignment spreads the matrix evenly across the available DPUs as taught by Gómez-Luna (Section 5.1.1, p. 20). As per claim 14, the combination of references above teaches wherein the interacting elements include the elements of the matrix that are combinable as part of a reduction computation of a general matrix-vector multiplication operation (Gómez-Luna section 4.2, p. 16 our PIM implementation of GEMV partitions the matrix across the DPUs available in the system; Seo [0058] a semiconductor memory device performs a matrix vector multiplication operation; Song [0095] the MAC circuit 222 may perform a multiplying calculation and an accumulative adding calculation for the first and second data DA1 and DA2). As per claim 15, the combination of references above teaches wherein the element parameters include a row of the matrix (Gómez-Luna section 4.2, p. 16 set of rows i assigned to DPU i; Song [0236] a row address RA, a bank address BA, a column address CA) and a column of the matrix associated with the element (Song [0236] a column address CA and [0238] the set value VAL may be set by the address generator 3220 to correspond to the column address of a position in which the weight data are stored). As per claim 16, the combination of references above teaches wherein the layout parameters include: a first number of bank columns allocated to the matrix in the one or more banks of the multiple processing-in-memory units (Seo [0058] a DRAM array may be configured to store 256 bits of data corresponding to at least one row of a matrix); a second number of lanes included in each of the multiple processing-in-memory units (Seo [0058] the 16 bits of data corresponding to the 16 DRAM cells may correspond to each element of the matrix, and may be transferred to a corresponding MAC 42); a third number of the multiple processing-in-memory units (Gómez-Luna section 4.2, p. 16 partitions the matrix across the DPUs available in the system); a fourth number of matrix columns in the matrix (Gómez-Luna section 4.2, p. 16, the per-DPU row size following from the matrix dimensions; Song [0162] the weight data W1.1~W1.128 arrayed in the first row R1 of the weight matrix may be stored into the first row ROW0 of the first memory bank BK0 in the first channel CH0; Seo [0058], the row of the matrix spanning the bank columns); and a base type size of the elements in the matrix (Seo [0058] the 16 bits of data corresponding to the 16 DRAM cells may correspond to each element of the matrix). Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Seo in view of Gómez-Luna in view of Islam. As per claim 17, Alsop in view of Song in view of Seo in view of Gómez-Luna disclose the claimed invention as detailed above for claim 13. Song further teaches wherein the host processor is further configured to generate a memory address for the access request based on the processing-in-memory unit and the offset ([0232] the address generator 3220 may receive a base address ADDR_B and an offset signal OFFSET from the host 3300. The address generator 3220 may generate a restored address ADDR_RE having a restored address map state with change of a column address in the base address ADDR_B in response to the offset signal OFFSET), the memory address generated using a physical address map ([0236] the base address ADDR_B is mapped in order of the row address RA, the bank address BA, the column address CA, and the channel address CHA), the element of the matrix being accessed based on the memory address ([0240] the weight data transmitted from the memory banks BK0~BK15 to the MAC operators MAC0~MAC15 may be selected by the second column address CA2 included in the restored address ADDR_RE). Alsop in view of Song in view of Seo in view of Gómez-Luna do not explicitly teach: that includes one or more mappings that assign bit positions of the memory address to different components of the memory. However, Islam teaches: one or more mappings that assign bit positions of the memory address to different components of the memory ([0034] FIG. 3 depicts how physical address bits are mapped for indexing inside the PIM-enabled memory using address interleaving memory mapping. For example, bit 0 represents a channel number. Bits 1 and 2 represent a bank number. Bits 3-5 represent a column number, Bits 6-7 represent a row number; [0042] the same bits of physical address as depicted in FIG. 3 are used for row and column addressing, but to decode the bits, different equations are utilized by a memory controller). Alsop, Song, Seo, Gómez-Luna, and Islam are all concerned with placing and addressing the operand data of processing-in-memory units and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Seo in view of Gómez-Luna in view of Islam because it would provide for a physical address map that assigns designated bits of the address to the channel, the bank, the column, and the row of the memory, so that elements accessed together land in the same row of one bank. Motivation would achieve superior PIM performance and energy efficiency since the mapping reduces the number of row activations as taught by Islam ([0025]). Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Song in view of Seo in view of Gómez-Luna in view of Islam in view of Gupta. As per claim 18, Alsop in view of Song in view of Seo in view of Gómez-Luna in view of Islam disclose the claimed invention as detailed above for claims 13 and 17, and further teach wherein the processing-in-memory unit and the offset are indicated by one or more numerical identifiers (Alsop [0026] the partition ID for a given compute task may be determined by some function of a subset of the address bits in the compute task data and [0034] a base address and index offset, rather than a full address, may be supplied in the address information; Song [0236] a row address RA, a bank address BA, a column address CA, and a channel address CHA), in accordance with a routing protocol corresponding to a mapping of the physical address map specified for a workload that includes the access request (Islam [0080] the memory controller recognizes the NoS information that is included with each such memory request and decodes the physical memory address per IBFS or ICFS and [0075] NoS is provided statically as per user preference at system startup time (e.g. via the basic input/output system) or dynamically to achieve flexibility). Alsop in view of Song in view of Seo in view of Gómez-Luna in view of Islam do not explicitly teach: to generate the memory address, the host processor is configured to route source bits of the one or more numerical identifiers to the bit positions of the memory address. However, Gupta teaches: to generate the memory address, the host processor is configured to route source bits of the one or more numerical identifiers to the bit positions of the memory address (abstract any address bit provided by the processor can be routed to any bank, row, or column bit, and can be used to generate any rank bit; col. 9:1-2 FIG. 6 shows address bit routing circuitry 48, which can route any address bit to any bank, row, or column bit; col. 9:24-26 if rank bit 0 is "1", then address bit 0 is routed to the output of multiplexer 52 indicated by the routing stored in register 50). Alsop, Song, Seo, Gómez-Luna, Islam, and Gupta are all concerned with mapping data onto the bank, row, and column organization of memory and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Song in view of Seo in view of Gómez-Luna in view of Islam in view of Gupta because it would provide for routing circuitry that can route any address bit from the processor to any bank, row, or column bit. Motivation would improve system adaptability because the routing supports a variety of memory sizes and interleaving schemes as taught by Gupta (Abstract). Claim(s) 19 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alsop in view of Islam in view of Gupta. As per claim 19, Alsop primarily teaches the invention as claimed including: a method, comprising: receiving, by a host processor, an access request of a workload to access an element of a data structure stored in memory ([0019] in operation, threads push compute task data to the AMAT mechanism before the compute tasks are executed and [0030] the thread determines one or more next operations to be performed on data stored at a particular memory address and generates compute task data that specifies the operation(s), a memory address, and a value); the access request including numerical identifiers of a processing-in-memory unit that is configured to access a bank where the element is stored ([0017] one or more addresses or index values; [0018] the AMAT partition corresponding to the memory module containing data that will be processed by the task according to an input index parameter; [0022] each PIM-enabled memory element includes one or more banks and a corresponding PIM execution unit configured to execute PIM commands; [0026] the partition ID for a given compute task may be determined by some function of a subset of the address bits in the compute task data); and an offset of the element relative to other elements of the data structure ([0034] a base address and index offset, rather than a full address, may be supplied in the address information); and accessing, by the host processor, the element of the data structure based on the memory address ([0045] the tasking manager 148 constructs a fully valid PIM command using the compute task data, e.g., using the address, operation and values specified by the compute task data and the ID of the PIM execution unit that corresponds to the source partition and [0046] PIM commands for different partitions may be generated and issued to their respective PIM execution units in parallel). In response to the access request, Alsop teaches calculating the target address from a base address and an index offset and identifying the PIM execution unit from bits of that address ([0034] the target address is calculated based upon the base address and index offset and [0026] the partition ID for a given compute task may be determined by some function of a subset of the address bits in the compute task data). Alsop's step generates the address arithmetically and derives the unit from the resulting bits but does not describe the generation as a routing of identifier bits to assigned positions. Calculating an address arithmetically and deriving the unit from its bits is not generating the memory address using a routing protocol indicating how source bits of the numerical identifiers are routed to bit positions of the memory address. Alsop therefore does not explicitly teach: generating, by the host processor, a memory address for the access request using a routing protocol indicating how source bits of the numerical identifiers are routed to bit positions of the memory address assigned to respective components of the memory. However, Islam teaches: bit positions of the memory address assigned to respective components of the memory ([0034] FIG. 3 depicts how physical address bits are mapped for indexing inside the PIM-enabled memory using address interleaving memory mapping. For example, bit 0 represents a channel number. Bits 1 and 2 represent a bank number. Bits 3-5 represent a column number, Bits 6-7 represent a row number). Alsop and Islam are both concerned with physical address mapping for processing-in-memory operands and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Islam because it would provide for a physical address map that assigns designated bits of the address to the channel, the bank, the column, and the row of the memory, so that elements accessed together land in the same row of one bank. Motivation would achieve superior PIM performance and energy efficiency since the mapping reduces the number of row activations as taught by Islam ([0025]). Alsop in view of Islam do not explicitly teach: generating, by the host processor, a memory address for the access request using a routing protocol indicating how source bits of the numerical identifiers are routed to the bit positions of the memory address. However, Gupta teaches: generating a memory address using a routing protocol indicating how source bits are routed to the bit positions of the memory address (abstract any address bit provided by the processor can be routed to any bank, row, or column bit, and can be used to generate any rank bit; col. 9:1-2 FIG. 6 shows address bit routing circuitry 48, which can route any address bit to any bank, row, or column bit; col. 9:24-26 if rank bit 0 is "1", then address bit 0 is routed to the output of multiplexer 52 indicated by the routing stored in register 50). Alsop, Islam, and Gupta are all concerned with mapping processor addresses onto the bank, row, and column organization of memory and are therefore combinable/modifiable. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Alsop in view of Islam in view of Gupta because it would provide for routing circuitry that can route any address bit from the processor to any bank, row, or column bit. Motivation would improve system adaptability because the routing supports a variety of memory sizes and interleaving schemes as taught by Gupta (Abstract). As per claim 20, the combination of references above teaches wherein the routing protocol is implemented in hardware of the host processor that is reconfigurable to account for different mappings of bit positions of the memory address to corresponding components of the memory (Gupta abstract any address bit provided by the processor can be routed to any bank, row, or column bit and col. 9:1-2 FIG. 6 shows address bit routing circuitry 48, which can route any address bit to any bank, row, or column bit; col. 9:30-31 multiplexer 52 has an output for each possible bank, row, and column bit), and generating the memory address includes: receiving a mapping associated with the workload (Islam [0075] NoS is provided statically as per user preference at system startup time (e.g. via the basic input/output system) or dynamically to achieve flexibility and [0080] the NoS information for each data structure/page must be tracked and communicated along with a physical memory address for any memory access request to the memory controller); updating machine status registers of the host processor to specify the routing protocol corresponding to the mapping (Islam [0081] four possible approaches are utilized: (i) Instruction-based approach and (ii) Page Table Entry (PTE)-based approach, (iii) configuration register approach, and (iv) a mode register based approach; Gupta col. 9:12-13 register 50 is configured when the computer system is initialized, as described above), and reconfiguring the hardware to implement the routing protocol as specified by the machine status registers (Gupta col. 9:24-26 if rank bit 0 is "1", then address bit 0 is routed to the output of multiplexer 52 indicated by the routing stored in register 50). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kessler (US 6,567,900) discloses that “for both striped and contiguous addressing, the address space includes a processor identification field to identify the processor where the associated memory resides, together with an offset indicating where in memory the address is located” (Abstract), which relates to the claimed generating of a memory address based on a processing-in-memory unit identifier and an offset. Desai et al. (US 10,628,308) discloses that “the MMU is configured to receive a request for a memory allocation request for one or more memory pages ... and, in response, select one of a plurality of interleave granularities for the one or more memory pages” (Abstract), which relates to the claimed receiving of a mapping associated with a workload for generating the memory address. Kwon et al., System Architecture and Software Stack for GDDR6-AiM, Hot Chips 34 (August 2022), discloses that “the AiM runtime library reshapes the matrix and stores it into AiMs,” the store commands carrying channel mask, bank, row, and column fields, which relates to the claimed storing of elements of a matrix at locations in the memory that map to the processing-in-memory units. Examiner has cited particular columns/paragraphs/sections and line numbers in the references applied and not relied upon to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THANH NGO whose telephone number is (571)270-3019. The examiner can normally be reached M-F 9am to 6pm ET, first F of bi-week off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Vital can be reached at (571)272-4215. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /T.N./Examiner, Art Unit 2198 /PIERRE VITAL/Supervisory Patent Examiner, Art Unit 2198
Read full office action

Prosecution Timeline

Mar 28, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month