Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
As per the instant application having Application No. 19/060,889, the amendment filed on 6/25/2026 is herein acknowledged. Claims 1-3, 5, 7-9 and 11-14 have been amended. Claims 1-20 are pending.
In response to this Office action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application.
Examiner cites columns and line numbers in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by prior art or disclosed by the examiner.
REJECTIONS BASED ON PRIOR ART
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5 and 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Kalamatianos et al. (US 2023/0077933) in view of Dutu et al. (US 2024/0220107) and Puthoor et al. (US 2023/0195645).
1. A method comprising: determining a candidate pseudo channel (PC) to which a processing in memory (PIM) instruction is assignable among a plurality of PCs based on an idle state of a PC; [Kalamatianos teaches “[0014] An implementation of supporting PIM execution in a multiprocessing environment also includes determining an availability of resources of the PIM device to support execution of the PIM instructions. Based on the availability, the method includes providing, to the first thread, a grant response indicating that access to the PIM device by the first thread is granted…In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.“
“[0039] The work scheduler 160 can be a single logic block tracking PIM resource usage across all memory channels (i.e., DRAM channels). In other examples, the work scheduler can be logic physically distributed (address interleaved in a similar manner that DRAM channels are) among different physical partitions. The flow described above works for a centralized work scheduler 160 implementation with one queue, whereas a distributed work scheduler 160 implementation (with a local queue per work scheduler 160 block) requires each processor core to track the grant and response status from all physical partitions of the work scheduler 160. For example, for an SoC with 128 memory channels, the core would need a 128-wide bit vector for tracking grant/response status of each physical partition. Only when all physical partitions of the work scheduler 160 grant access to the thread is the thread allowed to dispatch PIM instructions to the PIM devices of all DRAM channels. In cases where a processor core itself supports multithreading, the processor core must track grant response statuses for each hardware context separately.” Where as availability is tracked for all channels/channel partitions, candidate channels are identified as available or not currently executing operations, which corresponds to channels being idle]; however, Kalamatianos does not expressly refer to the channels/channel partitions as pseudo channels
determining a target PC set corresponding to a PIM operation based on the candidate PC; [Kalamatianos teaches “[0042] Yet another example implementation uses a vertical multithreaded dispatch policy in which access to PIM execution units is granted to threads that dispatch work to fixed partitions of memory channels (as opposed to all memory channels). Consider an example where four threads T0, T1, T2, and T3 have been granted access to PIM execution units each using a fixed 2-channel partition. Threads T0 and T1 are dispatching PIM instructions to channels 0 and 1 only, where the PIM resources in channels 0 and 1 are shared by threads T0 and T1. Threads T2 and T3 are dispatching PIM instructions to channels 30 and 31 only, where the PIM resources in channels 30 and 31 are shared by threads T2 and T3. In this implementation, the channel partition size (i.e., 2) is the same across all threads. A physically distributed implementation of a work scheduler must ensure a table per channel partition. Moreover, if the work scheduler 160 is physically distributed, each processor core must track grant/response status for every channel partition. A centralized work scheduler 160 implementation must track reserved PIM resources per thread and per channel partition.
[0043] In another implementation, the size of the memory channel partition varies per thread. Consider an example where T0 can dispatch PIM instructions to all 32 channels, while T1 dispatches work to a 2-channel partition (e.g., channels 0 and 1) and T2 dispatches PIM instructions to a different 4-channel partition (e.g., channels 2-5). A physically distributed work scheduler 160 must ensure a table per minimum size channel partition while the processor core must track grant/response status for the minimum channel partition supported. A centralized work scheduler implementation must be able to track reserved PIM resources per thread and per minimum size channel partition.”]
allocating data and the PIM instruction to the target PC set, for the target PC set; and…, wherein two or more target PCs included in the target PC set perform the PIM operation in parallel based on the data and the PIM instruction [Kalamatianos teaches “In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.” (par. 0014) where “FIG. 5 sets forth a method of supporting PIM execution in a multiprocessing environment in which multiple threads are executing concurrently according to implementations of the present disclosure. In implementations in which multiple threads execute concurrently and share PIM execution resources,” (par. 0066)].
With respect to the limitations while the PIM operation is performed, determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation, [Kalamatianos teaches ““[0038] Consider, for example, that every processor thread issues a start of kernel command before it starts dispatching PIM instructions to the work scheduler 160. The work scheduler 160 grants access to threads based on PIM resource availability and the PIM resource requirements specified in the start of kernel commands. The work scheduler only provides a grant response to the threads that have been granted access to PIM execution units. That is, the work scheduler only grants access to threads once resources have been reserved. All other threads wait for a response and do not dispatch any PIM instructions while waiting. The threads that are not waiting eventually issue an end of kernel command to the work scheduler 160 when they have completed dispatching a set of PIM instructions. The work scheduler 160 then releases the PIM resources for that thread, reserves resources for one or more threads that are pending in the queue, and grants access to those threads once the resources are reserved. This process continues until all threads have been granted access to the PIM execution units and dispatched all of their PIM instructions.” Where threads are dispatched to selected memory channels according to the reserved resources (pars. 0042-0043)], thus teaching allocating determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation but does not expressly disclose doing so while the PIM operation is performed.
With respect to the channels/channel partitions as pseudo channels, Dutu teaches [teaches “[0029] A third arbiter (e.g., a third arbitration stage) of the arbitration system then schedules an execution order for the requests output by the second arbiter. In implementations where the memory controller is tasked with scheduling requests for a memory channel allocated into multiple pseudo-channels (e.g., two or more pseudo-channels), the first and second arbiters are configured to perform their functionality for each pseudo-channel simultaneously. In such scenarios where the memory channel is allocated into multiple pseudo-channels, the third arbiter schedules requests output by the second arbiter (e.g., the priority winner for each of the multiple pseudo-channels) in a round-robin manner… [0035] In some aspects, the techniques described herein relate to a system, wherein the memory controller is associated with a channel in the memory and the channel in the memory is allocated into two or more pseudo-channels.”].
Kalamatianos and Dutu are analogous art because they are from the same field of endeavor of memory access and control.
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify Kalamatianos to have the channels/channel partitions as implemented as a plurality of pseudo channels in the manner taught by Dutu since doing so would provide the benefits of providing flexibility of design and allowing for “adaptive scheduling of memory requests and processing-in-memory requests is described” (par. 0020).
With respect to while the PIM operation is performed determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation, Puthoor teaches [“[0044]… Consider, as an example, that process 172 has allocated a virtual address space for a configuration context of the PIM device 181 and the PIM driver has mapped that configuration address space to the process 172 and assigned physical pages for receiving operands of PIM instructions to the process 172. Consider also that process 174, while the virtual address space is allocated to the process 172, makes a call to allocate a second virtual address space that stores a second PIM configuration context of the process 174. The PIM driver, having another available PIM device 183 allocates the second virtual address space to the second process and maps the physical address space of the second PIM device 183 to the second virtual address space only if the configuration and orchestration space of the second PIM device 183 is not mapped to another process's virtual address space. The driver 124 can then program the second PIM device's 183 configuration registers according to the second configuration context of the process 174. In this way, multiple processes can concurrently access different PIM drivers or PIM driver partitions while maintaining process isolation.” “[0060] FIG. 5 sets forth a flow chart illustrating process isolation for a PIM device using virtualization in which multiple different processes utilize different PIM resources simultaneously. While the examples above generally describe assigning ownership of a PIM device to a single process, some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
Kalamatianos, Dutu and Puthoor are analogous art because they are from the same field of endeavor of memory access and control.
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify the combination Kalamatianos and Dutu to include determining next PIM resources (such as the target PC set of the combination of Kalamatianos and Dutu) corresponding to a next PIM operation and allocating data and a PIM instruction for the next operation as taught by Puthoor since doing so would provide the benefits of [“[0060]… some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
Therefore, it would have been obvious to combine Kalamatianos and Dutu with Puthoor for the benefit of creating a storage system/method to obtain the invention as specified in claim 1.
2. The method of claim 1, wherein the next PIM operation independent of the PIM operation is independent of the PIM operation [Kalamatianos teaches “[0038] Consider, for example, that every processor thread issues a start of kernel command before it starts dispatching PIM instructions to the work scheduler 160. The work scheduler 160 grants access to threads based on PIM resource availability and the PIM resource requirements specified in the start of kernel commands. The work scheduler only provides a grant response to the threads that have been granted access to PIM execution units. That is, the work scheduler only grants access to threads once resources have been reserved. All other threads wait for a response and do not dispatch any PIM instructions while waiting. The threads that are not waiting eventually issue an end of kernel command to the work scheduler 160 when they have completed dispatching a set of PIM instructions. The work scheduler 160 then releases the PIM resources for that thread, reserves resources for one or more threads that are pending in the queue, and grants access to those threads once the resources are reserved. This process continues until all threads have been granted access to the PIM execution units and dispatched all of their PIM instructions.” Where threads are dispatched to selected memory channels according to the reserved resources (pars. 0042-0043). Dutu teaches “[0126] The selected requests are then ordered for multiple pseudo-channels (block 614). The arbitration system 124, for instance, causes the third arbiter 210 to maintain a priority winner queue 230 that stores the priority winner 228 output by the second arbiter 208 for each of one or more pseudo-channels of a channel in memory 110 to which the memory controller 114 is assigned. In implementations where the priority winner queue 230 includes priority winners 228 for multiple pseudo-channels, the third arbiter 210 outputs individual ones of the priority winner 228 as a scheduled request 212 by cycling through different pseudo-channel priority winners in a round-robin selection process. After outputting a scheduled request 212 for each memory pseudo-channel, operation of the procedure 600 continues by returning to block 602.” Where each of the operations considered independent. Additionally, Puthoor teaches “[0044]… while the virtual address space is allocated to the process 172, makes a call to allocate a second virtual address space that stores a second PIM configuration context of the process 174. The PIM driver, having another available PIM device 183 allocates the second virtual address space to the second process and maps the physical address space of the second PIM device 183 to the second virtual address space only if the configuration and orchestration space of the second PIM device 183 is not mapped to another process's virtual address space. The driver 124 can then program the second PIM device's 183 configuration registers according to the second configuration context of the process 174. In this way, multiple processes can concurrently access different PIM drivers or PIM driver partitions while maintaining process isolation.”].
3. The method of claim 2, wherein the allocating of the data and the PIM instruction for the next PIM operation comprises: determining the next target PC set for the next PIM operation such that the PIM operation shares resources with the next PIM operation; and allocating the data and the PIM instruction for the next PIM operation [Kalamatianos teaches “[0011] Multithreaded applications executing in PIM require the sharing of limited PIM resources among the threads running PIM code simultaneously” “[0041] Another example implementation uses a horizontal multithreaded dispatch policy in which multiple threads are allowed to dispatch work to all memory channels concurrently, so long as enough PIM execution unit resources are available. For example, two threads executing on two different processors can share the resources of each PIM unit as long as each PIM unit has enough resources to support execution of the PIM instructions of both threads. The work scheduler 160 in such an implementation tracks PIM execution unit resource utilization for all threads that have been granted access. A table in which PIM unit resources per thread are tracked can be sized to allow all or a subset of hardware contexts in the host device 130 to dispatch work to the PIM execution units.” “PIM resources in channels 0 and 1 are shared by threads T0 and T1… PIM resources in channels 30 and 31 are shared by threads T2 and T3” (par. 0042), Puthoor teaches [“[0044]… Consider, as an example, that process 172 has allocated a virtual address space for a configuration context of the PIM device 181 and the PIM driver has mapped that configuration address space to the process 172 and assigned physical pages for receiving operands of PIM instructions to the process 172. Consider also that process 174, while the virtual address space is allocated to the process 172, makes a call to allocate a second virtual address space that stores a second PIM configuration context of the process 174. The PIM driver, having another available PIM device 183 allocates the second virtual address space to the second process and maps the physical address space of the second PIM device 183 to the second virtual address space only if the configuration and orchestration space of the second PIM device 183 is not mapped to another process's virtual address space. The driver 124 can then program the second PIM device's 183 configuration registers according to the second configuration context of the process 174. In this way, multiple processes can concurrently access different PIM drivers or PIM driver partitions while maintaining process isolation.” “[0060] FIG. 5 sets forth a flow chart illustrating process isolation for a PIM device using virtualization in which multiple different processes utilize different PIM resources simultaneously. While the examples above generally describe assigning ownership of a PIM device to a single process, some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
4. The method of claim 1, wherein the determining of the candidate PC comprises determining a PC in an idle state among the plurality of PCs as the candidate PC [Kalamatianos teaches “[0014] An implementation of supporting PIM execution in a multiprocessing environment also includes determining an availability of resources of the PIM device to support execution of the PIM instructions. Based on the availability, the method includes providing, to the first thread, a grant response indicating that access to the PIM device by the first thread is granted…In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.“
“[0039] The work scheduler 160 can be a single logic block tracking PIM resource usage across all memory channels (i.e., DRAM channels). In other examples, the work scheduler can be logic physically distributed (address interleaved in a similar manner that DRAM channels are) among different physical partitions. The flow described above works for a centralized work scheduler 160 implementation with one queue, whereas a distributed work scheduler 160 implementation (with a local queue per work scheduler 160 block) requires each processor core to track the grant and response status from all physical partitions of the work scheduler 160. For example, for an SoC with 128 memory channels, the core would need a 128-wide bit vector for tracking grant/response status of each physical partition. Only when all physical partitions of the work scheduler 160 grant access to the thread is the thread allowed to dispatch PIM instructions to the PIM devices of all DRAM channels. In cases where a processor core itself supports multithreading, the processor core must track grant response statuses for each hardware context separately.” Where as availability is tracked for all channels/channel partitions, candidate channels are identified as available or not currently executing operations, which corresponds to channels being idle. See pars. 0041-0043 for channel grant/response status as well per channel partition tracking].
5. The method of claim 1, further comprising updating an allocation state of the data and the PIM instruction for the two or more target PCs [Kalamatianos teaches “A table in which PIM unit resources per thread are tracked can be sized to allow all or a subset of hardware contexts in the host device 130 to dispatch work to the PIM execution units” (par. 0041) “dispatch policy in which access to PIM execution units is granted to threads that dispatch work to fixed partitions of memory channels (as opposed to all memory channels)… the channel partition size (i.e., 2) is the same across all threads. A physically distributed implementation of a work scheduler must ensure a table per channel partition… each processor must track grant/response status for every channel partition. A centralized work scheduler 160 implementation must track reserved PIM resources per thread and per channel partition” (par. 0042) “[0043] In another implementation, the size of the memory channel partition varies per thread... A physically distributed work scheduler 160 must ensure a table per minimum size channel partition while the processor core must track grant/response status for the minimum channel partition supported. A centralized work scheduler implementation must be able to track reserved PIM resources per thread and per minimum size channel partition.” “[0060] FIG. 5 sets forth a flow chart illustrating process isolation for a PIM device using virtualization in which multiple different processes utilize different PIM resources simultaneously. While the examples above generally describe assigning ownership of a PIM device to a single process, some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
9. An accelerator comprising: [Kalamatianos teaches system 100 including processing-in-memory PIM, fig. 1 and related text, where “ PIM may also include so-called processing-near-memory implementations and other accelerator architectures.” (par. 0002)]
a memory comprising a plurality of pseudo channels (PCs); [Kalamatianos teaches “In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.” (par. 0014) “[0034] In the example system of FIG. 1, the host device 130 also includes a memory controller 140 that is shared by the processor cores 102, 104, 106, 108 for accessing various memory channels coupling the processor 132 to the memory device 180.” (fig. 1 and related text)] but Kalamatianos does not expressly refer to the channels as pseudo channels
and a processor configured to: determine a candidate PC to which a processing in memory (PIM) instruction is assignable among the plurality of PCs based on an idle state of a PC; [Kalamatianos teaches Host Device 130 including Work Scheduler 160 where ““[0014] An implementation of supporting PIM execution in a multiprocessing environment also includes determining an availability of resources of the PIM device to support execution of the PIM instructions. Based on the availability, the method includes providing, to the first thread, a grant response indicating that access to the PIM device by the first thread is granted…In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.“
“[0039] The work scheduler 160 can be a single logic block tracking PIM resource usage across all memory channels (i.e., DRAM channels). In other examples, the work scheduler can be logic physically distributed (address interleaved in a similar manner that DRAM channels are) among different physical partitions. The flow described above works for a centralized work scheduler 160 implementation with one queue, whereas a distributed work scheduler 160 implementation (with a local queue per work scheduler 160 block) requires each processor core to track the grant and response status from all physical partitions of the work scheduler 160. For example, for an SoC with 128 memory channels, the core would need a 128-wide bit vector for tracking grant/response status of each physical partition. Only when all physical partitions of the work scheduler 160 grant access to the thread is the thread allowed to dispatch PIM instructions to the PIM devices of all DRAM channels. In cases where a processor core itself supports multithreading, the processor core must track grant response statuses for each hardware context separately.” Where as availability is tracked for all channels/channel partitions, candidate channels are identified as available or not currently executing operations, which corresponds to channels being idle]
determine a target PC set comprising one or more of the plurality of PCs and corresponding to a PIM operation, based on the candidate PC; [Kalamatianos teaches “[0042] Yet another example implementation uses a vertical multithreaded dispatch policy in which access to PIM execution units is granted to threads that dispatch work to fixed partitions of memory channels (as opposed to all memory channels). Consider an example where four threads T0, T1, T2, and T3 have been granted access to PIM execution units each using a fixed 2-channel partition. Threads T0 and T1 are dispatching PIM instructions to channels 0 and 1 only, where the PIM resources in channels 0 and 1 are shared by threads T0 and T1. Threads T2 and T3 are dispatching PIM instructions to channels 30 and 31 only, where the PIM resources in channels 30 and 31 are shared by threads T2 and T3. In this implementation, the channel partition size (i.e., 2) is the same across all threads. A physically distributed implementation of a work scheduler must ensure a table per channel partition. Moreover, if the work scheduler 160 is physically distributed, each processor core must track grant/response status for every channel partition. A centralized work scheduler 160 implementation must track reserved PIM resources per thread and per channel partition.
[0043] In another implementation, the size of the memory channel partition varies per thread. Consider an example where T0 can dispatch PIM instructions to all 32 channels, while T1 dispatches work to a 2-channel partition (e.g., channels 0 and 1) and T2 dispatches PIM instructions to a different 4-channel partition (e.g., channels 2-5). A physically distributed work scheduler 160 must ensure a table per minimum size channel partition while the processor core must track grant/response status for the minimum channel partition supported. A centralized work scheduler implementation must be able to track reserved PIM resources per thread and per minimum size channel partition.”]
allocate data and the PIM instruction to the target PC set, for the target PC set; and…, wherein two or more target PCs included in the target PC set perform the PIM operation in parallel based on the data and the PIM instruction [Kalamatianos teaches “In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.” (par. 0014; see pars. 0042-0043) where “FIG. 5 sets forth a method of supporting PIM execution in a multiprocessing environment in which multiple threads are executing concurrently according to implementations of the present disclosure. In in implementations in which multiple threads execute concurrently and share PIM execution resources,” (par. 0066)].
Regarding while the PIM operation is performed, determine a next target PC set corresponding to a next PIM operation and allocate data and a PIM instruction for the next PIM operation [Kalamatianos teaches ““[0038] Consider, for example, that every processor thread issues a start of kernel command before it starts dispatching PIM instructions to the work scheduler 160. The work scheduler 160 grants access to threads based on PIM resource availability and the PIM resource requirements specified in the start of kernel commands. The work scheduler only provides a grant response to the threads that have been granted access to PIM execution units. That is, the work scheduler only grants access to threads once resources have been reserved. All other threads wait for a response and do not dispatch any PIM instructions while waiting. The threads that are not waiting eventually issue an end of kernel command to the work scheduler 160 when they have completed dispatching a set of PIM instructions. The work scheduler 160 then releases the PIM resources for that thread, reserves resources for one or more threads that are pending in the queue, and grants access to those threads once the resources are reserved. This process continues until all threads have been granted access to the PIM execution units and dispatched all of their PIM instructions.” Where threads are dispatched to selected memory channels according to the reserved resources (pars. 0042-0043)], thus teaching allocating determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation but does not expressly disclose doing so while the PIM operation is performed.
With respect to the channels/channel partitions as pseudo channels, Dutu teaches [“[0029] A third arbiter (e.g., a third arbitration stage) of the arbitration system then schedules an execution order for the requests output by the second arbiter. In implementations where the memory controller is tasked with scheduling requests for a memory channel allocated into multiple pseudo-channels (e.g., two or more pseudo-channels), the first and second arbiters are configured to perform their functionality for each pseudo-channel simultaneously. In such scenarios where the memory channel is allocated into multiple pseudo-channels, the third arbiter schedules requests output by the second arbiter (e.g., the priority winner for each of the multiple pseudo-channels) in a round-robin manner… [0035] In some aspects, the techniques described herein relate to a system, wherein the memory controller is associated with a channel in the memory and the channel in the memory is allocated into two or more pseudo-channels.”].
Kalamatianos and Dutu are analogous art because they are from the same field of endeavor of memory access and control.
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify Kalamatianos to have the channels/channel partitions as implemented as a plurality of pseudo channels in the manner taught by Dutu since doing so would provide the benefits of providing flexibility of design and allowing for “adaptive scheduling of memory requests and processing-in-memory requests is described” (par. 0020).
With respect to while the PIM operation is performed determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation, Puthoor teaches [“[0044]… Consider, as an example, that process 172 has allocated a virtual address space for a configuration context of the PIM device 181 and the PIM driver has mapped that configuration address space to the process 172 and assigned physical pages for receiving operands of PIM instructions to the process 172. Consider also that process 174, while the virtual address space is allocated to the process 172, makes a call to allocate a second virtual address space that stores a second PIM configuration context of the process 174. The PIM driver, having another available PIM device 183 allocates the second virtual address space to the second process and maps the physical address space of the second PIM device 183 to the second virtual address space only if the configuration and orchestration space of the second PIM device 183 is not mapped to another process's virtual address space. The driver 124 can then program the second PIM device's 183 configuration registers according to the second configuration context of the process 174. In this way, multiple processes can concurrently access different PIM drivers or PIM driver partitions while maintaining process isolation.” “[0060] FIG. 5 sets forth a flow chart illustrating process isolation for a PIM device using virtualization in which multiple different processes utilize different PIM resources simultaneously. While the examples above generally describe assigning ownership of a PIM device to a single process, some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
Kalamatianos, Dutu and Puthoor are analogous art because they are from the same field of endeavor of memory access and control.
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify the combination Kalamatianos and Dutu to include determining next PIM resources (such as the target PC set of the combination of Kalamatianos and Dutu) corresponding to a next PIM operation and allocating data and a PIM instruction for the next operation as taught by Puthoor since doing so would provide the benefits of [“[0060]… some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
Therefore, it would have been obvious to combine Kalamatianos and Dutu with Puthoor for the benefit of creating a storage system/method to obtain the invention as specified in claim 9.
10. An electronic device comprising: a host processor configured to provide the PIM instruction to the accelerator of claim 9; and the accelerator of claim 9 [Kalamatianos teaches system 100 including processing-in-memory PIM, fig. 1 and related text, where “ PIM may also include so-called processing-near-memory implementations and other accelerator architectures.” (par. 0002) “To that end, providing enough space to hold all PIM architectural registers for every hardware context in a multicore processor can result in a significant space and power overhead for a memory device or accelerator implementing PIM logic. Additionally, resource sharing or virtualization within the PIM device can be a difficult task.” (par. 0011) and host device 130 providing instruction to PIM devices corresponding to the accelerator architecture] but the memory device 180 including PIM is not expressly defined as an accelerator in a manner that a host provides instructions to the accelerator; however, regarding these limitations, Puthoor teaches [“0049] In the example of FIG. 2, the execution unit 150 is a component of a PIM device 280 that is implemented in a processing-near-memory (PNM) fashion. For example, the PIM device 280 can be a memory accelerator that is used to execute memory-intensive operations that have been offloaded to by the host processor 132 to the accelerator. The host processor 132 and the PIM device 280 are both coupled to the same memory 220. The host processor 132 provides PIM instructions to the PIM device 280 through the memory controller 140, which the execution unit 150 of the PIM device 280 performs on data stored in the memory 220. “].
Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Kalamatianos et al. (US 2023/0077933) in view of Dutu et al. (US 2024/0220107) and Puthoor et al. (US 20230195645) as applied in the rejection of claim 1, and further in view of Kanayama et al. (US 2023/0418772).
7. The method of claim 1, wherein the allocating of the data and the PIM instruction to the target PC set comprises: inputting the data and the PIM instruction to a port for any one target PC among the two or more target PCs; and … allocating the data and the PIM instruction to each of the two or more target PCs [Kalamatianos teaches “[0042] Yet another example implementation uses a vertical multithreaded dispatch policy in which access to PIM execution units is granted to threads that dispatch work to fixed partitions of memory channels (as opposed to all memory channels). Consider an example where four threads T0, T1, T2, and T3 have been granted access to PIM execution units each using a fixed 2-channel partition. Threads T0 and T1 are dispatching PIM instructions to channels 0 and 1 only, where the PIM resources in channels 0 and 1 are shared by threads T0 and T1. Threads T2 and T3 are dispatching PIM instructions to channels 30 and 31 only, where the PIM resources in channels 30 and 31 are shared by threads T2 and T3. In this implementation, the channel partition size (i.e., 2) is the same across all threads. A physically distributed implementation of a work scheduler must ensure a table per channel partition. Moreover, if the work scheduler 160 is physically distributed, each processor core must track grant/response status for every channel partition. A centralized work scheduler 160 implementation must track reserved PIM resources per thread and per channel partition… [0043] In another implementation, the size of the memory channel partition varies per thread. Consider an example where T0 can dispatch PIM instructions to all 32 channels, while T1 dispatches work to a 2-channel partition (e.g., channels 0 and 1) and T2 dispatches PIM instructions to a different 4-channel partition (e.g., channels 2-5). A physically distributed work scheduler 160 must ensure a table per minimum size channel partition while the processor core must track grant/response status for the minimum channel partition supported. A centralized work scheduler implementation must be able to track reserved PIM resources per thread and per minimum size channel partition.” Where each of the channels connecting to the memory and PIM devices includes a port. Dutu teaches pseudo channels (pars. 0029 and 0035, 0126). Puthoor teaches (pars. 0044, 0060)] but the combination of Kalamatianos, Dutu and Puthoor does not expressly disclose dividing the data and the PIM instruction into numbers corresponding to the one or more target PCs and; however, regarding these limitations, Kanayama teaches [“An important feature of newer HBM memories is known as pseudo channel mode. Pseudo channel mode divides a channel into two individual subchannels that operate semi-independently. The pseudo-channels share the command bus, but execute commands individually. “ (par. 0002)]
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify Kalamatianos, Dutu and Puthoor to include as taught by Kanayama since doing so would provide the benefits of [allowing for higher speed operation (par. 0002)].
Therefore, it would have been obvious to combine Kalamatianos, Dutu and Puthoor with Kanayama for the benefit of creating a storage system/method to obtain the invention as specified in claim 7.
8. The method of claim 7, wherein the inputting of the data and the PIM instruction comprises inputting the data and the PIM instruction to a port of a target PC having a lowest index among the one or more target PCs [Kalamatianos teaches “[0042] Yet another example implementation uses a vertical multithreaded dispatch policy in which access to PIM execution units is granted to threads that dispatch work to fixed partitions of memory channels (as opposed to all memory channels). Consider an example where four threads T0, T1, T2, and T3 have been granted access to PIM execution units each using a fixed 2-channel partition. Threads T0 and T1 are dispatching PIM instructions to channels 0 and 1 only, where the PIM resources in channels 0 and 1 are shared by threads T0 and T1. Threads T2 and T3 are dispatching PIM instructions to channels 30 and 31 only, where the PIM resources in channels 30 and 31 are shared by threads T2 and T3. In this implementation, the channel partition size (i.e., 2) is the same across all threads. A physically distributed implementation of a work scheduler must ensure a table per channel partition. Moreover, if the work scheduler 160 is physically distributed, each processor core must track grant/response status for every channel partition. A centralized work scheduler 160 implementation must track reserved PIM resources per thread and per channel partition… [0043] In another implementation, the size of the memory channel partition varies per thread. Consider an example where T0 can dispatch PIM instructions to all 32 channels, while T1 dispatches work to a 2-channel partition (e.g., channels 0 and 1) and T2 dispatches PIM instructions to a different 4-channel partition (e.g., channels 2-5). A physically distributed work scheduler 160 must ensure a table per minimum size channel partition while the processor core must track grant/response status for the minimum channel partition supported. A centralized work scheduler implementation must be able to track reserved PIM resources per thread and per minimum size channel partition.” Where each of the channels connecting to the memory and PIM devices includes a port. Dutu teaches pseudo channels (pars. 0029 and 0035) where “[0126] The selected requests are then ordered for multiple pseudo-channels (block 614). The arbitration system 124, for instance, causes the third arbiter 210 to maintain a priority winner queue 230 that stores the priority winner 228 output by the second arbiter 208 for each of one or more pseudo-channels of a channel in memory 110 to which the memory controller 114 is assigned. In implementations where the priority winner queue 230 includes priority winners 228 for multiple pseudo-channels, the third arbiter 210 outputs individual ones of the priority winner 228 as a scheduled request 212 by cycling through different pseudo-channel priority winners in a round-robin selection process. After outputting a scheduled request 212 for each memory pseudo-channel, operation of the procedure 600 continues by returning to block 602.”)] thus cycling through the pseudo channels includes selecting a first pseudo channel and then other but the combination does not expressly disclose a lowest index among the channels/pseudo channels; however, It would have been obvious to one having ordinary skill in the art at the time the invention was made to have the channel selected based on a lowest index or number or any value for the channel, since it has been held that discovering an optimum value of a result effective variable involves only routine skill in the art. In re Boesch, 617 F.2d 272, 205 USPQ 215 (CCPA 1980) and doing so would at least provide the benefits of labeling/identifying the selected pseudo channel.
Claims 11-16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kalamatianos et al. (US 2023/0077933) in view of Puthoor et al. (US 20230195645) and Dutu et al. (US 2024/0220107).
11. An electronic device comprising: a host processor configured to provide a processing in memory (PIM) instruction to an accelerator; and the accelerator configured to: [Kalamatianos teaches system 100 including processing-in-memory PIM, fig. 1 and related text, where “ PIM may also include so-called processing-near-memory implementations and other accelerator architectures.” (par. 0002) “To that end, providing enough space to hold all PIM architectural registers for every hardware context in a multicore processor can result in a significant space and power overhead for a memory device or accelerator implementing PIM logic. Additionally, resource sharing or virtualization within the PIM device can be a difficult task.” (par. 0011) and host device 130 providing instruction to PIM devices corresponding to the accelerator architecture] but the memory device 180 including PIM is not expressly defined as an accelerator where the host device provides instructions to the accelerator
determine a candidate pseudo channel (PC) to which the PIM instruction is assignable among a plurality of PCs based on an idle state of a PC; [Kalamatianos teaches “[0014] An implementation of supporting PIM execution in a multiprocessing environment also includes determining an availability of resources of the PIM device to support execution of the PIM instructions. Based on the availability, the method includes providing, to the first thread, a grant response indicating that access to the PIM device by the first thread is granted…In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.“
“[0039] The work scheduler 160 can be a single logic block tracking PIM resource usage across all memory channels (i.e., DRAM channels). In other examples, the work scheduler can be logic physically distributed (address interleaved in a similar manner that DRAM channels are) among different physical partitions. The flow described above works for a centralized work scheduler 160 implementation with one queue, whereas a distributed work scheduler 160 implementation (with a local queue per work scheduler 160 block) requires each processor core to track the grant and response status from all physical partitions of the work scheduler 160. For example, for an SoC with 128 memory channels, the core would need a 128-wide bit vector for tracking grant/response status of each physical partition. Only when all physical partitions of the work scheduler 160 grant access to the thread is the thread allowed to dispatch PIM instructions to the PIM devices of all DRAM channels. In cases where a processor core itself supports multithreading, the processor core must track grant response statuses for each hardware context separately.” Where as availability is tracked for all channels/channel partitions, candidate channels are identified as available or not currently executing operations, which corresponds to channels being idle]; however, Kalamatianos does not expressly refer to the channels/channel partitions as pseudo channels
determine a target PC set corresponding to a PIM operation based on the candidate PC; [Kalamatianos teaches “[0042] Yet another example implementation uses a vertical multithreaded dispatch policy in which access to PIM execution units is granted to threads that dispatch work to fixed partitions of memory channels (as opposed to all memory channels). Consider an example where four threads T0, T1, T2, and T3 have been granted access to PIM execution units each using a fixed 2-channel partition. Threads T0 and T1 are dispatching PIM instructions to channels 0 and 1 only, where the PIM resources in channels 0 and 1 are shared by threads T0 and T1. Threads T2 and T3 are dispatching PIM instructions to channels 30 and 31 only, where the PIM resources in channels 30 and 31 are shared by threads T2 and T3. In this implementation, the channel partition size (i.e., 2) is the same across all threads. A physically distributed implementation of a work scheduler must ensure a table per channel partition. Moreover, if the work scheduler 160 is physically distributed, each processor core must track grant/response status for every channel partition. A centralized work scheduler 160 implementation must track reserved PIM resources per thread and per channel partition.
[0043] In another implementation, the size of the memory channel partition varies per thread. Consider an example where T0 can dispatch PIM instructions to all 32 channels, while T1 dispatches work to a 2-channel partition (e.g., channels 0 and 1) and T2 dispatches PIM instructions to a different 4-channel partition (e.g., channels 2-5). A physically distributed work scheduler 160 must ensure a table per minimum size channel partition while the processor core must track grant/response status for the minimum channel partition supported. A centralized work scheduler implementation must be able to track reserved PIM resources per thread and per minimum size channel partition.”]
allocate data and the PIM instruction to the target PC set, for the target PC set; perform the PIM operation using the target PC set to which the data and the PIM instruction is allocated; and [Kalamatianos teaches “In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.” (par. 0014; see pars. 0042-0043) where “FIG. 5 sets forth a method of supporting PIM execution in a multiprocessing environment in which multiple threads are executing concurrently according to implementations of the present disclosure. In in implementations in which multiple threads execute concurrently and share PIM execution resources,” (par. 0066)]
while the PIM operation is performed, determine a next target PC set corresponding to a next PIM operation and allocate data and a PIM instruction for the next PIM operation [Kalamatianos teaches “[0038] Consider, for example, that every processor thread issues a start of kernel command before it starts dispatching PIM instructions to the work scheduler 160. The work scheduler 160 grants access to threads based on PIM resource availability and the PIM resource requirements specified in the start of kernel commands. The work scheduler only provides a grant response to the threads that have been granted access to PIM execution units. That is, the work scheduler only grants access to threads once resources have been reserved. All other threads wait for a response and do not dispatch any PIM instructions while waiting. The threads that are not waiting eventually issue an end of kernel command to the work scheduler 160 when they have completed dispatching a set of PIM instructions. The work scheduler 160 then releases the PIM resources for that thread, reserves resources for one or more threads that are pending in the queue, and grants access to those threads once the resources are reserved. This process continues until all threads have been granted access to the PIM execution units and dispatched all of their PIM instructions.” Where threads are dispatched to selected memory channels according to the reserved resources (pars. 0042-0043)], thus teaching allocating determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation but does not expressly disclose doing so while the PIM operation is performed.
Regarding the PIM array being implemented as an accelerator where the host device provides instructions to the accelerator, Puthoor teaches [0049] In the example of FIG. 2, the execution unit 150 is a component of a PIM device 280 that is implemented in a processing-near-memory (PNM) fashion. For example, the PIM device 280 can be a memory accelerator that is used to execute memory-intensive operations that have been offloaded to by the host processor 132 to the accelerator. The host processor 132 and the PIM device 280 are both coupled to the same memory 220. The host processor 132 provides PIM instructions to the PIM device 280 through the memory controller 140, which the execution unit 150 of the PIM device 280 performs on data stored in the memory 220.]
while the PIM operation is performed determining a next target PC set corresponding to a next PIM operation and allocating data and a PIM instruction for the next PIM operation, Puthoor teaches [“[0044]… Consider, as an example, that process 172 has allocated a virtual address space for a configuration context of the PIM device 181 and the PIM driver has mapped that configuration address space to the process 172 and assigned physical pages for receiving operands of PIM instructions to the process 172. Consider also that process 174, while the virtual address space is allocated to the process 172, makes a call to allocate a second virtual address space that stores a second PIM configuration context of the process 174. The PIM driver, having another available PIM device 183 allocates the second virtual address space to the second process and maps the physical address space of the second PIM device 183 to the second virtual address space only if the configuration and orchestration space of the second PIM device 183 is not mapped to another process's virtual address space. The driver 124 can then program the second PIM device's 183 configuration registers according to the second configuration context of the process 174. In this way, multiple processes can concurrently access different PIM drivers or PIM driver partitions while maintaining process isolation.” “[0060] FIG. 5 sets forth a flow chart illustrating process isolation for a PIM device using virtualization in which multiple different processes utilize different PIM resources simultaneously. While the examples above generally describe assigning ownership of a PIM device to a single process, some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify Kalamatianos to implement the PIM device as a PIM accelerator where the host device provides instructions to the accelerator and modify Kalamatianos to include determining next PIM resources (such as the target PC set of the combination of Kalamatianos) corresponding to a next PIM operation and allocating data and a PIM instruction for the next operation as taught by Puthoor since doing so would provide the benefits of [“[0060]… some PIM devices may include multiple logical or physical partitions that can be separately assigned to different processes and utilized by those process simultaneously without conflict or security concerns. Additionally, a system that includes multiple different PIM devices may assign out those resources to separate processes for parallel ownership.”].
The combination of Kalamatianos and Puthoor does not expressly define the channels/channel partitions as pseudo channels; however, regarding these limitations, Dutu teaches [“[0029] A third arbiter (e.g., a third arbitration stage) of the arbitration system then schedules an execution order for the requests output by the second arbiter. In implementations where the memory controller is tasked with scheduling requests for a memory channel allocated into multiple pseudo-channels (e.g., two or more pseudo-channels), the first and second arbiters are configured to perform their functionality for each pseudo-channel simultaneously. In such scenarios where the memory channel is allocated into multiple pseudo-channels, the third arbiter schedules requests output by the second arbiter (e.g., the priority winner for each of the multiple pseudo-channels) in a round-robin manner… [0035] In some aspects, the techniques described herein relate to a system, wherein the memory controller is associated with a channel in the memory and the channel in the memory is allocated into two or more pseudo-channels.”].
Kalamatianos, Puthoor and Dutu are analogous art because they are from the same field of endeavor of memory access and control.
Before the effective filing date of the claimed inventions, it would have been obvious to a person of ordinary skill in the art to modify the combination of Kalamatianos and Puthoor to have the channels/channel partitions as implemented as a plurality of pseudo channels in the manner taught by Dutu since doing so would provide the benefits of providing flexibility of design and allowing for “adaptive scheduling of memory requests and processing-in-memory requests is described” (par. 0020).
Therefore, it would have been obvious to combine Kalamatianos with Puthoor and Dutu for the benefit of creating a storage system/method to obtain the invention as specified in claim 11.
12. The electronic device of claim 11, wherein two or more target PCs included in the target PC set are configured to perform the PIM operation in parallel based on the data and the PIM instruction [Kalamatianos teaches “In an implementation, the first thread dispatches the plurality of PIM instructions to a set of memory channels concurrently with at least one second thread dispatching PIM instructions to that set of memory channels. Also, in an implementation, the first thread dispatches the plurality of PIM instructions to a first partition of memory channels concurrently with at least one second thread dispatching PIM instructions to a second partition of memory channels.” (par. 0014) where “FIG. 5 sets forth a method of supporting PIM execution in a multiprocessing environment in which multiple threads are executing concurrently according to implementations of the present disclosure. In in implementations in which multiple threads execute concurrently and share PIM execution resources,” (par. 0066). Puthoor teaches pars. 0044, 0060].
13. The electronic device of claim 11, wherein the next PIM operation is independent of the PIM operation [The rationale in the rejection of claim 2 is herein incorporated].
14. The electronic device of claim 13, wherein, for the allocating of the data and the PIM instruction for the next PIM operation, the accelerator is configured to: determine the target PC set for the next PIM operation such that the PIM operation shares resources with the next PIM operation; and allocate the data and the PIM instruction for the next PIM operation [The rationale in the rejection of claim 3 is herein incorporated].
15. The electronic device of claim 11, wherein, for the determining of the candidate PC, the accelerator is configured to determine a PC in an idle state among the plurality of PCs as the candidate PC [The rationale in the rejection of claim 4 is herein incorporated].
16. The electronic device of claim 11, wherein the accelerator is configured to update an allocation state of the data and the PIM instruction for one or more target PCs [The rationale in the rejection of claim 5 is herein incorporated].
20. The electronic device of claim 11, wherein the accelerator comprises a processor implementing control logic, and the processor comprises any one or any combination of any two or more of a field- programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a central processing unit (CPU), a graphics processing unit (GPU), and a neural processing unit (NPU) [Zhou teaches “[0030] FIG. 2A illustrates an exemplary neural network accelerator architecture 200, consistent with embodiments of the present disclosure. In the context of this disclosure, a neural network accelerator may also be referred to as a machine learning accelerator or deep learning accelerator. In some embodiments, accelerator architecture 200 may be referred to as a neural network processing unit (NPU) architecture 200. As shown in FIG. 2A, accelerator architecture 200 can include a PIM accelerator 210, an interface 212, and the like. It is appreciated that, PIM accelerator 210 can perform algorithmic operations based on communicated data.”].
Claims 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Kalamatianos et al. (US 2023/0077933) in view of Puthoor et al. (US 20230195645) and Dutu et al. (US 2024/0220107) as applied in the rejection of claim 11 above, and further in view of Kanayama et al. (US 2023/0418772).
18. The electronic device of claim 11, wherein, for the allocating of the data and the PIM instruction to the target PC set, the accelerator is configured to: receive the data and the PIM instruction as input through a port for any one target PC among one or more target PCs; and divide the data and the PIM instruction into numbers corresponding to the one or more target PCs and allocate the data and the PIM instruction to each of the one or more target PCs [The rationale in the rejection of claim 7 is herein incorporated but the combination of Kalamatianos, Puthoor, Dutu and Kanayama applies to claim 18 instead of the combination of Kalamatianos, Dutu, Purthoor and Kanayama].
19. The electronic device of claim 18, wherein, for the inputting of the data and the PIM instruction, the accelerator is configured to input the data and the PIM instruction to a port of a target PC having a lowest index among the one or more target PCs [The rationale in the rejection of claim 8 is herein incorporated but the combination of Kalamatianos, Puthoor, Dutu and Kanayama applies to claim 19 instead of the combination of Kalamatianos, Dutu, Puthoor and Kanayama].
RELEVANT ART CITED BY THE EXAMINER
The following prior art made of record and not relied upon is cited to establish the level of skill in the applicant’s art and those arts considered reasonably pertinent to applicant’s disclosure. See MPEP 707.05(c).
Puthoor et al. (US 2023/0195375) teaches “Process isolation for a PIM device through exclusive locking includes receiving, from a process, a call requesting ownership of a PIM device. The request includes one or more PIM configuration parameters. The exclusive locking technique also includes granting the process ownership of the PIM device responsive to determining that ownership is available. The PIM device is configured according to the PIM configuration parameters.” (Abstract).
ACKNOWLEDGEMENT OF ISSUES RAISED BY APPLICANT
Response to Amendment
Applicant’s amendments filed on 6/25/2026 have overcome the 35 USC 112 rejections; thus, these rejections are herein withdrawn.
Applicant's arguments filed on 6/25/2026 with respect to the 35 USC 103 rejections have been considered but are moot in view of the new ground(s) of rejection.
CLOSING COMMENTS
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
a. STATUS OF CLAIMS IN THE APPLICATION
a(1) CLAIMS REJECTED IN THE APPLICATION
Per the instant office action, claims 1-5, 7-16 and 18-20 have received a first action on the merits and are subject of a first action non-final.
a(2) ALLOWABLE SUBJECT MATTER
Per the instant office action, claims 6 and 17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
6. The method of claim 1, wherein the determining of the target PC set comprises determining the target PC set based on a level value input from a processor of an accelerator and predetermined tree logic.
17. The electronic device of claim 11, wherein, for the determining of the target PC set, the accelerator is configured to determine the target PC set based on a level value input from a processor of the accelerator and predetermined tree logic.
b. DIRECTION OF FUTURE CORRESPONDENCES
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YAIMA RIGOL whose telephone number is (571)272-1232. The examiner can normally be reached Monday-Friday 9:00AM-5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jared I. Rutz can be reached on (571) 272-5535. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
September 8, 2026
/YAIMA RIGOL/
Primary Examiner, Art Unit 2135