DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Action is in response to communications filed 08/14/2025.
Claims 1-20 are pending.
Claims 1-20 are rejected.
Priority
Applicant’s priority claim to foreign document KR10-2025-0022900 filed 02/21/2025 is herein acknowledged.
It is noted, however, that applicant has not filed a certified copy of the application as required by 37 CFR 1.55. The Examiner reminds Applicant to file the certified copy as required in order to properly realize the claim for priority.
Information Disclosure Statement
As required by M.P.E.P. 609(C), the applicant’s submission of the Information Disclosure Statement dated 08/14/2025 is acknowledged by the examiner and the cited references have been considered in the examination of the claims now pending. As required by M.P.E.P 609 C(2), a copy of the PTOL-1449 initialed and dated by the examiner is attached to the instant office action.
Drawings
The applicant’s drawings submitted on 08/14/2025 are acceptable for examination purposes.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103(a) are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, 9-8, 12-16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Marcovitch et al. (US 2021/0406179) in view of Zhang et al. (US 2026/0003681).
Regarding claim 1, Marcovitch discloses, in the italicized portions, a storage device, comprising: a shared memory storing first data received from a first computing node and second data received from a second computing node ([0023] In some embodiments, the work-request initiators in the group provide their inputs by modifying a value in a shared memory location that is accessible over the network. The shared memory location may reside, for example, in a memory of the network device, in a memory of the compute node performing the distributed operation, or in any other memory that is accessible to the work-request initiators and to the network device.); and a memory controller configured to: perform an instruction on the first data and the second data ([0037] In the example of FIG. 1, network device 24 comprises processing hardware (H/W) 36, a memory 40, and a controller 44 that runs suitable software (S/W). [0038] Typically, controller 44 holds a definition of the operation to be performed. For example, when network device 24 comprises a network adapter of a certain compute node 28, controller 44 may receive the definition of the operation from the CPU of this compute node.); store third data, which is a result of performing the instruction, in the shared memory; and transmit the third data to the first computing node and the second computing node. Herein Marcovitch discloses a shared memory device as part of a network of connected compute nodes wherein the memory of the shared memory device is targeted by processing operations which stores inputs from the plurality of compute nodes and the controller of the shared memory device executes an operation utilizing the input data. Marcovitch does not explicitly address the how the shared memory device manages storing data from the result of the executed operation and transmitting the result data back to the compute nodes. Regarding these limitations, Zhang discloses in Paragraphs [0046-47] “[0046] The shared memory 210 may be shared by multiple devices including the management processor 230 and the N PEs 250.sub.k's (k=1, . . . , N). It may include a shared static random-access memory (SRAM) 212 and an HBM 214… It may have buffered input/output interfaces to allow access from multiple devices… The shared memory controller 220 controls the shared memory 210 including the SRAM and HBM control such as read/write controls, row and column addresses, pre-charge control, and bank select. [0047] The management processor 230 performs the management functions for the shared memory 210 and the processing operations within itself and the PEs 250.sub.k's (k=1, . . . , N). It may communicate with one or more PEs 250.sub.k's via the bus 240 and/or the communication channel 270.” Herein Zhang discloses the structure of the shared memory which comprises a buffered interface to temporarily store inputs and outputs before operation execution and result transfer back to involved compute devices. In this manner, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the shared memory to temporarily store result data prior to transmitting the data to respective nodes in order to provide the known functionality in the art of buffered inputs and outputs for controlling data flow between components (Zhang [0029]). Marcovitch and Zhang are analogous art because they are from the same field of endeavor of managing connected processing elements to a shared memory device.
Regarding claim 2, Marcovitch and Zhang in combination further disclose the storage device of claim 1, wherein the first data is application data corresponding to the first computing node, and wherein the second data is application data corresponding to the second computing node (Marcovitch [0019] In the present context, the term “distributed operation” refers to any operation whose execution depends on inputs from a plurality of entities, e.g., software processes and/or compute nodes. The entities that provide inputs to a distributed operation are referred to herein as “work-request initiators” (WRIs). A WRI may comprise, for example, a remote compute node, a local process, a thread or other entity, and/or an update issued by a network device. The work-request initiators of a given distributed operation may reside on different compute nodes and/or share the same compute node. [0027] The network device monitors whether the value in the shared memory location renders the condition true. When the condition is met, the network device triggers execution of the distributed operation.). Herein Marcovitch discloses the network device waits for each respective node to transmit data to the shared memory before executing the operation which requires data from each involved node.
Regarding claim 6, Marcovitch and Zhang in combination further disclose the storage device of claim 1, wherein the memory controller performs an operation including at least one of Max, Min, Sum, Multiply, AND, OR, Bitwise AND, or Bitwise OR operation in the instruction (Marcovitch [0089] FIG. 6 is a diagram that schematically illustrates memory-based synchronization of a reduction operation, in accordance with an embodiment of the present invention. In the present example, multiple work-request initiators 124A-124Z provide inputs to two separate distributed summation operations, using RDMA WRITE commands. The inputs to the first summation operation are denoted A0-Z0. The inputs to the second summation operation are denoted A1-Z1. The total number of work-request initiators is denoted numTargets.). Herein Marcovitch identifies performing a summation operation as part of the reduction operation involving the inputs from the initiators.
Regarding claim 7, Marcovitch discloses, in the italicized portions, a computing system, comprising: a plurality of computing nodes; and a storage device configured to: receive input data including data corresponding to each of the plurality of computing nodes from the plurality of computing nodes ([0023] and [0031]); perform an instruction on the input data ([0037-38]); and transmit output data which is a result of performing the instruction to the plurality of computing nodes. Herein Marcovitch discloses a shared memory device as part of a network of connected compute nodes wherein the memory of the shared memory device is targeted by processing operations which stores inputs from the plurality of compute nodes and the controller of the shared memory device executes an operation utilizing the input data. Marcovitch does not explicitly address the how the shared memory device transmits the result data back to the compute nodes. Regarding this limitation, Zhang discloses in Paragraphs [0046-47] the structure of the shared memory which comprises a buffered interface to temporarily store inputs and outputs before operation execution and result transfer back to involved compute devices. Claim 7 is rejected on a similar basis as claim 1.
Regarding claim 8, Marcovitch and Zhang in combination further disclose the computing system of claim 7, wherein the input data includes application data corresponding to each of the plurality of computing nodes (Marcovitch [0019] and [0027]). Claim 8 is rejected on a similar basis as claim 2.
Regarding claim 12, Marcovitch and Zhang in combination further disclose the computing system of claim 7, wherein the storage device includes a shared memory shared by the plurality of computing nodes and storing the input data (Marcovitch [0023]). Herein Marcovitch identifies the network device as comprising a shared memory accessible by the plurality of nodes.
Regarding claim 13, Marcovitch and Zhang in combination further disclose the computing system of claim 7, wherein the storage device performs an operation including at least one of Max, Min, Sum, Multiply, AND, OR, Bitwise AND, or Bitwise OR operation in the instruction (Marcovitch [0089]). Claim 13 is rejected on a similar basis as claim 6.
Regarding claim 14, Marcovitch discloses, in the italicized portions, a method of operating a storage device including a shared memory, the method comprising: receiving first data from a first computing node and second data from a second computing node; storing the first data and the second data in the shared memory ([0023]); performing an instruction on the first data and the second data ([0037-38]); and storing third data, which is a result of performing the instruction, in the shared memory. Herein Marcovitch discloses a shared memory device as part of a network of connected compute nodes wherein the memory of the shared memory device is targeted by processing operations which stores inputs from the plurality of compute nodes and the controller of the shared memory device executes an operation utilizing the input data. Marcovitch does not explicitly address the how the shared memory device stores the result data of the instruction. Regarding this limitation, Zhang discloses in Paragraphs [0046-47] the structure of the shared memory which comprises a buffered interface to temporarily store inputs and outputs before operation execution and result transfer back to involved compute devices. Claim 14 is rejected on a similar basis as claim 1.
Regarding claim 15, Marcovitch and Zhang in combination further disclose the method of claim 14, further comprising transmitting the third data to the first computing node and the second computing node (Zhang [0046-47]). Herein Zhang discloses the buffers are used to temporarily store data prior to transmitting the data to respective nodes.
Regarding claim 16, Marcovitch and Zhang in combination further disclose the method of claim 14, wherein the first data is application data corresponding to the first computing node, and wherein the second data is application data corresponding to the second computing node (Marcovitch [0019] and [0027]). Claim 8 is rejected on a similar basis as claim 2.
Regarding claim 20, Marcovitch and Zhang in combination further disclose the method of claim 14, wherein performing the instruction includes performing an operation including at least one of Max, Min, Sum, Multiply, AND, OR, Bitwise AND, or Bitwise OR operation on the first data and the second data (Marcovitch [0089]). Claim 20 is rejected on a similar basis as claim 6.
Claims 3, 9, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Marcovitch in view of Zhang and further in view of Paul et al. (US 2022/0004488).
Regarding claim 3, Marcovitch and Zhang do not explicitly disclose the storage device of claim 1, wherein the memory controller communicates with the first computing node and the second computing node via a Compute Express Link (CXL) interface. Regarding this aspect of the limitation, Paul discloses in Paragraphs [0019-20] “[0019] The node can communicate with other devices in the pooled memory architecture (e.g., with shared memory controller 215, networking controller 220, other nodes, etc.) using one or more protocols, including Compute Express Link (CXL) (an interconnect technology for removable high-bandwidth devices, such as GPU-based compute accelerators, in a data-center environment), PCIe, QPI, Ethernet, among other examples. [0020] SMC 215 may include logic for handling load/store requests of nodes 210a-210n. Load/store requests can be received by the SMC 215 over links (such as CXL) connecting the nodes 210a-210n to the SMC 215.” Herein Paul discloses a shared memory architecture between a plurality of processing units over a Compute Express Link (CXL) structure. It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize a CXL structure for shared memory pooling in order to reduce latency costs and improve memory utilization among a plurality of nodes (Paul [0026]). Marcovitch, Zhang, and Paul are analogous art because they are from the same field of endeavor of managing connected processing elements to a shared memory device.
Regarding claim 9, Marcovitch and Zhang do not explicitly disclose the computing system of claim 7, wherein the storage device communicates with the plurality of computing nodes via a Compute Express Link (CXL) interface. Regarding this aspect of the limitation, Paul discloses in Paragraphs [0019-20] a shared memory architecture between a plurality of processing units over a Compute Express Link (CXL) structure. Claim 9 is rejected on a similar basis as claim 3.
Regarding claim 17, Marcovitch and Zhang do not explicitly disclose the method of claim 14, wherein the storage device communicates with the first computing node and the second computing node via a Compute Express Link (CXL) interface. Regarding this aspect of the limitation, Paul discloses in Paragraphs [0019-20] a shared memory architecture between a plurality of processing units over a Compute Express Link (CXL) structure. Claim 17 is rejected on a similar basis as claim 3.
Claims 4, 10, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Marcovitch in view of Zhang and further in view of Archer et al. (US 2011/0239003).
Regarding claim 4, Marcovitch and Zhang do not explicitly disclose the storage device of claim 1, wherein the memory controller performs the instruction based on a Message Passing Interface (MPI) protocol. Regarding this aspect of the limitation, Archer discloses in Paragraphs [0020-21] and [0043] “[0020] Within each compute node (102) of FIG. 1, a host computer (110) and one or more accelerators (104) are adapted to one another for data communications by a system level message passing module (`SLMPM`) (146) and by two or more data communications fabrics (106, 107) of at least two different fabric types. An SLMPM (146) is a module or library of computer program instructions that exposes an application programming interface (`API`) to user-level applications for carrying out message-based data communications between the host computer (110) and the accelerator (104). Examples of message-based data communications libraries that may be improved for use as an SLMPM according to embodiments of the present invention include: [0021] the Message Passing Interface or `MPI,` an industry standard interface in two versions. [0043] Monitoring data communications performance for a plurality of data communications modes may also include monitoring utilization of a shared memory space (158). In the example of FIG. 2, shared memory space (158) is allocated in RAM (140) of the accelerator. Utilization is the proportion of the allocated shared memory space to which data has been stored for sending to a target device and has not yet been read or received by the target device, monitored by tracking the writes and reads to and from the allocated shared memory.” Herein Archer discloses a shared memory architecture between a plurality of processing units communicating via message-based data communication libraries including MPI. It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize a MPI structure for shared memory pooling in order to utilize the well-known communication standard in the art (Archer [0024]) as the claim does not otherwise distinguish any particular purpose for utilizing the protocol which would distinguish over the known use. Marcovitch, Zhang, and Archer are analogous art because they are from the same field of endeavor of managing connected processing elements to a shared memory device.
Regarding claim 10, Marcovitch and Zhang do not explicitly disclose the computing system of claim 7, wherein the storage device performs the instruction based on a Message Passing Interface (MPI) protocol. Regarding this aspect of the limitation, Archer discloses in Paragraphs [0020-21] and [0043] a shared memory architecture between a plurality of processing units communicating via message-based data communication libraries including MPI. Claim 10 is rejected on a similar basis as claim 4.
Regarding claim 18, Marcovitch and Zhang do not explicitly disclose the method of claim 14, wherein the instruction is performed based on a Message Passing Interface (MPI) protocol. Regarding this aspect of the limitation, Archer discloses in Paragraphs [0020-21] and [0043] a shared memory architecture between a plurality of processing units communicating via message-based data communication libraries including MPI. Claim 18 is rejected on a similar basis as claim 4.
Claims 5, 11, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Marcovitch in view of Zhang and further in view of Raj et al. (US 2021/0224213).
Regarding claim 5, Marcovitch and Zhang do not explicitly disclose the storage device of claim 1, wherein the memory controller includes a Near Data Processing (NDP) engine performing the instruction on the first data and the second data. Regarding this aspect of the limitation, Raj discloses in Paragraphs [0059] and [0081] “[0059] In some examples, for logic flow 900 at block 920 an NDP such as NDP 522-1 may be discoverable by a device driver for an OS such as an OS executed by one or more elements of processor 501. Discovering NDP 522-1 may include recognizing that NDP 522-1 is an accelerator resource of memory controller 520-1 to facilitate near data processing of data primary stored to HBM 530-1. [0081] According to some examples, memory system 1330 may include a controller 1332 and a memory 1334. For these examples, circuitry resident at or located at controller 1332 may be included in a near data processor and may execute at least some processing operations or logic for apparatus 1000 based on instructions included in a storage media that includes storage medium 1200.” Herein Raj discloses a shared memory architecture between a plurality of processing units including a near data processor to execute at least some processing operations instead of primary cores in the computing system. It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize near data processors to offload tasks from the processor cores in view of data locality to improve system throughput (Raj [0029]). Marcovitch, Zhang, and Raj are analogous art because they are from the same field of endeavor of managing connected processing elements to a shared memory device.
Regarding claim 11, Marcovitch and Zhang do not explicitly disclose the computing system of claim 7, wherein the storage device includes a Near Data Processing (NDP) engine performing the instruction on the input data. Regarding this aspect of the limitation, Raj discloses in Paragraphs [0059] and [0081] a shared memory architecture between a plurality of processing units including a near data processor to execute at least some processing operations instead of primary cores in the computing system. Claim 11 is rejected on a similar basis as claim 5.
Regarding claim 19, Marcovitch and Zhang do not explicitly disclose the method of claim 14, wherein the instruction is performed by a Near Data Processing (NDP) engine of the storage device. Regarding this aspect of the limitation, Raj discloses in Paragraphs [0059] and [0081] a shared memory architecture between a plurality of processing units including a near data processor to execute at least some processing operations instead of primary cores in the computing system. Claim 19 is rejected on a similar basis as claim 5.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Liu et al. (US 2024/0362165) – Paragraph [0007] wherein managing access to a shared memory pool is discussed.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER J YOON whose telephone number is (408)918-7629. The examiner can normally be reached on Monday-Friday 8am-3pm ET. The examiner’s email is alexander.yoon2@uspto.gov.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jared Rutz can be reached on 571-272-5535. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER YOON/
Examiner, Art Unit 2135
/JARED I RUTZ/Supervisory Patent Examiner, Art Unit 2135