DETAILED ACTION
Status of Claims
This action is in reply to the communication filed on 06/05/2026.
Claims 1, 2, 5, 6, 9-13, and 19-20 have been amended.
Claims 1-20 are currently pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments filed 06/05/2026 with respect to the rejections under 35 USC § 102/103 have been considered but are not persuasive and/or are moot in view of the new grounds of rejection as described below.
On pg. 7-8 of the Remarks, Applicant essentially argues (emphasis original):
None of the cited references teach or suggest all of the limitations of amended claim 1. For example, Rajbhandari discloses "a powerful offload mechanism called the infinity offload engine which can offload all of the partitioned model states to CPU or NVMe memory, or keep them on the GPU based on the memory requirements." (P. 7, 2ⁿᵈ column). (Emphasis added). Rajbhandari further discloses that "[w]e only offload when the aggregated GPU memory is not sufficient," and that "[f]irst we offload optimizer-states and gradients to the fastest memory with enough capacity,"
…
Rajbhandari does not appear to disclose a "host computing device having a host memory and a processing circuit" that "provide[s] the first data from the host memory to an acceleration device based on determining the first attribute," and that also "provide[s] the second data from the host memory to a second device different from the acceleration device," as is now required by claim 1. (Emphasis added). Thus, claim 1 is in condition for allowance.
Examiner respectfully disagrees that Rajbhandari does not disclose “provid[ing] the first data from the host memory to an acceleration device based on determining the first attribute”. Rajbhandari specifically discloses that initial data and previously offloaded data are transferred from CPU memory to the GPU (accelerator) at the appropriate times of the ML training process, e.g. Rajbhandari pg. 7, emphasis added: “Partitioned parameters are moved from slow memory to GPU and then collected to form the full layer…Activation checkpoints are moved to CPU memory during forward pass, and moved back to GPU one layer at a time before the backward pass on the corresponding layer…ZeRO-Infinity can decrease the frequency of activation checkpoints as well as effectively overlap the communication of activation checkpoints both to and from CPU memory with the forward and backward computation on the GPU.”
Applicant’s remaining arguments regarding “provide the second data from the host memory to a second device different from the first acceleration device”, are moot in view of the new grounds of rejection.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 5-9, and 12-19 are rejected under 35 U.S.C. 103 as being unpatentable over Rajbhandari et al. (“ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning”, 11/2021, SC ’21, ACM conference publication1) in view of Kim et al. (US 2024/0127056 A1).
Claims 1 and 12:
Rajbhandari discloses the limitations as shown in the following rejections
A system comprising: a host computing device having host memory (CPU memory); and a processing circuit (CPU with offload engine) in communication with the host memory(pg. 7, Fig. 4; pg. 9, § 6.3; pg. 14).
the processing circuit being configured to: identify first data associated with a task (training) associated with a machine learning program (model); determine a first attribute (e.g. current pass, current operation of training) associated with the first data; provide the first data from the host memory to an acceleration device (GPU) based on determining the first attribute (e.g. parameters/gradients needed for current pass); identify second data corresponding to the task associated with the machine learning program; determine a second attribute (e.g. activation checkpoint type, optimizer state) associated with the second data; and provide the second data…to a second device (e.g. NVMe SSD) different from the acceleration device (see at least pg. 2, col. 1, last 2 para.; pg. 6, § 5.1.1; pg. 7, Fig. 4, description; pg. 7-8, § 5.2.3 and 5.3; pg. 9, § 7.1, para. 4) disclosing a system for parallel training of machine learning models using a heterogeneous memory architecture, including partitioning model and residual state data of the training process are partitioned and determining which portions should be transferred to GPU (acceleration device) memory from CPU memory and which should be transferred to CPU memory or NVMe solid-state drive (SSD) (second device) based on the type of data (attributes) and the stage of the training process; for example parameters needed (first data) for a current operation forward pass are stored at GPU while activations (second data) are offloaded to CPU and/or NVMe storage. Exemplary quotations:
“ZeRO-Infinity partitions [model states such as parameter, gradients and optimizer states] to fully leverage the aggregate memory across all data parallel processes. During the training, ZeRO-Infinity uses communication collectives to gather only those model states that are needed at the moment...ZeRO-Infinity offloads them to CPU or NVMe memory, moving them between GPU, CPU and NVMe as needed, allowing ZeRO-Infinity to fully leverage CPU and NVMe memory in addition to the GPU memory” (pg. 2, col. 1).
“Communication for the backward pass of the first layer is depicted. Partitioned parameters are moved from slow memory to GPU and then collected to form the full layer. After gradients are computed, they are aggregated, re-partitioned, and then offloaded to slow memory…Activation checkpoints…are moved to CPU memory during forward pass, and moved back to GPU one layer at a time before the backward pass on the corresponding layer… ZeRO-Infinity can decrease the frequency of activation checkpoints as well as effectively overlap the communication of activation checkpoints both to and from CPU memory with the forward and backward computation on the GPU” (pg. 7).
“First we offload optimizer states and gradients to the fastest memory with enough capacity, since it gives the largest memory savings with the least communication overhead. Next, between parameters and activation checkpoints, if only one needs to be offloaded to CPU memory, we empirically chose to offload the one that gives better performance. When both need to be offloaded, activation checkpoints are offloaded to CPU and parameters are offloaded to the fastest memory with enough capacity” (pg. 9).
wherein the machine learning program is configured to generate a result based on an input (see at least pg. 3, § 2, first para.; pg. 6, § 5.1.3; pg. 7, Fig. 4, description; pg. 7, col. 2;.
Rajbhandari discloses the ZeRO-Infinity system moves ML task data “between GPU, CPU and NVMe as needed” (pg. 2, col. 1) and specifically discloses: “The offload engine…allows ZeRO-Infinity to overlap NVMe to CPU reads with CPU to NVMe writes” (pg. 7, col. 2, emphasis added). But Rajbhandari’s disclosure focuses on ML task data movement between GPUs and CPU/NVMe storage devices and does not clearly anticipate provide the second data from the host memory to a second device different from the first acceleration device.
Kim, however, discloses (FIG. 1B and 2; ¶0045, 0048, 0054-0055) an analogous method and system for accelerating distributed ML training in heterogeneous GPU/CPU clusters configured to identify second data (input data) corresponding to the task associated with the machine learning program and provide the second data from host memory (host CPU DRAM) to a second device (e.g. local storage/computational storage) different from the acceleration device (GPU).
It would have been obvious to one of ordinary skill in the art prior to the filing date of the invention to modify Rajbhandari to utilize local storage/computational storage for ML task input data preprocessing as taught by Kim to accelerate the training data preprocessing to maintain model accuracy while ensuring low end-to-end latency and high-energy efficiency (Kim ¶0039-0041, 0053) .
Claims 2, 3, 13, and 14:
The combination of Rajbhandari/Kim discloses the limitations as shown in the rejections above. Rajbhandari further discloses wherein the acceleration device includes at least one of a graphics processing unit, tensor processing unit, or co-processor…wherein the second device includes a solid-state drive (Rajbhandari: NVMe; Kim: SSD) (see at least pg. 1, Abstract; pg. 6, § 5.1.1 and 5.1.2; pg. 9, § 7.1, para. 4; pg. 14: “Relevant hardware details…2x NVME SSD Samsung 960GB”). See also Kim FIG. 1B and 2; ¶0008, 0037, 0045, 0048)
Claims 5, 6, 15, and 16:
The combination of Rajbhandari/Kim discloses the limitations as shown in the rejections above. Rajbhandari further discloses wherein the first data is stored in the host memory and includes at least one of an optimizer state, gradient, or weight computed for the machine learning program…wherein the second data includes at least one of an activation…or checkpoint data of the machine learning program (see at least pg. 6, § 5.1.1 and 5.1.2; pg. 7, Fig. 4, description; pg. 7-8, § 5.3; pg. 9, § 7.1, para. 4). See also Kim FIG. 1B and 2; ¶0008, 0037, 0045, 0048) wherein the second data includes an…input batch… of the machine learning program.
Claims 7, 8, 17, and 18:
The combination of Rajbhandari/Kim discloses the limitations as shown in the rejections above. Rajbhandari further discloses wherein the first attribute includes a state (e.g. current pass and/or current operation of training) of a computing logic for training the machine learning program…wherein the second attribute includes a data type (e.g. input data or activation or checkpoint data type) (see at least pg. 6, § 5.1.1 and 5.1.2; pg. 7, Fig. 4, description; pg. 7-8, § 5.3; pg. 9, § 7.1, para. 4). See also Kim FIG. 1B and 2; ¶0008, 0037, 0045, 0048.
Claims 9 and 19:
The combination of Rajbhandari/Kim discloses the limitations as shown in the rejections above. Rajbhandari further discloses wherein the acceleration device is configured to perform a computation (operation/module) based on the first data, generate third data (e.g. gradients, activations) based on the computation, and transfer the third data, wherein the processing circuit is configured to: receive the third data; and store the third data in the host memory (CPU memory), (see at least pg. 7, § 5.2.2; pg. 7-8, § 5.3; pg. 7, Fig. 4, description: “After gradients are computed, they are aggregated, repartitoned, and then offloaded to slow memory.)”
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Rajbhandari in view of Kim in further view of Wang et al. (“Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems”, 2022).
Claim 4:
The combination of Rajbhandari/Kim discloses the limitations as shown in the rejections above. Rajbhandari/Kim does not specifically disclose the processing circuit is configured to operate using a cache coherent protocol.
Wang, however, discloses analogous heterogeneous system for parallel training of ML models comprising GPUs and disaggregated memory devices coordinated by a host CPU (processing circuit) that is configured to operate using a cache coherent protocol (see at least pg. 129; pg. 126, Abstract: “we propose COARSE, a disaggregated memory extension for distributed DL training. COARSE is built on modern cache-coherent interconnect (CCI) protocols…to allow low-latency and parallel access to training data and model parameters shared among worker GPUs.”
It would have been obvious to one of ordinary skill in the art prior to the filing date of the invention to modify Rajbhandari/Kim to employ CCI protocols as taught by Wang because:
“CCI is beneficial to data parallel training in two aspects: (i) in data parallel training, cross-device communication is a major overhead; this communication can take advantage of CCI low-latency memory access to improve the parameter synchronization performance. (ii) the parameter synchronization operations block the training procedure and take up GPU computing resources; while using CCI, GPU can work with memory device processors coherently to offload these synchronization operations to memory device processors, thus improving the GPU utilization and reducing the communication overhead.” (Wang pg. 129)
Claims 10, 11, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Rajbhandari in view of Kim in further view of Chang et al. (US 2019/0311257 A1).
Claims 10, 11, and 20:
The combination of Rajbhandari/Kim discloses the limitations as shown in the rejections above. Rajbhandari/Kim does not specifically disclose receive a signal indicative of a state of the acceleration device; and update a parameter associated with the task…wherein the state of the acceleration device includes available memory of the acceleration device, and the parameter includes a batch size of training data for training the machine learning program.
Chang, however, discloses analogous heterogeneous system for parallel training of ML models by a plurality of GPU work servers (acceleration device) and including an advisor server (processing circuit) that is configured to receive a signal indicative of a state of the acceleration device; and update a parameter associated with the task…wherein the state of the acceleration device includes available memory of the acceleration device, and the parameter includes a batch size of training data for training the machine learning program (see at least ¶0006, 0039-0042, 0051-0052, 0056). Exemplary quotation: “Advisor server 140 may also determine a batch size for each work server…controller 132 may also report throughput parameters (e.g., idle time, duration spent processing the batch, free GPU memory, etc.) to advisor server 140. Advisor server 140 may then update a batch size for work server” (¶0052).
It would have been obvious to one of ordinary skill in the art prior to the filing date of the invention to modify Rajbhandari/Kim to dynamically update batch sizes as taught by Chang to prevent model mis-training and to increase the overall speed of the training process (¶0002-0003, 0055-0056).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure:
“Efficient Utilization of Heterogeneous Compute and Memory Systems” divides model training tasks between CPUs and GPUs based on the compute patterns and hierarchically utilizes various memory types, such as HBM, DRAM, and SCMs.
“STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model Training” extends GPU memory for model training using host memory and SSD.
US 20230259747 A1, US 11704572 B1, US 20230229899 A1 disclose training systems with heterogenous storage.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry of a general nature or relating to the status of this application or concerning this communication or earlier communications from the Examiner should be directed to Paul Mills whose telephone number is 571-270-5482. The Examiner can normally be reached on Monday-Friday 11:00am-8:00pm. If attempts to reach the examiner by telephone are unsuccessful, the Examiner’s supervisor, April Blair can be reached at 571-270-1014.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/P. M./
Paul Mills
08/28/2026
/APRIL Y BLAIR/Supervisory Patent Examiner, Art Unit 2196
1 For clarity of record, Examiner notes the reference Rajbhandari cited in the rejection refers to the reference provided with the 03/09/2026 NFOA, which is a different version than the paper with same title cited in the IDS dated 02/10/2025.