Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the arguments do not apply to any of the references being used in the current rejection.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Arur et al. (US 2025/0225369, hereinafter Arur) in view of Kaitha (US 2023/0011315, hereinafter Kaitha).
Regarding claim 1, Arur discloses
A method for managing workload placement, the method comprising (fig. 1-6):
obtaining, by a workload placement service, a request for assigning a model adaptation workload to one of a plurality of production environments (paragraphs [0020], [0060]: an orchestration service and scheduler (i.e., the workload placement service) that receives new ML models for training and assigns them to various computing resource groups/datacenters (i.e., production environments)) based on completion time (paragraph [0031]: The training constraints may further include … an expected training time (e.g., a specified time deadline); paragraph [0071]: the objectives may include minimizing total training time by selecting data centers with required hardware being available);
in response to the request:
performing an initial workload placement to assign the model adaptation workload to a first production environment of the plurality of production environments (paragraphs [0067], [0077]: the scheduler computes and generates a deployment plan to deploy these ML models to various computing resource groups (e.g., assigning a model to a first datacenter, DC1);
wherein the initial workload placement is performed using a workload placement model that inputs a parameter to be optimized and a type of the model adaptation workload and outputs the first production environment as a selection (paragraphs [0070]-[0073]: the scheduler applies a multi-objective genetic algorithm (i.e., workload placement model) that evaluates model requirements/architecture (i.e., type of the model adaptation workload) and optimization objectives (i.e., parameter to be optimized, such as minimizing training time) to output a datacenter selection);
after performing the initial workload placement (paragraphs [0067], [0077]: the scheduler computes and generates a deployment plan to deploy these ML models to various computing resource groups (e.g., assigning a model to a first datacenter, DC1), monitoring:
execution of the model adaptation workload on the first production
environment, and
performance of computing resource in the plurality of production environments to obtain telemetry data associated with the execution and the performance (paragraphs [0059], [0074]: the system continuously monitors the computing resources, hardware capabilities, and changing conditions across the datacenters; paragraph [0080]: The training service is further configured to monitor all ML model training jobs and track their iteration progress … Performance metrics like accuracy and loss are pushed back to the orchestration service);
performing a completion time analysis using the telemetry data to generate a placement recommendation (paragraph [0060], [0074]: the scheduler periodically reruns its algorithm during active training jobs to account for changing conditions; paragraph [0072]: The scheduler 310 further estimates training times based on model requirements, hardware mapping, and data center capabilities; paragraphs [0070]-[0074]: the scheduler evaluates placements to generate an updated deployment plan (the placement recommendation) that meets the optimization objectives (which includes minimizing training time);
making a determination that the placement recommendation specifies a second production environment of the plurality of production environments (paragraph [0077]: Two ML models (e.g., codependent models) are then to be migrated (transferred) to the second computing resource group 420b for continued training at the second computing resource group 420b); and
wherein the computing resources comprise: a number of graphics processing units (GPUs), central processing units (CPUs), memory components, storage capability, and network bandwidth (paragraph [0059]: the asset inventory includes the number of GPUs, TPUs, CPUs, available storage, memory resources, and network bandwidth);
based on the determination, initiating deployment of the model adaptation
workload to the second production environment (paragraph [0081]: when notified (e.g., instructions in the deployment plan), the training service pauses an ML model training job, checkpoints it, and transfer the ML model training job to the newly assigned computing resource group (migration of ML model); paragraph [0083]: when the orchestration service 410 instructs to migrate the neural network 426 (that uses training data set stored the a training data set storage 428) to the second computing resource group 420b … training can resume from the last point rather than restarting).
Auru does not explicitly disclose wherein the second production environment comprises more computing resources than the first production environment. Kaitha discloses wherein the second production environment comprises more computing resources than the first production environment (paragraph [0048]: if the test… indicates that system 152 is running without memory resources capable of performing operations within an allotted timeframe, the test infrastructure may reapportion memory… to support the continued operation; paragraph [0057]: Environment engine 325 included in augmented decisioning engine 250 may be used to configure, automatically, test infrastructure 130 based on a type of metrics data retrieved by data engine 315 from production environment 140. The type of metrics data retrieved from production environment 140 may correspond to one or more of CPU usage, memory usage, other system overhead limitations, network downlink rate, network uplink rate, other network bandwidth limitations, application logs, overall speed, responsiveness, or stability of a system in production environment 140 or production environment 280 (e.g., production infrastructure 150 and/or any other computing system/resource that may be included therein); paragraph [0058]: test engine 330 … configured to manage, allocate, or otherwise reapportion any combination of the computing clusters, servers, databases, applications, or other computing resources included in test infrastructure 130 to production infrastructure 150 based on the recommended configuration of production infrastructure 150; Note: Kaitha discloses monitoring the performance metrics of a production infrastructure against expected performance values (e.g., speed, responsiveness, allotted timeframe) and, based on the comparison, generating a recommendation to allocate or reapportion additional computing resources (e.g., computing clusters, servers, memory) to the infrastructure to support the continued operation of the system and meet the expected performance metrics. Specifically, Kaitha teaches that if a system lacks the resources to perform operations within an allotted timeframe, additional resources are allocated to overcome the deficit).
In other words, Kaitha’s initial, under-resourced infrastructure corresponds to the claimed “first production environment,” and Kaitha’s reconfigured infrastructure, which has been allocated additional computing clusters, servers, and memory to meet the allotted timeframe, corresponds to the claimed “second production environment comprising more computing resources.”
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the workload migration method of Arur to include migrating the workload to a second production environment that comprises more computing resources than the first production environment, applying the resource-scaling principles taught by Kaitha. The motivation would have been to ensure that strict completion-time deadlines and SLA are met (Kaitha paragraph [0039]).
Regarding claim 8 referring to claim 1, Arur discloses A non-transitory computer readable medium comprising computer readable program code, which when executed by a computer processor enables the computer processor to perform a method for managing information handling systems, the method comprising: … (paragraph [0111]).
Regarding claim 15 referring to claim 1, Arur discloses A system, comprising: a processor; and memory including instructions, which when executed by the processor, perform a method comprising: … (FIG. 6).
Regarding claims 2, 9, and 16, Arur discloses
wherein the model adaptation workload comprises performing a parameter-efficient fine-tuning (PEFT) process on a training workload (paragraph [0016]: Parameter Efficient Fine Tuning (PEFT) and Low-Rank Adaptation (LoRA)), and wherein the training workload comprises a training of a generative artificial intelligence (AI) model using a training dataset (paragraph [0011]: generative pre-trained transformers (GPT) and large language models (LLMs); paragraphs [0031], [0079]: training dataset).
Regarding claims 3, 10, and 17, Arur discloses
wherein the generative AI model is utilized by a front-end environment to obtain an inferencing payload (paragraph [0011]: AI/ML models are deployed to help with enterprise tasks and gain insights; paragraphs [0026]-[0027]: deploying ML models for hosting and computing functions).
Regarding claims 4 and 11, Arur discloses
wherein the completion time analysis is further based on causal variables associated with the completion time (paragraph [0072]: estimating training times based on hardware mapping and data center capabilities).
Regarding claims 5, 12, and 18, Arur discloses
wherein the causal variables comprise at least one of: clock speed of a graphics processing unit (GPU) of the first production environment, a number of GPUs used for the model adaptation workload in the first production environment, a second number of GPUs available in the second production environment, and latency between GPUs executing the model adaptation workload (paragraphs [0015], [0032], [0059]: tracking the number of GPUs/TPUs available; paragraph [0034] Insufficient network capacity between GPUs causes delays that slow down training. The network constraints may further include the network bandwidth needed for migrating the ML models to a different datacenter).
Regarding claims 6, 13, and 19, Arur discloses
wherein the first production environment is a computing device of an on-premise environment (paragraphs [0024], [0027]: geographically remote enterprise sites and local profilers).
Regarding claims 7, 14, and 20, Arur discloses
wherein the first production environment is a computing device of a cloud environment (paragraphs [0051], [0063]: the service resides in one or more cloud(s)).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Nia et al. (US 2022/0188645) discloses evaluating and adapting machine learning models by generating counterfactuals to explain and adjust model behavior, which relates generally to the claimed model adaptation workload (paragraphs [0004]-[0005], [0034]) .
Mehrotra et al. (US 10,467,000) discloses monitoring the execution and performance of an application to obtain usage logs (i.e., telemetry data) and performing an analysis on that data to generate a migration or deployment recommendation, which relates to the claimed steps of monitoring execution to obtain telemetry data and generating a placement recommendation (col. 6, line 34-col. 7 line 47).
Cheng et al. (US 2025/0117699) discloses utilizing machine learning models to determine the optimal device placement of machine learning operations across CPUs and GPUs based on time-based parameters such as latency and throughput, which relates to the claimed workload placement model and the computing resources comprising CPUs and GPUs (paragraphs [0003]-[0006], [0073]-[0074]).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in [0037] CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to [0037] CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SISLEY N. KIM whose telephone number is (571)270-7832. The examiner can normally be reached M-F 11:30AM -7:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, April Y. Blair can be reached on (571)270-1014. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SISLEY N KIM/Primary Examiner, Art Unit 2196 9/12/2026