Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claim Objections
Claim 17 is objected to because of the following informalities:
Claim 17 recites “One” in line 2. It should be “one”. Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
1. Claims 1-15, 17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kurkure et al., IDS, U.S Patent Application Publication No.2024/0362052 (“Kurkure”) in view of QIU et al, IDS, CN118377598 (English translated) (“QIU”) further in view of Kuperman et al., IDS, U.S Patent No.12236193 (“Kuperman”) further in view of Dhakal et al., U.S Patent No.12346252 (“Dhakal”)
Regarding independent claim 1, Kurkure teaches a computer-implemented method (Fig.1) comprising:
analyzing an inference request to determine a set of parameters of execution corresponding to the inference request(see at least [0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: [0021] 1. A set of coefficients corresponding to the compute slices and memory slices requested by each VM i for i=1, . . . , M; [0022] 2. a set of decision variables that indicate, among other things, whether a given VM i is placed on a given GPU j for j=1, . . . , N per the problem solution; [0023] 3. a set of constraints that ensures, among other things, that (a) each (placed) VM i is placed on a single GPU j with sufficient compute and memory resources to satisfy the VM's requested requirements, (b) the total allocated compute and memory resources on each GPU j do not exceed its maximum capacity, and (c) each VM i is placed on at most one GPU j; and [0024] 4. an objective function corresponding to the total number of GPUs used.”; [0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”);
a set of multi-instance Graphical Processing Units (GPUs) (MIGs), a MIG in the set of MIGs comprising a set of slices of a corresponding GPU (set of MIG slices) (see at least [0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.; [0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”);
causing, by sending a set of instructions to a controller associated with the MIG, the controller to modify a second amount of the computing resource available to a MIG slice in the set of MIG slices (see at least [0008] Embodiments of the present disclosure are directed to techniques for optimally placing clients on GPUs under the MIG model. As used herein, the phrase “placing a client on a GPU” refers to the act of allocating portions of the resources of the GPU for use by that client, typically in accordance with the client's requirements. Once placed in this manner, the client can consume the GPU resources allocated to it over the course of its execution.”; [0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available.”; [0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.”); and
scheduling the inference request to execute using the first amount of computing resource at the MIG slice (see at least [0013] By way of example, FIG. 2 depicts a scenario 200 in which a GPU 202 is partitioned into three MIG instances 204 (1)-(3). These MIG instances are in turn assigned to VMs 206 (1)-(3) respectively, and thus these VMs are said to be “placed” on GPU 202. Each MIG instance 204 includes a separate execution path through GPU 202 that includes a dedicated portion of GPU 202's compute resources (reference numeral 208) and a dedicated portion of GPU 202's memory resources (reference numeral 210). This advantageously ensures that the GPU workload of each VM 206 runs with predictable quality of service (e.g., throughput, latency, etc.) and prevents one VM from impacting the work or scheduling of another. [0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: [0021] 1. A set of coefficients corresponding to the compute slices and memory slices requested by each VM i for i=1, . . . , M; [0022] 2. a set of decision variables that indicate, among other things, whether a given VM i is placed on a given GPU j for j=1, . . . , N per the problem solution; [0023] 3. a set of constraints that ensures, among other things, that (a) each (placed) VM i is placed on a single GPU j with sufficient compute and memory resources to satisfy the VM's requested requirements, (b) the total allocated compute and memory resources on each GPU j do not exceed its maximum capacity, and (c) each VM i is placed on at most one GPU j; and [0024] 4. an objective function corresponding to the total number of GPUs used. [0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU. [0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”) Kurkure is understood to be silent on the remaining limitations of claim 1.
In the same field endeavor, QIU teaches analyzing an inference request to determine a set of parameters of execution corresponding to the inference request (see at least [n0022] The multi-instance GPU inference service job resource performance prediction module also performs a small amount of online testing and evaluation, updates the model, and tries to fit the actual performance of the inference service job as closely as possible: the serialization model optimization algorithm determines the resource parameter combination to be tested in the next round, starts the corresponding instance through the inference service job image, and the test evaluation client sends a request to the instance to monitor the request processing time, resource utilization, and CPU utilization information; the serialization model optimization algorithm updates and corrects the resource performance prediction model for multi-instance GPUs after multiple iterations and a small amount of online testing and evaluation, and finally predicts the performance of the inference service job under various resource parameter combinations using the resource performance prediction model for multi-instance GPUs. Finally, the algorithm server returns the prediction results to the user and stores them in the database.”);
computing, a first amount of a computing resource that will be needed to produce and store the output (see at least [n0012] The multi-instance GPU elastic resource management module communicates with each agent in the cluster and sends information to the multi-instance GPU inference service job management module to manage the lifecycle of all inference service jobs in the cluster. The inference service job controller manages all state changes during the lifecycle. Furthermore, the multi-instance GPU inference service job management module communicates with the multi-instance GPU inference service job resource performance prediction module and the inference service job online scheduling module to perform calculations. [n0015]Initialization of the inference service job: Based on the service level target of the inference service job and the relationship between job resources and performance, calculate the minimum number of GPU instances required for the inference service job to meet the service level target, determine the number of Pods accordingly, create the required number of Pods, update the status of the inference service job to Pending, and thus set a globally unique job ID for the job.[n0020] The multi-instance GPU inference service job resource performance prediction module uses a resource performance prediction model for multi-instance GPUs. It incorporates the structural features of the inference model, namely FLOPs, into the construction of the prediction model. When a user's performance prediction request is received, the module first queries the database to see if there is historical data for the inference service job. If so, it returns the result directly. Then, it parses the inference model through the model structure parsing component to obtain the model's structural information. This component integrates the tensorflow.profiler and thop libraries, which are used to parse tensorflow and pytorch models, respectively., [n0022] The multi-instance GPU inference service job resource performance prediction module also performs a small amount of online testing and evaluation, updates the model, and tries to fit the actual performance of the inference service job as closely as possible: the serialization model optimization algorithm determines the resource parameter combination to be tested in the next round, starts the corresponding instance through the inference service job image, and the test evaluation client sends a request to the instance to monitor the request processing time, resource utilization, and CPU utilization information; the serialization model optimization algorithm updates and corrects the resource performance prediction model for multi-instance GPUs after multiple iterations and a small amount of online testing and evaluation, and finally predicts the performance of the inference service job under various resource parameter combinations using the resource performance prediction model for multi-instance GPUs. Finally, the algorithm server returns the prediction results to the user and stores them in the database.”; [n0061] First, a resource performance prediction algorithm for multi-instance GPUs was designed. The resource performance prediction model adopts the random forest regression model. The computational cost of the model, namely the number of floating point operations (FLOPs), represents the number of floating point operations required to run the model once, which can be used to characterize the complexity of the deep learning model structure.This invention incorporates the structural features of the inference model, namely FLOPs, into the construction of the prediction model and completes the pre-training of the regression model. When a new job is submitted, the Sequential Model-based Optimization (SMBO) technique is also used.);the output comprising a set of results produced while processing the inference request by executing using a set of multi-instance Graphical Processing Units (GPUs) (MIGs) (see at least [n0005] Question 1: How to characterize the relationship between job execution efficiency and GPU fragmentation types? MIG can divide the GPU into 5 different GPU slices, each with different computing and memory resources. Based on these 5 GPU partitioning methods, MIG can configure a single GPU with 19 different resource partitioning schemes. The execution efficiency of deep learning inference jobs varies depending on the partitioning scheme. Furthermore, the processing latency of inference jobs also varies greatly depending on the batch size and CPU configuration. Finding the relationship between a massive set of configuration parameters and job execution efficiency has become a thorny issue. As shown below, there are 180 different parameter combinations in a multi-instance GPU scenario.”);
causing, by sending a set of instructions to a controller associated with the MIG, the controller to modify a second amount of the computing resource available to a MIG slice (see at least [n0012] The multi-instance GPU elastic resource management module communicates with each agent in the cluster and sends information to the multi-instance GPU inference service job management module to manage the lifecycle of all inference service jobs in the cluster. The inference service job controller manages all state changes during the lifecycle. Furthermore, the multi-instance GPU inference service job management module communicates with the multi-instance GPU inference service job resource performance prediction module and the inference service job online scheduling module to perform calculations. [n0017] Running, the inference service job is running normally: Check if the inference service job has been deleted. If so, delete all Pods for the job, reclaim the GPU instance resources allocated to the Pods, check the status of all Pods for the inference service job. If all are Running, the job continues to be in the Running state. If any Pod has been deleted, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. If any Pod is in the Failed state, delete that Pod, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. Check if the throughput requirements of the inference service job have changed. If there are changes, trigger the elastic resource allocation process. The inference service job controller adopts a two-phase operation: first add GPU instances, then remove GPU instances. First, it interacts with the elastic resource management module to determine the GPU instances to be added, creates Pods and completes the injection of GPU resources, and updates the status of the inference service job to Add. If no GPU instances are needed, no Pods are created, but the job status is still updated to Add. [n0019] Delete, the job resource elastic allocation and reclamation phase; the controller communicates with the elastic resource management module to determine the list of Pods to be deleted in the current phase, deletes these Pods, reclaims the corresponding GPU instance resources, updates the inference service job Pod list maintained by the controller, modifies the current actual throughput of the job, and finally updates the job status to Running.”); and scheduling the inference request to execute using the first amount of computing resource at the MIG slice (see at least [n0029] After the scheduling process begins, the QueueSort plugin sorts the nodes based on a scoring strategy that considers their GPU throughput contribution. Then, the Filter plugin filters out nodes that do not meet the optimal resource combination and selects GPU instances of nodes with the second-best resource combination to inject into the PostFilter plugin. Next, the Score plugin scores the fragmentation index, and the PreBind plugin injects GPU instance resources into the Pod to complete resource binding, thus ending the scheduling process. [n0030] The QueueSort plugin provides a sorting function to sort the scheduling order of a set of Pods to be scheduled. The core functionality is implemented through the Less(Pod1,Pod2) method. First, based on the completion rate of the inference service job to which the two Pods belong and the throughput information under the optimal GPU instance, the scoring strategy is applied to calculate the priority of the Pod with the larger value. Second, if the calculation results are the same, the two Pods are considered to have the same scheduling priority, and they will be sorted according to their creation time, with the Pod created earlier having a higher priority.”; [n0069] Fragindex represents the fragmentation index, GPUInstanceNum represents the number of GPU instances, RequestComputeInstanceNum represents the number of compute instances required, and IDLEComputeInstanceNum represents the number of currently idle compute instances on the GPU. [n0070] The closer the index is to 1000, the more severe the fragmentation. The algorithm aims to allocate GPU instances from GPUs with more severe fragmentation, thereby reducing the number of resource fragments. [n0071] The QueueSort plugin provides sorting functions to sort the scheduling order of a set of Pods to be scheduled. Its core functionality is implemented through the Less(Pod1,Pod2) method, which is used to compare the priorities of two Pods. In implementing this plugin, the caller rewrote the logic of the Less(Pod1,Pod2) method based on the characteristics of the inference service job. First, based on the completion rate of the inference service jobs to which the two Pods belong and the throughput information under the optimal GPU instance, the scoring strategy is applied to calculate the priority of the Pod with the larger value. Second, if the calculation results are the same, the two Pods are considered to have the same scheduling priority, and they will be sorted according to their creation time, with the Pod created earlier having a higher priority.”)
Therefore, it would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to modify the method of specifying a number of GPU compute slices and a number of GPU memory slices of Kurkure with a cluster resource scheduling for a multi-instance GPU of QIU because this modification would improve resource utilization, and solve job scheduling and execution issues under multi-instance GPUs ([n0009] of QIU) Kurkure and QIU are understood to be silent on the remaining limitations of claim 1.
In the same field of endeavor, Kuperman teaches computing, for a Large Language Model (LLM), a first amount the LLM output, the LLM output while processing the inference request by executing the LLM using a set of multi-instance Graphical Processing Units (GPUs) (MIGs) (see at least col.3, lines 42-54 “Additionally, in some embodiments, the AI automation system is further configured to time-slice GPUs and/or provide multi-instance GPU (MIG) device support. When an LLM model is deployed on a private Kubernetes (K8s) cluster, GPUs are utilized to their full extent. For example, if an entity runs the Llama 2 model on their private cluster on a Spot VM, there might be scenarios where the AI automation system determines that the GPU is not being fully utilized, such as at 70% of its capacity. In such cases, the AI automation system can split the GPU into two or more virtual GPUs and run two or more models side by side on the same GPU, resulting in nearly 100% utilization of the resource.”; col.4, lines 1-12 “he AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140. In some embodiments, the AI automation system 110 causes a proxy application 122 to be deployed on each entity system 120. This proxy application 122 is configured to collect prompts generated by the entity system 120 and pass them to the AI automation system 110. Upon receiving a prompt, the AI automation system 110 determines the performance metrics of the multiple LLM platforms 130, 140. It then selects an LLM platform from the plurality of LLM platforms 130, 140 based on their performance metrics.”, col.4, liens 22-34 “In some embodiments, the AI automation system 110 sends the selected LLM platform to the proxy application 122, causing the proxy application 122 to send the prompt to the selected LLM platform. Upon receiving the prompt, the selected LLM platform generates a response based on the prompt and sends the response back to the proxy application 122. Alternatively, the AI automation system 110 sends the prompt to the selected LLM platform. Upon receiving the prompt, the LLM platform 130 or 140 generates a response based on the prompt and sends this response back to the AI automation system 110, which in turn passes the response to the proxy application 122.”; Fig.6)
Therefore, it would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to modify the method of specifying a number of GPU compute slices and a number of GPU memory slices of Kurkure and QIU with applying large language model of Kuperman because this modification would generate a sufficiently good response with the lowest cost for processing the request (col.1, lines 62-67-col.2,lines 1-3 0f Kuperman) Kurkure, QIU and Kuperman are understood to be silent on the remaining limitations of claim 1.
In the same field of endeavor, Dhakal teaches computing, for a Large Language Model (LLM), first amount of a computing resource, the LLM output comprising a set of intermediate results produced while processing the inference request by executing the LLM (see at least col.1, lines 51-67-col.2, lines 1-7 “During the inference phase, an LLM can process an input sequence (e.g., a sentence) token by token, with each token representing a word or sub-word. As the model progresses through the input sequence, it computes intermediate representations (e.g., key-value pairs) for each token based on its context and the surrounding tokens. When the model needs to generate or predict subsequent tokens in a sequence, it can benefit from reusing previously computed representations. For example, when generating the next word in a sentence, the model can leverage the representations of preceding words to inform its prediction. To increase the inference efficiency, the model can store the intermediate key-value pairs in a cache, referred to as a key-value (KV) cache. These key-value pairs correspond to the keys and values used in the self-attention mechanism to compute attention scores for each token. The KV cache can be dynamically updated to incorporate the latest intermediate representations as the model generates new tokens or processes additional user queries to make sure that the model maintains accurate context information throughout the inference process. Caching the intermediate key-value pairs allows the model to efficiently reuse previously computed attention scores and context representations as it progresses through the sequence.”; col.3, lines 64-67-col.4. lines 1 -17 “Once a copy of the entire KV cache associated with a particular user has been transferred to the remote storage node, the original or local KV cache can be deleted from the GPU's memory to free up the memory space for other users' applications. For example, after smart NIC 116 completes the transfer of all updates of the KV cache to the remote storage node, it can notify the GPU to delete the KV cache from its memory. However, the KV cache may need to be transferred back to the GPU memory when a subsequent query from the same user is received, such that the GPU can take advantage of the intermediate key-value pairs to expedite the inference process. According to some aspects, a smart NIC can include a query parser that can use match-action tables to determine the memory location of the KV cache associated with each user. The smart NIC on the compute node (e.g., smart NIC 116) can then send a KV-cache-transfer request to the remote storage node, requesting the KV cache. The smart NIC on the storage node (e.g., smart NIC 124) can retrieve available portions of the requested KV cache from its SSDs in response to the cache-transfer request.”)
Therefore, it would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to modify the method of specifying a number of GPU compute slices and a number of GPU memory slices of Kurkure, QIU and Kuperman with computing intermediate representations (e.g., key-value pairs) for each token of LLM as seen in Dhakal because this modification would benefit from reusing previously computed representations (col.1, lines 51-67-col.2, lines 1-7 of Dhakal).
Thus, the combination of Kurkure, QIU, Kuperman and Dhakal teaches a computer-implemented method comprising: analyzing an inference request to determine a set of parameters of execution corresponding to the inference request; computing, for a Large Language Model (LLM), a first amount of a computing resource that will be needed to produce and store the LLM output, the LLM output comprising a set of intermediate results produced while processing the inference request by executing the LLM using a set of multi-instance Graphical Processing Units (GPUs) (MIGs), a MIG in the set of MIGs comprising a set of slices of a corresponding GPU (set of MIG slices); causing, by sending a set of instructions to a controller associated with the MIG, the controller to modify a second amount of the computing resource available to a MIG slice in the set of MIG slices; and scheduling the inference request to execute using the first amount of computing resource at the MIG slice.
Regarding claim 2, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, further comprising: identifying an attention layer in the LLM, wherein the set of intermediate results is an output of the attention layer (see at least Dhakal col.1, lines 51-67-col.2, lines 1-7 “During the inference phase, an LLM can process an input sequence (e.g., a sentence) token by token, with each token representing a word or sub-word. As the model progresses through the input sequence, it computes intermediate representations (e.g., key-value pairs) for each token based on its context and the surrounding tokens. When the model needs to generate or predict subsequent tokens in a sequence, it can benefit from reusing previously computed representations. For example, when generating the next word in a sentence, the model can leverage the representations of preceding words to inform its prediction. To increase the inference efficiency, the model can store the intermediate key-value pairs in a cache, referred to as a key-value (KV) cache. These key-value pairs correspond to the keys and values used in the self-attention mechanism to compute attention scores for each token. The KV cache can be dynamically updated to incorporate the latest intermediate representations as the model generates new tokens or processes additional user queries to make sure that the model maintains accurate context information throughout the inference process. Caching the intermediate key-value pairs allows the model to efficiently reuse previously computed attention scores and context representations as it progresses through the sequence.”; col.4, lines 48-60In addition to using a flag, other messaging mechanisms can also be used to trigger smart NIC 206 to automatically transfer the KV cache. The LLM can include many attention layers, and each layer can generate a set of KV vectors to be added to the KV cache. In one example, instead of the per-token transfer scheme, the KV cache can also be transferred each time an LLM layer finishes computing and generates a set of KV vectors. In such a situation, a predetermined number of triggers (which corresponds to the number of LLM layers) can be registered or set, with each trigger corresponding to an LLM layer to allow for KV cache transfer each time an LLM layer finishes processing the user query.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 3, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, wherein the set of intermediate results is a set of Key-Value pairs at an intermediate layer in a neural network of the LLM (see at least Dhakal col.1, lines 51-67-col.2, lines 1-7 “During the inference phase, an LLM can process an input sequence (e.g., a sentence) token by token, with each token representing a word or sub-word. As the model progresses through the input sequence, it computes intermediate representations (e.g., key-value pairs) for each token based on its context and the surrounding tokens. When the model needs to generate or predict subsequent tokens in a sequence, it can benefit from reusing previously computed representations. For example, when generating the next word in a sentence, the model can leverage the representations of preceding words to inform its prediction. To increase the inference efficiency, the model can store the intermediate key-value pairs in a cache, referred to as a key-value (KV) cache. These key-value pairs correspond to the keys and values used in the self-attention mechanism to compute attention scores for each token. The KV cache can be dynamically updated to incorporate the latest intermediate representations as the model generates new tokens or processes additional user queries to make sure that the model maintains accurate context information throughout the inference process. Caching the intermediate key-value pairs allows the model to efficiently reuse previously computed attention scores and context representations as it progresses through the sequence.”col.6, lines 61-67-col.7, lines 1-18 “FIG. 3 presents a flowchart illustrating the operation process of a smart network interface controller (NIC) of a compute node, according to one aspect of the instant application. In this example, the smart NIC is part of the compute node that executes the LLM to infer a query from a user. The compute node can also include an LLM accelerator, which can include a GPU. The GPU can be used to accelerate the LLM inference. During inference, intermediate KV pairs (i.e., key and value vectors) can be stored in the GPU's memory as a KV cache. During operation, the smart NIC can monitor a key-value cache associated with the LLM (operation 302). According to some aspects, the smart NIC can monitor the key-value cache by reading one or more memory locations in the GPU's memory. A thread executing on the GPU may update flags stored in those memory locations. In one example, when the LLM generates a new token, a flag corresponding to the new token generation can be updated (e.g., a token count may be incremented). In another example, when a particular attention layer of the LLM finishes computation, a flag corresponding to that particular layer can be updated. Other message-passing mechanisms can also be used to allow the smart NIC to monitor the state of the key-value cache in the GPU's memory (referred to as a local key-value cache). In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 4, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, further comprising: analyzing, to extract a set of parameters of environment, a computing environment of the LLM, the computing environment comprising the MIGs; and using, in the computing, at least a subset of the set of parameters of environment (see Kurkure: [0008] Embodiments of the present disclosure are directed to techniques for optimally placing clients on GPUs under the MIG model. As used herein, the phrase “placing a client on a GPU” refers to the act of allocating portions of the resources of the GPU for use by that client, typically in accordance with the client's requirements. Once placed in this manner, the client can consume the GPU resources allocated to it over the course of its execution. [0027] FIG. 3 depicts a flowchart 300 that provides additional details regarding the processing that may be performed by VIM server 102 for optimally placing a set of M MIG-enabled VMs on a set of N GPUs using MIG-aware placement algorithm 114 of FIG. 1 according to certain embodiments. [0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU. [0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”; [n0073] of QIU “The PostFilter plugin is only invoked if the Filter plugin fails to filter out any schedulable worker nodes. First, PostFilter will sort the feasible GPU instance sizes by throughput on each unit of the inference service job across different GPU instances based on the performance of the job across different GPU instances. Next, the available free resources are searched in the multi-level GPU instance queue according to the list of GPU instance sizes. When there are multiple available resources for the same GPU instance size, the fragmentation index is calculated according to the formula, and the worker node with a higher degree of fragmentation is selected, thereby completing the selection of worker nodes and the allocation of GPU instance resources. In Kubernetes, if a Pod needs to use GPU resources, it can do so by mounting driver files or injecting environment variables. However, this method of allocating GPU resources needs to be specified through a declared API when the Pod is created. Therefore, this invention uses ConfigMap resources to assist the Pod in injecting GPU instances. When a Pod is created, a configuration file is attached to the Pod via a ConfigMap. During the scheduling phase, the configuration file is modified to achieve flexible binding of GPU resources.”; Kuperman: col.3, lines 42-54 “Additionally, in some embodiments, the AI automation system is further configured to time-slice GPUs and/or provide multi-instance GPU (MIG) device support. When an LLM model is deployed on a private Kubernetes (K8s) cluster, GPUs are utilized to their full extent. For example, if an entity runs the Llama 2 model on their private cluster on a Spot VM, there might be scenarios where the AI automation system determines that the GPU is not being fully utilized, such as at 70% of its capacity. In such cases, the AI automation system can split the GPU into two or more virtual GPUs and run two or more models side by side on the same GPU, resulting in nearly 100% utilization of the resource.”; col.4, lines 1-12 “he AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140. In some embodiments, the AI automation system 110 causes a proxy application 122 to be deployed on each entity system 120. This proxy application 122 is configured to collect prompts generated by the entity system 120 and pass them to the AI automation system 110. Upon receiving a prompt, the AI automation system 110 determines the performance metrics of the multiple LLM platforms 130, 140. It then selects an LLM platform from the plurality of LLM platforms 130, 140 based on their performance metrics.”, col.4, liens 22-34 “In some embodiments, the AI automation system 110 sends the selected LLM platform to the proxy application 122, causing the proxy application 122 to send the prompt to the selected LLM platform. Upon receiving the prompt, the selected LLM platform generates a response based on the prompt and sends the response back to the proxy application 122. Alternatively, the AI automation system 110 sends the prompt to the selected LLM platform. Upon receiving the prompt, the LLM platform 130 or 140 generates a response based on the prompt and sends this response back to the AI automation system 110, which in turn passes the response to the proxy application 122.”; col.4, lines 61-67-col.5,lines 1-5 “FIGS. 2A and 2B illustrate examples outputs of two different LLMs based on a same prompt, “why is the sky blue?” FIG. 2A illustrates an example output 200A of a 7B parameter LLM model, which is an open-source model. FIG. 2B illustrates an example output 200B of ChatGPT 4.0 model, which is a commercial proprietary model. As illustrated, the output of the two models are different but contain similar information. Depending on the purpose of the application, even though the output of ChatGPT4.0 is better, the output of the 7B parameter LLM model may be good enough; therefore, the AI automation system 110 may select the 7B parameter LLM model for the prompt.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 5, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 4, further comprising: detecting a change in a parameter in the parameters of environment (see at least Kurkure:[0008] Embodiments of the present disclosure are directed to techniques for optimally placing clients on GPUs under the MIG model. As used herein, the phrase “placing a client on a GPU” refers to the act of allocating portions of the resources of the GPU for use by that client, typically in accordance with the client's requirements. Once placed in this manner, the client can consume the GPU resources allocated to it over the course of its execution.; [0014] A GPU that supports MIG is composed of a number of compute slices and a number of memory slices, where each compute slice is a disjoint subset of the GPU's total compute resources and each memory slice is a disjoint subset of the GPU's total memory resources. The specific number of compute slices and memory slices will vary depending on the GPU model. For example, the A100 GPU mentioned earlier is composed of seven compute slices (each comprising 1/7 of its 6192 processing cores) and eight memory slices (each comprising ⅛ of its 40 GB of VRAM). These slices can be combined in various permutations in the form of MIG profiles, which are policies that a MIG-enabled client can select to define its GPU compute and memory requirements. 0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available.”; [n0017] of QIU “Running, the inference service job is running normally: Check if the inference service job has been deleted. If so, delete all Pods for the job, reclaim the GPU instance resources allocated to the Pods, check the status of all Pods for the inference service job. If all are Running, the job continues to be in the Running state. If any Pod has been deleted, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. If any Pod is in the Failed state, delete that Pod, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. Check if the throughput requirements of the inference service job have changed. If there are changes, trigger the elastic resource allocation process. The inference service job controller adopts a two-phase operation: first add GPU instances, then remove GPU instances. First, it interacts with the elastic resource management module to determine the GPU instances to be added, creates Pods and completes the injection of GPU resources, and updates the status of the inference service job to Add. If no GPU instances are needed, no Pods are created, but the job status is still updated to Add.”);
recomputing, using a changed value of the parameter in the parameters of environment, the first amount of the computing resource to form a third amount of the computing resource (see at least Kurkure: [0042] ; QIU [n0022] The multi-instance GPU inference service job resource performance prediction module also performs a small amount of online testing and evaluation, updates the model, and tries to fit the actual performance of the inference service job as closely as possible: the serialization model optimization algorithm determines the resource parameter combination to be tested in the next round, starts the corresponding instance through the inference service job image, and the test evaluation client sends a request to the instance to monitor the request processing time, resource utilization, and CPU utilization information; the serialization model optimization algorithm updates and corrects the resource performance prediction model for multi-instance GPUs after multiple iterations and a small amount of online testing and evaluation, and finally predicts the performance of the inference service job under various resource parameter combinations using the resource performance prediction model for multi-instance GPUs. Finally, the algorithm server returns the prediction results to the user and stores them in the database.”; col.2, lines 8-30 of Dhakal “The KV cache may occupy a large amount of memory on the Graphics Processing Unit (GPU) accelerators that perform the inference tasks. The enormous memory requirement of the KV cache may prevent the GPU from running other applications. To free up the GPU memory, some approaches place the KV cache into the memory of the CPU. Considering the high cost associated with the dense CPU double data rate (DDR) memory, some approaches place the KV cache in non-volatile storage devices (e.g., the solid-state drive (SSD)) attached to a remote storage node. However, transferring large amounts of data between the GPU and the remote storage node may add significant network overhead and increase latency. More specifically, such data transfers often require the involvement of the Central Processing Unit (CPU) on the compute node, which is needed to extract the data from the GPU and transfer it to the remote storage node using the network/communication stack on the compute node. Similarly, the CPU on the storage node is needed to receive data through the network/communication stack and write it to the attached SSD. The involvement of the CPUs can add latency and increase the idle time of the GPU as it is waiting for the data to be sent to or received from the remote SSDs”) ; and
causing, by sending a second set of instructions to the controller, the controller to modify the first amount of the computing resource available to the MIG slice to a third amount (see at least Kurkure [0008] Embodiments of the present disclosure are directed to techniques for optimally placing clients on GPUs under the MIG model. As used herein, the phrase “placing a client on a GPU” refers to the act of allocating portions of the resources of the GPU for use by that client, typically in accordance with the client's requirements. Once placed in this manner, the client can consume the GPU resources allocated to it over the course of its execution.; [0014] A GPU that supports MIG is composed of a number of compute slices and a number of memory slices, where each compute slice is a disjoint subset of the GPU's total compute resources and each memory slice is a disjoint subset of the GPU's total memory resources. The specific number of compute slices and memory slices will vary depending on the GPU model. For example, the A100 GPU mentioned earlier is composed of seven compute slices (each comprising 1/7 of its 6192 processing cores) and eight memory slices (each comprising ⅛ of its 40 GB of VRAM). These slices can be combined in various permutations in the form of MIG profiles, which are policies that a MIG-enabled client can select to define its GPU compute and memory requirements. 0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available.”; [n0012] of QIU “The multi-instance GPU elastic resource management module communicates with each agent in the cluster and sends information to the multi-instance GPU inference service job management module to manage the lifecycle of all inference service jobs in the cluster. The inference service job controller manages all state changes during the lifecycle. Furthermore, the multi-instance GPU inference service job management module communicates with the multi-instance GPU inference service job resource performance prediction module and the inference service job online scheduling module to perform calculations. [n0017] Running, the inference service job is running normally: Check if the inference service job has been deleted. If so, delete all Pods for the job, reclaim the GPU instance resources allocated to the Pods, check the status of all Pods for the inference service job. If all are Running, the job continues to be in the Running state. If any Pod has been deleted, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. If any Pod is in the Failed state, delete that Pod, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. Check if the throughput requirements of the inference service job have changed. If there are changes, trigger the elastic resource allocation process. The inference service job controller adopts a two-phase operation: first add GPU instances, then remove GPU instances. First, it interacts with the elastic resource management module to determine the GPU instances to be added, creates Pods and completes the injection of GPU resources, and updates the status of the inference service job to Add. If no GPU instances are needed, no Pods are created, but the job status is still updated to Add. [n0019] Delete, the job resource elastic allocation and reclamation phase; the controller communicates with the elastic resource management module to determine the list of Pods to be deleted in the current phase, deletes these Pods, reclaims the corresponding GPU instance resources, updates the inference service job Pod list maintained by the controller, modifies the current actual throughput of the job, and finally updates the job status to Running.” ) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 6, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 5, wherein the modification of the amount to the third amount occurs while the MIG slice is processing the inference request by suspending the processing of the inference request (see at least Kurkure [0008] Embodiments of the present disclosure are directed to techniques for optimally placing clients on GPUs under the MIG model. As used herein, the phrase “placing a client on a GPU” refers to the act of allocating portions of the resources of the GPU for use by that client, typically in accordance with the client's requirements. Once placed in this manner, the client can consume the GPU resources allocated to it over the course of its execution.; [0014] A GPU that supports MIG is composed of a number of compute slices and a number of memory slices, where each compute slice is a disjoint subset of the GPU's total compute resources and each memory slice is a disjoint subset of the GPU's total memory resources. The specific number of compute slices and memory slices will vary depending on the GPU model. For example, the A100 GPU mentioned earlier is composed of seven compute slices (each comprising 1/7 of its 6192 processing cores) and eight memory slices (each comprising ⅛ of its 40 GB of VRAM). These slices can be combined in various permutations in the form of MIG profiles, which are policies that a MIG-enabled client can select to define its GPU compute and memory requirements. [0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available.”;[0042]; [n0056] of QIU” Pending: The inference service job has not been scheduled. First, check if all created Pods have been scheduled. If any have not been scheduled, the inference service job will remain in a pending state. Next, it calculates whether the throughput of the created and scheduled Pods meets the service level target. If it does not, it calculates the minimum number of GPU instances required based on the difference between the current throughput and the expected throughput, and continues to create Pods based on this value. If all created Pods are scheduled successfully and the total throughput meets the service level target, then the status of the inference service job will be updated to Running. [n0057] Running: The inference service is running normally. Check if the inference service job has been deleted. If so, delete all Pods for that job and reclaim the GPU instance resources allocated to the Pods. Check the status of all Pods in the inference service job. If all are Running, the job remains in the Running state. If any Pod has been deleted, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. If any Pod is in the Failed state, delete that Pod, reclaim the GPU instance resources allocated to that Pod, update the inference service job Pod list maintained by the controller, and update the job status to Pending. The system checks whether the throughput requirements of the inference service job have changed. If there are changes, it triggers the elastic resource allocation process. To ensure that the service level target of the inference service job is met during the elastic resource allocation process, the controller adopts a two-stage operation: first, it adds GPU instances, and then it removes GPU instances to ensure that the resources allocated by the scheduling system are always no less than the job's requirements. To this end, the controller first interacts with the elastic resource management module to determine the GPU instances to be added, creates a Pod and completes the injection of GPU resources, and updates the status of the inference service job to Add. If no GPU instances are needed, no Pod is created, but the job status is still updated to Add) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 7, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 5, wherein the change in the parameter is a change in a performance of the LLM in the computing environment (Kuperman: col.3, lines 42-54 “Additionally, in some embodiments, the AI automation system is further configured to time-slice GPUs and/or provide multi-instance GPU (MIG) device support. When an LLM model is deployed on a private Kubernetes (K8s) cluster, GPUs are utilized to their full extent. For example, if an entity runs the Llama 2 model on their private cluster on a Spot VM, there might be scenarios where the AI automation system determines that the GPU is not being fully utilized, such as at 70% of its capacity. In such cases, the AI automation system can split the GPU into two or more virtual GPUs and run two or more models side by side on the same GPU, resulting in nearly 100% utilization of the resource.”; col.4, lines 1-12 “he AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140. In some embodiments, the AI automation system 110 causes a proxy application 122 to be deployed on each entity system 120. This proxy application 122 is configured to collect prompts generated by the entity system 120 and pass them to the AI automation system 110. Upon receiving a prompt, the AI automation system 110 determines the performance metrics of the multiple LLM platforms 130, 140. It then selects an LLM platform from the plurality of LLM platforms 130, 140 based on their performance metrics.”, col.4, lines 22-34 “In some embodiments, the AI automation system 110 sends the selected LLM platform to the proxy application 122, causing the proxy application 122 to send the prompt to the selected LLM platform. Upon receiving the prompt, the selected LLM platform generates a response based on the prompt and sends the response back to the proxy application 122. Alternatively, the AI automation system 110 sends the prompt to the selected LLM platform. Upon receiving the prompt, the LLM platform 130 or 140 generates a response based on the prompt and sends this response back to the AI automation system 110, which in turn passes the response to the proxy application 122.”; col.4, lines 61-67-col.5,lines 1-5 “FIGS. 2A and 2B illustrate examples outputs of two different LLMs based on a same prompt, “why is the sky blue?” FIG. 2A illustrates an example output 200A of a 7B parameter LLM model, which is an open-source model. FIG. 2B illustrates an example output 200B of ChatGPT 4.0 model, which is a commercial proprietary model. As illustrated, the output of the two models are different but contain similar information. Depending on the purpose of the application, even though the output of ChatGPT4.0 is better, the output of the 7B parameter LLM model may be good enough; therefore, the AI automation system 110 may select the 7B parameter LLM model for the prompt.”.) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 8, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 5, wherein the change in the parameter is a change in a number of active users using the computing environment (see at least Kurkure [0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available. 0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: [0021] 1. A set of coefficients corresponding to the compute slices and memory slices requested by each VM i for i=1, . . . , M; [0022] 2. a set of decision variables that indicate, among other things, whether a given VM i is placed on a given GPU j for j=1, . . . , N per the problem solution; [0023] 3. a set of constraints that ensures, among other things, that (a) each (placed) VM i is placed on a single GPU j with sufficient compute and memory resources to satisfy the VM's requested requirements, (b) the total allocated compute and memory resources on each GPU j do not exceed its maximum capacity, and (c) each VM i is placed on at most one GPU j; and [0024] 4. an objective function corresponding to the total number of GPUs used. 0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.”; [0042]; [n0022] of QIU “The multi-instance GPU inference service job resource performance prediction module also performs a small amount of online testing and evaluation, updates the model, and tries to fit the actual performance of the inference service job as closely as possible: the serialization model optimization algorithm determines the resource parameter combination to be tested in the next round, starts the corresponding instance through the inference service job image, and the test evaluation client sends a request to the instance to monitor the request processing time, resource utilization, and CPU utilization information; the serialization model optimization algorithm updates and corrects the resource performance prediction model for multi-instance GPUs after multiple iterations and a small amount of online testing and evaluation, and finally predicts the performance of the inference service job under various resource parameter combinations using the resource performance prediction model for multi-instance GPUs. Finally, the algorithm server returns the prediction results to the user and stores them in the database.” ) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 9, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 5, wherein the change in the parameter is a change in rate of requests being directed at the LLM in the computing environment (see at least Kuperman col.12, lines 62- col.12, lines 3 “In some embodiments, the AI automation system 110 is further configured to determine a GPU utilization rate at a private Kubernetes that host an open source LLM. Responsive to determining that the GPU utilization rate is lower than a predetermined threshold, e.g., 70%, the AI automation system 110 causes the GPU to be divided into a plurality of virtual GPUs, and causes each of the plurality of virtual GPUs to execute a separate instance of the open source LLM, such that the GPU utilization rate can be increased to near 100%; ol.3, lines 42-54 “Additionally, in some embodiments, the AI automation system is further configured to time-slice GPUs and/or provide multi-instance GPU (MIG) device support. When an LLM model is deployed on a private Kubernetes (K8s) cluster, GPUs are utilized to their full extent. For example, if an entity runs the Llama 2 model on their private cluster on a Spot VM, there might be scenarios where the AI automation system determines that the GPU is not being fully utilized, such as at 70% of its capacity. In such cases, the AI automation system can split the GPU into two or more virtual GPUs and run two or more models side by side on the same GPU, resulting in nearly 100% utilization of the resource.”; col.4, lines 1-12 “he AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140. In some embodiments, the AI automation system 110 causes a proxy application 122 to be deployed on each entity system 120. This proxy application 122 is configured to collect prompts generated by the entity system 120 and pass them to the AI automation system 110. Upon receiving a prompt, the AI automation system 110 determines the performance metrics of the multiple LLM platforms 130, 140. It then selects an LLM platform from the plurality of LLM platforms 130, 140 based on their performance metrics.”, col.4, lines 22-34 “In some embodiments, the AI automation system 110 sends the selected LLM platform to the proxy application 122, causing the proxy application 122 to send the prompt to the selected LLM platform. Upon receiving the prompt, the selected LLM platform generates a response based on the prompt and sends the response back to the proxy application 122. Alternatively, the AI automation system 110 sends the prompt to the selected LLM platform. Upon receiving the prompt, the LLM platform 130 or 140 generates a response based on the prompt and sends this response back to the AI automation system 110, which in turn passes the response to the proxy application 122.”; Fig. 6) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 10, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 5, wherein the change in the parameter is a change in a utilization of at least one of the MIG slices (see at least Kurkure [0008];[0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available. 0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: [0021] 1. A set of coefficients corresponding to the compute slices and memory slices requested by each VM i for i=1, . . . , M; [0022] 2. a set of decision variables that indicate, among other things, whether a given VM i is placed on a given GPU j for j=1, . . . , N per the problem solution; [0023] 3. a set of constraints that ensures, among other things, that (a) each (placed) VM i is placed on a single GPU j with sufficient compute and memory resources to satisfy the VM's requested requirements, (b) the total allocated compute and memory resources on each GPU j do not exceed its maximum capacity, and (c) each VM i is placed on at most one GPU j; and [0024] 4. an objective function corresponding to the total number of GPUs used. 0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.”; [0042]; [n0022] of QIU “The multi-instance GPU inference service job resource performance prediction module also performs a small amount of online testing and evaluation, updates the model, and tries to fit the actual performance of the inference service job as closely as possible: the serialization model optimization algorithm determines the resource parameter combination to be tested in the next round, starts the corresponding instance through the inference service job image, and the test evaluation client sends a request to the instance to monitor the request processing time, resource utilization, and CPU utilization information; the serialization model optimization algorithm updates and corrects the resource performance prediction model for multi-instance GPUs after multiple iterations and a small amount of online testing and evaluation, and finally predicts the performance of the inference service job under various resource parameter combinations using the resource performance prediction model for multi-instance GPUs. Finally, the algorithm server returns the prediction results to the user and stores them in the database.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 11, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, wherein the computing resource is a memory available to the MIG slice in the set of MIG slices (see at least [0011] of Kurkure “ Host cluster 104 comprises a plurality of host systems 106, each running in software a hypervisor 108 that provides an execution environment for one or more VMs 110. As known in the art, a VM is a virtual representation of a physical computer system with its own virtual CPU(s), virtual storage, virtual GPU(s), etc. Each host system 106 also includes hardware components that are provisioned for use by VMs 110 via hypervisor 108. These hardware components include, among other things, a physical GPU 112. Although not shown in FIG. 1, GPU 112 comprises a set of compute resources (e.g., processing cores, copy engines, hardware encoders/decoders, etc.) and a set of memory resources (e.g., video RAM (VRAM), caches, memory controllers, etc.). For example, the Nvidia Ampere A100 GPU includes 6192 processing cores and 40 gigabytes (GB) of VRAM. n0003]Multi-Instance GPU (MIG) is a new hardware feature added by NVIDIA to GPUs such as the A100 and H100. This technology supports the spatial division of a GPU into up to 7 GPU shards, each with independent computing and memory resources, enabling true parallel execution of jobs.” [0015] For instance, the table below depicts an example list of MIG profiles available for the A100 GPU:TABLE-US-00001 TABLE 1 Number of Profile Number of Number of Instances Name Compute Slices Memory Slices Available MIG 1g.5gb 1 (out of 7 total) 1 (out of 8 total) 7 MIG 2g.10gb 2 (out of 7 total) 2 (out of 8 total) 3 MIG 3g.20gb 3 (out of 7 total) 4 (out of 8 total) 2 MIG 4g.20gb 4 (out of 7 total) 4 (out of 8 total) 1 MIG 7g.40gb 7 (out of 7 total) 8 (out of 8 total) 1; [0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available. [n0004] of QIU “However, MIG technology also increases the difficulty of cluster resource scheduling, bringing three problems: [n0005] Question 1: How to characterize the relationship between job execution efficiency and GPU fragmentation types? MIG can divide the GPU into 5 different GPU slices, each with different computing and memory resources. Based on these 5 GPU partitioning methods, MIG can configure a single GPU with 19 different resource partitioning schemes. The execution efficiency of deep learning inference jobs varies depending on the partitioning scheme. Furthermore, the processing latency of inference jobs also varies greatly depending on the batch size and CPU configuration. Finding the relationship between a massive set of configuration parameters and job execution efficiency has become a thorny issue. As shown below, there are 180 different parameter combinations in a multi-instance GPU scenario.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 12, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 11, wherein the memory is one of a cache memory and a primary memory ([0011] of Kurkure “ Host cluster 104 comprises a plurality of host systems 106, each running in software a hypervisor 108 that provides an execution environment for one or more VMs 110. As known in the art, a VM is a virtual representation of a physical computer system with its own virtual CPU(s), virtual storage, virtual GPU(s), etc. Each host system 106 also includes hardware components that are provisioned for use by VMs 110 via hypervisor 108. These hardware components include, among other things, a physical GPU 112. Although not shown in FIG. 1, GPU 112 comprises a set of compute resources (e.g., processing cores, copy engines, hardware encoders/decoders, etc.) and a set of memory resources (e.g., video RAM (VRAM), caches, memory controllers, etc.). For example, the Nvidia Ampere A100 GPU includes 6192 processing cores and 40 gigabytes (GB) of VRAM.”; n0003] of QIU “Multi-Instance GPU (MIG) is a new hardware feature added by NVIDIA to GPUs such as the A100 and H100. This technology supports the spatial division of a GPU into up to 7 GPU shards, each with independent computing and memory resources, enabling true parallel execution of jobs.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 13, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, wherein the causing occurs while the MIG is processing another request (see at least Kurkure [0008];[0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available. [0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: 0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.”; [[0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”; [n0062] of QIU “Upon receiving a user's performance prediction request, the system first queries the database to check if the inference service job has historical data. If so, it returns the result directly. Next, the inference model is parsed using the model structure parsing component to obtain the model's structural information. This component integrates the tensorflow.profiler and thop libraries, which are used to parse tensorflow and pytorch models, respectively. Subsequently, various resource parameter combinations and model structure information are used as inputs to predict job performance using a pre-trained random forest regression model. To prevent issues such as model aging and low generalization, the prediction module will conduct a small number of online tests to evaluate and update the model, aiming to fit the actual performance of the inference service as closely as possible. In this stage, the serialization model optimization algorithm determines the combination of resource parameters to be tested in the next round, the corresponding instance is launched through the inference service job image, and the test evaluation client sends a request to the instance to monitor information such as request processing time, resource utilization, and CPU utilization. The serialization model optimization algorithm is evaluated through multiple iterations and a small number of online tests. The random forest regression model is updated and corrected, and finally the random forest regression model is used to predict the performance of the inference service job under various resource parameter combinations. The algorithm server then returns the prediction results to the user and stores them in the database.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 14, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, wherein the set of instructions comprises an instruction to allocate the first amount by deallocating a third amount of computing resource from an existing amount of the computing resource already configured in the MIG slice (see at least Kurkure [0008];[0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available. [0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: 0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.”; [[0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”; [n0029] of QIU “After the scheduling process begins, the QueueSort plugin sorts the nodes based on a scoring strategy that considers their GPU throughput contribution. Then, the Filter plugin filters out nodes that do not meet the optimal resource combination and selects GPU instances of nodes with the second-best resource combination to inject into the PostFilter plugin. Next, the Score plugin scores the fragmentation index, and the PreBind plugin injects GPU instance resources into the Pod to complete resource binding, thus ending the scheduling process. [n0030] The QueueSort plugin provides a sorting function to sort the scheduling order of a set of Pods to be scheduled. The core functionality is implemented through the Less(Pod1,Pod2) method. First, based on the completion rate of the inference service job to which the two Pods belong and the throughput information under the optimal GPU instance, the scoring strategy is applied to calculate the priority of the Pod with the larger value. Second, if the calculation results are the same, the two Pods are considered to have the same scheduling priority, and they will be sorted according to their creation time, with the Pod created earlier having a higher priority. [n0031] The Filter plugin provides node filtering functionality, filtering out worker nodes that cannot run the scheduled Pod. In this stage, the scheduler will filter worker nodes based on the best resource configuration combination of the inference service job to which the Pod belongs. The Filter plugin will filter out all worker nodes that do not meet the optimal number of CPU cores and GPU instance size for the job.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding claim 15, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 14, wherein the deallocating occurs while the MIG is processing another request (see at least Kurkure [0008];[0017] At the time of provisioning a VM in host cluster 104, the creator/user of the VM can submit a provisioning request to VIM server 102 with a selection of a MIG profile that is appropriate for the VM's GPU workload, or in other words a MIG profile that specifies a sufficient number of compute and memory slices to meet the VM's requirements. Such a VM is referred to as a MIG-enabled VM. In response, VIM server 102 will place the VM on a GPU in host cluster 104 that has the number of compute and memory slices specified in the selected MIG profile free/unallocated, assuming such a GPU is available. [0020] At a high level, algorithm 114 involves formulating the placement optimization problem as an integer linear programming (ILP) problem. For example, given M VMs to be placed (each associated with a MIG profile specifying a number of compute slices and a number of memory slices requested by the VM) and N GPUs that can serve as placement targets, algorithm 114 can define an ILP problem that includes: 0028] Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on. It is assumed that each of these requests are fractional requests, or in other words includes a MIG profile that specifies a fraction of the total compute and memory slices of a given GPU. This is because a non-fractional request can be fulfilled by simply placing the VM corresponding to that request on the entirety of a single GPU.”; [[0042] Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”; [n0029] of QIU “After the scheduling process begins, the QueueSort plugin sorts the nodes based on a scoring strategy that considers their GPU throughput contribution. Then, the Filter plugin filters out nodes that do not meet the optimal resource combination and selects GPU instances of nodes with the second-best resource combination to inject into the PostFilter plugin. Next, the Score plugin scores the fragmentation index, and the PreBind plugin injects GPU instance resources into the Pod to complete resource binding, thus ending the scheduling process. [n0030] The QueueSort plugin provides a sorting function to sort the scheduling order of a set of Pods to be scheduled. The core functionality is implemented through the Less(Pod1,Pod2) method. First, based on the completion rate of the inference service job to which the two Pods belong and the throughput information under the optimal GPU instance, the scoring strategy is applied to calculate the priority of the Pod with the larger value. Second, if the calculation results are the same, the two Pods are considered to have the same scheduling priority, and they will be sorted according to their creation time, with the Pod created earlier having a higher priority. [n0031] The Filter plugin provides node filtering functionality, filtering out worker nodes that cannot run the scheduled Pod. In this stage, the scheduler will filter worker nodes based on the best resource configuration combination of the inference service job to which the Pod belongs. The Filter plugin will filter out all worker nodes that do not meet the optimal number of CPU cores and GPU instance size for the job.”) In addition, the same motivation is used as the rejection for claim 1.
Regarding independent claim 17, Kurkure teaches a computer program product comprising: One or more computer readable storage media; and program instructions stored on the one or more storage media and configured to perform operations ([0046] Further, one or more embodiments can relate to a device or an apparatus for performing the foregoing operations. The apparatus can be specially constructed for specific required purposes, or it can be a generic computer system comprising one or more general purpose processors (e.g., Intel or AMD x86 processors) selectively activated or configured by program code stored in the computer system. In particular, various generic computer systems may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus to perform the required operations. The various embodiments described herein can be practiced with other computer system configurations including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. [0047] Yet further, one or more embodiments can be implemented as one or more computer programs or as one or more computer program modules embodied in one or more non-transitory computer readable storage media. The term non-transitory computer readable storage medium refers to any storage device, based on any existing or subsequently developed technology, that can store data and/or computer programs in a non-transitory state for access by a computer system. Examples of non-transitory computer readable media include a hard drive, network attached storage (NAS), read-only memory, random-access memory, flash-based nonvolatile memory (e.g., a flash memory card or a solid state disk), persistent memory, NVMe device, a CD (Compact Disc) (e.g., CD-ROM, CD-R, CD-RW, etc.), a DVD (Digital Versatile Disc), a magnetic tape, and other optical and non-optical data storage devices. The non-transitory computer readable media can also be distributed over a network coupled computer system so that the computer readable code is stored and executed in a distributed fashion.”) comprising: Remaining limitations of claim 17 is similar scope to claim 1 and therefore rejected under the same rational.
Regarding independent claim 20, Kurkure teaches a computer system comprising a processor and one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by the processor to cause the processor to perform operations ([0046] Further, one or more embodiments can relate to a device or an apparatus for performing the foregoing operations. The apparatus can be specially constructed for specific required purposes, or it can be a generic computer system comprising one or more general purpose processors (e.g., Intel or AMD x86 processors) selectively activated or configured by program code stored in the computer system. In particular, various generic computer systems may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus to perform the required operations. The various embodiments described herein can be practiced with other computer system configurations including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. [0047] Yet further, one or more embodiments can be implemented as one or more computer programs or as one or more computer program modules embodied in one or more non-transitory computer readable storage media. The term non-transitory computer readable storage medium refers to any storage device, based on any existing or subsequently developed technology, that can store data and/or computer programs in a non-transitory state for access by a computer system. Examples of non-transitory computer readable media include a hard drive, network attached storage (NAS), read-only memory, random-access memory, flash-based nonvolatile memory (e.g., a flash memory card or a solid state disk), persistent memory, NVMe device, a CD (Compact Disc) (e.g., CD-ROM, CD-R, CD-RW, etc.), a DVD (Digital Versatile Disc), a magnetic tape, and other optical and non-optical data storage devices. The non-transitory computer readable media can also be distributed over a network coupled computer system so that the computer readable code is stored and executed in a distributed fashion.”) comprising: Remaining limitations of claim 17 is similar scope to claim 1 and therefore rejected under the same rational.
2. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Kurkure et al., IDS, U.S Patent Application Publication No.2024/0362052 (“Kurkure”) in view of QIU et al, IDS, CN118377598 (English translated) (“QIU”) further in view of Kuperman et al., IDS, U.S Patent No.12236193 (“Kuperman”) further in view of Dhakal et al., U.S Patent No.12346252 (“Dhakal”) further in view of YANOSY, JR. et al., U.S Patent Application Publication No.2024/0296352 (“YANOSY”)
Regarding claim 16, Kurkure, QIU, Kuperman and Dhakal teach the computer-implemented method of claim 1, Kurkure, QIU, Kuperman and Dhakal are understood to be silent on the remaining limitations of claim 16.
In the same field of endeavor, YASOSY teaches wherein the LLM is a transformer-type model ([0098] An example of a transformer type machine learning model suitable for the system of the present invention can include, for example, large language models (LLMs). The large language models can be configured to understand and generate human language by learning patterns and relationships from vast amounts of input textual data. The model configuration can include setting selected hyperparameters, including the number of layers, hidden units per layer, attention mechanisms, and other architectural details. The LLMs can utilize deep learning techniques, particularly the foregoing transformer architectures, to process and generate text. The models can be pre-trained and trained on massive data corpora (e.g., text corpora, image corpora, and the like) and can perform tasks such as text generation, language translation, text summarization, image generation, sentiment analysis, and the like. The LLMs can include, by simple way of example, generative artificial intelligence (AI) or machine learning models. The generative artificial intelligence (AI) model refers to a computational system designed to create new and original data based on patterns and information learned from existing datasets. The generative AI model can employ selected machine learning techniques to generate content, such as text, images, audio, or other forms of media or data, that closely resembles the input data but is not an exact replication. The generative AI models can leverage neural networks and probabilistic methods to produce outputs that exhibit creativity and diversity while maintaining coherence with the input data distribution.”)
Therefore, it would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to modify the method of specifying a number of GPU compute slices and a number of GPU memory slices of Kurkure, QIU, Kuperman and Dhakal with LLM is a transformer-type model as seen in YASOSY because this modification would understand and generate human language by learning patterns and relationships from vast amounts of input textual data ([0098] of YASOSY)
3. Claims 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kurkure et al., IDS, U.S Patent Application Publication No.2024/0362052 (“Kurkure”) in view of QIU et al, IDS, CN118377598 (English translated) (“QIU”) further in view of Kuperman et al., IDS, U.S Patent No.12236193 (“Kuperman”) further in view of Dhakal et al., U.S Patent No.12346252 (“Dhakal”) further in view of Sui et al., U.S Patent Application Publication No.20220398250 (“Sui”)
Regarding claim 18, Kurkure, QIU, Kuperman and Dhakal teach the computer program product of claim 17, Kurkure, QIU, Kuperman and Dhakal are understood to be silent on the remaining limitations of claim 18.
In the same field of endeavor, Sui teaches wherein the stored program instructions are stored in a computer readable storage device in a data processing system, and wherein the stored program instructions are transferred over a network from a remote data processing system ([0046] Furthermore, the illustrative embodiments may be implemented with respect to any type of data, data source, or access to a data source over a data network. Any type of data storage device may provide the data to an embodiment of the invention, either locally at a data processing system or over a data network, within the scope of the invention. Where an embodiment is described using a mobile device, any type of data storage device suitable for use with the mobile device may provide the data to such embodiment, either locally at the mobile device or over a data network, within the scope of the illustrative embodiments.”; [0123] Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present invention.”; claim 14)
Therefore, in combination of Kurkure, QIU, Kuperman and Dhakal, it would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to modify the method of specifying a number of GPU compute slices and a number of GPU memory slices of Kurkure with transferring the stored program instructions over a network from a remote data processing system as seen in Sui because this modification would access to a data source over a data network ([0046] of Sui)
Regarding claim 19, Kurkure, QIU, Kuperman and Dhakal teach the computer program product of claim 17, Kurkure, QIU, Kuperman and Dhakal are understood to be silent on the remaining limitations of claim 19.
In the same field of endeavor, Sui teaches wherein the stored program instructions are stored in a computer readable storage device in a server data processing system, and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system([0122] Computer readable program instructions described herein can be downloaded to respective computing/processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and/or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and/or edge servers. A network adapter card or network interface in each computing/processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing/processing device. [0123] Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present invention.”) further comprising:
program instructions to meter use of the program instructions associated with the request; and program instructions to generate an invoice based on the metered use ([0072] In one example, management layer 80 may provide the functions described below. Resource provisioning 81 provides dynamic procurement of computing resources and other resources that are utilized to perform tasks within the cloud computing environment. Metering and Pricing 82 provide cost tracking as resources are utilized within the cloud computing environment, and billing or invoicing for consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management such that required service levels are met. Service Level Agreement (SLA) planning and fulfillment 85 provide pre-arrangement for, and procurement of, cloud computing resources for which a future requirement is anticipated in accordance with an SLA.0128] Embodiments of the present invention may also be delivered as part of a service engagement with a client corporation, nonprofit organization, government entity, internal organizational structure, or the like. Aspects of these embodiments may include configuring a computer system to perform, and deploying software, hardware, and web services that implement, some or all of the methods described herein. Aspects of these embodiments may also include analyzing the client's operations, creating recommendations responsive to the analysis, building systems that implement portions of the recommendations, integrating the systems into existing processes and infrastructure, metering use of the systems, allocating expenses to users of the systems, and billing for use of the systems. Although the above embodiments of present invention each have been described by stating their individual advantages, respectively, present invention is not limited to a particular combination thereof. To the contrary, such embodiments may also be combined in any way and number according to the intended deployment of present invention without losing their beneficial effects.”; claim 15)
Therefore, in combination of Kurkure, QIU, Kuperman and Dhakal, it would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to modify the method of specifying a number of GPU compute slices and a number of GPU memory slices of Kurkure with generating an invoice based on the metered use as seen in Sui because this modification would provide billing or invoicing for consumption of these resources ([0072] of Sui)
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SARAH LE whose telephone number is (571)270-7842. The examiner can normally be reached Monday: 8AM-4:30PM EST, Tuesday: 8 AM-3:30PM EST, Wednesday: 8AM-2:30PM EST, Thursday and Friday off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SARAH LE/Primary Examiner, Art Unit 2614