DETAILED ACTION
This communication is in response to the application filed on November 22, 2024 in which claims 1-20 are pending in the application. Claims 1, 8, and 15 are in independent form.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on January 9, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities:
[0041] recites "(i) the entire GPU of at least one GPU of the one or more GPUs is used for processing LLM inference requests and (2) at least one GPU of the one or more GPUs is partitioned into multiple MIG devices for processing LLM inference requests." The enumerator "(2)" should recite "(ii)," consistent with the "(i)" opening the sentence and with the enumeration at [0042].
[0046] recites "the power consumed (e.g., in watts) by the multiple GPU for processing the LLM inference requests" and [0049] recites "a desired amount of power to be consumed by the multiple GPU." The GPUs are consistently designated "the one or more GPUs" elsewhere (e.g., [0045], [0048]). Both passages should recite "the one or more GPUs."
[0047] recites "a desired request rate of LLM inference requests submitted to the multiple the one or more GPUs to be processed by the one or more GPUs." Consistent with [0043], [0044], and [0048], LLM inference requests are submitted to the multiple LLMs and processed by the one or more GPUs, and the passage should apparently recite "submitted to the multiple LLMs to be processed by the one or more GPUs."
[0103] recites "(or one or more additional memory devices such as read only memory device 96)." Reference numeral 96 designates the input data (FIG. 5), and the read only memory device is designated 98 ([0104], FIG. 5). The passage should recite "read only memory device 98."
[0112] recites "even though it is not shown in a cloud in FIG. 1." Computer 101 appears in FIG. 6, and FIG. 1 depicts the GPU 10 partitioned into MIG devices. The passage should recite "FIG. 6."
[0126] recites "(not separately shown in Figure 1)" and "private and public clouds 106 are programmed and configured." The computing environment is depicted in FIG. 6, the private cloud is designated 106, and the public cloud is designated 105. The passages should recite "FIG. 6" and "private and public clouds 106 and 105."
The following minor typographical errors should also be corrected: [0042] recites "the entire GPU of at the least one GPU" (should recite "of the at least one GPU"), "such that the least one GPU is partitioned" (should recite "the at least one GPU"), and "for processing the ther LLM inference requests" (should recite "the other LLM inference requests"); [0050] recites "Embodiment of the present invention calculate" (should recite "Embodiments of the present invention calculate"); [0067] recites "Steps 215-290 is an iterative process" (should recite "are an iterative process"); [0072] and [0076] recite "the multiple LLM requests" (should recite "the multiple LLM inference requests," consistent with the remainder of the disclosure).
Appropriate correction is required.
Claim Objections
Claims 2, 7, 9, 14, and 16 are objected to because of the following informalities:
As per claims 7 and 14, the claims recite "changing from a first mixed GPU-MIG state to a pure MIG state." A pure MIG state is already recited in claims 5 and 12, from which claims 7 and 14 respectively depend. The claims should recite "to the pure MIG state."
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 15-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 15 recites "to implement a method for graphics processing unit (GPU) resource allocation for processing large language model (LLM) inference requests, said a computer system, inputs of target request rate, target token throughput, and power budget;". Words appear to have been omitted after "said". The passage recites no act, so it is unclear whether claim 15 requires receiving the inputs that the later calculating limitation uses. It is also unclear whether "a computer system" following "said" requires a second computer system in addition to the one recited in the preamble. For examination purposes, the passage is interpreted as reciting "said method comprising: receiving, by the one or more processors of the computer system, inputs of target request rate, target token throughput, and power budget;", consistent with claims 1 and 8.
Claims 16-20, which are dependent on claim 15, are similarly rejected.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 5, 6, 8-10, 12, 13, 15-17, 19, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Stojkovic et al. ("DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency," arXiv:2408.00741, August 1, 2024) (hereinafter Stojkovic) in view of "Horizontal Pod Autoscaling," Kubernetes documentation, version 1.28 (hereinafter Kubernetes) in view of Patel et al. ("POLCA: Power Oversubscription in LLM Cloud Providers," arXiv:2308.12908, August 24, 2023) (hereinafter Patel) in view of Duluk et al. (US 2023/0288471 A1) (hereinafter Duluk).
As per claim 1, Stojkovic primarily teaches the invention as claimed including:
A method for graphics processing unit (GPU) resource allocation for processing large language model (LLM) inference requests, said method comprising:
performing, by the one or more processors, an iterative process, wherein each iteration of the iterative process (Section IV-B, at every epoch (e.g., 30 minutes), the cluster manager computes the minimal number of nodes per pool that can support the load of a given request type and Section IV, periodically re-evaluates how many pools and how many model instances per pool are needed based on the system load; The examiner notes that each epoch repeats the same taking of the load and recomputing of the allocation, which makes every epoch one iteration of the claimed iterative process) comprises:
monitoring metrics of current request rate, current token throughput (Section IV-D, when an instance manager detects that its queue is building up, indicating that the rate of request processing is lower than the rate of request arrival, it triggers an emergency event; Section III-A, a medium system load of 2000 tokens per second (TPS); Section IV-B, the manager uses a load predictor to forecast the incoming load, PL, for each request type based on its historic data; The examiner notes that the rate of request processing is how much of that work the cluster completes per unit time, and Stojkovic states that load in tokens per second per Section III-A, so the monitored processing rate is the claimed current token throughput. The instance manager can detect that one rate has fallen below the other only by holding current readings of both, so the two metrics are read while the cluster runs, which is the claimed monitoring).
Stojkovic does not explicitly teach:
receiving, by one or more processors of a computer system, inputs of target request rate, target token throughput, and power budget;
and current power consumption with respect to LLM inference requests submitted to multiple LLMs and processed by one or more GPUs;
calculating, from the monitored metrics, parameters of traffic factor (TF), throughput factor (TPF), power factor (PF), GPU scaling factor (GSF), and MIG scaling factor (MSF) via TF = current request rate / target request rate, TPF = target token throughput / the current token throughput, PF = power budget / the current power consumption; and
converting (i) Multi-Instance GPU (MIG) service to GPU service or (ii) the GPU service to the MIG service, said converting being implemented as a function of the GSF, the MSF, and the PF.
However, Kubernetes teaches:
receiving, by one or more processors of a computer system, inputs of target request rate, target token throughput (Walkthrough, you can introduce additional metrics to use when autoscaling; Walkthrough, metric: name: requests-per-second ... target: type: Value value: 10k; Walkthrough, your HorizontalPodAutoscaler would attempt to ensure that each pod was ... serving 1000 packets per second, and that all pods behind the main-route Ingress were serving a total of 10000 requests per second; Walkthrough, if you provide multiple such metric blocks, the HorizontalPodAutoscaler will consider each metric in turn; The examiner notes that a further block for the token throughput carries a received input of a target token throughput, and Stojkovic states that metric in tokens per second per Section III-A);
calculating, from the monitored metrics (How does a HorizontalPodAutoscaler work, once during each period, the controller manager queries the resource utilization against the metrics specified in each HorizontalPodAutoscaler definition and obtains the metrics from either the resource metrics API (for per-pod resource metrics), or the custom metrics API (for all other metrics); Algorithm details, desiredReplicas = ceil[currentReplicas * ( currentMetricValue / desiredMetricValue )]; Algorithm details, the control plane skips any scaling action if the ratio is sufficiently close to 1.0 (within a globally-configurable tolerance, 0.1 by default); Walkthrough, if you provide multiple such metric blocks, the HorizontalPodAutoscaler will consider each metric in turn), parameters of
traffic factor (TF) via TF = current request rate / target request rate (Algorithm details, if the current metric value is 200m, and the desired value is 100m, the number of replicas will be doubled, since 200.0 / 100.0 == 2.0 and if the current value is instead 50m, you'll halve the number of replicas, since 50.0 / 100.0 == 0.5),
throughput factor (TPF), power factor (PF) via TPF = target token throughput / the current token throughput, PF = power budget / the current power consumption (Walkthrough, if you provide multiple such metric blocks, the HorizontalPodAutoscaler will consider each metric in turn; Walkthrough, metric: name: packets-per-second target: type: AverageValue averageValue: 1k and metric: name: requests-per-second ... target: type: Value value: 10k; Algorithm details, desiredReplicas = ceil[currentReplicas * ( currentMetricValue / desiredMetricValue )]; The examiner notes that a metric block for the token throughput and a metric block for the power draw each carry that metric's target, and the controller runs the cited division on each block in turn. The recited TPF and PF are that same comparison of a metric against its target, written with the target and the budget on top.),
GPU scaling factor (GSF), and MIG scaling factor (MSF) (Algorithm details, if the current metric value is 200m, and the desired value is 100m, the number of replicas will be doubled, since 200.0 / 100.0 == 2.0 and if the current value is instead 50m, you'll halve the number of replicas, since 50.0 / 100.0 == 0.5; Algorithm details, the control plane skips any scaling action if the ratio is sufficiently close to 1.0 (within a globally-configurable tolerance, 0.1 by default); The examiner notes that claim 1 recites a formula only for the TF, the TPF, and the PF, and therefore gives the GSF and the MSF their broadest reasonable interpretation as the scale-up and scale-down decision factors calculated from the monitored metrics. Scaling up is the change toward the GPU service and scaling down is the change toward the MIG service, each factor named for its direction. The cited tolerance shows the calculated factor is the value the controller acts on, scaling when it departs from one and holding while it stays close. The factor on which the cited example doubles the capacity, 2.0, is a calculated GSF, and the factor on which the cited example halves the capacity, written with the target on top as 100.0 / 50.0 == 2.0, is a calculated MSF).
Stojkovic and Kubernetes are both concerned with setting the amount of allocated computing resource from monitored metrics measured against target values and are therefore combinable/modifiable.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Stojkovic to include targets for the request rate and the token throughput provided by the operator, and a factor for each monitored metric calculated by dividing the metric’s current value by its target, as taught by Kubernetes, in order to measure how far the running cluster sits from each of its targets.
Motivation would scale the allocated GPU resources to match the demand because the scaling action is taken from the ratio of each current metric value to its target and is skipped while the ratio stays within a tolerance of one, as taught by Kubernetes (Algorithm details).
Stojkovic in view of Kubernetes do not explicitly teach:
and power budget;
and current power consumption with respect to LLM inference requests submitted to multiple LLMs and processed by one or more GPUs;
However, Patel teaches:
and power budget (Section 1, datacenters are deployed with a fixed power budget, based on generators and contracts with the utility companies and Section 5.1, when T2 is breached, we start by frequency capping all the low-priority workloads);
The examiner notes that the act of receiving the inputs is shown by Kubernetes above, and in the combination the budget is provided to the controller the same way as the other two targets.
monitoring metrics of current power consumption (Section 2.2, we run NVIDIA DCGM with a 100 ms interval to track utilization, power consumption, temperature, hardware activity, and other performance counters on each GPU and Section 5.2, the power manager running at rack-level receives frequent telemetry from the PDU)
with respect to LLM inference requests (Section 3.2, Table 2 shows the normalized aggregate power consumption patterns of LLM training and inference clusters at a large cloud provider and we consider an interactive inference cluster (i.e., where users make inference requests and expect rapid responses))
submitted to multiple LLMs (Figure 4, GPU power usage timeseries for multiple inference models and Section 2.2, we consider popular opensource LLM models varying structures and sizes: Encoder (RoBERTa [27]), Decoder (GPT-NeoX, OPT, BLOOM [47]), and Encoder+Decoder (Flan-T5 [10]) transformer models)
and processed by one or more GPUs (Section 2.2, on each GPU).
Stojkovic, Kubernetes, and Patel are all concerned with managing the resources of a serving cluster from monitored operating values measured against set levels and are therefore combinable/modifiable.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Stojkovic in view of Kubernetes to include a reading of the power drawn by each GPU against a fixed power budget for the cluster, as taught by Patel, in order to know how much power the cluster is drawing and how much it may draw.
Motivation would increase the serving capacity available from the same power supply because clusters use only part of their budgeted power, and the unused share safely supports more servers when the draw is watched against the budget, as taught by Patel (Abstract).
Stojkovic in view of Kubernetes in view of Patel do not explicitly teach:
converting (i) Multi-Instance GPU (MIG) service to GPU service or (ii) the GPU service to the MIG service, said converting being implemented as a function of the GSF, the MSF, and the PF.
However, Duluk teaches:
converting (i) Multi-Instance GPU (MIG) service to GPU service or (ii) the GPU service to the MIG service ([0148], NVIDIA previously introduced a Multiple Instance GPU ("MIG") feature that allows a GPU to be spatially subdivided into multiple smaller GPU Instances, each GPU Instance of which can be running a different instance of an operating system (or separate containers under one OS); [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances using virtual GPUs; [0166], avoid a full reset that may be needed to reconfigure a GPU chip while allowing unused portions of the chip to be turned off when not needed and turned back on when needed, and also to dynamically reconfigure hardware partitions of the GPU so different numbers of users and/or applications can make use of differently sized hardware/processing partitions depending upon need; The examiner notes that the transition from GPU Instances to the full GPU is the claimed converting of the MIG service to the GPU service, and the reverse transition is the claimed converting of the GPU service to the MIG service),
said converting being implemented as a function of the GSF, the MSF, and the PF ([0166], depending upon need; The examiner notes that in the combination the need is stated by the calculated factors, the change toward the full GPU taken when the factors call for more capacity and the change toward MIG devices when they call for less or when the power draw reaches the trigger value of Patel's thresholds above).
Stojkovic, Kubernetes, Patel, and Duluk are all concerned with matching computing hardware capacity to the workloads it serves and are therefore combinable/modifiable.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Stojkovic in view of Kubernetes in view of Patel to include MIG transitions of each GPU between a full GPU and sets of GPU Instances with the direction of each transition being selected by the calculated factors, as taught by Duluk, in order to grow or shrink each GPU’s serving capacity without adding or removing hardware.
Motivation would improve the utilization of the GPU hardware because unused portions of the chip are turned off when not needed and turned back on when needed, differently sized hardware partitions serve different numbers of users and applications, as taught by Duluk ([0166]).
As per claim 2, Stojkovic in view of Kubernetes in view of Patel in view of Duluk discloses the claimed invention as detailed above for claim 1, and further teaches wherein during one iteration of the iterative process, the one iteration comprises:
determining that GSF > 1 (Kubernetes Algorithm details, if the current metric value is 200m, and the desired value is 100m, the number of replicas will be doubled, since 200.0 / 100.0 == 2.0 and skips any scaling action if the ratio is sufficiently close to 1.0; The examiner notes that the worked example doubles the capacity because dividing the current value by the target yields 2.0, so the grow action is taken when the computed factor exceeds one, which is the claimed determining that GSF > 1),
PF > 1 (Patel Section 1, LLM inference clusters utilize only up to 80% of the provisioned power, making them excellent candidates for power oversubscription; The examiner notes that using at most 80% of the provisioned power means the cluster draws less than its budget, and the budget divided by a smaller draw is greater than one, so a cluster with that headroom is one in which the PF exceeds one),
and the one or more GPUs are not in a pure GPU state Duluk [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances)
and in response, converting, by the one or more processors, the MIG service to the GPU service (Duluk [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances).
As per claim 3, Stojkovic in view of Kubernetes in view of Patel in view of Duluk disclose the claimed invention as detailed above for claims 1 and 2, and further teaches wherein said converting results in a current GPU-MIG state changing from a pure MIG state to the pure GPU state or to a mixed GPU-MIG state (Duluk [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances using virtual GPUs; The examiner notes that with eight 1/8 GPU Instances every serving resource is a MIG device, the pure MIG state as construed above, and with the full GPU the entire GPU serves, the pure GPU state. The transition from the eight instances to the full GPU is therefore the claimed change from a pure MIG state to the pure GPU state).
As per claim 5, Stojkovic in view of Kubernetes in view of Patel in view of Duluk disclose the claimed invention as detailed above for claim 1, and further teaches wherein during one iteration of the iterative process, the one iteration comprises:
determining that the MSF > 1 or the PF < 1 (Kubernetes, Algorithm details, if the current value is instead 50m, you'll halve the number of replicas, since 50.0 / 100.0 == 0.5; The examiner notes that the worked example halves the capacity because dividing the current value by the target yields 0.5, so the shrink action is taken when the current value has fallen below its target. Written with the target on top, the order addressed at claim 1, the factor of that example reads 100.0 / 50.0 == 2.0, the claimed determining that the MSF is greater than one; Patel, Section 5.1, upon reaching T1, we set all the low-priority workloads to the base frequency and when T2 is breached, we start by frequency capping all the low-priority workloads; The examiner notes that the capping actions fire when the draw reaches a threshold's power value, so at the trigger the consumption has met the budgeted level and the budget divided by the draw, the PF, is less than one),
and the one or more GPUs are not in a pure GPU state (Duluk [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances)
and in response, converting, by the one or more processors, the MIG service to the GPU service (Duluk [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances).
As per claim 6, Stojkovic in view of Kubernetes in view of Patel in view of Duluk disclose the claimed invention as detailed above for claims 1 and 6, and further teaches wherein said converting results in a current GPU-MIG state changing from a pure GPU state to the pure MIG state or to a mixed GPU-MIG state (Duluk [0167], FIG. 14 shows such MIG transitions between a full GPU, two half GPU Instances, four quarter GPU Instances, and eight 1/8 GPU Instances using virtual GPUs; The examiner notes that with the full GPU the entire GPU serves, the pure GPU state as construed above, and with eight 1/8 GPU Instances every serving resource is a MIG device, the pure MIG state. The transition from the full GPU to the eight instances is therefore the claimed change from a pure GPU state to the pure MIG state).
As per claim 8, it has similar limitations as claim 1 and is therefore rejected using the same rationale. Duluk further teaches a computer program product, comprising one or more computer readable hardware storage devices having computer readable program code stored therein, said program code containing instructions executable by one or more processors of a computer system to implement a method ([0036], the NVIDIA virtual GPU (vGPU) software is installed at a virtualization layer along with a hypervisor and each node can support multiple Virtual Machines (VMs), where each VM runs its own instance of an Operating System (OS); The examiner notes that installed software is program code held on the node's storage devices and executed through its processors, so the vGPU software and the operating systems are stored program code executed by one or more processors).
As per claim 9, it has similar limitations as claim 2 and is therefore rejected using the same rationale.
As per claim 10, it has similar limitations as claim 3 and is therefore rejected using the same rationale.
As per claim 12, it has similar limitations as claim 5 and is therefore rejected using the same rationale.
As per claim 13, it has similar limitations as claim 6 and is therefore rejected using the same rationale.
As per claim 15, it has similar limitations as claim 1 and is therefore rejected using the same rationale. Duluk further teaches a computer system, comprising one or more processors, one or more memories, and one or more computer readable hardware storage devices ([0028], an overall system may include any number of such multi-GPC processing units 202 and associated memories 204 that are coupled to a host CPU via a memory bridge 105 and [0036], the NVIDIA virtual GPU (vGPU) software is installed at a virtualization layer along with a hypervisor).
As per claim 16, it has similar limitations as claim 2 and is therefore rejected using the same rationale.
As per claim 17, it has similar limitations as claim 3 and is therefore rejected using the same rationale.
As per claim 19, it has similar limitations as claim 5 and is therefore rejected using the same rationale.
As per claim 20, it has similar limitations as claim 6 and is therefore rejected using the same rationale.
Claim(s) 4, 7, 11, 14, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Stojkovic in view of Kubernetes in view of Patel in view of Duluk in view of Tan et al. ("Serving DNN Models with Multi-Instance GPUs: A Case of the Reconfigurable Machine Scheduling Problem," arXiv:2109.11067, September 18, 2021) (hereinafter Tan).
As per claim 4, Stojkovic in view of Kubernetes in view of Patel in view of Duluk disclose the claimed invention as detailed above for claims 1 and 2 but do not explicitly teach:
said converting results in a current GPU-MIG state changing from a first mixed GPU-MIG state to the pure GPU state or to a second mixed GPU-MIG state.
However, Tan teaches:
said converting results in a current GPU-MIG state changing from a first mixed GPU-MIG state to the pure GPU state or to a second mixed GPU-MIG state (Section 1, MIG supports partial reconfiguration [45]: a subset of a GPU's instances can be repartitioned on-the-fly, without affecting other working instances on the same GPU; Section 3, first, GPUs are partitioned into homogeneous instances (either 1/7 or 7/7 instances) ... Second, GPUs are partitioned to heterogeneous instances (a mix of multiple instance sizes); Section 1, two of the 7 instances in A100 (which we call 1/7 instances) can merge to a 2/7 instance with twice the resources; The examiner notes that Tan serves models on a fleet of MIG-partitioned GPUs and repartitions the fleet a subset at a time, the resources outside the subset continuing to serve unchanged per Section 1. A fleet with some GPUs serving whole and some partitioned into MIG devices is in the claimed mixed GPU-MIG state as construed above. Converting one subset under the combination therefore moves the fleet from that first mixed state to a second mixed state, and converting the last partitioned GPUs moves it to the pure GPU state).
Stojkovic, Kubernetes, Patel, Duluk, and Tan are all concerned with adjusting the computing capacity that serves a changing workload and are therefore combinable/modifiable.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Stojkovic in view of Kubernetes in view of Patel in view of Duluk to include partial reconfiguration of the GPUs one subset at a time, with the remaining GPUs continuing to serve their LLM inference requests while the subset changes between MIG devices and entire GPUs, as taught by Tan, in order to keep the cluster serving during every conversion.
Motivation would reduce the number of GPUs needed to sustain a given throughput because a subset of a GPU's instances is repartitioned on the fly without affecting other working instances, as taught by Tan (Abstract and Section 1).
As per claim 7, Stojkovic in view of Kubernetes in view of Patel in view of Duluk disclose the claimed invention as detailed above for claims 1 and 5 but do not explicitly teach:
said converting results in a current GPU-MIG state changing from a first mixed GPU-MIG state to a pure MIG state or to a second mixed GPU-MIG state.
However, Tan teaches:
said converting results in a current GPU-MIG state changing from a first mixed GPU-MIG state to a pure MIG state or to a second mixed GPU-MIG state (Section 1, MIG supports partial reconfiguration [45]: a subset of a GPU's instances can be repartitioned on-the-fly, without affecting other working instances on the same GPU and Section 3, first, GPUs are partitioned into homogeneous instances (either 1/7 or 7/7 instances) ... Second, GPUs are partitioned to heterogeneous instances (a mix of multiple instance sizes)).
Stojkovic, Kubernetes, Patel, Duluk, and Tan are all concerned with adjusting the computing capacity that serves a changing workload and are therefore combinable/modifiable.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Stojkovic in view of Kubernetes in view of Patel in view of Duluk to include partial reconfiguration of the GPUs one subset at a time, with the remaining GPUs continuing to serve their LLM inference requests while the subset changes between MIG devices and entire GPUs, as taught by Tan, in order to keep the cluster serving during every conversion.
Motivation would reduce the number of GPUs needed to sustain a given throughput because a subset of a GPU's instances is repartitioned on the fly without affecting other working instances, as taught by Tan (Abstract and Section 1).
As per claim 11, it has similar limitations as claim 4 and is therefore rejected using the same rationale.
As per claim 14, it has similar limitations as claim 7 and is therefore rejected using the same rationale.
As per claim 18, it has similar limitations as claim 4 and is therefore rejected using the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Roche et al. (US 2023/0236887) discloses a method for allocating graphics processing unit partitions for a computer vision environment ... generating an optimal GPU partition allocation ... using a GPU partition model (abstract), which relates to the claimed GPU resource allocation by partitioning of the one or more GPUs.
Lewis et al. (US 2017/0339196) discloses in response to receipt of a notification from a third service, a scaling policy specified by a customer ... is obtained ... a request is submitted to a second service to scale the resource (abstract), which relates to the claimed scaling of computing resources toward configured target values.
Aman et al. (US 5,473,773) discloses a workload manager creates goal control data ... calculating a performance index for each class; selecting a receiver class to receive improved service based on the relative performance indexes (abstract), which relates to the claimed calculating of parameters comparing current metrics against target values.
Examiner has cited particular columns/paragraphs/sections and line numbers in the references applied and not relied upon to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THANH NGO whose telephone number is (571)270-3019. The examiner can normally be reached M-F 9am to 6pm ET, first F of biweek off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Vital can be reached at (571)272-4215. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.N./Examiner, Art Unit 2198
/PIERRE VITAL/Supervisory Patent Examiner, Art Unit 2198