Prosecution Insights
Last updated: August 17, 2026
Application No. 18/732,130

METHOD FOR GPU RESOURCE MANAGEMENT USING REINFORCEMENT LEARNING AND APPARATUS USING THE SAME

Non-Final OA §103
Filed
Jun 03, 2024
Priority
Nov 27, 2023 — RE 10-2023-0166527
Examiner
SWIFT, CHARLES M
Art Unit
Tech Center
Assignee
Electronics and Telecommunications Research Institute
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
722 granted / 891 resolved
+21.0% vs TC avg
Strong +22% interview lift
Without
With
+21.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
36 currently pending
Career history
936
Total Applications
across all art units

Statute-Specific Performance

§101
11.3%
-28.7% vs TC avg
§103
56.7%
+16.7% vs TC avg
§102
16.5%
-23.5% vs TC avg
§112
6.2%
-33.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 891 resolved cases

Office Action

§103
DETAILED ACTION This office action is in response to application filed on 6/3/2024. Claims 1 – 20 are pending. Priority is claimed to Korean application KR10-2023-0166527 (filed on 11/27/2023). Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 11 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kurkure et al (US 20240362052, hereinafter Kurkure), and in view of Zad Tootaghaj et al (US 20230281052, hereinafter Zad Tootaghaj). As per claim 1, Kurkure discloses: A method for Graphics Processing Unit (GPU) resource management using reinforcement learning, performed by an apparatus for GPU resource management, comprising: deriving a Multi-Instance GPU (MIG) instance configuration that meets a Service Level Objective (SLO) condition assigned to a workload by utilizing a pretrained reinforcement learning model; (Kurkure [0028]: “Starting with step 302, VIM server 102 can receive requests for placing the M VMs on the N GPUs, where each request includes a MIG profile specifying the number of compute slices and the number of memory slices requested by the corresponding VM. For example, a first request for a first VM may include a MIG profile that specifies two compute slices and two memory slices, a second request for a second VM may include a MIG profile that specifies three compute slices and two memory slices, and so on.”; [0029] – [0041]: “Upon receiving these requests, VIM server 102 can proceed with formulating an ILP problem for computing an optimal placement of the M VMs on the N GPUs… Once the foregoing problem components are created/defined, VIM server 102 can generate a solution to the ILP problem using any ILP solver known in the art (e.g., Gurobi optimizer, etc.) (step 314). This solution will include values for the decision variables that minimize the objective function while satisfying the constraints, thereby resulting in an optimal placement of the M VMs on the N GPUs.”; [0013]: “Each MIG instance 204 includes a separate execution path through GPU 202 that includes a dedicated portion of GPU 202's compute resources (reference numeral 208) and a dedicated portion of GPU 202's memory resources (reference numeral 210). This advantageously ensures that the GPU workload of each VM 206 runs with predictable quality of service (e.g., throughput, latency, etc.) and prevents one VM from impacting the work or scheduling of another.”) and reorganizing MIG resources of a GPU device to correspond to the workload by transferring the MIG instance configuration to the GPU device. (Kurkure [0042]: “Finally, at step 316, VIM server 102 can proceed with placing the VMs in accordance with the solution generated at 314. This process can include, e.g., creating or updating metadata associated with each VM to indicate the GPU on which it is placed and the VM's MIG profile. With this metadata in place, upon being powered on, the VM will be able to access a MIG instance of the GPU with the resources specified in its MIG profile.”) Kurkure did not explicitly disclose: Wherein the workload further comprises a request rate; However, Zad Tootaghaj teaches: Wherein the workload further comprises a request rate; (Zad Tootaghaj [0016]) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Zad Tootaghaj into that of Kurkure in order to have the workload further comprises a request rate. Kurkure [0013[ teaches some examples to the quality of service, one of ordinary skill in the art can easily see that other form of constraints or conditions can be imposed on the workload as well, such as request rate, that can dictate the SLO of the workload, applicants have thus merely claimed the combination of known parts in the field to achieve predictable results of having request rate being a part of the workload’s requirement, and is therefore rejected under 35 UCS 103. As per claim 11, the combination of Kurkure and Zad Tootaghaj further teach: The method of claim 1, wherein the apparatus runs in a backend process of a cloud environment. (Kurkure [0009]) As per claim 12, it is the apparatus variant of claim 1 and is therefore rejected under the same rationale. (Kurkure figure 1) Claim(s) 2 – 4 and 13 – 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kurkure and Zad Tootaghaj, and further in view of Padmanbha Iyer et al (US 20230342278, hereinafter Padmanabha Iyer). As per claim 2, the combination of Kurkure and Zad Tootaghaj did not teach: The method of claim 1, wherein deriving the MIG instance configuration includes setting a maximum batch size that does not violate a latency constraint defined in the SLO condition; and measuring throughput of the workload depending on the maximum batch size. However, Padmanbha Iyer teaches: The method of claim 1, wherein deriving the MIG instance configuration includes setting a maximum batch size that does not violate a latency constraint defined in the SLO condition; and measuring throughput of the workload depending on the maximum batch size. (Padmanabha Iyer [0040] and [0067]) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Padmanabha Iyer into that of Kurkure and Zad Tootaghaj in order to derive the MIG instance configuration includes setting a maximum batch size that does not violate a latency constraint defined in the SLO condition; and measuring throughput of the workload depending on the maximum batch size. Kurkure [0013] teaches some examples to the quality of service that can be used to influence the placements of the VMs, one of ordinary skill in the art can easily see that other form of constraints or conditions can be imposed on the workload as well, such as maximum batch size and latency, that can dictate the SLO of the workload, applicants have thus merely claimed the combination of known parts in the field to achieve predictable results of having request rate being a part of the workload’s requirement, and is therefore rejected under 35 UCS 103. As per claim 3, the combination of Kurkure, Zad Tootaghaj and Padmanabha Iyer further teach: The method of claim 2, wherein setting the maximum batch size comprises setting a batch size that makes a sum of batching latency and inference latency closest to the latency constraint without exceeding the latency constraint as the maximum batch size. (Padmanabha Iyer [0040] and [0067]) As per claim 4, the combination of Kurkure, Zad Tootaghaj and Padmanabha Iyer further teach: The method of claim 2, wherein the maximum batch size is calculated by performing curve fitting on latency data of a specific batch size for each instance of the MIG instance configuration. (Padmanabha Iyer [0030]) As per claim 13, it is the apparatus variant of claim 2 and is therefore rejected under the same rationale. As per claim 14, it is the apparatus variant of claim 3 and is therefore rejected under the same rationale. As per claim 15, it is the apparatus variant of claim 4 and is therefore rejected under the same rationale. Claim(s) 5 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kurkure, Zad Tootaghaj and Padmanbha Iyer, and further in view of Wang et al (US 20210037113, hereinafter Wang). As per claim 5, the combination of Kurkure, Zad Tootaghaj and Padmanbha Iyer did not teach: The method of claim 2, further comprising: performing, by the apparatus, evaluation about whether the MIG instance configuration is a configuration that allocates a minimum amount of the MIG resources while meeting the SLO condition and the request rate. However, Wang teaches: The method of claim 2, further comprising: performing, by the apparatus, evaluation about whether the MIG instance configuration is a configuration that allocates a minimum amount of the MIG resources while meeting the SLO condition and the request rate. (Wang figure 8 and [0151]) It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Wang into that of Kurkure, Zad Tootaghaj and Padmanabha Iyer in order to perform evaluation about whether the MIG instance configuration is a configuration that allocates a minimum amount of the MIG resources while meeting the SLO condition and the request rate. Using the minimum amount of resource required to meet the SLO would prevent overprovisioning of the resource and improves the overall placement efficient and is therefore rejected under 35 USC 103. As per claim 16, it is the apparatus variant of claim 5 and is therefore rejected under the same rationale. Allowable Subject Matter Claims 6 – 10 and 17 – 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chen et al (US 20250077289) teaches “optimizing graphics-processing unit (GPU) utilization and a system thereof. The method includes the following steps: receiving an application workload which is to be executed on the GPU; predicting GPU resource requirements for the application workload; scheduling the application workload according to the prediction of the GPU resource requirements; dynamically allocating and deallocating GPU resources based on the prediction of the GPU resource requirements for the application workload; and executing the application workload on the GPU.”; Cho et al (US 20230089925) teaches “Architectures and techniques for managing heterogeneous sets of physical GPUs. Functionality information is collected for one or more physical GPUs with a GPU device manager coupled with a heterogeneous set of physical GPUs. At least one of the physical GPUs is to be managed as multiple virtual GPUs based on the collected functionality information with the GPU device manager. Each of the physical GPUs is classified as either a single physical GPU or as one or more virtual GPUs with the device manager. Traffic representing processing jobs to be processed is received by at least a subset of the physical GPUs via a gateway programmed by a traffic manager. The GPU application to process received processing jobs scheduled by and distributed into the scheduled GPU application with a GPU scheduler communicatively coupled with the traffic manager and with the GPU device manager.”; Bernat (US 20210271517) teaches “determine multiple configurations of hardware resources to perform a workload associated with a workload request in a subsequent stage based on a pre-processing operation associated with the workload request and at least one service level agreement (SLA) parameter associated with the workload request. In some examples, an executable binary is associated with the workload request and execution of the executable binary performs the pre-processing operation. In some examples, the circuitry is to store the multiple configurations of hardware resources to perform a workload associated with the workload request in a subsequent stage, wherein the multiple configurations of hardware resources are available for access by one or more accelerator devices to perform the workload.”. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHARLES M SWIFT whose telephone number is (571)270-7756. The examiner can normally be reached Monday - Friday: 9:30 AM - 7PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, April Blair can be reached at 5712701014. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CHARLES M SWIFT/Primary Examiner, Art Unit 2196
Read full office action

Prosecution Timeline

Jun 03, 2024
Application Filed
Jul 27, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705094
GRAPH COMPUTING APPARATUS, PROCESSING METHOD, AND RELATED DEVICE
3y 5m to grant Granted Aug 11, 2026
Patent 12699588
DISTRIBUTED COMMUNICATION BETWEEN DEVICES
3y 10m to grant Granted Aug 04, 2026
Patent 12693905
WORKLOAD MIGRATION RECOMMENDATIONS IN HETEROGENEOUS WORKSPACE ENVIRONMENTS
4y 11m to grant Granted Jul 28, 2026
Patent 12693897
Methods and Apparatus for Supporting Application Mobility in Multi-Access Edge Computing Platform Architectures
3y 2m to grant Granted Jul 28, 2026
Patent 12681741
VIRTUAL VOLUME PLACEMENT BASED ON ACTIVITY LEVEL
4y 2m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+21.7%)
3y 0m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 891 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month