Prosecution Insights
Last updated: October 02, 2026
Application No. 18/334,723

System and Method for Token-based Graphics Processing Unit (GPU) Utilization

Final Rejection §103
Filed
Jun 14, 2023
Examiner
NGUYEN, BRANDON A
Art Unit
2195
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
19 currently pending
Career history
19
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments and amendments appear to overcome the 101 rejection, and it is hereby rescinded. Applicant’s arguments with respect to claim 1 has been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 5. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Maschmeyer et al. Pub No. US-20240311192-A1 with a priority date to Provisional Application No. 63/501,854 (hereafter Maschmeyer) in view of Dutta et al. Pub. No. US 2023/0251785 A1 (hereafter Dutta), GU et al. Pub. No. WO 2024/073087 A1 (priority to 9/29/2022, hereafter Gu) and … 6. Regarding claim 1, Maschmeyer teaches “A computer-implemented method, executed on a computing device, comprising: generating, at an artificial intelligence (AI) model, processed workload data by processing workload data of a workload on a processing unit; ([0061] – [0063] teaches a LLM generating outputs based on user input (prompts) using processing units. [0073] teaches that the tokenization of prompts may be done by the LLM such that processed workload data is generated). Maschmeyer does not explicitly teach a simulation engine to determine available cache. Dutta teaches a simulation engine such that it teaches the limitation simulating, at a simulation engine, performance of the processing unit based on the processed workload data to determine a maximum number of … cache blocks available for the workload data; ([0049-0052] teaches providing a workload profile to a simulation engine to simulate the execution of a workload in order to determine an amount of additional headroom of a storage system such that the method may be used to determine the amount of available storage for the workload based on workload data). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to apply the teachings of Dutta with the invention of Maschmeyer to teach as evidence, that processed workload data may be simulated in order to determine an available amount of storage to execute the workload. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of improving provisioning decisions such as data reduction (Dutta [0076]). The combination does not explicitly teach that the determined resource are KV cache blocks. Gu teaches transformer-based models using KV caches such that it teaches the limitation “… KV cache blocks available for the workload data ([0073] teaches that KV caches can be used to store computational results. [0016-0020 teach of inferencing using tokens in order to produce embeddings such that the inferred tokens are then stored as a result in the KV caches)”. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to apply the teachings of Gu to the combination of Maschmeyer and Gu to teach as evidence, the amount of storage available to store computational results may be a KV cache. A person having ordinary skill in the art would have been motivated to make this combination in order to execute future inference runs with more efficiency (Gu [0073]). The combination now teaches “determining a token utilization for the workload data based upon, the maximum number of KV cache blocks available for the workload data (Maschmeyer [0074] teaches that total resource usage parameter may be the token count, tokens used in generating a response. Maschmeyer [0016] teaches that the total resource usage parameter is computed based on the prompt resource usage parameter. Maschmeyer [0140] teaches that the prompt resource usage parameter is a computed token count associated with the prompt such that when the prompt is simulated in the engine taught by Dutta, it determines an available storage amount for the computed token count. Also the predicted response usage parameter and complexity measure mentioned in Maschmeyer [0140] are associated with the maximum resource capacity, such that when combined with the prompt resource usage parameter, it represents the total available resource capacity that can be allocated to the user such that it determines the total token count, hence the token utilization based upon a simulated prompt, predicted response token count based on the prompt, and the remaining storage used as a complexity measure. As Dutta only discloses that the system determines a respective amount of storage used by the workload, it would be obvious for one of ordinary skill to apply the total token count calculation in determining the total amount of storage used from simulating a workload, in this case the workload being a prompt run on a processing unit); and allocating processing unit resources for the AI model based upon, the token utilization (Dutta [0020] teaches selecting a storage system based upon the simulations of the workload, wherein the results of the simulation may result in a total token count as described in Maschmeyer, ).” 7. Regarding claim 2, wherein the combination Maschmeyer teaches “The computer-implemented method of claim 1, wherein processing the workload data includes mirroring a plurality of requests received for processing by the AI model on the processing unit to the simulation engine. ([0133]-[0135] teaches a resource prediction model estimating resources for a given prompt such that it teaches mirroring the requests onto a simulation engine).” 8. Regarding claim 3, wherein the combination Maschmeyer implicitly teaches “The computer-implemented method of claim 1, wherein determining the token utilization includes determining a processing unit memory utilization limit. ([0073]-[0075] teaches that the computer systems’ memory and computing power dictates the token limit such that it the memory utilization limit must be determined before determining the token limit).” 9. Regarding claim 4, wherein the combination Maschmeyer teaches “The computer-implemented method of claim 1, wherein determining the token utilization includes determining a processing unit computing utilization limit. ([0073]-[0075] also implicitly teaches that similar to claim 3).” 10. Regarding claim 5, wherein the combination Maschmeyer teaches “The computer-implemented method of claim 1, wherein determining the maximum number of KV cache blocks available for the workload data includes converting the maximum number of KV cache blocks available into a number of tokens available. ([0074]-[0076] teaches total resource usage represents a token count with respect to available tokens remaining).” 11. Regarding claim 6, wherein the combination Maschmeyer teaches “The computer-implemented method of claim 5, wherein determining the token utilization includes determining a number of processing tokens ([0139]-[0140] and [00001] teaches the equation used to determine total utilization. Also [0073] teaches “For example, an LLM may have a token limit of 4097 tokens which can be shared between a prompt and a response (e.g., a prompt that uses 4000 tokens will limit the response to a maximum of 97 tokens). [0055]-[0056] teaches how input tokens are broken down into processing tokens).” 12. Regarding claim 7, wherein the combination Maschmeyer teaches “The computer-implemented method of claim 6, wherein determining the token utilization includes determining a performance configuration for the workload data based upon, at least in part, the number of tokens available and the number of processing tokens ([0140]-[0150] and Fig. 3A-3D, 4 teach generating a visual representation of the configuration and recommendations/warnings based on user input and the calculated token utilization).” 13. Regarding claim 8, wherein the combination Maschmeyer teaches “The computer-implemented method of claim 7, wherein allocating the processing unit resources for the AI model includes allocating processing unit resources for the AI model using the performance configuration. ([0140]-[0145] teaches a visual representation of total resources used based on token utilization such that it is already using the performance configuration to allocate the resources as soon as the user continues with their prompt).” 14. Claim 9 is similar to claims 1 and 2, and is therefore rejected for similar reasons. Claim 9 is directed towards “A computing system (Maschmeyer [0023] computing system) comprising: a memory; and a processor configured to generate, at an artificial intelligence (AI) model, processed workload data by processing workload data of a workload on a graphics processing unit (GPU) (Maschmeyer [0064] processor may be a GPU) including mirroring a plurality of requests of the workload received for processing by the AI model on the GPU to a simulation engine, to determine a maximum number of key-value (KV) cache blocks available for the workload data, to determine a token utilization for the workload data based upon, the maximum number of KV cache blocks available for the workload data, and to allocate GPU resources for the AI model based upon, at least in part, the token utilization.” 15. Claim 10 is similar to claim 3, therefore is rejected for similar reasons. 16. Claim 11 is similar to claim 4, therefore is rejected for similar reasons. 17. Claim 12 is similar to claim 5, therefore is rejected for similar reasons. 18. Claim 13 is similar to claim 6, therefore is rejected for similar reasons. 19. Claim 14 is similar to claim 7, therefore is rejected for similar reasons. 20. Claim 15 is similar to claim 1, therefore is rejected for similar reasons. Claim 15 is directed towards “A computer program product (Maschmeyer [0188]) residing on a non-transitory computer readable medium (Maschmeyer [0023]) having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising: processing, at an artificial intelligence (AI) model, processed workload data of a workload on a graphics processing unit (GPU); simulating, at a simulation engine, performance of the processing unit based on the processed workload data to determine a maximum number of key-value (KV) cache blocks available for the workload data; converting the maximum number of KV cache blocks available into a maximum number of tokens available; determining a token utilization for the workload data based upon, at least in part, the number of tokens available; and allocating GPU resources for the AI model based upon, at least in part, the token utilization.” 21. Claim 16 is similar to claim 2, therefore is rejected for similar reasons. 22. Claim 17 is similar to claims 3 and 10, and are therefore rejected for similar reasons. 23. Claim 18 is similar to claims 4 and 11, and are therefore rejected for similar reasons. 24. Claim 19 is similar to claims 6 and 13, and are therefore rejected for similar reasons. 25. Claim 20 is similar to claims 7 and 14, and are therefore rejected for similar reasons. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRANDON A NGUYEN whose telephone number is (571)272-6074. The examiner can normally be reached Mon-Fri (10am-6pm). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BRANDON NGUYEN/Examiner, Art Unit 2195 /PIERRE VITAL/Supervisory Patent Examiner, Art Unit 2198
Read full office action

Prosecution Timeline

Jun 14, 2023
Application Filed
Jan 22, 2026
Non-Final Rejection mailed — §103
Apr 20, 2026
Applicant Interview (Telephonic)
Apr 20, 2026
Examiner Interview Summary
Apr 22, 2026
Response Filed
Jul 16, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month