Prosecution Insights
Last updated: August 17, 2026
Application No. 18/652,420

REQUEST THROTTLING IN MODEL-AS-A-SERVICE PLATFORM

Non-Final OA §101§103§112
Filed
May 01, 2024
Examiner
RIGGINS, ARI FAITH COLEMA
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
50%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
2 granted / 4 resolved
-10.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
20 currently pending
Career history
40
Total Applications
across all art units

Statute-Specific Performance

§101
27.1%
-12.9% vs TC avg
§103
43.6%
+3.6% vs TC avg
§102
8.4%
-31.6% vs TC avg
§112
20.9%
-19.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office Action is in response to claims filed on 05/01/2024. Claims 1-20 are pending. Claim Objections Claims 1-16 are objected to because of the following informalities: The limitation “and a throttling service that that limits” in claim 1 should read “and a throttling service requested number of input tokens”. The limitation “throttling customer endpoint requests to directed to instances” in claim 10 should read “throttling customer endpoint requests . Appropriate correction is required. Claims 2-9 and 11-16 depend, directly or indirectly, from objected to claims and do not resolve the deficiencies thereof and are therefore objected to for at least the same reasons. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a metric standardizer” of claims 1-9 which is accompanied by the functional language of “a metric standardizer that determines” in claim 1, “wherein the metric standardizer is further configured to:” in claim 2. Claims 2-9 depend, directly or indirectly, from the claims and thus inherit the invocation of 112(f) and are therefore included in the following claim interpretation. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. A review of the specification reveals the corresponding structure of metric standardizer as merely general purpose computers. “One or more applications 512 (e.g., LLMs, the throttling service 204, the metric standardizer 314 or the autoscaler 312) are loaded in the memory device(s) 504 and executed on the operating system510 by the processor unit(s) 502” [¶ 101]. MPEP § 2181(II)(B) states: “For a computer- implemented 35 U.S.C. 112(f) claim limitation, the specification must disclose an algorithm for performing the claimed specific computer function, or else the claim is indefinite under 35 U.S.C. 112(b)”. The specification fails to disclose the computer + algorithm and, as such, is rejected under 35 U.S.C. § 112(a) and (b) below. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-9 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. “When a claim containing a computer-implemented 35 U.S.C. 112(f) claim limitation is found to be indefinite under 35 U.S.C. 112(b) for failure to disclose sufficient corresponding structure (e.g., the computer and the algorithm) in the specification that performs the entire claimed function, it will also lack written description under 35 U.S.C. 112(a)” [MPEP § 2181(II)(B)]. The claims have been found to be indefinite under 35 U.S.C. 112(b), as referenced below, and thus lack written description under 35 U.S.C. 112(a). The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim limitations “a metric standardizer that determines” in claim 1, “wherein the metric standardizer is further configured to: receive” in claim 2. invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The specification is devoid of adequate structure to perform the claimed functions as no algorithm is disclosed for performing the specific computer functions, as referenced above in the 112(f) analysis. There is no disclosure of any particular structure, either explicitly or inherently, to perform the above function of execute. The use of the terms “determines” and “receive” is not adequate structure for performing the function because it does not describe a particular structure for performing the function. In accordance with MPEP § 2181(II)(B), the specification must disclose corresponding computer + algorithm for computer-implemented means-plus-function limitations. No such computer + algorithm is disclosed. See also “Claim Interpretation” under 35 U.S.C. § 112(f) above. Claims 2-9 depend, directly or indirectly, from rejected claims and do not resolve the deficiencies thereof and are therefore rejected for at least the same reasons. Therefore, the claims are indefinite and are rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 17-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the claim is directed to “one or more computer-readable storage media” which is a signal per se. “For example, the BRI of machine readable media can encompass non-statutory transitory forms of signal transmission, such as a propagating electrical or electromagnetic signal per se. See In re Nuijten, 500 F.3d 1346, 84 USPQ2d 1495 (Fed. Cir. 2007). When the BRI encompasses transitory forms of signal transmission, a rejection under 35 U.S.C. 101 as failing to claim statutory subject matter would be appropriate. Thus, a claim to a computer readable medium that can be a compact disc or a carrier wave covers a non-statutory embodiment and therefore should be rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. See, e.g., Mentor Graphics v. EVE-USA, Inc., 851 F.3d at 1294-95, 112 USPQ2d at 1134 (claims to a "machine-readable medium" were non-statutory, because their scope encompassed both statutory random-access memory and non-statutory carrier waves)” [MPEP§ 2106.03(II)]. The specification does not define “computer-readable storage media” as non-transitory but only mentions: “Tangible computer-readable storage media excludes intangible and transitory communications signals and includes volatile and nonvolatile, removable, and non-removable storage media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data” [¶ 104]. So “computer-readable storage media” may include at least non-transitory type and transitory type. Proposed amendment to overcome the U.S.C. 101 rejection: “non-transitory computer-readable storage media” or “tangible computer-readable storage media”. Claims 17-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, is directed to that judicial exception, an abstract idea, as it has not been integrated into practical application and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below. Step 1: Claims 17-20 are directed to one or more computer-readable storage media which is signal per se. Therefore, “Are the claims to a process, machine, manufacture or composition of matter?” No. In order to evaluate the Step 2A inquiry “Is the claim directed to a law of nature, a natural phenomenon or an abstract idea?” we must determine, at Step 2A Prong 1, whether the claim recites a law of nature, a natural phenomenon or an abstract idea and further whether the claim recites additional elements that integrate the judicial exception into a practical application. Step 2A Prong 1: Claim 17: The limitations of “determining, based on identity of the target LLM instance, an applicable subset of the model-specific benchmark metrics;”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can observe an identity of a target LLM and model-specific benchmark metrics and, based on these observations, can mentally determine an applicable subset of the model-specific benchmark metrics. Further the limitations of “using the applicable subset of the model-specific benchmark metrics to derive a model-agnostic estimated utilization associated with the LLM processing task, the model-agnostic estimated utilization being defined as a quantity of units of a model-agnostic unit type;”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can observe a subset of model-specific benchmark metrics and, based on these observations, can mentally derive a model-agnostic estimated utilization as a quantity of units of a model-agnostic unit type. This may also be done with pen and paper. Further the limitations of “and determining, by the throttling service, whether to grant or deny the LLM processing task based on the model-agnostic estimated utilization for the LLM processing task”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can observe a model-agnostic estimated utilization for a LLM processing task and, based on these observations, can mentally determine whether to grant or deny the LLM processing task. Therefore, Yes, claim 17 recites a judicial exception. Step 2A Prong 2: Claim 17: The judicial exception is not integrated into a practical application. In particular, the claims recite additional element recitations of “One or more computer-readable storage media encoding computer-executable instructions for executing a computer process for throttling customer requests to access instances of different large language models (LLMs) deployed within a model-as-a-service platform, the computer process comprising:”, which are merely recitations of generic computing components (see MPEP § 2106.05(f)) which does not integrate a judicial exception into practical application. Further, the claims recite additional element recitations of “storing model-specific benchmark metrics that define relationships between resource utilization and token processing according to different model-specific tokenization schemes utilized by different large language models;”, “receiving, from a throttling service, token-based job metrics identifying a target LLM instance, a requested number of input tokens, and a requested number of output tokens associated with an LLM processing task;”, and “and providing the model-agnostic estimated utilization to the throttling service;” which are merely recitations of data storage, reception, and transmission which are insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Therefore, “Do the claims recite additional elements that integrate the judicial exception into a practical application? No, these additional elements do not integrate the abstract idea into a practical application and they do not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea. After having evaluated the inquires set forth in Steps 2A Prong 1 and 2, it has been concluded that claim 17 not only recites a judicial exception but that the claims are directed to the judicial exception as the judicial exception has not been integrated into practical application. Step 2B: Claim 17: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than generic computing components/functions and insignificant extra solution activity which do not amount to significantly more than the abstract idea. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Therefore, “Do the claims recite additional elements that amount to significantly more than the judicial exception? No, these additional elements, alone or in combination, do not amount to significantly more than the judicial exception. Having concluded analysis within the provided framework, Claim 17 does not recite patent eligible subject matter under 35 U.S.C. § 101. With regard to claim 18, the claim recites additional abstract idea recitations of “wherein determining whether to grant or deny the LLM processing task is based on a customer-allotted quota that limits a number of requests a given client compute platform can submit to the model-as-a-service platform, the customer-allotted quota defined as a quantity units of the model-agnostic unit type”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can observe a customer allotted quota that is defined as a quantity units of the model-agnostic unit type and, based on these observations, can mentally determine whether to grant or deny the LLM processing task. Further, claim 18 does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 18 also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 18 does not recite patent eligible subject matter under 35 U.S.C. § 101. With regard to claim 19, the claim recites additional element recitations of “wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization”, which are merely recitations of technological environment/field of use (see MPEP § 2106.05(h)) which does not integrate a judicial exception into practical application. Further, claim 19 does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 19 also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 19 does not recite patent eligible subject matter under 35 U.S.C. § 101. With regard to claim 20, the claim recites additional abstract idea recitations of “determining, by the throttling service, whether to grant or deny the LLM processing task based on a sum of the model-agnostic estimated utilization and the current utilization of the client compute platform”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can observe a model-agnostic estimated utilization and the current utilization of the client compute platform and, based on these observations, can mentally determine whether to grant or deny the LLM processing task based on a mental calculation of the sum of the model-agnostic estimated utilization and the current utilization of the client compute platform. This may also be done with pencil and paper. Further, the claim recites additional element recitations of “wherein the computer process further comprises: querying, by the throttling service, a platform-level database to determine a current utilization of a client compute platform associated with the LLM processing task;”, which are merely recitations of data transmission which is insignificant extra solution activity (see MPEP §2106.05(g)) which does not integrate a judicial exception into practical application. Further, the insignificant extra solution activity is well-understood, routine, and conventional in the art. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network…iv. Storing and retrieving information in memory” [MPEP§ 2106.05(d)(II)]. Further, claim 20 does not recite any further additional elements and for the same reasons as above with regard to integration into practical application and whether additional elements amount to significantly more, claim 20 also fails both Step 2A prong 2, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 20 does not recite patent eligible subject matter under 35 U.S.C. § 101. Therefore, Claims 17-20 do not recite patent eligible subject matter under U.S.C. §101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3 and 8-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kuperman et al. US 2025/0292023 A1 (hereafter Kuperman) in view of Jain et al. US 12,147,513 B1 (hereafter Jain) in view of Hu et al. US 2024/0086249 A1 (hereafter Hu). With regard to claim 1, Kuperman teaches: A model-as-a-service platform including: “The AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140 … The one or more LLM SaaS platform(s) 130 may include (but are not limited to) non-open source LLMs, such as GPT-3, GPT-4, which are provided through APIs of a SaaS platform. The one or more open source LLM platform(s) 140 may include (but are not limited to) open source LLMs deployed on private Kubernetes (K8s)” [Kuperman ¶ 26-27]. a shared resource pool of graphics processing units (GPUs) shared (used) among model pools executing instances of different large language models (LLMs), “Additionally, in some embodiments, the AI automation system is further configured to time-slice GPUs and/or provide multi-instance GPU (MIG) device support. When an LLM model is deployed on a private Kubernetes (K8s) cluster, GPUs are utilized to their full extent. For example, if an entity runs the Llama 2 model on their private cluster on a Spot VM, there might be scenarios where the AI automation system determines that the GPU is not being fully utilized, such as at 70% of its capacity. In such cases, the AI automation system can split the GPU into two or more virtual GPUs and run two or more models side by side on the same GPU, resulting in nearly 100% utilization of the resource” [Kuperman ¶ 24]. the different LLMs generating and processing text according to different model-specific tokenization schemes; “The token count is a number of tokens in a prompt, which include individual units of text, such as words, subwords, or special tokens. The token count corresponds to complexity and specificity of the input given to the LLM. Also, a token count for a prompt depends on how the prompt is tokenized by a specific LLM being used. If the LLM tokenizes words and punctuation marks, the token count might be one number. However, if the LLM uses subword tokenization or includes special tokens, the token count would be a different number” [Kuperman ¶ 39]. “Some LLMs may support multimodal inputs, such as image input, audio input, or video input. When multimodal inputs are received, other tokenization methods may be applied” [Kuperman ¶ 42]. a metric standardizer that determines a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; “The pricing models for LLM services typically depend on the level of usage, which can be quantified in various ways, such as the number of API requests, the volume of text processed (measured in tokens or characters), or the amount of compute time utilized” [Kuperman ¶ 5]. the units of the customer-allotted quota being redeemable in exchange for compute tasks performed by the different LLMs. “LLMs are able to perform a variety of tasks that were traditionally considered challenging for computers. These tasks include generating human-like text, translating languages, summarizing long pieces of content, answering questions, and more” [Kuperman ¶ 3]. “Cost per million tokens is an example of a measure of cost-effectiveness when using Large Language Model (LLM) tooling in cloud environments, such as Software as a Service (SaaS) environments. Another example measure of cost-effectiveness, aside from SaaS environments, could be the cost per million API requests” [Kuperman ¶ 29]. Kuperman fails to teach a metric standardizer that determines a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; and a throttling service that that limits a number requests a user can concurrently submit to the instances of the different LLMs based on a customer-allotted quota of the units of the model-agnostic unit type, the units of the customer-allotted quota being redeemable in exchange for compute tasks. However, Jain teaches: a metric standardizer that determines a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; “Following input validation, the performance engine 118 can evaluate the performance of LLMs to route the prompt to an appropriate LLM (e.g., large language model(s) 410)” [Jain Col. 10 Lines 14-16]. “In some implementations, the data generation platform evaluates the performance requirements associated with the prompt generate a recommendation for a suitable LLM for the received prompt (e.g., to improve the efficiency of system resource use)” [Jain Col. 4 Lines 14-19]. “As such, the data generation platform 102 can dynamically select LLMs for processing output generation requests on the basis of estimated system resources and associated limitations or requirements, thereby improving the efficiency, flexibility, and robustness of the associated development pipeline” [Jain Col. 16 Lines 29-34]. and a throttling service that that limits a number requests a user can concurrently submit to the instances of the different LLMs based on a customer-allotted quota of the units of the model-agnostic unit type, “The data generation platform 102 can determine that the estimated resource metric value (e.g., the estimated cost) does not satisfy the threshold metric value. For example, the data generation platform 102 determines that the estimated cost is greater than or equal to the threshold cost. Based on this determination, the performance evaluation 408 can determine to prevent provision of the associated prompt to one or more LLMs of the LLMs 410” [Jain Col. 16 Lines 16-23]. “Furthermore, the threshold metric can depend on the user (e.g., the user device 402a of FIG. 4) and/or the service 402b. For example, the threshold metric includes resource allotments that are specific to particular users of the data generation platform 102” [Jain Col. 16 Lines 4-8]. “For example, the access control engine 114 determines to throttle the bandwidth associated with receiving outputs from the LLMs (e.g., by specifying a number of responses per unit time that are allowed to be transmitted to the given user). As such, the access control engine 114 can control the system-wide performance by limiting the assignment of system resources to particular users” [Jain Col. 12 Lines 30-37]. the units of the customer-allotted quota being redeemable in exchange for compute tasks performed by the different LLMs. “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request.” [Jain Col. 15 Lines 40-49]. Jain is considered to be analogous to the claimed invention because it is in the same field of workload prediction. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman to incorporate the teachings of Jain and include: a metric standardizer that determines a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; and a throttling service that that limits a number requests a user can concurrently submit to the instances of the different LLMs based on a customer-allotted quota of the units of the model-agnostic unit type, the units of the customer-allotted quota being redeemable in exchange for compute tasks. Doing so would allow for improved flexibility and control over the system resources. “By doing so, the data generation platform 102 provides improved flexibility and/or control over the use of system resources (including memory, computational, and/or financial resources), enabling optimization of the associated development pipeline” [Jain Col. 22 Lines 2-6]. Kuperman in view of Jain fails to explicitly teach a shared resource pool of graphics processing units (GPUs) shared among model pools executing instances of different large language models (LLMs). However, Hu teaches a shared resource pool of graphics processing units (GPUs) shared among model pools executing instances of different large language models (LLMs), “Cloud computing is a form of network-based computing (e.g., Internet-based computing) that enables access to shared pools of configurable computing resources and higher-level services that can be rapidly provisioned with minimal management effort, often over the Internet” [Hu ¶ 3]. “According to an example aspect of the present disclosure, there is provided a method for training a plurality of models using a cloud computing resource pool comprising a plurality of nodes. Each node comprises a plurality of processor devices” [Hu ¶ 26]. “A node is most commonly a cluster of 8 graphical processing units (GPUs), but may be some other number of GPUs (e.g., 2, 4, or 16) and/or other processor devices such as central processing units (CPUs), tensor processing units (TPUs), neural processing units (NPUs), or other hardware artificial intelligence (AI) accelerators” [Hu ¶ 6]. Hu is considered to be analogous to the claimed invention because it is in the same field of indexing schemes relating to resource pools. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman in view of Jain to incorporate the teachings of Hu and include: a shared resource pool of graphics processing units (GPUs) shared among model pools executing instances of different large language models (LLMs). Doing so would allow for the rapid provisioning of resources with minimal management needed. “Cloud computing is a form of network-based computing (e.g., Internet-based computing) that enables access to shared pools of configurable computing resources and higher-level services that can be rapidly provisioned with minimal management effort, often over the Internet” [Hu ¶ 3]. With regard to claim 2, Kuperman in view of Jain in view of Hu teaches the model-as-a-service platform of claim 1, as referenced above. Kuperman further teaches: wherein the metric standardizer is further configured to: receive, as input, token-based job metrics for the incoming customer-requested LLM processing task, “Monitoring prompt token counts can help ensure that the input provided to the LLM is within the acceptable limits of the LLM's input capacity and does not exceed any constraints set by an application or environment in which the LLM is being used” [Kuperman ¶ 39]. “The data store 410 may store collected historical prompts and other data associated with the historical prompts, such as (but not limited to) corresponding responses generated by LLMs, a token count associated with the prompt, a cost associated with the prompt, etc” [Kuperman ¶ 38]. a request number of input tokens, and a requested number of output tokens; “It then determines whether the request should be reconstructed to reduce a token count, which includes input and/or output token counts” [Kuperman ¶ 7]. and determine the model-agnostic estimated utilization for the incoming customer-requested LLM processing task based on stored model-specific benchmark metrics “Notably, the costs of LLM services are often directly or indirectly related to token counts. FIG. 6 illustrates an example cost per thousand tokens for a plurality of LLMs, such as GPT-4 Turbo 610, GPT-4 620, and GPT-3.5 Turbo 630. As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens” [Kuperman ¶ 43]. that define relationships between resource utilization and token processing according to the different model-specific tokenization schemes. “As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens. Therefore, monitoring token counts can also help identify strategies to restructure certain types of prompts to reduce the token count” [Kuperman ¶ 43]. “Further, large token counts (prompt and/or response) may consume more computational resources and increase inference time, especially in online or real-time applications. As such, monitoring prompt or response token counts can help manage resource usage effectively” [Kuperman ¶ 40]. “The AI automation system 110 described herein collects and analyzes data associated with prompts to, and responses from, Large Language Models (LLMs), including associated metadata. The AI automation system 110 uses this collected data to train machine learning models. These machine-learning models are trained to intelligently select and route prompts from applications to different LLMs to achieve desired performance metrics and reduce operational costs. Furthermore, the AI automation system 110 is capable of dynamically adjusting computational resources, such as GPU utilization, to further enhance operational efficiency” [Kuperman ¶ 75]. Kuperman fails to teach the token-based job metrics including a target model instance, and determine the model-agnostic estimated utilization for the incoming customer-requested LLM processing task based on stored model-specific benchmark metrics. However, Jain teaches: the token-based job metrics including a target model instance, “In some implementations, the user provides an indication of a desired LLM to be used to generate the resulting output, such as through the specification of a natural language generation (NLG) engine or architecture” [Jain Col. 3 Lines 10-14]. and determine the model-agnostic estimated utilization for the incoming customer-requested LLM processing task “Based on the selected LLM, the data generation platform can determine a set of performance metrics (and/or corresponding values) associated with processing the requested prompt via the selected LLM” [Jain Col. 3 Lines 17-20]. “To illustrate, the performance engine 118 can estimate performance metric values associated with processing a given prompt with a selected LLM (e.g., an estimated cost or memory usage). By doing so, the performance engine 118 can determine whether to allow access to a given LLM by a user, based on the user's requested output and the associated estimated system effects” [Jain Col. 7 Lines 28-35]. With regard to claim 3, Kuperman in view of Jain in view of Hu teaches the model-as-a-service platform of claim 1, as referenced above. Kuperman further teaches wherein the model-agnostic unit type is a logical unit of GPU capacity “In some embodiments, the AI automation system 110 is further configured to determine a GPU utilization rate at a private Kubernetes that host an open source LLM” [Kuperman ¶ 70]. Kuperman fails to teach wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization. However, Jain teaches wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request” [Jain Col. 15 Lines 39-49]. “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. With regard to claim 8, Kuperman in view of Jain in view of Hu teaches The model-as-a-service platform of claim 1, as referenced above. Kuperman fails to teach wherein the throttling service grants the incoming customer-requested LLM processing task in response to determining that a sum of the model-agnostic estimated utilization and a determined current utilization associated with a source of the incoming customer-requested LLM processing task is less than the customer-allotted quota. However, Jain teaches wherein the throttling service grants the incoming customer-requested LLM processing task in response to determining that a sum of the model-agnostic estimated utilization and a determined current utilization associated with a source of the incoming customer-requested LLM processing task is less than the customer-allotted quota. “Based on determining that the cost is greater than a threshold cost (e.g., a remaining budget within the user's allotment), the data generation platform 102 can determine not to provide the prompt to the LLM. Additionally or alternatively, the data generation platform 102 can determine that the cost is less than the threshold cost and, in response to this determination, proceed to provide the prompt to the LLM” [Jain Col. 21-22 Lines 62-67, 1-2]. With regard to claim 9, Kuperman in view of Jain in view of Hu teaches the model-as-a-service platform of claim 1, as referenced above. Kuperman further teaches wherein the different LLMs include one or more multimodal LLMs that tokenize input strings representing image, audio, or video data. “Some LLMs may support multimodal inputs, such as image input, audio input, or video input. When multimodal inputs are received, other tokenization methods may be applied” [Kuperman ¶ 42]. Claim(s) 4-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kuperman et al. US 2025/0292023 A1 (hereafter Kuperman) in view of Jain et al. US 12,147,513 B1 (hereafter Jain) in view of Hu et al. US 2024/0086249 A1 (hereafter Hu) in view of Jayatunga US 2025/0259045 A1 (hereafter Jayatunga). With regard to claim 4, Kuperman in view of Jain in view of Hu teaches the model-as-a-service platform of claim 1, as referenced above. Kuperman fails to teach wherein the model-agnostic unit type is a unit of token throughput representing a quantity of tokens that varies based on identify of a target LLM for the incoming customer-requested LLM processing task. However, Jain teaches wherein the model-agnostic unit type is a unit of token throughput representing a quantity of tokens that varies based on identify of a target LLM for the incoming customer-requested LLM processing task “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. Kuperman in view of Jain in view of Hu fails to explicitly teach a quantity of tokens that varies based on identify of a target LLM and GPU architecture supporting the target LLM. However, Jayatunga teaches a quantity of tokens that varies based on identify of a target LLM and GPU architecture supporting the target LLM. “For example, the number of tokens per unit of time may be measured. If the latency falls within a particular range (e.g. the average tokens per second measured over a rolling time window drops below a certain threshold), then a switch may be made to a second LLM. The second LLM may be a different LLM from the first LLM, or the first LLM and the second LLM may be the same model, e.g. different hardware instances of the same model accessed via different endpoints” [Jayatunga ¶ 5]. “The two models might actually be the same model, e.g. the same model parameters and code, just two different instances. For example, one may be a replica of the other, each operating in parallel, e.g. on two different computers (e.g. servers) and/or on two different specialized processing circuits (e.g. two different graphical processing units (GPUs)). For example, both the first generative model 532 and second generative model 534 might both be GPT-4™, just different hardware instances, e.g. separately executing in parallel on different hardware and/or on different computers and/or on different specialized processing circuits (e.g. different GPUs) and/or at two different network locations … In another variation, the first generative model 532 and the second generative model 534 may be different models, e.g. different model architectures. For example, the first generative model 532 may be GPT and the second generative model 534 may be BERT” [Jayatunga ¶ 64]. Jayatunga is considered to be analogous to the claimed invention because it is in the same field of resource capping. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman in view of Jain in view of Hu to incorporate the teachings of Jayatunga and include: a quantity of tokens that varies based on identify of a target LLM and GPU architecture supporting the target LLM. Doing so would allow for the measurement a use of latency values in the scheduling process. “To address this issue, the latency of the symbols (e.g. tokens or words or Unicode characters, etc.) received from the first LLM may be measured” [Jayatunga ¶ 5]. With regard to claim 5, Kuperman in view of Jain in view of Hu teaches the model-as-a-service platform of claim 4, as referenced above. Kuperman further teaches deployed within the GPU architecture. “Additionally, in some embodiments, the AI automation system is further configured to time-slice GPUs and/or provide multi-instance GPU (MIG) device support. When an LLM model is deployed on a private Kubernetes (K8s) cluster, GPUs are utilized to their full extent” [Kuperman ¶ 24]. Kuperman fails to teach wherein the model-agnostic unit type is defined to equal a fraction of an observed max utilization of the target LLM. However, Jain teaches wherein the model-agnostic unit type is defined to equal a fraction of an observed max utilization of the target LLM “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. Claim(s) 6-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kuperman et al. US 2025/0292023 A1 (hereafter Kuperman) in view of Jain et al. US 12,147,513 B1 (hereafter Jain) in view of Hu et al. US 2024/0086249 A1 (hereafter Hu) in view of Maschmeyer US 2024/0311192 A1 (hereafter Maschmeyer) in view of Nakura US 2022/0131610 A1 (hereafter Nakura). With regard to claim 6, Kuperman in view of Jain in view of Hu teaches the model-as-a-service platform of claim 2, as referenced above. Kuperman further teaches wherein the metric standardizer determines the model-agnostic estimated utilization for the incoming customer-requested LLM processing task “The pricing models for LLM services typically depend on the level of usage, which can be quantified in various ways, such as the number of API requests, the volume of text processed (measured in tokens or characters), or the amount of compute time utilized” [Kuperman ¶ 5]. Kuperman fails to teach wherein the metric standardizer determines the model-agnostic estimated utilization for the incoming customer-requested LLM processing task. However, Jain teaches wherein the metric standardizer determines the model-agnostic estimated utilization for the incoming customer-requested LLM processing task “To illustrate, the performance engine 118 can estimate performance metric values associated with processing a given prompt with a selected LLM (e.g., an estimated cost or memory usage). By doing so, the performance engine 118 can determine whether to allow access to a given LLM by a user, based on the user's requested output and the associated estimated system effects” [Jain Col. 7 Lines 28-35]. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request.” [Jain Col. 15 Lines 40-49]. Kuperman in view of Jain in view of Hu fails to teach a max utilization for a LLM and GPU architecture corresponding to the target model instance. However, Maschmeyer teaches a max utilization for a LLM and GPU architecture corresponding to the target model instance. “In examples, the total resource usage parameter is a representation of the total resource usage parameter with respect to a maximum resource capacity for the LLM” [Maschmeyer ¶ 150]. “Notably, a remote language model may employ a plurality of processors (e.g., hardware processors such as, for example, processors of cooperating computer systems). Indeed, processing of inputs by an LLM may be computationally expensive/may involve a large number of operations (e.g., many instructions may be executed/large data structures may be accessed from memory) and providing output in a required time frame (e.g., real-time or near real-time) may require the use of a plurality of processors/cooperating computing devices as discussed above” [Maschmeyer ¶ 61]. Maschmeyer is considered to be analogous to the claimed invention because it is in the same field of resource capping. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman in view of Jain in view of Hu to incorporate the teachings of Maschmeyer and include: a max utilization for a LLM and GPU architecture corresponding to the target model instance. Doing so would allow for the enforcement of resource usage limits. “The predicted total resource usage parameter may be communicated to a user device in a manner that enables a user of the user device to adjust the user input in real-time to comply with the resource usage limit” [Maschmeyer ¶ 69]. Kuperman in view of Jain in view of Hu in view of Maschmeyer fails to teach by identifying a relevant stored probability distribution modeling a max utilization. However, Nakura teaches by identifying a relevant stored probability distribution modeling a max utilization “In addition, while the OLT 1 calculates a maximum usage rate in each window and informs the controller 4 of the maximum usage rate as resource information in the present embodiment, the OLT 1 may generate distribution of probability of occurrence of a certain usage rate, and provide information of the probability distribution as resource information” [Nakura ¶ 72]. “In the controller 4, the resource managing unit 41 receives and holds resource information input from the resource information generating unit 16 of the OLT 1” [Nakura ¶ 44]. Nakura is considered to be analogous to the claimed invention because it is in the same field of resource capping. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman in view of Jain in view of Hu in view of Maschmeyer to incorporate the teachings of Nakura and include: by identifying a relevant stored probability distribution modeling a max utilization. Doing so would allow for the further efficiency and user cost savings. “This can reduce the burden of cost on users of a pay-per-use network service in which charges change depending on network use amounts, and improve the use efficiency of the whole network” [Nakura ¶ 73]. With regard to claim 7, Kuperman in view of Jain in view of Hu in view of Maschmeyer in view of Nakura teaches the model-as-a-service platform of claim 6, as referenced above. Kuperman further teaches while the LLM processes workloads with input and output prompt characteristics determined to be similar to those of the incoming customer-requested LLM processing task. “In some embodiments, the system is also configured to apply a similarity model to the prompt, identifying a set of historical requests similar to the current one” [Kuperman ¶ 7]. “The data store 410 may store collected historical prompts and other data associated with the historical prompts, such as (but not limited to) corresponding responses generated by LLMs, a token count associated with the prompt, a cost associated with the prompt, etc. The prompts analysis module 420 is configured to analyze historical prompts and associated data” [Kuperman, 38-39]. Kuperman fails to teach throughput of the LLM. However, Jain teaches throughput of the LLM “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. Kuperman in view of Jain in view of Hu in view of Maschmeyer fails to teach wherein the relevant stored probability distribution is generated by modeling. However, Nakura teaches wherein the relevant stored probability distribution is generated by modeling “In addition, while the OLT 1 calculates a maximum usage rate in each window and informs the controller 4 of the maximum usage rate as resource information in the present embodiment, the OLT 1 may generate distribution of probability of occurrence of a certain usage rate, and provide information of the probability distribution as resource information” [Nakura ¶ 72]. Claim(s) 10-12, and 16-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kuperman et al. US 2025/0292023 A1 (hereafter Kuperman) in view of Jain et al. US 12,147,513 B1 (hereafter Jain). With regard to claim 10, Kuperman teaches: A method of throttling (sending) customer endpoint requests to directed to instances of different large language models (LLMs) in a model-as-a-service platform, “The AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140 … The one or more LLM SaaS platform(s) 130 may include (but are not limited to) non-open source LLMs, such as GPT-3, GPT-4, which are provided through APIs of a SaaS platform. The one or more open source LLM platform(s) 140 may include (but are not limited to) open source LLMs deployed on private Kubernetes (K8s)” [Kuperman ¶ 26-27]. the method comprising: determining a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; “The pricing models for LLM services typically depend on the level of usage, which can be quantified in various ways, such as the number of API requests, the volume of text processed (measured in tokens or characters), or the amount of compute time utilized” [Kuperman ¶ 5]. the units of the customer-allotted quota being redeemable in exchange for compute tasks performed by the different LLMs. “LLMs are able to perform a variety of tasks that were traditionally considered challenging for computers. These tasks include generating human-like text, translating languages, summarizing long pieces of content, answering questions, and more” [Kuperman ¶ 3]. “Cost per million tokens is an example of a measure of cost-effectiveness when using Large Language Model (LLM) tooling in cloud environments, such as Software as a Service (SaaS) environments. Another example measure of cost-effectiveness, aside from SaaS environments, could be the cost per million API requests” [Kuperman ¶ 29]. Kuperman fails to teach A method of throttling customer endpoint requests to directed to instances of different large language models (LLMs) … the method comprising: determining a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; and limiting a number requests a user can concurrently submit to the instances of the different LLMs based on a customer-allotted quota of the units of the model-agnostic unit type, the units of the customer-allotted quota being redeemable in exchange for compute tasks performed by the different LLMs. However, Jain teaches: A method of throttling customer endpoint requests to directed to instances of different large language models (LLMs) “For example, the access control engine 114 determines to throttle the bandwidth associated with receiving outputs from the LLMs (e.g., by specifying a number of responses per unit time that are allowed to be transmitted to the given user). As such, the access control engine 114 can control the system-wide performance by limiting the assignment of system resources to particular users” [Jain Col. 12 Lines 30-37]. the method comprising: determining a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; “Following input validation, the performance engine 118 can evaluate the performance of LLMs to route the prompt to an appropriate LLM (e.g., large language model(s) 410)” [Jain Col. 10 Lines 14-16]. “In some implementations, the data generation platform evaluates the performance requirements associated with the prompt generate a recommendation for a suitable LLM for the received prompt (e.g., to improve the efficiency of system resource use)” [Jain Col. 4 Lines 14-19]. “As such, the data generation platform 102 can dynamically select LLMs for processing output generation requests on the basis of estimated system resources and associated limitations or requirements, thereby improving the efficiency, flexibility, and robustness of the associated development pipeline” [Jain Col. 16 Lines 29-34]. and limiting a number requests a user can concurrently submit to the instances of the different LLMs based on a customer-allotted quota of the units of the model-agnostic unit type, “The data generation platform 102 can determine that the estimated resource metric value (e.g., the estimated cost) does not satisfy the threshold metric value. For example, the data generation platform 102 determines that the estimated cost is greater than or equal to the threshold cost. Based on this determination, the performance evaluation 408 can determine to prevent provision of the associated prompt to one or more LLMs of the LLMs 410” [Jain Col. 16 Lines 16-23]. “Furthermore, the threshold metric can depend on the user (e.g., the user device 402a of FIG. 4) and/or the service 402b. For example, the threshold metric includes resource allotments that are specific to particular users of the data generation platform 102” [Jain Col. 16 Lines 4-8]. “For example, the access control engine 114 determines to throttle the bandwidth associated with receiving outputs from the LLMs (e.g., by specifying a number of responses per unit time that are allowed to be transmitted to the given user). As such, the access control engine 114 can control the system-wide performance by limiting the assignment of system resources to particular users” [Jain Col. 12 Lines 30-37]. the units of the customer-allotted quota being redeemable in exchange for compute tasks performed by the different LLMs. “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request.” [Jain Col. 15 Lines 40-49]. Jain is considered to be analogous to the claimed invention because it is in the same field of workload prediction. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman to incorporate the teachings of Jain and include: A method of throttling customer endpoint requests to directed to instances of different large language models (LLMs) … the method comprising: determining a model-agnostic estimated utilization for an incoming customer-requested LLM processing task as a quantity of units of a model-agnostic unit type; and limiting a number requests a user can concurrently submit to the instances of the different LLMs based on a customer-allotted quota of the units of the model-agnostic unit type, the units of the customer-allotted quota being redeemable in exchange for compute tasks performed by the different LLMs. Doing so would allow for improved flexibility and control over the system resources. “By doing so, the data generation platform 102 provides improved flexibility and/or control over the use of system resources (including memory, computational, and/or financial resources), enabling optimization of the associated development pipeline” [Jain Col. 22 Lines 2-6]. With regard to claim 11, Kuperman in view of Jain teaches the method of claim 10, as referenced above. Kuperman further teaches: further comprising: receiving, as input, token-based job metrics for the incoming customer-requested LLM processing task, “Monitoring prompt token counts can help ensure that the input provided to the LLM is within the acceptable limits of the LLM's input capacity and does not exceed any constraints set by an application or environment in which the LLM is being used” [Kuperman ¶ 39]. “The data store 410 may store collected historical prompts and other data associated with the historical prompts, such as (but not limited to) corresponding responses generated by LLMs, a token count associated with the prompt, a cost associated with the prompt, etc” [Kuperman ¶ 38]. a request number of input tokens, and a requested number of output tokens; “It then determines whether the request should be reconstructed to reduce a token count, which includes input and/or output token counts” [Kuperman ¶ 7]. and determining the model-agnostic estimated utilization for the incoming customer-requested LLM processing task based on stored model-specific benchmark metrics “Notably, the costs of LLM services are often directly or indirectly related to token counts. FIG. 6 illustrates an example cost per thousand tokens for a plurality of LLMs, such as GPT-4 Turbo 610, GPT-4 620, and GPT-3.5 Turbo 630. As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens” [Kuperman ¶ 43]. that define relationships between resource utilization and token processing according to the different model-specific tokenization schemes utilized by the instances of the different LLMs. “As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens. Therefore, monitoring token counts can also help identify strategies to restructure certain types of prompts to reduce the token count” [Kuperman ¶ 43]. “Further, large token counts (prompt and/or response) may consume more computational resources and increase inference time, especially in online or real-time applications. As such, monitoring prompt or response token counts can help manage resource usage effectively” [Kuperman ¶ 40]. “The AI automation system 110 described herein collects and analyzes data associated with prompts to, and responses from, Large Language Models (LLMs), including associated metadata. The AI automation system 110 uses this collected data to train machine learning models. These machine-learning models are trained to intelligently select and route prompts from applications to different LLMs to achieve desired performance metrics and reduce operational costs. Furthermore, the AI automation system 110 is capable of dynamically adjusting computational resources, such as GPU utilization, to further enhance operational efficiency” [Kuperman ¶ 75]. Kuperman fails to teach the token-based job metrics including a target model instance, and determining the model-agnostic estimated utilization for the incoming customer-requested LLM processing task based on stored model-specific benchmark metrics. However, Jain teaches: the token-based job metrics including a target model instance, “In some implementations, the user provides an indication of a desired LLM to be used to generate the resulting output, such as through the specification of a natural language generation (NLG) engine or architecture” [Jain Col. 3 Lines 10-14]. and determining the model-agnostic estimated utilization for the incoming customer-requested LLM processing task “Based on the selected LLM, the data generation platform can determine a set of performance metrics (and/or corresponding values) associated with processing the requested prompt via the selected LLM” [Jain Col. 3 Lines 17-20]. “To illustrate, the performance engine 118 can estimate performance metric values associated with processing a given prompt with a selected LLM (e.g., an estimated cost or memory usage). By doing so, the performance engine 118 can determine whether to allow access to a given LLM by a user, based on the user's requested output and the associated estimated system effects” [Jain Col. 7 Lines 28-35]. With regard to claim 12, Kuperman in view of Jain in view of Hu teaches the method of claim 10, as referenced above. Kuperman further teaches wherein the model-agnostic unit type is a logical unit of GPU capacity “In some embodiments, the AI automation system 110 is further configured to determine a GPU utilization rate at a private Kubernetes that host an open source LLM” [Kuperman ¶ 70]. Kuperman fails to teach wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization. However, Jain teaches wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request” [Jain Col. 15 Lines 39-49]. “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. With regard to claim 16, Kuperman in view of Jain teaches The method of claim 10, as referenced above. Kuperman fails to teach further comprising: granting the incoming customer-requested LLM processing task in response to determining that a sum of the model-agnostic estimated utilization and a determined current utilization associated with a source of the incoming customer-requested LLM processing task is less than the customer-allotted quota. However, Jain teaches further comprising: granting the incoming customer-requested LLM processing task in response to determining that a sum of the model-agnostic estimated utilization and a determined current utilization associated with a source of the incoming customer-requested LLM processing task is less than the customer-allotted quota. “Based on determining that the cost is greater than a threshold cost (e.g., a remaining budget within the user's allotment), the data generation platform 102 can determine not to provide the prompt to the LLM. Additionally or alternatively, the data generation platform 102 can determine that the cost is less than the threshold cost and, in response to this determination, proceed to provide the prompt to the LLM” [Jain Col. 21-22 Lines 62-67, 1-2]. With regard to claim 17, Kuperman teaches: One or more computer-readable storage media encoding computer-executable instructions for executing a computer process for throttling (sending) customer requests to access instances of different large language models (LLMs) deployed within a model-as-a-service platform, the computer process comprising: “In the embodiment shown in FIG. 11, the storage device 1108 is a non-transitory computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 1106 holds instructions and data used by the processor 1102” [Kuperman ¶ 73]. “The AI automation system 110 helps entity systems 120 intelligently select and route prompts to different LLM platforms 130, 140 … The one or more LLM SaaS platform(s) 130 may include (but are not limited to) non-open source LLMs, such as GPT-3, GPT-4, which are provided through APIs of a SaaS platform. The one or more open source LLM platform(s) 140 may include (but are not limited to) open source LLMs deployed on private Kubernetes (K8s)” [Kuperman ¶ 26-27]. storing model-specific benchmark metrics that define relationships between resource utilization and token processing according to different model-specific tokenization schemes utilized by different large language models; “As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens. Therefore, monitoring token counts can also help identify strategies to restructure certain types of prompts to reduce the token count” [Kuperman ¶ 43]. “Further, large token counts (prompt and/or response) may consume more computational resources and increase inference time, especially in online or real-time applications. As such, monitoring prompt or response token counts can help manage resource usage effectively” [Kuperman ¶ 40]. “The AI automation system 110 described herein collects and analyzes data associated with prompts to, and responses from, Large Language Models (LLMs), including associated metadata. The AI automation system 110 uses this collected data to train machine learning models. These machine-learning models are trained to intelligently select and route prompts from applications to different LLMs to achieve desired performance metrics and reduce operational costs. Furthermore, the AI automation system 110 is capable of dynamically adjusting computational resources, such as GPU utilization, to further enhance operational efficiency” [Kuperman ¶ 75]. “Notably, the costs of LLM services are often directly or indirectly related to token counts. FIG. 6 illustrates an example cost per thousand tokens for a plurality of LLMs, such as GPT-4 Turbo 610, GPT-4 620, and GPT-3.5 Turbo 630. As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens” [Kuperman ¶ 43]. receiving, from a throttling service, token-based job metrics “Monitoring prompt token counts can help ensure that the input provided to the LLM is within the acceptable limits of the LLM's input capacity and does not exceed any constraints set by an application or environment in which the LLM is being used” [Kuperman ¶ 39]. “The data store 410 may store collected historical prompts and other data associated with the historical prompts, such as (but not limited to) corresponding responses generated by LLMs, a token count associated with the prompt, a cost associated with the prompt, etc” [Kuperman ¶ 38]. a requested number of input tokens, and a requested number of output tokens associated with an LLM processing task; “It then determines whether the request should be reconstructed to reduce a token count, which includes input and/or output token counts” [Kuperman ¶ 7]. determining, based on identity of the target LLM instance, an applicable subset of the model-specific benchmark metrics; “The token count is a number of tokens in a prompt, which include individual units of text, such as words, subwords, or special tokens. The token count corresponds to complexity and specificity of the input given to the LLM. Also, a token count for a prompt depends on how the prompt is tokenized by a specific LLM being used. If the LLM tokenizes words and punctuation marks, the token count might be one number. However, if the LLM uses subword tokenization or includes special tokens, the token count would be a different number” [Kuperman ¶ 39]. “Notably, the costs of LLM services are often directly or indirectly related to token counts. FIG. 6 illustrates an example cost per thousand tokens for a plurality of LLMs, such as GPT-4 Turbo 610, GPT-4 620, and GPT-3.5 Turbo 630. As illustrated, the cost per thousand tokens are based on both a number of input tokens and output tokens, and different LLMs charge different amounts per thousand input/output tokens” [Kuperman ¶ 43]. using the applicable subset of the model-specific benchmark metrics to derive a model-agnostic estimated utilization associated with the LLM processing task, “The pricing models for LLM services typically depend on the level of usage, which can be quantified in various ways, such as the number of API requests, the volume of text processed (measured in tokens or characters), or the amount of compute time utilized” [Kuperman ¶ 5]. “LLMs are able to perform a variety of tasks that were traditionally considered challenging for computers. These tasks include generating human-like text, translating languages, summarizing long pieces of content, answering questions, and more” [Kuperman ¶ 3]. “Cost per million tokens is an example of a measure of cost-effectiveness when using Large Language Model (LLM) tooling in cloud environments, such as Software as a Service (SaaS) environments. Another example measure of cost-effectiveness, aside from SaaS environments, could be the cost per million API requests” [Kuperman ¶ 29]. Kuperman fails to teach throttling customer requests to access instances of different large language models (LLMs) … identifying a target LLM instance, derive a model-agnostic estimated utilization associated with the LLM processing task, the model-agnostic estimated utilization being defined as a quantity of units of a model-agnostic unit type; and providing the model-agnostic estimated utilization to the throttling service; and determining, by the throttling service, whether to grant or deny the LLM processing task based on the model-agnostic estimated utilization for the LLM processing task. However, Jain teaches: throttling customer requests to access instances of different large language models (LLMs) “For example, the access control engine 114 determines to throttle the bandwidth associated with receiving outputs from the LLMs (e.g., by specifying a number of responses per unit time that are allowed to be transmitted to the given user). As such, the access control engine 114 can control the system-wide performance by limiting the assignment of system resources to particular users” [Jain Col. 12 Lines 30-37]. identifying a target LLM instance, “In some implementations, the user provides an indication of a desired LLM to be used to generate the resulting output, such as through the specification of a natural language generation (NLG) engine or architecture” [Jain Col. 3 Lines 10-14]. derive a model-agnostic estimated utilization associated with the LLM processing task, the model-agnostic estimated utilization being defined as a quantity of units of a model-agnostic unit type; “Following input validation, the performance engine 118 can evaluate the performance of LLMs to route the prompt to an appropriate LLM (e.g., large language model(s) 410)” [Jain Col. 10 Lines 14-16]. “In some implementations, the data generation platform evaluates the performance requirements associated with the prompt generate a recommendation for a suitable LLM for the received prompt (e.g., to improve the efficiency of system resource use)” [Jain Col. 4 Lines 14-19]. “As such, the data generation platform 102 can dynamically select LLMs for processing output generation requests on the basis of estimated system resources and associated limitations or requirements, thereby improving the efficiency, flexibility, and robustness of the associated development pipeline” [Jain Col. 16 Lines 29-34]. and providing the model-agnostic estimated utilization to the throttling service; “The data generation platform 102 can determine that the estimated resource metric value (e.g., the estimated cost) does not satisfy the threshold metric value. For example, the data generation platform 102 determines that the estimated cost is greater than or equal to the threshold cost. Based on this determination, the performance evaluation 408 can determine to prevent provision of the associated prompt to one or more LLMs of the LLMs 410” [Jain Col. 16 Lines 16-23]. “Furthermore, the threshold metric can depend on the user (e.g., the user device 402a of FIG. 4) and/or the service 402b. For example, the threshold metric includes resource allotments that are specific to particular users of the data generation platform 102” [Jain Col. 16 Lines 4-8]. “For example, the access control engine 114 determines to throttle the bandwidth associated with receiving outputs from the LLMs (e.g., by specifying a number of responses per unit time that are allowed to be transmitted to the given user). As such, the access control engine 114 can control the system-wide performance by limiting the assignment of system resources to particular users” [Jain Col. 12 Lines 30-37]. and determining, by the throttling service, whether to grant or deny the LLM processing task based on the model-agnostic estimated utilization for the LLM processing task. “Based on determining that the cost is greater than a threshold cost (e.g., a remaining budget within the user's allotment), the data generation platform 102 can determine not to provide the prompt to the LLM. Additionally or alternatively, the data generation platform 102 can determine that the cost is less than the threshold cost and, in response to this determination, proceed to provide the prompt to the LLM” [Jain Col. 21-22 Lines 62-67, 1-2]. Jain is considered to be analogous to the claimed invention because it is in the same field of workload prediction. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman to incorporate the teachings of Jain and include: throttling customer requests to access instances of different large language models (LLMs) … identifying a target LLM instance, derive a model-agnostic estimated utilization associated with the LLM processing task, the model-agnostic estimated utilization being defined as a quantity of units of a model-agnostic unit type; and providing the model-agnostic estimated utilization to the throttling service; and determining, by the throttling service, whether to grant or deny the LLM processing task based on the model-agnostic estimated utilization for the LLM processing task. Doing so would allow for improved flexibility and control over the system resources. “By doing so, the data generation platform 102 provides improved flexibility and/or control over the use of system resources (including memory, computational, and/or financial resources), enabling optimization of the associated development pipeline” [Jain Col. 22 Lines 2-6]. With regard to claim 18, Kuperman in view of Jain teaches the one or more computer-readable storage media of claim 17, as referenced above. Kuperman fails to teach wherein determining whether to grant or deny the LLM processing task is based on a customer-allotted quota that limits a number of requests a given client compute platform can submit to the model-as-a-service platform, the customer-allotted quota defined as a quantity units of the model-agnostic unit type. However, Jain teaches wherein determining whether to grant or deny the LLM processing task is based on a customer-allotted quota that limits a number of requests a given client compute platform can submit to the model-as-a-service platform, the customer-allotted quota defined as a quantity units of the model-agnostic unit type. “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. “Based on determining that the cost is greater than a threshold cost (e.g., a remaining budget within the user's allotment), the data generation platform 102 can determine not to provide the prompt to the LLM. Additionally or alternatively, the data generation platform 102 can determine that the cost is less than the threshold cost and, in response to this determination, proceed to provide the prompt to the LLM” [Jain Col. 21-22 Lines 62-67, 1-2]. With regard to claim 19, Kuperman in view of Jain teaches the one or more computer-readable storage media of claim 17, as referenced above. Kuperman further teaches wherein the model-agnostic unit type is a logical unit of GPU capacity “In some embodiments, the AI automation system 110 is further configured to determine a GPU utilization rate at a private Kubernetes that host an open source LLM” [Kuperman ¶ 70]. Kuperman fails to teach wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization. However, Jain teaches wherein the model-agnostic unit type is a logical unit of GPU capacity that facilitates direct comparison of memory utilization across the different LLMs without unitary conversion or normalization. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request” [Jain Col. 15 Lines 39-49]. “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. With regard to claim 20, Kuperman in view of Jain teaches the one or more computer-readable storage media of claim 17, as referenced above. Kuperman further teaches wherein the computer process further comprises: querying, by the throttling service, a platform-level database to determine “In some embodiments, the schema may also include an SQL statement to obtain a dataset or constraints of a database. For example, in some embodiments, each historical prompt in the set includes a dataset that can be queried from a database” [Kuperman ¶ 67]. Kuperman fails to explicitly teach a current utilization of a client compute platform associated with the LLM processing task; determining, by the throttling service, whether to grant or deny the LLM processing task based on a sum of the model-agnostic estimated utilization and the current utilization of the client compute platform. However, Jain teaches: a current utilization of a client compute platform associated with the LLM processing task; “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. determining, by the throttling service, whether to grant or deny the LLM processing task based on a sum of the model-agnostic estimated utilization and the current utilization of the client compute platform. “Based on determining that the cost is greater than a threshold cost (e.g., a remaining budget within the user's allotment), the data generation platform 102 can determine not to provide the prompt to the LLM. Additionally or alternatively, the data generation platform 102 can determine that the cost is less than the threshold cost and, in response to this determination, proceed to provide the prompt to the LLM” [Jain Col. 21-22 Lines 62-67, 1-2]. Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kuperman et al. US 2025/0292023 A1 (hereafter Kuperman) in view of Jain et al. US 12,147,513 B1 (hereafter Jain) in view of Jayatunga US 2025/0259045 A1 (hereafter Jayatunga). With regard to claim 13, Kuperman in view of Jain in view of Hu teaches the method of claim 10, as referenced above. Kuperman fails to teach wherein the model-agnostic unit type is a unit of token throughput representing a quantity of tokens that varies based on identify of a target LLM for the incoming customer-requested LLM processing task. However, Jain teaches wherein the model-agnostic unit type is a unit of token throughput representing a quantity of tokens that varies based on identify of a target LLM for the incoming customer-requested LLM processing task “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. Kuperman in view of Jain in view of Hu fails to explicitly teach a quantity of tokens that varies based on identify of a target LLM and GPU architecture supporting the target LLM. However, Jayatunga teaches a quantity of tokens that varies based on identify of a target LLM and GPU architecture supporting the target LLM. “For example, the number of tokens per unit of time may be measured. If the latency falls within a particular range (e.g. the average tokens per second measured over a rolling time window drops below a certain threshold), then a switch may be made to a second LLM. The second LLM may be a different LLM from the first LLM, or the first LLM and the second LLM may be the same model, e.g. different hardware instances of the same model accessed via different endpoints” [Jayatunga ¶ 5]. “The two models might actually be the same model, e.g. the same model parameters and code, just two different instances. For example, one may be a replica of the other, each operating in parallel, e.g. on two different computers (e.g. servers) and/or on two different specialized processing circuits (e.g. two different graphical processing units (GPUs)). For example, both the first generative model 532 and second generative model 534 might both be GPT-4™, just different hardware instances, e.g. separately executing in parallel on different hardware and/or on different computers and/or on different specialized processing circuits (e.g. different GPUs) and/or at two different network locations … In another variation, the first generative model 532 and the second generative model 534 may be different models, e.g. different model architectures. For example, the first generative model 532 may be GPT and the second generative model 534 may be BERT” [Jayatunga ¶ 64]. Jayatunga is considered to be analogous to the claimed invention because it is in the same field of resource capping. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman in view of Jain to incorporate the teachings of Jayatunga and include: a quantity of tokens that varies based on identify of a target LLM and GPU architecture supporting the target LLM. Doing so would allow for the measurement a use of latency values in the scheduling process. “To address this issue, the latency of the symbols (e.g. tokens or words or Unicode characters, etc.) received from the first LLM may be measured” [Jayatunga ¶ 5]. Claim(s) 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kuperman et al. US 2025/0292023 A1 (hereafter Kuperman) in view of Jain et al. US 12,147,513 B1 (hereafter Jain) in view of Maschmeyer US 2024/0311192 A1 (hereafter Maschmeyer). With regard to claim 14, Kuperman in view of Jain in view of Hu teaches the method of claim 11, as referenced above. Kuperman further teaches determining the model-agnostic estimated utilization for the incoming customer-requested LLM processing task “The pricing models for LLM services typically depend on the level of usage, which can be quantified in various ways, such as the number of API requests, the volume of text processed (measured in tokens or characters), or the amount of compute time utilized” [Kuperman ¶ 5]. Kuperman fails to teach determining the model-agnostic estimated utilization for the incoming customer-requested LLM processing task by identifying a relevant stored dataset modeling a max utilization for a LLM. However, Jain teaches determining the model-agnostic estimated utilization for the incoming customer-requested LLM processing task “To illustrate, the performance engine 118 can estimate performance metric values associated with processing a given prompt with a selected LLM (e.g., an estimated cost or memory usage). By doing so, the performance engine 118 can determine whether to allow access to a given LLM by a user, based on the user's requested output and the associated estimated system effects” [Jain Col. 7 Lines 28-35]. “Additionally or alternatively, the performance metric includes an indication of computational resources associated with processing the request (e.g., a bandwidth required to transmit the prompt to the LLM and/or execute the associated API call). The indication of computational resources can include memory requirements (e.g., a storage size associated with the prompt or the estimated output storage size) for processing the prompt. Additionally or alternatively, the estimated resource requirement includes an estimate of CPU processing speeds and/or time associated with the processing the request.” [Jain Col. 15 Lines 40-49]. by identifying a relevant stored dataset modeling a max utilization for a LLM “A performance metric can include an indication of an estimated resource use, such as a monetary cost (e.g., cost metric) associated with an API call to the requested LLM. For example, referring to FIG. 5, the performance engine 118 determines an estimated resource use 510, such as a monetary cost, for processing the prompt with the selected LLM” [Jain Col. 15 Lines 30-36]. “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. Kuperman in view of Jain in view of Hu fails to teach a max utilization for a LLM and GPU architecture corresponding to the target model instance. However, Maschmeyer teaches a max utilization for a LLM and GPU architecture corresponding to the target model instance. “In examples, the total resource usage parameter is a representation of the total resource usage parameter with respect to a maximum resource capacity for the LLM” [Maschmeyer ¶ 150]. “Notably, a remote language model may employ a plurality of processors (e.g., hardware processors such as, for example, processors of cooperating computer systems). Indeed, processing of inputs by an LLM may be computationally expensive/may involve a large number of operations (e.g., many instructions may be executed/large data structures may be accessed from memory) and providing output in a required time frame (e.g., real-time or near real-time) may require the use of a plurality of processors/cooperating computing devices as discussed above” [Maschmeyer ¶ 61]. Maschmeyer is considered to be analogous to the claimed invention because it is in the same field of resource capping. Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kuperman in view of Jain to incorporate the teachings of Maschmeyer and include: a max utilization for a LLM and GPU architecture corresponding to the target model instance. Doing so would allow for the enforcement of resource usage limits. “The predicted total resource usage parameter may be communicated to a user device in a manner that enables a user of the user device to adjust the user input in real-time to comply with the resource usage limit” [Maschmeyer ¶ 69]. With regard to claim 15, Kuperman in view of Jain in view of Maschmeyer teaches the method of claim 14, as referenced above. Kuperman further teaches the LLM deployed in the GPU architecture when the LLM is processing a set of workloads with input prompt characteristics and output prompt characteristics identified as similar to those of the incoming customer-requested LLM processing task. “In some embodiments, the system is also configured to apply a similarity model to the prompt, identifying a set of historical requests similar to the current one” [Kuperman ¶ 7]. “The data store 410 may store collected historical prompts and other data associated with the historical prompts, such as (but not limited to) corresponding responses generated by LLMs, a token count associated with the prompt, a cost associated with the prompt, etc. The prompts analysis module 420 is configured to analyze historical prompts and associated data” [Kuperman, 38-39]. Kuperman fails to teach wherein the relevant stored dataset describes token throughput observed for the LLM. However, Jain teaches wherein the relevant stored dataset describes token throughput observed for the LLM “In some implementations, the GUI 500 (e.g., as shown in FIG. 5) includes an indication of the user's allotment and a running indication of resource usage (e.g., as a percentage of the user's allotment)” [Jain Col. 16 Lines 8-11]. “The performance metric can include a maximum number of requests per minute (e.g., a throughput). In some implementations, the performance metric includes a number of prompt tokens per request (e.g., a number of words, phrases, or other natural language or numerical units within the request)” [Jain Col. 15 Lines 53-56]. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARI F RIGGINS whose telephone number is (571)272-2772. The examiner can normally be reached Monday-Friday 7:00AM-4:30PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at (571) 272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /A.F.R./Examiner, Art Unit 2197 /JOANNE G MACASIANO/Examiner, Art Unit 2197
Read full office action

Prosecution Timeline

May 01, 2024
Application Filed
Jul 23, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675316
USING MULTIPLE QUOTA TREES IN RESOURCE SCHEDULING
4y 6m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
50%
Grant Probability
99%
With Interview (+100.0%)
3y 7m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month