Prosecution Insights
Last updated: August 18, 2026
Application No. 18/749,630

Execution Methods of a Machine Learning Model

Final Rejection §101
Filed
Jun 21, 2024
Priority
Aug 02, 2023 — provisional 63/517,122
Examiner
PATEL, SHREYANS A
Art Unit
2659
Tech Center
2600 — Communications
Assignee
MediaTek Inc.
OA Round
2 (Final)
89%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
364 granted / 411 resolved
+26.6% vs TC avg
Moderate +8% lift
Without
With
+8.5%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
34 currently pending
Career history
457
Total Applications
across all art units

Statute-Specific Performance

§101
26.4%
-13.6% vs TC avg
§103
41.1%
+1.1% vs TC avg
§102
21.4%
-18.6% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 411 resolved cases

Office Action

§101
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments/Amendments Applicant's amendments with respect to Claim Objections of claims 2, 4, 12 and 14 have been considered and found persuasive, and the rejection has been withdrawn. Applicant's arguments with respect to 35 U.S.C. 112 rejection of claims 1 and 11 have been considered and found persuasive, and the rejection has been withdrawn. Applicant's arguments with respect to 35 U.S.C. 101 Abstract Idea in regards to claims 1 and 11 have been considered, however are not found to be persuasive due to the following reasons. Examiner respectfully disagrees with Applicant’s arguments because Applicant’s amendments do not overcome the 101 rejection because claims 1 and 11 remine directed to the abstract idea of mathematical calculations and data processing performed by a machine learning model. The claims recite generating outputs, creating and using a BoS cache, performing model quantization, selecting a next token, and generating subsequent outputs based on pervious outputs or input content. These are mathematical operations used to process data and execute a predictive algorithm. While Applicant argues that the Office Action overgeneralizes the claims, considering the claims as a whole does not change its focus, which remains the execution of mathematical computations for token prediction and cache management rather than a technological improvement to a computer or other technology. Under Step 2A, prong two, Applicant’s argument that the claimed BoS cache and quantization technique improve machine learning inference is not persuasive because the claimed improvement is directed only to the operation of the mathematical model itself, not to an improvement in computer technology. The additional limitations – such as generating the BoS cache before inference, excluding the BoS token from quantization, using floating-point data for the BoS cache, and using the BoS cache to indicate the beginning of the token sequence – describe how the machine learning algorithm processes data more efficiently or accurately. They do not recite a specific improvement to processor architecture, memory hardware, cache hardware, or another technological component. Instead, the claims merely use generic processors to execute mathematical operations and data manipulation, which does not integrate the abstract idea into a practical application. Under Step 2B, Applicant’s claims that the claimed execution technique improves inference accuracy and avoids activation outliers does not provide an inventive concept. The claimed features merely refine the mathematical rules used by the machine learning model by choosing how to process a BoS token, how to quantize model data, and how to reuse cached information during inference. These are improvements to the algorithm itself rather than to computer functionality. The claims do not recite any unconventional hardware, specialized processor, or other technological mechanism that transforms the abstract idea into patent eligible subject matter. When considered individually and as an ordered combination, the additional elements amount to only implementing an abstract mathematical algorithm on generic computer components. Therefore, the claims stand rejected. Applicant's amendments with respect to 35 U.S.C. 103 rejection of claims 1 and 11 have been considered and found persuasive, and the rejection has been withdrawn. See detailed reason for allowance below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 4-11 and 14-20 are rejected under 35 U.S.C. 101. Claims 1 and 11 are directed to an abstract idea of reciting generating outputs, computing and maintaining a “BoS cache,” applying quantization, and selecting a next token based on computed outputs or input content. These elements recite mental operations and mathematical concepts (data transformations, algorithmic selection, quantization math, token prediction), which are judicial exceptions to patentable subject matter. See MPEP §2106.02(II) (mental processes) and MPEP §2106.04(a) (mathematical concepts); Alice Corp. v. CLS Bank Int’l, 573 U.S. 208 (2014); Electric Power Group v. Alstom, 830 F.3d 1350 (Fed. Cir. 2016). In particular, because the claim language does not require any specific machine implemented data structures or hardware and recites high level steps of computing and selecting tokens (operations that can be expressed as mental steps or pen and paper calculations), the claim reads on performance in the human mind or by generic computation steps. Thus the claim recites a judicially recognized abstract idea comprised of mental processes and mathematical concepts. The claims merely applies the abstract idea using generic computer components and conventional machine-learning techniques, such as “a machine learning model,” “quantized model,” “cache,” and “inference,” without specifying any particular model architecture, data structure, memory organization, or hardware-level optimization. The use of quantization and caching is described at a high level of abstraction and reflects well-understood practices in machine learning for efficiency, rather than a specific improvement to computer functionality itself. The claim does not recite how quantization is technically performed, how the cache is structured in memory, or how these steps improve processing at a hardware or system level. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device. Dependent claims 4-10 and 14-20 further recite an abstract idea and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known in machine learning models. Claims 4 and 14: an abstract mathematical classification/quantization algorithm on data. Claims 5 and 15: an abstract data clustering rule based on numeric similarity. Claims 6 and 16: an abstract rule-based parameter selection for model execution. Claims 7 and 17: an abstract choice of numeric representation for performing the model’s calculations. Claims 8 and 18: an abstract model configuration choice without a specific technical improvement. Claims 9 and 19: an abstract mathematical number format conversion for model parameters. Claim 10: an abstract functional description of a prediction algorithm (next token depends on prior tokens). Claim 20: an abstract data sequence definition/formatting limitation. Allowable Subject Matter Claims 1, 4-11 and 14-20 are allowed if the Applicant can overcome the 101 Abstract Idea. The following is a statement of reasons for the indication of allowable subject matter: Yan et al. (US 2022/0318601) in view of Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with Single GPU”; July 3, 2023): Claims 1 and 11, Yan teaches an execution method of a machine learning model ([Abstract] attention mechanism 102), comprising: generating output and a begin of sentence (BoS) cache of a BoS token using the machine learning model ([Figs. 1-2] [0031-0032] [0038-0039] a token that designates a start of a sequence; upon predicating an output token, the decoder system adds that token to the end of the sentence that is fed as input information into the decoder system; the decoder system is tasked; the decoder system is tasked with responsibility of caching the head-specific key information and the head-specific value information during self-attention and the input sequence explicitly begins with, meaning the cache includes the BoS token’s state) during the inference, input the next token following the BoS token as a first input token ([0038] “Jack”; example input sequence shows the BoS token followed by the next token) wherein the next token is based on the output of the Bos token or based on an input content ([0038] “Jack” (given input token after BoS); output-based case (decoder output information predicts one or more candidate tokens that follow). The difference between the prior art and the claimed invention is that Yan does not explicitly teach before or after performing model quantization on the machine learning model to generate a quantized model; and executing inference based on the quantized model, and the BoS cache into the quantized model to generate output and cache of the next token. Sheng teaches before or after performing model quantization on the machine learning model to generate a quantized model ([5.] [Group-wise Quantization] 4-bit means using group-wise quantization to compress both weights and KV cache into 4-bit integers); and executing inference based on the quantized model ([3.] [Generative Inference] inference using the quantized representations: prefill and decoding stages operate with compressed (quantized) weight and KV cache), and the BoS cache into the quantized model to generate output and cache of the next token ([3.] [Generative Inference] [Section 3. Equations] during the decode phase, the inference computation need to update the KV cache (Concat, the decode equations) indicating use of previously generated cache when processing the next token). Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Yan with teachings of Sheng by modifying the resource-efficient attention in a neural network as taught by Yan to include before or after performing model quantization on the machine learning model to generate a quantized model; and executing inference based on the quantized model, and the BoS cache into the quantized model to generate output and cache of the next token for the benefit of producing high-throughput LLM inference using limited resources (Sheng [Abstract]). The difference between the prior art and the claimed invention is that Yan nor Sheng explicitly teach the output and the BoS cache of the BoS token is generated by using the machine learning model which is based on float data type before inference is executed; the BoS token does not participate in the model quantization; and the BoS cache is used as information to indicate the BoS token is at the beginning of all tokens. Therefore, it would not have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Yan and Sheng to include the output and the BoS cache of the BoS token is generated by using the machine learning model which is based on float data type before inference is executed; the BoS token does not participate in the model quantization; and the BoS cache is used as information to indicate the BoS token is at the beginning of all tokens. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHREYANS A. PATEL Primary Examiner Art Unit 2653 /SHREYANS A PATEL/ Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jun 21, 2024
Application Filed
Feb 24, 2026
Non-Final Rejection mailed — §101
May 25, 2026
Response Filed
Jul 23, 2026
Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12659658
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
2y 10m to grant Granted Jun 16, 2026
Patent 12646496
METHODS AND SYSTEMS OF TEXT-CONDITIONED AUDIO-VISUAL SPEECH GENERATION WITH MULTI-MODAL LATENT DIFFUSION MODELS
2y 2m to grant Granted Jun 02, 2026
Patent 12608559
METHOD AND SYSTEM FOR ENHANCING A MUTIMODAL INPUT CONTENT
3y 0m to grant Granted Apr 21, 2026
Patent 12609128
METHOD FOR IMPROVING FAR-FIELD SPEECH INTERACTION PERFORMANCE, AND FAR-FIELD SPEECH INTERACTION SYSTEM
2y 0m to grant Granted Apr 21, 2026
Patent 12586597
ENHANCED AUDIO FILE GENERATOR
3y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
89%
Grant Probability
97%
With Interview (+8.5%)
2y 0m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 411 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month