Prosecution Insights
Last updated: October 02, 2026
Application No. 19/057,428

MEMORY-EFFICIENT DECODING WITH KV CACHE COMPRESSION FOR LARGE LANGUAGE MODELS

Non-Final OA §101§103
Filed
Feb 19, 2025
Examiner
BLANKENAGEL, BRYAN S
Art Unit
2658
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
262 granted / 390 resolved
+5.2% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
23 currently pending
Career history
418
Total Applications
across all art units

Statute-Specific Performance

§101
25.1%
-14.9% vs TC avg
§103
50.7%
+10.7% vs TC avg
§102
11.8%
-28.2% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 390 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: Fig. 4 elements 440, 445. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. Using the subject matter eligibility test from page 74621 of the Federal Register Notice titled “2014 Interim Guidance on Patent Subject Matter Eligibility,” a two-step process is performed. Under step 1, the claims are analyzed to determine if the claim is directed to a process, machine, article of manufacture, or composition of matter. In this case, claims 1-10 are directed to a method, which is a process; claims 11-15 are directed to a computer program product, which is a machine or an article of manufacture; and claims 16-20 are directed to a system, which is a machine or an article of manufacture. Step 2A (part 1 of the Mayo test), using the guidance from pages 50-57 of the Federal Register Vol. 84 No. 4 from Monday, January 7, 2019, requires applying a two-prong inquiry. In Prong One, examiners evaluate whether the claim recites a judicial exception, determining if the claim is directed to a law of nature, a natural phenomenon, or an abstract idea. In this case, claim 1 recites performing compression and decompression, which are mathematical calculations. In Prong Two, examiners evaluate whether the judicial exception is integrated into a practical application that imposes a meaningful limit on the judicial exception. In this case, additional elements of processor, memory, and storage medium are generic computing components, and do not integrate the abstract ideas into a practical application. Step 2B (part 2 of the Mayo test) requires analyzing the claims to determine if they recite additional elements that amount to significantly more than the judicial exception. In this case, the claims do not include additional elements that are sufficient to amount to significantly more than the abstract idea itself. Regarding claims 1, 11, and 16, compressing and decompressing data is a mathematical calculation, which is an abstract idea. Additional limitations of processor, memory, and storage medium are generic computing components, and do not integrate the abstract ideas into a practical application or constitute significantly more. Regarding claims 2-3, 7, 13, and 18, the limitations are further clarifications of the above abstract ideas. Regarding claims 4, 12, and 17, evicting from a cache is a mathematical calculation, such as deleting a value, which is an abstract idea without integration into a practical application and without significantly more. Regarding claim 5, deleting a value is a mathematical calculation, which is an abstract idea without integration into a practical application and without significantly more. Regarding claim 6, comparing a value to a threshold is a mathematical calculation, which is an abstract idea without integration into a practical application and without significantly more. Regarding claims 8-9, 14, and 19, combining values along an axis and performing singular value decomposition are mathematical calculations, which is an abstract idea without integration into a practical application and without significantly more. Regarding claims 10, 15, and 20, decompressing is a mathematical calculation, which is an abstract idea without integration into a practical application and without significantly more. The limitations of the claims, taken alone, do not amount to significantly more than the above-identified judicial exception (the abstract idea). Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements individually. Applicable case law cited in the Federal Register includes, but is not limited to: Alice Corp., 134 S. Ct. at 2355-56, Digitech Image Tech., LLC v. Electronics for Imaging, Inc., 758 F.3d 1344 (Fed. Cir. 2014), Benson, 409 U.S. at 63. See "Preliminary Examination Instructions in view of the Supreme Court Decision in Alice Corporation Pty. Ltd. v. CLS Bank International, et al.," dated June 25, 2014, and the Federal Register notice titled "2014 Interim Guidance on Patent Subject Matter Eligibility" (79 FR 74618). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-7, 10-13, 15-18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jha et al. (US 2025/0124105 A1), hereinafter referred to as Jha, in view of Martini et al. (US 2026/0017323 A1), hereinafter referred to as Martini. Regarding claim 1, Jha teaches: A method comprising: performing, during a decoding phase performed by a large language model (LLM), multi- head compression of key-value data in a cache (para [0045-46], where an encoder-decoder utilized multiple attention heads, and para [0031], where KV cache compression is performed to increase LLM inference); and wherein the LLM comprises a decoder-based transformer model that is configured to generate an output based on an input by performing a tokenization phase, a prefill phase, the decoding phase, and a detokenization phase (Fig. 1 elements 104, 114, para [0037-39], where an input is tokenized, and the output is detokenized, and para [0070-72], where prefill and decoding are performed). Jha does not teach: performing partial decompression of the key-value data during the decoding phase, Martini teaches: performing partial decompression of the key-value data during the decoding phase (para [0053], where the keys and values associated with the combination of the outputs of the transformer and adapter are decompressed), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jha by performing the compression and decompression of Martini (Martini para [0053]) on the KV data of Jha (Jha para [0031]), in order to minimize or reduce the size of the cache (Martini para [0037]). Regarding claim 2, Jha in view of Martini teaches: The method of claim 1, wherein the cache is in a memory in a computing device (Jha Fig. 16 elements 904, 1604, 1600, para [0143-145], where the KV cache is in memory of a computing device). Regarding claim 3, Jha in view of Martini teaches: The method of claim 1, wherein the cache is in a memory in one or more graphics processing units (GPUs) (Jha para [0053], where the KV cache is allocated in GPU memory). Regarding claim 4, Jha in view of Martini teaches: The method of claim 1, further comprising evicting from the cache, and during the decoding phase, one or more tokens of a plurality of tokens based on scaled accumulated attention scores (Jha para [0083], where KV tokens are evicted from the cache during compression based on the attention weights, and para [0099], where the weights are summed and normalized). Regarding claim 5, Jha in view of Martini teaches: The method of claim 4, wherein: the cache includes the key-value data in the form of key vectors and value vectors associated with respective ones of the plurality of tokens (Jha para [0066], where the cache stores key and value tensors); and the evicting comprises deleting from the cache the key vectors and value vectors associated with the one or more tokens of the plurality of tokens (Jha para [0026], where the evicting means the tensors are not stored or kept in the KV cache). Regarding claim 6, Jha in view of Martini teaches: The method of claim 4, further comprising performing the evicting based on determining that a size of the data stored in the cache exceeds a threshold value (Jha para [0066], where the cache runs out of blocks to store the tensors, and the block allocation requests eviction). Regarding claim 7, Jha in view of Martini teaches: The method of claim 4, wherein the evicting is based on respective positions of the plurality of tokens (Jha para [0038], where the token positions within the sequences contribute to the attention, which is used for eviction). Regarding claim 10, Jha in view of Martini teaches: The method of claim 1, wherein the performing partial decompression comprises decompressing a subset of the key-value data based on variance of attention scores associated with respective ones of a plurality of tokens (Martini para [0053], where the keys and values associated with the combination of the outputs of the transformer and adapter are decompressed, and Jha para [0083], where the attention weights are used to determine importance). Regarding claim 11, Jha teaches: A computer program product comprising: one or more computer-readable storage media (para [0143], where computer-readable storage media is used); and program instructions stored on the one or more computer-readable storage media to perform operations (para [0144], where computer-readable storage media store instructions) comprising: performing, during a decoding phase performed by a large language model (LLM), multi-head compression of key-value data in a cache (para [0045-46], where an encoder-decoder utilized multiple attention heads, and para [0031], where KV cache compression is performed to increase LLM inference); and wherein the LLM comprises a decoder-based transformer model that is configured to generate an output based on an input by performing a tokenization phase, a prefill phase, the decoding phase, and a detokenization phase (Fig. 1 elements 104, 114, para [0037-39], where an input is tokenized, and the output is detokenized, and para [0070-72], where prefill and decoding are performed). Jha does not teach: performing partial decompression of the key-value data during the decoding phase, Martini teaches: performing partial decompression of the key-value data during the decoding phase (para [0053], where the keys and values associated with the combination of the outputs of the transformer and adapter are decompressed), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jha by performing the compression and decompression of Martini (Martini para [0053]) on the KV data of Jha (Jha para [0031]), in order to minimize or reduce the size of the cache (Martini para [0037]). Regarding claim 12, Jha in view of Martini teaches: The computer program product of claim 11, further comprising evicting from the cache, and during the decoding phase, one or more tokens of a plurality of tokens based on scaled accumulated attention scores (Jha para [0083], where KV tokens are evicted from the cache during compression based on the attention weights, and para [0099], where the weights are summed and normalized). Regarding claim 13, Jha in view of Martini teaches: The computer program product of claim 12, wherein the evicting is based on respective positions of the plurality of tokens in a recency window (Jha para [0038], where the token positions within the sequences contribute to the attention, which is used for eviction). Regarding claim 15, Jha in view of Martini teaches: The computer program product of claim 11, wherein the performing partial decompression comprises decompressing a subset of the key-value data based on variance of attention scores associated with respective ones of a plurality of tokens (Martini para [0053], where the keys and values associated with the combination of the outputs of the transformer and adapter are decompressed, and Jha para [0083], where the attention weights are used to determine importance). Regarding claim 16, Jha teaches: A computer system comprising: a processor set (para [0142], where processing devices are used); one or more computer-readable storage media (para [0143], where computer-readable storage media is used); and program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations (para [0144], where computer-readable storage media store instructions) comprising: performing, during a decoding phase performed by a large language model (LLM), multi-head compression of key-value data in a cache (para [0045-46], where an encoder-decoder utilized multiple attention heads, and para [0031], where KV cache compression is performed to increase LLM inference); and wherein the LLM comprises a decoder-based transformer model that is configured to generate an output based on an input by performing a tokenization phase, a prefill phase, the decoding phase, and a detokenization phase (Fig. 1 elements 104, 114, para [0037-39], where an input is tokenized, and the output is detokenized, and para [0070-72], where prefill and decoding are performed). Jha does not teach: performing partial decompression of the key-value data during the decoding phase, Martini teaches: performing partial decompression of the key-value data during the decoding phase (para [0053], where the keys and values associated with the combination of the outputs of the transformer and adapter are decompressed), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jha by performing the compression and decompression of Martini (Martini para [0053]) on the KV data of Jha (Jha para [0031]), in order to minimize or reduce the size of the cache (Martini para [0037]). Regarding claim 17, Jha in view of Martini teaches: The computer system of claim 16, further comprising evicting from the cache, and during the decoding phase, one or more tokens of a plurality of tokens based on scaled accumulated attention scores (Jha para [0083], where KV tokens are evicted from the cache during compression based on the attention weights, and para [0099], where the weights are summed and normalized). Regarding claim 18, Jha in view of Martini teaches: The computer system of claim 17, wherein the evicting is based on respective positions of the plurality of tokens in a recency window (Jha para [0038], where the token positions within the sequences contribute to the attention, which is used for eviction). Regarding claim 20, Jha in view of Martini teaches: The computer system of claim 16, wherein the performing partial decompression comprises decompressing a subset of the key-value data based on variance of attention scores associated with respective ones of a plurality of tokens (Martini para [0053], where the keys and values associated with the combination of the outputs of the transformer and adapter are decompressed, and Jha para [0083], where the attention weights are used to determine importance). Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jha, in view of Martini, and further in view of Behnam et al. (US 2026/0220041 A1), hereinafter referred to a Behnam. Regarding claim 8, Jha in view of Martini teaches: The method of claim 1, Jha in view of Martini does not teach: wherein the performing multi-head compression comprises compressing information shared across heads by combining values in the cache along a head axis. Behnam teaches: wherein the performing multi-head compression comprises compressing information shared across heads by combining values in the cache along a head axis (para [0034], where compression involves accumulating per-group rather than per-head). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jah in view of Martini by using the compression method of Behnam (Behnam para [0034]) in the compression of Jha in view of Martini (Jah para [0031]), in order to avoid redundant storage of the same KV tokens (Behnam para [0034]). Claim(s) 9, 14, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Jha, in view of Martini, and Behnam, and further in view of Zhang et al. (Zhang, R., Wang, K., Liu, L., Wang, S., Cheng, H., Zhang, C., & Shen, Y. (2024). Lorc: Low-rank compression for llms kv cache with a progressive compression strategy. arXiv preprint arXiv:2410.03111.), hereinafter referred to as Zhang. Regarding claim 9, Jha in view of Martini and Behnam teaches: The method of claim 8 Jha in view of Martini and Behnam does not teach: wherein the performing multi-head compression further comprises reducing respective matrices in the cache into components by performing singular value decomposition. Zhang teaches: wherein the performing multi-head compression further comprises reducing respective matrices in the cache into components by performing singular value decomposition (pages 4-5, section 4, 4.1, where SVD is performed on the weight matrices). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jah in view of Martini and Behnam by using the decomposition of Zhang (Zhang pages 4-5, section 4, 4.1) in the compression of Jha in view of Martini and Behnam (Jah para [0031]), in order to leverage low-rank properties of the weight matrices (Zhang page 4 section 4.1). Regarding claim 14, Jha in view of Martini teaches: The computer program product of claim 11, wherein the performing multi-head compression comprises: Jha in view of Martini does not teach: compressing information shared across heads by combining values in the cache along a head axis; and reducing respective matrices into components by performing singular value decomposition. Behnam teaches: compressing information shared across heads by combining values in the cache along a head axis (para [0034], where compression involves accumulating per-group rather than per-head); and It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jah in view of Martini by using the compression method of Behnam (Behnam para [0034]) in the compression of Jha in view of Martini (Jah para [0031]), in order to avoid redundant storage of the same KV tokens (Behnam para [0034]). Zhang teaches: reducing respective matrices into components by performing singular value decomposition (pages 4-5, section 4, 4.1, where SVD is performed on the weight matrices). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jah in view of Martini and Behnam by using the decomposition of Zhang (Zhang pages 4-5, section 4, 4.1) in the compression of Jha in view of Martini and Behnam (Jah para [0031]), in order to leverage low-rank properties of the weight matrices (Zhang page 4 section 4.1). Regarding claim 19, Jha in view of Martini teaches: The computer system of claim 16, wherein the performing multi-head compression comprises: Jha in view of Martini does not teach: compressing information shared across heads by combining values in the cache along a head axis; and reducing respective matrices into components by performing singular value decomposition. Behnam teaches: compressing information shared across heads by combining values in the cache along a head axis (para [0034], where compression involves accumulating per-group rather than per-head); and It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jah in view of Martini by using the compression method of Behnam (Behnam para [0034]) in the compression of Jha in view of Martini (Jah para [0031]), in order to avoid redundant storage of the same KV tokens (Behnam para [0034]). Zhang teaches: reducing respective matrices into components by performing singular value decomposition (pages 4-5, section 4, 4.1, where SVD is performed on the weight matrices). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jah in view of Martini and Behnam by using the decomposition of Zhang (Zhang pages 4-5, section 4, 4.1) in the compression of Jha in view of Martini and Behnam (Jah para [0031]), in order to leverage low-rank properties of the weight matrices (Zhang page 4 section 4.1). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 2026/0228532 A1 para [0053] teaches improving LLM performance by compressing, storing, and using KV cache data generating during input prompt processing. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRYAN S BLANKENAGEL whose telephone number is (571)270-0685. The examiner can normally be reached 8:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BRYAN S BLANKENAGEL/Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Feb 19, 2025
Application Filed
Aug 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724962
PRE-TRAINING LANGUAGE MODELS USING NATURAL LANGUAGE EXPRESSIONS EXTRACTED FROM STRUCTURED DATABASES
4y 3m to grant Granted Sep 01, 2026
Patent 12718826
AUDIO SIGNAL ENCODING AND DECODING METHOD AND APPARATUS
2y 7m to grant Granted Aug 25, 2026
Patent 12711947
METHOD FOR TRAINING A NEURAL NETWORK AND A DATA PROCESSING DEVICE
2y 8m to grant Granted Aug 18, 2026
Patent 12711969
METHODS AND APPARATUS FOR SUPPLEMENTING PARTIALLY READABLE AND/OR INACCURATE CODES IN MEDIA
2y 2m to grant Granted Aug 18, 2026
Patent 12711983
METHOD OF DETECTING SPEECH AND SPEECH DETECTOR FOR LOW SIGNAL-TO-NOISE RATIOS
2y 1m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+33.3%)
2y 8m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 390 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month