Prosecution Insights
Last updated: October 01, 2026
Application No. 19/228,714

SYSTEMS AND METHODS FOR COMPRESSING, DECOMPRESSING, AND PROCESSING DATA FOR USE BY MACHINE LEARNING MODELS

Non-Final OA §102§103
Filed
Jun 04, 2025
Priority
Jun 04, 2024 — provisional 63/655,968 +2 more
Examiner
SADLER, NATHAN
Art Unit
2139
Tech Center
2100 — Computer Architecture & Software
Assignee
Meta Platforms Technologies LLC
OA Round
1 (Non-Final)
71%
Grant Probability
Favorable
1-2
OA Rounds
1y 7m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 71% — above average
71%
Career Allowance Rate
481 granted / 679 resolved
+15.8% vs TC avg
Strong +26% interview lift
Without
With
+26.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
21 currently pending
Career history
712
Total Applications
across all art units

Statute-Specific Performance

§101
6.8%
-33.2% vs TC avg
§103
50.6%
+10.6% vs TC avg
§102
19.6%
-20.4% vs TC avg
§112
19.0%
-21.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 679 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event a determination of the status of the application as subject to AIA 35 U.S.C. 102, 103, and 112 (or as subject to pre-AIA 35 U.S.C. 102, 103, and 112) is incorrect, any correction of the statutory basis for a rejection will not be considered a new ground of rejection if the prior art relied upon and/or the rationale supporting the rejection, would be the same under either status. Notice of Claim Interpretation Claims in this application are not interpreted under 35 U.S.C. 112(f) unless otherwise noted in an office action. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference characters not mentioned in the description: 900-1400. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Color photographs and color drawings are not accepted in utility applications unless a petition filed under 37 CFR 1.84(a)(2) is granted. Any such petition must be accompanied by the appropriate fee set forth in 37 CFR 1.17(h), one set of color drawings or color photographs, as appropriate, if submitted via the USPTO patent electronic filing system or three sets of color drawings or color photographs, as appropriate, if not submitted via the via USPTO patent electronic filing system, and, unless already present, an amendment to include the following language as the first paragraph of the brief description of the drawings section of the specification: The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color photographs will be accepted if the conditions for accepting color drawings and black and white photographs have been satisfied. See 37 CFR 1.84(b)(2). Specification The disclosure is objected to because of the following informalities: “1502” should be --1802-- in paragraph 00120. Appropriate correction is required. Claim Rejections - 35 USC § 102 (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1 and 15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”). In regards to claims 1 and 15, Sheng teaches a non-transitory, computer-readable storage medium including executable instructions (“FlexGen is implemented on top of PyTorch”, page 7, paragraph 9) that, when executed by one or more processors (See Table 1, page 7), cause the one or more processors to perform or cause performance of: receiving a tensor (“Given a tensor”, page 7, paragraph 1); compressing the tensor by applying a compression scheme to values of the tensor to form a compressed tensor (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); storing the compressed tensor into a tensor cache (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); reading the compressed tensor from the tensor cache (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); decompressing the compressed tensor by applying a decompression scheme to values of the compressed tensor to form a decompressed tensor (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); and forwarding the decompressed tensor to a compute unit (“The tensors are stored in the quantized format and converted back to FP16 before computation.”, page 7, paragraph 2). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”) in view of Loh (US 2020/0195273). In regards to claims 2 and 16, Sheng teaches claims 1 and 15. Sheng fails to teach that the compression scheme is indicated by a compression flag in an instruction for storing the tensor. Loh teaches that the compression scheme is indicated by a compression flag in an instruction for storing the tensor (“When the processor executes a lossy store instruction, the value in the registers are recompressed by the compression unit 105, and these compressed values are stored in the cache.”, paragraph 0024) in order to support applications that need to exactly reproduce stored information (paragraph 0001). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Sheng with Loh such that the compression scheme is indicated by a compression flag in an instruction for storing the tensor in order to support applications that need to exactly reproduce stored information (id.). Claims 3-6, 8, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”) in view of Vakili et al. (“Fast and Low-Cost Approximate Multiplier for FPGAs using Dynamic Reconfiguration”). In regards to claims 3 and 17, Sheng further teaches that the compression scheme corresponds to quantizing an input into a quantized format having a reduced bit size (“For each group, we compute the min and max of the group elements and quantize each element x into b-bit integers by xquant = round(x−min/max−min × (2b − 1)). The tensors are stored in the quantized format and converted back to FP16 before computation.”, page 7, paragraphs 1-2). Sheng fails to teach quantizing an integer into a quantized floating-point format. Vakili teaches quantizing an integer into a quantized floating-point format (“For an INT8 case study, illustration of the design detail of a highly optimized encoder circuit to convert INT8 to 8-bit floating point format”, page 2, paragraph 2, bullet 3) thereby “reducing significantly the number of required LUTs in FPGAs” (page 2, paragraph 2). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Shen with Vakili to include quantizing an integer into a quantized floating-point format thereby “reducing significantly the number of required LUTs in FPGAs” (id.). In regards to claims 4 and 18, Vakili further teaches that the quantized floating-point format retains a sign bit of the integer (“ PNG media_image1.png 21 82 media_image1.png Greyscale ”, page 2, equation 4). In regards to claims 5 and 19, Vakili further teaches that the quantized floating-point format includes an exponent portion corresponding to a position of a non-zero most significant bit (MSB) of the integer (“ PNG media_image2.png 39 286 media_image2.png Greyscale ”, page 2, equation 4). In regards to claims 6 and 20, Sheng further teaches that the compression scheme includes determining the exponent portion for each integer value (“For each group, we compute the min and max of the group elements and quantize each element x into b-bit integers”, page 7, paragraph 1). In regards to claim 8, Vakili further teaches that the quantized floating-point format includes a mantissa portion corresponding to a value of a non-zero most significant bit (MSB) of the integer (“The mantissa is a segment of mntBW bits from |X|, with the leftmost bit in |X| that contains ’1’ being the most significant bit.”, page 2, paragraph 6). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”) in view of Vakili et al. (“Fast and Low-Cost Approximate Multiplier for FPGAs using Dynamic Reconfiguration”) and Ghaffari et al. (US 2025/0200277). In regards to claim 7, Sheng in view of Vakili teaches claim 5. Sheng in view of Vakili fails to teach that the compression scheme includes determining the exponent portion for a group of integer values. Ghaffari teaches that the compression scheme includes determining the exponent portion for a group of integer values (“In this example, the processor 110 can be configured to determine, in each column, a group of the values having a groupsize<m, and determine the maximum exponent 404 in each group.”, paragraph 0091) which requires less space (paragraph 0100). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Sheng with Vakili and Ghaffari such that the compression scheme includes determining the exponent portion for a group of integer values which requires less space (id.). Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”) in view of Liu et al. (US 2022/0383081). In regards to claim 9, Sheng teaches a system, comprising: memory including one or more programs (“FlexGen is implemented on top of PyTorch”, page 7, paragraph 9) that are configured to be executed by one or more processors (See Table 1, page 7), the one or more programs including instructions for: receiving a tensor (“Given a tensor”, page 7, paragraph 1); compressing the tensor by applying a compression scheme to values of the tensor to form a compressed tensor (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); storing the compressed tensor into a tensor cache (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); reading the compressed tensor from the tensor cache (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); decompressing the compressed tensor by applying a decompression scheme to values of the compressed tensor to form a decompressed tensor (“The tensors are stored in the quantized format and converted back to FP16 before computation. Since both the weights and KV cache consume a significant amount of memory, we compress both to 4 bits with a group size of 64.”, page 7, paragraph 2); and forwarding the decompressed tensor to a compute unit (“The tensors are stored in the quantized format and converted back to FP16 before computation.”, page 7, paragraph 2). Sheng fails to teach a wearable device; and wherein the one or more processors are in communication with the wearable device. Liu teaches a wearable device (“To achieve low latency and high energy efficiency for always-accessible user experiences, AR/VR hardware needs to reduce data movement cost between different modules, and needs to have a small form factor due to area and size constraints in wearable or portable devices.”, paragraph 0053); and wherein the one or more processors (“In some embodiments, console 110 may include a processor and a non-transitory computer-readable storage medium storing instructions executable by the processor.”, paragraph 0077) are in communication with the wearable device (“In some embodiments, console 110 may provide content to near-eye display 120 for presentation to the user in accordance with information received from one or more of external imaging device 150, near-eye display 120, and input/output interface 140.”, paragraph 0076) “[t]o achieve low latency and high energy efficiency for always-accessible user experiences” (paragraph 0053). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Sheng with Liu to include a wearable device; and wherein the one or more processors are in communication with the wearable device “[t]o achieve low latency and high energy efficiency for always-accessible user experiences” (id.). Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”) in view of Liu et al. (US 2022/0383081) and Loh (US 2020/0195273). In regards to claim 10, Sheng in view of Liu teaches claim 9. Sheng in view of Liu fails to teach that the compression scheme is indicated by a compression flag in an instruction for storing the tensor. Loh teaches that the compression scheme is indicated by a compression flag in an instruction for storing the tensor (“When the processor executes a lossy store instruction, the value in the registers are recompressed by the compression unit 105, and these compressed values are stored in the cache.”, paragraph 0024) in order to support applications that need to exactly reproduce stored information (paragraph 0001). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Sheng with Liu and Loh such that the compression scheme is indicated by a compression flag in an instruction for storing the tensor in order to support applications that need to exactly reproduce stored information (id.). Claims 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over Sheng et al. (“FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU”) in view of Liu et al. (US 2022/0383081) and Vakili et al. (“Fast and Low-Cost Approximate Multiplier for FPGAs using Dynamic Reconfiguration”). In regards to claim 11, Sheng further teaches that the compression scheme corresponds to quantizing an input into a quantized format having a reduced bit size (“For each group, we compute the min and max of the group elements and quantize each element x into b-bit integers by xquant = round(x−min/max−min × (2b − 1)). The tensors are stored in the quantized format and converted back to FP16 before computation.”, page 7, paragraphs 1-2). Sheng in view of Liu fails to teach quantizing an integer into a quantized floating-point format. Vakili teaches quantizing an integer into a quantized floating-point format (“For an INT8 case study, illustration of the design detail of a highly optimized encoder circuit to convert INT8 to 8-bit floating point format”, page 2, paragraph 2, bullet 3) thereby “reducing significantly the number of required LUTs in FPGAs” (page 2, paragraph 2). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Shen with Liu and Vakili to include quantizing an integer into a quantized floating-point format thereby “reducing significantly the number of required LUTs in FPGAs” (id.). In regards to claim 12, Vakili further teaches that the quantized floating-point format retains a sign bit of the integer (“ PNG media_image1.png 21 82 media_image1.png Greyscale ”, page 2, equation 4). In regards to claim 13, Vakili further teaches that the quantized floating-point format includes an exponent portion corresponding to a position of a non-zero most significant bit (MSB) of the integer (“ PNG media_image2.png 39 286 media_image2.png Greyscale ”, page 2, equation 4). In regards to claim 14, Sheng further teaches that the compression scheme includes determining the exponent portion for each integer value (“For each group, we compute the min and max of the group elements and quantize each element x into b-bit integers”, page 7, paragraph 1). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Shao (US 2020/0042189) teaches compressing data for the L2 cache and decompressing it. Wu (US 11,461,625) teaches decompress compressed tensor data from the tensor cache. Gobriel (US 2025/0061316) teaches different quantization levels for KV cache pages. Guo (WO 2025/185502) as supported by its translation; its CN application, Guo (CN 120597956); and the translation of the CN application teaches storing compressed K or V data in a cache and performing inverse quantization processing. Ma (WO 2025/185500) as supported by its translation; its CN application, Ma (CN 120597957); and the translation of the CN application teaches decompressing KV cache data. Wang (WO 2025/222855) as supported by its translation; its CN application, Wang (CN 120849529); and the translation of the CN application teaches compressing key and value vectors. Kundu (WO 2025/244669) teaches quantizing key and value tensors to store in a compressed KV cache. Zhao et al. ("Atom: Low-Bit Quantization for Efficient and Accurate LLM Serving") teaches dequantizing the output of the KV cache. Zhang et al. ("Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache") teaches a quantized KV cache. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NATHAN SADLER whose telephone number is (571)270-7699. The examiner can normally be reached Monday - Friday 8am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Reginald Bragdon can be reached at (571)272-4204. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Nathan Sadler/Primary Examiner, Art Unit 2139 9 June 2026
Read full office action

Prosecution Timeline

Jun 04, 2025
Application Filed
Jun 11, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724722
PROCESSOR AND MEMORY COMMUNICATION IN A STACKED MEMORY SYSTEM
2y 0m to grant Granted Sep 01, 2026
Patent 12688127
CACHE DATA PROCESSING METHOD, SYSTEM, APPARATUS, AND DEVICE, AND COMPUTER STORAGE MEDIUM
1y 1m to grant Granted Jul 21, 2026
Patent 12675411
CACHE OPTIMIZATION FOR A REMOTE STORAGE DEVICE
2y 6m to grant Granted Jul 07, 2026
Patent 12650935
SYSTEMS, METHODS, AND APPARATUS FOR CACHE OPERATION IN STORAGE DEVICES
2y 2m to grant Granted Jun 09, 2026
Patent 12645610
OBJECT-LEVEL METADATA LOCATOR
1y 10m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
71%
Grant Probability
97%
With Interview (+26.0%)
2y 11m (~1y 7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 679 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month