Prosecution Insights
Last updated: August 28, 2026
Application No. 18/504,016

FLEXIBLE MACHINE LEARNING MODEL COMPRESSION

Non-Final OA §102§103
Filed
Nov 07, 2023
Examiner
MEYER, JACQUELINE CHRISTINE
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
14 granted / 21 resolved
+6.7% vs TC avg
Strong +57% interview lift
Without
With
+57.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
16 currently pending
Career history
42
Total Applications
across all art units

Statute-Specific Performance

§101
28.1%
-11.9% vs TC avg
§103
50.0%
+10.0% vs TC avg
§102
10.2%
-29.8% vs TC avg
§112
10.7%
-29.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 21 resolved cases

Office Action

§102 §103
CTNF 18/504,016 CTNF 99760 DETAILED ACTION This nonfinal office action is responsive to claims filed on November 7, 2023. Claims 1-20 are pending. Claims 1, 10, and 12 are independent. Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Information Disclosure Statement The information disclosure statements (IDS) submitted on July 25, 2024 and April 28, 2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Objections 07-29-01 AIA Claim s 3, 4, 14, and 15 are objected to because of the following informalities: Claims 3 and 14 read “determining a number of exponent bits required for uniquely representing represent each value within the one or more dense ranges” but should read “determining a number of exponent bits required for uniquely representing represent each value within the one or more dense ranges.” Claims 4 and 15 read “determining a number of mantissa bits required for representing represent each value within the one or more dense ranges” but should read “determining a number of mantissa bits required for representing represent each value within the one or more dense ranges.” Appropriate correction is required. Claim Rejections - 35 USC § 102 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-15 AIA Claim s 1-6, 9-17, and 20 are rejected under 35 U.S.C. 102( a)(1 ) as being anticipated by Bordawekar et al. ( EFloat: Entropy-coded Floating Point Format for Compressing Vector Embedding Models ), hereinafter Bordawekar . Examiner’s note: Version 1 of Bordawekar was provided by applicant in the IDS dated 4/28/2025. Examiner is relying on version 3, a copy of which has been included. Regarding claim 1, Bordawekar teaches the method: obtaining trained values of a set of parameters for at least a portion of a machine learning model, the trained values represented in a first format; (Bordawekar, page 1, column 2, paragraph 3: “A trained vector embedding model is a snapshot of it’s weight matrices and consists of weight values represented in IEEE 32-bit single-precision floating point (FP32) format. T herefore, we focus on compression approaches that compress FP32 to low-precision floating point formats. In our approach, active rows of the compressed db2Vec model are uncompressed on-demand during inference computations and the original FP32 values are restored with smaller loss of accuracy than other low- precision methods, as we quantify in later sections.” – The weight values being compressed in the approach is indicative of the weight values first being obtained. Where 32-bit single-precision floating point format is a first format that the values are received in. ) identifying one or more dense ranges for the trained values, the one or more dense ranges collectively defining a range of values for which a density of the trained values within the range meets a minimum density criterion; (Bordawekar, page 4, column 2, paragraph 3: “Irrespective of the model type, we observe in Fig. 2 that exponent values cluster in a narrow range of values, and display a distinct histogram peak.” – The clusters in narrow range of values is analogous to the dense ranges where the histogram is used to identify these ranges. ) determining a least number of bits required to represent each trained value within the one or more dense ranges; (Bordawekar, page 6, column 1, paragraph 2: “Frequency of unique exponent values in the data set determine the coded-exponent widths which may vary between as small as 1-bit and some software configurable maximum, e.g., 8-bit (Figure 1(e,f,g)).” – The number of required bits being configurable is indicative of determining a least number of bits required. ) identifying, based at least on the determined least number of bits required to represent each trained value within the one or more dense ranges, a second format having a range that is smaller than a range of the first format; and (Bordawekar, page 2, column 1, paragraph 2: “Figure 1(e) presents the new low-precision FP format, EFloat (EFn), with a fixed total bit budget of 𝑛 bits, e.g., 𝑛 = 16. EF16 uses an entropy coded variable-width 𝑁 bit exponent and a variable-width 15 − 𝑁 bit significand (mantissa) with a total number of 16 bits, including the sign bit. The EF width is adjustable, e.g., EF12, permitting a tradeoff between compression ratio and accuracy, as we demonstrate in later sections.” – The second format being 16-bit is a smaller format than the first. With the variable width being adjustable this implies a range for the format that is smaller than that of the first. ) generating a compressed version of the at least a portion of the machine learning model by: (i) converting the trained values of the set of parameters within the one or more dense ranges from the first format to compressed values represented by the second format , and (Bordawekar, page 6, column 1, paragraph 2: “Thanks to the entropy coding of the Huffman algorithm, frequent exponent values are coded with fewer bits and infrequent exponents are coded with more bits as observed in Fig.9.” – The more frequent exponent values are those from the dense range and would thus be compressed (e.g., fewer bits). ) (ii) assigning predetermined values to each of the parameters in the compressed version of the at least a portion of the machine learning model whose trained values were not within the one or more dense ranges. (Bordawekar, page 5, column 2, last paragraph: “Then, one could use the same code-table any number times because the exponent frequencies hence the code-tables are practically identical across these models.” – The code-tables being able to be used any number of times implies that there are predetermined values for the parameters not in the dense ranges (e.g., the infrequent exponents noted above). ) Regarding claim 2, Bordawekar teaches the method of claim 1, as cited above. Bordawekar further teaches: wherein the first format is a floating-point format. (Bordawekar, page 8, column 1, paragraph 4: “Currently, both EFloat encoding and decoding are completely implemented in software that allows us to compress an FP32 value to an EFn format of variable size 𝑛 , and conversely, given a value in the EFn format, generate its FP32, BF16, or FP16 representations.”) Regarding claim 3, Bordawekar teaches the method of claim 1, as cited above. Bordawekar further teaches: determining a number of exponent bits required for uniquely representing represent each value within the one or more dense ranges. (Bordawekar, page 2, column 1, paragraph 4: “EFloat provides flexible variable-length reduced-bit representation of any floating point format (e.g., FP32, FP16) by using fewer exponent bits to map the same exponent range as the original value.” – Mapping the same exponent range as the original value implies that each value is uniquely represented. ) Regarding claim 4, Bordawekar teaches the method of claim 1, as cited above. Bordawekar further teaches: determining a number of mantissa bits required for representing represent each value within the one or more dense ranges. (Bordawekar, page 2, column 1, paragraph 4: “For a given bit budget (e.g., 16), EFloat provides more accurate representation of the FP32 values than BF16 and FP16 by using fewer exponent bits to capture the same range as before, and then using the remaining bits to increase significand precision.” – The significand being the remaining bits indicates the number of mantissa bits. ) Regarding claim 5, Bordawekar teaches the method of claim 1, as cited above. Bordawekar further teaches: selecting, from among multiple candidate formats that correspond respectively to different numbers of bits, a candidate format that corresponds to the determined least number of bits as the second format. (Bordawekar, page 7, column 2, paragraph 1: “Therefore, EFloat not only compresses the regular floats but for a given EF 𝑛 budget of 𝑛 -bits the software application can optimize the end-to-end precision and compression ratio by adjusting both the floating point bit-budget 𝑛 and the maximum coded-exponent width.” – the n-bits being optimized and adjusted indicates that there are multiple candidate formats that are being selected from where the floating point bit-budget is the least number of bits as the second format. ) Regarding claim 6, Bordawekar teaches the method of claim 1, as cited above. Bordawekar further teaches: generating a mapping from the trained values that each have the first format to the compressed values that each have the second format; and (Bordawekar, page 7, column 1, paragraph 2: “The Huffman algorithm builds probabilities based on frequency histogram of exponents and outputs a code-table mapping 8-bit exponents to variable-width coded-exponents.” – The probabilities based on the frequency histogram that outputs the code-table mapping is analogous to generating a mapping from the trained values from a first format to a compressed second format. ) converting the trained values in accordance with the mapping. (Bordawekar, page 7, column 1, paragraph 3: “With a worst case distribution, some coded exponent widths may not fit in an EF16 number (e.g., a 22 bit exponent). Therefore, we use the Length-Limiting variant of the Huffman algorithm to set a maximum coded-exponent width (Abali et al. [1])1.” - The length-limiting variant of the Huffman algorithm is analogous to converting the values in accordance with the mapping. ) Regarding claim 9, Bordawekar teaches the method of claim 1, as cited above. Bordawekar further teaches: wherein the neural network is a Transformer neural network, and wherein the first layer is one of an embedding layer, an attention layer, or a feed-forward layer of the Transformer neural network. (Bordawekar, page 2, column 2, bullet 4: “Since vector embedding models are used in a wide array of NLP architectures including transformers, in addition to db2Vec, the EFloat format can be used for a much wider (and more space consuming) class of NLP models.” And page 11, column 1, paragraph 2: “In particular, advent of new transformer-based NLP models (e.g., BERT and friends, T5, Megatron-LM, Open AI GPT-2/3) has highlighted the very high space and computational costs associated with these models [6, 22, 25, 26, 45, 53, 54, 59]. Given potential uses of NLP models in enterprise and consumer domains, a lot of attention is being devoted to compressing such models. The primary goal of these compression efforts is to reduce the size of a pre-trained model to enable its deployment in real world industrial applications that demand low memory footprint, low response times, and smaller computational and power budget during the inference phase.” – The first layer of a transformer model or BERT is typically the embedding layer. ) Regarding claim 10, claim 10 has all the same limitations of claim 1 which are taught by Bordawekar – see claim 1 above. Bordawekar additionally teaches: One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising: (Bordawekar, page 8, column 1, paragraph 4: “Currently, both EFloat encoding and decoding are completely implemented in software that allows us to compress an FP32 value to an EFn format of variable size 𝑛 , and conversely, given a value in the EFn format, generate its FP32, BF16, or FP16 representations… However, hardware support for EFloat decoding may be necessary so that the numerical values in the correct format are fed to the functional units with minimum delay. The EFloat decoding hardware may be implemented with a Static Random Access Memory (SRAM) based lookup table.”) Regarding claim 11, Bordawekar teaches the computer-readable storage medium of claim 10, as cited above. Claim 11 additionally has the same limitations of claim 9 which is taught by Bordawekar – see claim 9 above. Regarding claim 12, claim 12 has all the same limitations of claim 1 which are taught by Bordawekar – see claim 1 above. Bordawekar additionally teaches: one or more computers; and (Bordawekar, page 8, column 2, paragraph 2: “If the compute processor has many input ports, for example a systolic array such as the Google TPU [39], or a wide SIMD architecture [28, 29], many decoder tables will be necessary for parallel access.”) one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising: (Bordawekar, page 8, column 1, paragraph 4: “Currently, both EFloat encoding and decoding are completely implemented in software that allows us to compress an FP32 value to an EFn format of variable size 𝑛 , and conversely, given a value in the EFn format, generate its FP32, BF16, or FP16 representations… However, hardware support for EFloat decoding may be necessary so that the numerical values in the correct format are fed to the functional units with minimum delay. The EFloat decoding hardware may be implemented with a Static Random Access Memory (SRAM) based lookup table.”) Regarding claim 13, Bordawekar teaches the system of claim 12, as cited above. Claim 13 additionally has the same limitations of claim 2 which are taught by Bordawekar – see claim 2 above. Regarding claim 14, Bordawekar teaches the system of claim 12, as cited above. Claim 14 additionally has the same limitations of claim 3 which are taught by Bordawekar – see claim 3 above. Regarding claim 15, Bordawekar teaches the system of claim 12, as cited above. Claim 15 additionally has the same limitations of claim 4 which are taught by Bordawekar – see claim 4 above. Regarding claim 16, Bordawekar teaches the system of claim 12, as cited above. Claim 16 additionally has the same limitations of claim 5 which are taught by Bordawekar – see claim 5 above. Regarding claim 17, Bordawekar teaches the system of claim 12, as cited above. Claim 17 additionally has the same limitations of claim 6 which are taught by Bordawekar – see claim 6 above. Regarding claim 20, Bordawekar teaches the system of claim 12, as cited above. Claim 20 additionally has the same limitations of claim 9 which are taught by Bordawekar – see claim 9 above . Claim Rejections - 35 USC § 103 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 7, 8, 18, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bordawekar in view of Nagel et al. ( A White Paper on Neural Network Quantization ), hereinafter Nagel . Regarding claim 7, Bordawekar teaches the method of claim 6, as cited above. Bordawekar does not explicitly teach: wherein generating the mapping comprises generating a scale factor for each dense range. However, Nagel teaches: wherein generating the mapping comprises generating a scale factor for each dense range. (Nagel, page 6, paragraph 3: “The quantizer block implements the quantization function of equation (7) and each quantizer is defined by a set of quantization parameters (scale factor, zero-point, bit-width).” - The quantization function is analogous to the mapping that is taught by Bordawekar in claim 6 above. Where the quantization parameters including a scale factor indicates that the mapping generates a scale factor for the dense range. ) Nagel is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Bordawekar, which already teaches converting the trained values from a first format to a second format according to a mapping but does not explicitly teach that the mapping comprises generating a scale factor, to include the teachings of Nagel which does teach that the mapping comprises generating a scale factor which uses a finer granularity in order to improve performance. (Nagel, page 5, paragraph 3) Regarding claim 8, Bordawekar teaches the method of claim 1, as cited above. Bordawekar does not explicitly teach: wherein the minimum density criterion requires that at least a particular percentage of the set of parameters are within the one or more dense ranges. However, Nagel teaches: wherein the minimum density criterion requires that at least a particular percentage of the set of parameters are within the one or more dense ranges. (Nagel, page 9, paragraph 3: “Quantization range setting refers to the method of determining clipping thresholds of the quantization grid, qmin and qmax (see equation 7).” And page 10, paragraph 2: “Range setting for activation quantizers often requires some calibration data. If a layer has batch-normalized activations, the per-channel mean and standard deviation of the activations are equal to the learned batch normalization shift and scale parameters, respectively. These can then be used to find suitable parameters for activation quantizer as follows (Nagel et al., 2019): qmin = min(β−αγ) qmax = max(β+αγ) (18) (19) where β and γ are vectors of per-channel learned shift and scale parameters, and α > 0. Nagel et al. (2019) uses α = 6 so that only large outliers are clipped” – The qmin and qmax is indicative of the range which is already taught by Bordawekar above. The mean and standard deviation being used for the range indicates that a particular percentage of the set of parameters will fall within that range. ) Nagel is considered analogous to the claimed invention as it is in the same field of endeavor, machine learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified Bordawekar, which already teaches a dense range of the parameters using a minimum density criterion but does not explicitly teach that the minimum density criterion requires a particular percentage of the parameters, to include the teachings of Nagel which does teach that the minimum density criterion requires a particular percentage of the parameters in order to “find a good trade-off between clipping and rounding error.” (Nagel, page 10, paragraph 4) Regarding claim 18, Bordawekar teaches the system of claim 17, as cited above. Claim 18 additionally has the same limitations of claim 7 which are taught by Bordawekar and Nagel – see claim 7 above. Regarding claim 19, Bordawekar teaches the system of claim 12, as cited above. Claim 19 additionally has the same limitations of claim 8 which are taught by Bordawekar and Nagel – see claim 8 above . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Moshovos et al. (US20230334285) GadelRab et al. (US20210279635) Park et al. (US20210004663) Gajjala et al. ( Huffman Coding Based Encoding Techniques for Fast Distributed Deep Learning ) Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACQUELINE MEYER whose telephone number is (703)756-5676. The examiner can normally be reached M-F 8:00 am - 4:30 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571-272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.C.M./Examiner, Art Unit 2144 /TAMARA T KYLE/Supervisory Patent Examiner, Art Unit 2144 Application/Control Number: 18/504,016 Page 2 Art Unit: 2144 Application/Control Number: 18/504,016 Page 3 Art Unit: 2144 Application/Control Number: 18/504,016 Page 4 Art Unit: 2144 Application/Control Number: 18/504,016 Page 5 Art Unit: 2144 Application/Control Number: 18/504,016 Page 6 Art Unit: 2144 Application/Control Number: 18/504,016 Page 7 Art Unit: 2144 Application/Control Number: 18/504,016 Page 8 Art Unit: 2144 Application/Control Number: 18/504,016 Page 9 Art Unit: 2144 Application/Control Number: 18/504,016 Page 10 Art Unit: 2144 Application/Control Number: 18/504,016 Page 11 Art Unit: 2144 Application/Control Number: 18/504,016 Page 12 Art Unit: 2144 Application/Control Number: 18/504,016 Page 13 Art Unit: 2144
Read full office action

Prosecution Timeline

Nov 07, 2023
Application Filed
May 15, 2026
Non-Final Rejection mailed — §102, §103
Aug 06, 2026
Applicant Interview (Telephonic)
Aug 06, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12717933
METHOD AND SYSTEM FOR SECURING NEURAL NETWORK MODELS
4y 2m to grant Granted Aug 25, 2026
Patent 12718094
SYSTEM AND METHOD FOR CONTINUAL REFINABLE NETWORK
3y 8m to grant Granted Aug 25, 2026
Patent 12705498
SYSTEMS AND METHODS FOR FEDERATED VALIDATION OF MODELS
3y 8m to grant Granted Aug 11, 2026
Patent 12688412
TRAINING METHOD, STORAGE MEDIUM, AND TRAINING DEVICE
4y 4m to grant Granted Jul 21, 2026
Patent 12664421
MACHINE LEARNING MODELS WITH EFFICIENT FEATURE LEARNING
3y 11m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+57.1%)
3y 10m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 21 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month