Prosecution Insights
Last updated: October 02, 2026
Application No. 18/633,324

REDUCED PRECISION NEURAL FEDERATED LEARNING

Non-Final OA §102§103
Filed
Apr 11, 2024
Examiner
GONZALES, VINCENT
Art Unit
Tech Center
Assignee
ARM Limited
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
424 granted / 539 resolved
+18.7% vs TC avg
Moderate +11% lift
Without
With
+11.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
14 currently pending
Career history
557
Total Applications
across all art units

Statute-Specific Performance

§101
21.0%
-19.0% vs TC avg
§103
41.7%
+1.7% vs TC avg
§102
13.8%
-26.2% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 539 resolved cases

Office Action

§102 §103
Detailed Action This action is written in response to the application filed 11 April 2024. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Subject Matter Eligibility In determining whether the claims are subject matter eligible, the examiner has considered and applied guidance from MPEP § 2106. The examiner finds that the independent claims are directed to the practical application of improving the computational efficiency of a federated learning system by selectively reducing model precision for models distributed to some clients. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-2, 6-10, 12, and 16-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Abdelmoniem. Abdelmoniem, Ahmed M., and Marco Canini. "Towards mitigating device heterogeneity in federated learning via adaptive model quantization." Proceedings of the 1st Workshop on Machine Learning and Systems. 2021. Regarding claim 1, Abdelmoniem discloses a method comprising: iteratively training a neural network model comprising one or more weights in a reduced-precision format to produce a high-precision format neural network model by: P. 4, algorithm 1, (reproduced below) “run SGD for E epochs”. training the neural network comprising one or more reduced-precision format weights in a reduced-precision format compute unit; and P. 4, sec. 3: “Adaptive Quantized Federated Learning In real-world end-device settings, different bit-widths are supported by different devices [19, 20, 23, 27, 30]. This hard ware flexibility allows for the ability of configuring the model with different quantization levels to match the variety of hardware configurations of clients’ mobile or IoT devices. Therefore, it is natural to rethink the design of the FL server and allow the flexibility of choosing for the clients among different quantized versions of the model. The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” See also algorithm 1, reproduced below, especially line 5, “decide the quantization level Qk for each client k”. PNG media_image1.png 374 476 media_image1.png Greyscale updating the converted high-precision format neural network model based on the training; Also pp. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.) converting the trained high-precision format neural network model to the reduced-precision format to produce a trained reduced-precision format neural network model; and Algorithm 1 (reproduced above), line 5, “decide the quantization level Qk for each client k”. sending the trained reduced-precision format neural network model to an aggregating system. Algorithm 1 (reproduced above), line 6, “Server sends a customized quantized Qk(wt) to device k”. Regarding claim 2, Abdelmoniem discloses the further limitation comprising: receiving the neural network model comprising one or more weights received in a reduced-precision format from another device; Algorithm 1, line 10 “send the updated model Qk(wt+1k)) to the server” and line 12 “Server updates the model”. converting the received neural network model weights from the reduced-precision format to a high-precision format. PP. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.) Regarding claim 6, Abdelmoniem discloses the further limitation wherein the reduced-precision format comprises an FP8 format comprising an exponent and a mantissa, and the high-precision format comprises an FP32 or single-precision floating-point format. The Examiner notes that the system described throughout Abdelmoniem is implemented in TensorFlow, which can accommodate the Python, JavaScript, Java and C++. Floating point number formats in each of these languages comprise a significand (mantissa) and an exponent. PP. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.) P. 7, “Post-quantization training methods cannot support less than 8-bit values”. Regarding claim 7, Abdelmoniem discloses the further limitation wherein the neural network comprises part of a large language model or a machine vision model. P. 2, table 1, illustrating tasks including “next work prediction” using LSTM model and “image classification” using a CNN model. Regarding claim 8, Abdelmoniem discloses the further limitation wherein the method is performed on a plurality of federated learning devices, each operable to send results of training to the same aggregating system. See algorithm 1, reproduced supra, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”. Regarding claim 9, Abdelmoniem discloses a method comprising: receiving a plurality of trained neural network models from a respective plurality of remote devices in a reduced-precision format; and P. 4, sec. 3: “Adaptive Quantized Federated Learning In real-world end-device settings, different bit-widths are supported by different devices [19, 20, 23, 27, 30]. This hardware flexibility allows for the ability of configuring the model with different quantization levels to match the variety of hardware configurations of clients’ mobile or IoT devices. Therefore, it is natural to rethink the design of the FL server and allow the flexibility of choosing for the clients among different quantized versions of the model. The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” See also algorithm 1, reproduced below, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”. PNG media_image1.png 374 476 media_image1.png Greyscale aggregating the plurality of received trained neural network models to produce an aggregated trained neural network model in a high-precision format. Id. Also pp. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.) Regarding claim 10, Abdelmoniem discloses the further limitation comprising: receiving a neural network model comprising one or more weights in a high-precision format; P. 4, sec. 3: “To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” converting the neural network model weights from the high-precision format to a reduced-precision format to produce a converted reduced-precision format neural network model; and Id. “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round.” See also algorithm 1, reproduced supra. distributing the converted reduced-precision format neural network model to one or more remote devices for training using federated learning. Id. Regarding claim 12, Abdelmoniem discloses the further limitation comprising distributing the aggregated trained neural network in a high-precision format to one or more remote devices using a reduced-precision format. See algorithm 1, reproduced supra, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”. Regarding claim 16, Abdelmoniem discloses the further limitation wherein the reduced-precision format comprises an FP8 format comprising an exponent and a mantissa, and the high-precision format comprises an FP32 or single-precision floating-point format. The Examiner notes that the system described throughout Abdelmoniem is implemented in TensorFlow, which can accommodate the Python, JavaScript, Java and C++. Floating point number formats in each of these languages comprise a significand (mantissa) and an exponent. PP. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.) P. 7, “Post-quantization training methods cannot support less than 8-bit values”. Regarding claim 17, Abdelmoniem discloses the further limitation wherein the neural network comprises part of a large language model or a machine vision model. P. 2, table 1, illustrating tasks including “next work prediction” using LSTM model and “image classification” using a CNN model. Regarding claim 18, Abdelmoniem discloses a method, comprising: transmitting a reduced-precision format of a neural network to one or more remote devices for training using federated learning over an electronic communication network; P. 4, sec. 3: “Adaptive Quantized Federated Learning In real-world end-device settings, different bit-widths are supported by different devices [19, 20, 23, 27, 30]. This hard ware flexibility allows for the ability of configuring the model with different quantization levels to match the variety of hardware configurations of clients’ mobile or IoT devices. Therefore, it is natural to rethink the design of the FL server and allow the flexibility of choosing for the clients among different quantized versions of the model. The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” (Emphasis added.) See also algorithm 1, reproduced below, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”. PNG media_image1.png 374 476 media_image1.png Greyscale receiving from the electronic communication network one or more trained neural networks from the one or more remote devices in a reduced-precision format; and Algorithm 1, line 10 “send the updated model Qk(wt+1k)) to the server” and line 12 “Server updates the model”. aggregating the received trained neural networks in a high-precision format to produce an aggregated trained neural network. Id. Also pp. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.) Regarding claim 19, Abdelmoniem discloses the further limitation further comprising training the sent reduced-precision format neural network on the one or more remote devices using a reduced-precision compute module on at least one of the one or more remote devices. See algorithm 1, reproduced supra, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”. Regarding claim 20, Abdelmoniem discloses the further limitation comprising storing weights in a high-precision format on the one or more remote devices during training. P. 4, sec. 3: “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” (Emphasis added.) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. The following references are relied upon in the rejections below: Abdelmoniem (Abdelmoniem, Ahmed M., and Marco Canini. "Towards mitigating device heterogeneity in federated learning via adaptive model quantization." Proceedings of the 1st Workshop on Machine Learning and Systems. 2021.) Fang (US 2021/0133278 A1) Haase (US 2022/0393986 A1) Koker (US 2025/0292362 A1) Nadamuni Raghavan (US 2024/0062042 A1) Claims 3 is rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Haase. Regarding claim 3, Haase discloses the further limitation which Abdelmoniem does not disclose comprising updating the reduced-precision format weights based on the updated converted high-precision format neural network model by rounding the converted high-precision format neural network weights to reduced-precision format weights using nearest neighbor rounding. [0118] “The simplest quantization method rounds the neural network parameters t.sub.k 13 to the nearest reconstruction levels (also referred to as nearest neighbor quantization).” At the time of filing, it would have been obvious to a skilled machine learning engineer apply the rounding technique disclosed by Haase with the federated learning system of Abdelmoniem because this would provide for more realistic quantization then simple numeric rounding because rounded values are guaranteed to be identical to another existing value. Both disclosures pertain to neural network quantization. Claims 4 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Nadamuni Raghavan. Regarding claim 4, Nadamuni Raghavan discloses the following further limitation which Abdelmoniem does not disclose wherein converting the trained high-precision format neural network model to the reduced-precision format comprises rounding the trained high-precision format neural network weights to reduced-precision format weights using unbiased stochastic quantization. [0017] “FIG. 4 is a diagram illustrating unbiased stochastic quantization for nine bins, according to techniques of this disclosure.” [0074] “It should be noted that the stochastic quantization scheme is unbiased for any v∈[b.sub.i, b.sub.k] because each p.sub.ij(v) is unbiased and by linearity of expectation. The linearity of expectation states that the expected value of a sum of random variables is equal to the sum of the expected values of the random variables. This means that the expected value of the quantized value is equal to the sum of the probabilities of each bin multiplied by the value of the continuous value in that bin.” At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the stochastic quantization technique disclosed by Nadamuni Raghavan for the nearest neighbor quantization technique disclosed by Abdelmoniem because this would for provide for a low-sparsity sample with a high variance. See generally Nadamuni Raghavan [0072]. This substitution would yield predictable results. Regarding claim 11, Nadamuni Raghavan discloses the following further limitation which Abdelmoniem does not disclose wherein converting the neural network model weights from the high-precision format to a reduced-precision format comprises using unbiased stochastic quantization. [0017] “FIG. 4 is a diagram illustrating unbiased stochastic quantization for nine bins, according to techniques of this disclosure.” [0074] “It should be noted that the stochastic quantization scheme is unbiased for any v∈[b.sub.i, b.sub.k] because each p.sub.ij(v) is unbiased and by linearity of expectation. The linearity of expectation states that the expected value of a sum of random variables is equal to the sum of the expected values of the random variables. This means that the expected value of the quantized value is equal to the sum of the probabilities of each bin multiplied by the value of the continuous value in that bin.” At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the stochastic quantization technique disclosed by Nadamuni Raghavan for the nearest neighbor quantization technique disclosed by Abdelmoniem because this would for provide for a low-sparsity sample with a high variance. See generally Nadamuni Raghavan [0072]. This substitution would yield predictable results. Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Koker. Regarding claim 5, Koker discloses the following further limitation which Abdelmoniem does not disclose wherein the received neural network model further comprises one or more scale factors. [0339] “The techniques described herein enable quantization via a hierarchical scale factors for data associated with neural network models to reduce the memory required for training and/or inference for those models while limiting the accuracy loss due to quantization.“ At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the quantization technique disclosed by Koker for the quantization technique used in the Abdelmoniem system because this would constitute the simple substation of one technique for a comparable one. Tuning model architectures and hyper-parameters—eg different quantization types—is a common practice when preparing complex neural networks. Regarding claim 14, Koker discloses the following further limitation which Abdelmoniem does not disclose wherein aggregating the received trained neural network models further comprises mean squared error minimization of a scale factor. [0339] “The techniques described herein enable quantization via a hierarchical scale factors for data associated with neural network models to reduce the memory required for training and/or inference for those models while limiting the accuracy loss due to quantization.“ The obviousness analysis of claim 5 applies equally here. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Fang. Regarding claim 13, Fang discloses the following further limitation which Abdelmoniem does not disclose wherein aggregating the received trained neural network models comprises mean squared error minimization of weights of the received trained neural networks. [0008] “Dividing the quantization range may include locating a breakpoint for the first region and the second region. Locating the breakpoint may include determining a quantization error over at least a portion of the quantization range. Locating the breakpoint may include substantially minimizing the quantization error. Minimizing the quantization error may include formulating the quantization error as a function of a location of the breakpoint, formulating a first derivative of the function, and determining a value of the breakpoint that results in the first derivative being substantially zero. The value of the breakpoint that results in the first derivative being substantially zero may be determined using a binary search. The location of the breakpoint may be approximated using a regression. The quantization error may be substantially minimized using a grid search.” (Emphasis added.) At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the quantization technique disclosed by Fang for the quantization technique used in the Abdelmoniem system because this would constitute the simple substation of one technique for a comparable one. Tuning model architectures and hyper-parameters—eg different quantization types—is a common practice when preparing complex neural networks. Claims 15 is rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem, Koker and Fang. Regarding claim 15, Fang discloses the following further limitation which Abdelmoniem/Koker do not disclose wherein mean squared error minimization of a scale factor comprises performing a grid search of calculated errors using different scale factors. [0008] “Dividing the quantization range may include locating a breakpoint for the first region and the second region. Locating the breakpoint may include determining a quantization error over at least a portion of the quantization range. Locating the breakpoint may include substantially minimizing the quantization error. Minimizing the quantization error may include formulating the quantization error as a function of a location of the breakpoint, formulating a first derivative of the function, and determining a value of the breakpoint that results in the first derivative being substantially zero. The value of the breakpoint that results in the first derivative being substantially zero may be determined using a binary search. The location of the breakpoint may be approximated using a regression. The quantization error may be substantially minimized using a grid search.” (Emphasis added.) At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the quantization technique disclosed by Fang for the quantization technique used in the Abdelmoniem system because this would constitute the simple substation of one technique for a comparable one. Tuning model architectures and hyper-parameters—eg different quantization types—is a common practice when preparing complex neural networks. Additional Relevant Prior Art The following references were identified by the Examiner as being relevant to the disclosed invention, but are not relied upon in any rejection: Burger discloses a neural network system featuring error-minimizing quantization. See eg col. 4, line 57 et seq. (US 11,645,493 B2) Conclusion Information regarding the status of an application may be found at the USPTO Patent Center at https://patentcenter.uspto.gov. Any inquiry concerning this communication should be directed to Vincent Gonzales at (571) 270-3837. The examiner can normally be reached Monday-Friday 7 am to 4 pm MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092. /Vincent Gonzales/Primary Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Apr 11, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737626
JOINT INPUT PERTUBATION AND TEMPERATURE SCALING FOR NEURAL NETWORK CALIBRATION
3y 3m to grant Granted Sep 15, 2026
Patent 12705472
Prefetching Weights For Use In A Neural Network Processor
2y 9m to grant Granted Aug 11, 2026
Patent 12675989
FUSION MODEL TRAINING USING DISTANCE METRICS
2y 3m to grant Granted Jul 07, 2026
Patent 12651182
IDENTIFYING TRAITS OF PARTITIONED GROUP FROM IMBALANCED DATASET
4y 11m to grant Granted Jun 09, 2026
Patent 12639623
FAIR SELECTIVE CLASSIFICATION VIA A VARIATIONAL MUTUAL INFORMATION UPPER BOUND FOR IMPOSING SUFFICIENCY
4y 4m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
90%
With Interview (+11.2%)
3y 5m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 539 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month