Detailed Action
This action is written in response to the application filed 11 April 2024. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Subject Matter Eligibility
In determining whether the claims are subject matter eligible, the examiner has considered and applied guidance from MPEP § 2106. The examiner finds that the independent claims are directed to the practical application of improving the computational efficiency of a federated learning system by selectively reducing model precision for models distributed to some clients.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 6-10, 12, and 16-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Abdelmoniem.
Abdelmoniem, Ahmed M., and Marco Canini. "Towards mitigating device heterogeneity in federated learning via adaptive model quantization." Proceedings of the 1st Workshop on Machine Learning and Systems. 2021.
Regarding claim 1, Abdelmoniem discloses a method comprising:
iteratively training a neural network model comprising one or more weights in a reduced-precision format to produce a high-precision format neural network model by:
P. 4, algorithm 1, (reproduced below) “run SGD for E epochs”.
training the neural network comprising one or more reduced-precision format weights in a reduced-precision format compute unit; and
P. 4, sec. 3: “Adaptive Quantized Federated Learning In real-world end-device settings, different bit-widths are supported by different devices [19, 20, 23, 27, 30]. This hard ware flexibility allows for the ability of configuring the model with different quantization levels to match the variety of hardware configurations of clients’ mobile or IoT devices. Therefore, it is natural to rethink the design of the FL server and allow the flexibility of choosing for the clients among different quantized versions of the model. The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.”
See also algorithm 1, reproduced below, especially line 5, “decide the quantization level Qk for each client k”.
PNG
media_image1.png
374
476
media_image1.png
Greyscale
updating the converted high-precision format neural network model based on the training;
Also pp. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.)
converting the trained high-precision format neural network model to the reduced-precision format to produce a trained reduced-precision format neural network model; and
Algorithm 1 (reproduced above), line 5, “decide the quantization level Qk for each client k”.
sending the trained reduced-precision format neural network model to an aggregating system.
Algorithm 1 (reproduced above), line 6, “Server sends a customized quantized Qk(wt) to device k”.
Regarding claim 2, Abdelmoniem discloses the further limitation comprising:
receiving the neural network model comprising one or more weights received in a reduced-precision format from another device;
Algorithm 1, line 10 “send the updated model Qk(wt+1k)) to the server” and line 12 “Server updates the model”.
converting the received neural network model weights from the reduced-precision format to a high-precision format.
PP. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.)
Regarding claim 6, Abdelmoniem discloses the further limitation wherein the reduced-precision format comprises an FP8 format comprising an exponent and a mantissa, and the high-precision format comprises an FP32 or single-precision floating-point format.
The Examiner notes that the system described throughout Abdelmoniem is implemented in TensorFlow, which can accommodate the Python, JavaScript, Java and C++. Floating point number formats in each of these languages comprise a significand (mantissa) and an exponent.
PP. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.)
P. 7, “Post-quantization training methods cannot support less than 8-bit values”.
Regarding claim 7, Abdelmoniem discloses the further limitation wherein the neural network comprises part of a large language model or a machine vision model.
P. 2, table 1, illustrating tasks including “next work prediction” using LSTM model and “image classification” using a CNN model.
Regarding claim 8, Abdelmoniem discloses the further limitation wherein the method is performed on a plurality of federated learning devices, each operable to send results of training to the same aggregating system.
See algorithm 1, reproduced supra, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”.
Regarding claim 9, Abdelmoniem discloses a method comprising:
receiving a plurality of trained neural network models from a respective plurality of remote devices in a reduced-precision format; and
P. 4, sec. 3: “Adaptive Quantized Federated Learning In real-world end-device settings, different bit-widths are supported by different devices [19, 20, 23, 27, 30]. This hardware flexibility allows for the ability of configuring the model with different quantization levels to match the variety of hardware configurations of clients’ mobile or IoT devices. Therefore, it is natural to rethink the design of the FL server and allow the flexibility of choosing for the clients among different quantized versions of the model. The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.”
See also algorithm 1, reproduced below, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”.
PNG
media_image1.png
374
476
media_image1.png
Greyscale
aggregating the plurality of received trained neural network models to produce an aggregated trained neural network model in a high-precision format.
Id.
Also pp. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.)
Regarding claim 10, Abdelmoniem discloses the further limitation comprising:
receiving a neural network model comprising one or more weights in a high-precision format;
P. 4, sec. 3: “To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.”
converting the neural network model weights from the high-precision format to a reduced-precision format to produce a converted reduced-precision format neural network model; and
Id. “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round.”
See also algorithm 1, reproduced supra.
distributing the converted reduced-precision format neural network model to one or more remote devices for training using federated learning.
Id.
Regarding claim 12, Abdelmoniem discloses the further limitation comprising distributing the aggregated trained neural network in a high-precision format to one or more remote devices using a reduced-precision format.
See algorithm 1, reproduced supra, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”.
Regarding claim 16, Abdelmoniem discloses the further limitation wherein the reduced-precision format comprises an FP8 format comprising an exponent and a mantissa, and the high-precision format comprises an FP32 or single-precision floating-point format.
The Examiner notes that the system described throughout Abdelmoniem is implemented in TensorFlow, which can accommodate the Python, JavaScript, Java and C++. Floating point number formats in each of these languages comprise a significand (mantissa) and an exponent.
PP. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.)
P. 7, “Post-quantization training methods cannot support less than 8-bit values”.
Regarding claim 17, Abdelmoniem discloses the further limitation wherein the neural network comprises part of a large language model or a machine vision model.
P. 2, table 1, illustrating tasks including “next work prediction” using LSTM model and “image classification” using a CNN model.
Regarding claim 18, Abdelmoniem discloses a method, comprising:
transmitting a reduced-precision format of a neural network to one or more remote devices for training using federated learning over an electronic communication network;
P. 4, sec. 3: “Adaptive Quantized Federated Learning In real-world end-device settings, different bit-widths are supported by different devices [19, 20, 23, 27, 30]. This hard ware flexibility allows for the ability of configuring the model with different quantization levels to match the variety of hardware configurations of clients’ mobile or IoT devices. Therefore, it is natural to rethink the design of the FL server and allow the flexibility of choosing for the clients among different quantized versions of the model. The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” (Emphasis added.)
See also algorithm 1, reproduced below, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”.
PNG
media_image1.png
374
476
media_image1.png
Greyscale
receiving from the electronic communication network one or more trained neural networks from the one or more remote devices in a reduced-precision format; and
Algorithm 1, line 10 “send the updated model Qk(wt+1k)) to the server” and line 12 “Server updates the model”.
aggregating the received trained neural networks in a high-precision format to produce an aggregated trained neural network.
Id.
Also pp. 4-5, sec. 3.1, “Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” (Emphasis added.)
Regarding claim 19, Abdelmoniem discloses the further limitation further comprising training the sent reduced-precision format neural network on the one or more remote devices using a reduced-precision compute module on at least one of the one or more remote devices.
See algorithm 1, reproduced supra, especially line 5, “decide the quantization level Qk for each client k” and line 6, “Server sends a customized quantized Qk(wt) to device k”.
Regarding claim 20, Abdelmoniem discloses the further limitation comprising storing weights in a high-precision format on the one or more remote devices during training.
P. 4, sec. 3: “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round. To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.” (Emphasis added.)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
The following references are relied upon in the rejections below:
Abdelmoniem (Abdelmoniem, Ahmed M., and Marco Canini. "Towards mitigating device heterogeneity in federated learning via adaptive model quantization." Proceedings of the 1st Workshop on Machine Learning and Systems. 2021.)
Fang (US 2021/0133278 A1)
Haase (US 2022/0393986 A1)
Koker (US 2025/0292362 A1)
Nadamuni Raghavan (US 2024/0062042 A1)
Claims 3 is rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Haase.
Regarding claim 3, Haase discloses the further limitation which Abdelmoniem does not disclose comprising updating the reduced-precision format weights based on the updated converted high-precision format neural network model by rounding the converted high-precision format neural network weights to reduced-precision format weights using nearest neighbor rounding.
[0118] “The simplest quantization method rounds the neural network parameters t.sub.k 13 to the nearest reconstruction levels (also referred to as nearest neighbor quantization).”
At the time of filing, it would have been obvious to a skilled machine learning engineer apply the rounding technique disclosed by Haase with the federated learning system of Abdelmoniem because this would provide for more realistic quantization then simple numeric rounding because rounded values are guaranteed to be identical to another existing value. Both disclosures pertain to neural network quantization.
Claims 4 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Nadamuni Raghavan.
Regarding claim 4, Nadamuni Raghavan discloses the following further limitation which Abdelmoniem does not disclose wherein converting the trained high-precision format neural network model to the reduced-precision format comprises rounding the trained high-precision format neural network weights to reduced-precision format weights using unbiased stochastic quantization.
[0017] “FIG. 4 is a diagram illustrating unbiased stochastic quantization for nine bins, according to techniques of this disclosure.”
[0074] “It should be noted that the stochastic quantization scheme is unbiased for any v∈[b.sub.i, b.sub.k] because each p.sub.ij(v) is unbiased and by linearity of expectation. The linearity of expectation states that the expected value of a sum of random variables is equal to the sum of the expected values of the random variables. This means that the expected value of the quantized value is equal to the sum of the probabilities of each bin multiplied by the value of the continuous value in that bin.”
At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the stochastic quantization technique disclosed by Nadamuni Raghavan for the nearest neighbor quantization technique disclosed by Abdelmoniem because this would for provide for a low-sparsity sample with a high variance. See generally Nadamuni Raghavan [0072]. This substitution would yield predictable results.
Regarding claim 11, Nadamuni Raghavan discloses the following further limitation which Abdelmoniem does not disclose wherein converting the neural network model weights from the high-precision format to a reduced-precision format comprises using unbiased stochastic quantization.
[0017] “FIG. 4 is a diagram illustrating unbiased stochastic quantization for nine bins, according to techniques of this disclosure.”
[0074] “It should be noted that the stochastic quantization scheme is unbiased for any v∈[b.sub.i, b.sub.k] because each p.sub.ij(v) is unbiased and by linearity of expectation. The linearity of expectation states that the expected value of a sum of random variables is equal to the sum of the expected values of the random variables. This means that the expected value of the quantized value is equal to the sum of the probabilities of each bin multiplied by the value of the continuous value in that bin.”
At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the stochastic quantization technique disclosed by Nadamuni Raghavan for the nearest neighbor quantization technique disclosed by Abdelmoniem because this would for provide for a low-sparsity sample with a high variance. See generally Nadamuni Raghavan [0072]. This substitution would yield predictable results.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Koker.
Regarding claim 5, Koker discloses the following further limitation which Abdelmoniem does not disclose wherein the received neural network model further comprises one or more scale factors.
[0339] “The techniques described herein enable quantization via a hierarchical scale factors for data associated with neural network models to reduce the memory required for training and/or inference for those models while limiting the accuracy loss due to quantization.“
At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the quantization technique disclosed by Koker for the quantization technique used in the Abdelmoniem system because this would constitute the simple substation of one technique for a comparable one. Tuning model architectures and hyper-parameters—eg different quantization types—is a common practice when preparing complex neural networks.
Regarding claim 14, Koker discloses the following further limitation which Abdelmoniem does not disclose wherein aggregating the received trained neural network models further comprises mean squared error minimization of a scale factor.
[0339] “The techniques described herein enable quantization via a hierarchical scale factors for data associated with neural network models to reduce the memory required for training and/or inference for those models while limiting the accuracy loss due to quantization.“
The obviousness analysis of claim 5 applies equally here.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem and Fang.
Regarding claim 13, Fang discloses the following further limitation which Abdelmoniem does not disclose wherein aggregating the received trained neural network models comprises mean squared error minimization of weights of the received trained neural networks.
[0008] “Dividing the quantization range may include locating a breakpoint for the first region and the second region. Locating the breakpoint may include determining a quantization error over at least a portion of the quantization range. Locating the breakpoint may include substantially minimizing the quantization error. Minimizing the quantization error may include formulating the quantization error as a function of a location of the breakpoint, formulating a first derivative of the function, and determining a value of the breakpoint that results in the first derivative being substantially zero. The value of the breakpoint that results in the first derivative being substantially zero may be determined using a binary search. The location of the breakpoint may be approximated using a regression. The quantization error may be substantially minimized using a grid search.” (Emphasis added.)
At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the quantization technique disclosed by Fang for the quantization technique used in the Abdelmoniem system because this would constitute the simple substation of one technique for a comparable one. Tuning model architectures and hyper-parameters—eg different quantization types—is a common practice when preparing complex neural networks.
Claims 15 is rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem, Koker and Fang.
Regarding claim 15, Fang discloses the following further limitation which Abdelmoniem/Koker do not disclose wherein mean squared error minimization of a scale factor comprises performing a grid search of calculated errors using different scale factors.
[0008] “Dividing the quantization range may include locating a breakpoint for the first region and the second region. Locating the breakpoint may include determining a quantization error over at least a portion of the quantization range. Locating the breakpoint may include substantially minimizing the quantization error. Minimizing the quantization error may include formulating the quantization error as a function of a location of the breakpoint, formulating a first derivative of the function, and determining a value of the breakpoint that results in the first derivative being substantially zero. The value of the breakpoint that results in the first derivative being substantially zero may be determined using a binary search. The location of the breakpoint may be approximated using a regression. The quantization error may be substantially minimized using a grid search.” (Emphasis added.)
At the time of filing, it would have been obvious to a skilled machine learning engineer to substitute the quantization technique disclosed by Fang for the quantization technique used in the Abdelmoniem system because this would constitute the simple substation of one technique for a comparable one. Tuning model architectures and hyper-parameters—eg different quantization types—is a common practice when preparing complex neural networks.
Additional Relevant Prior Art
The following references were identified by the Examiner as being relevant to the disclosed invention, but are not relied upon in any rejection:
Burger discloses a neural network system featuring error-minimizing quantization. See eg col. 4, line 57 et seq. (US 11,645,493 B2)
Conclusion
Information regarding the status of an application may be found at the USPTO Patent Center at https://patentcenter.uspto.gov. Any inquiry concerning this communication should be directed to Vincent Gonzales at (571) 270-3837. The examiner can normally be reached Monday-Friday 7 am to 4 pm MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092.
/Vincent Gonzales/Primary Examiner, Art Unit 2124