DETAILED ACTION
Status of Claims
Claim(s) 1-20 are pending and are examined herein.
Claim(s) 1, 3, 6, 11, 13, 16, and 18 have been Amended.
Claim(s) 1-20 remain rejected under 35 U.S.C. § 103.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The amendment filed on August 03, 2026, has been entered. Claims 1-20 are pending in the application. Applicant’s amendments to the claims have been fully considered and are addressed in the rejections below.
Response to Arguments
Applicant's arguments with respect to the rejection under 35 U.S.C. § 103 filed on 08/03/2026 have been fully considered but they are not persuasive.
Applicant’s argument (Pp. 7-8 of the remark):
Applicant argues that Abdelmoniem fails to disclose the claim step as currently recited “quantizing, by a server, a global model, the global model being quantized at multiple different quantization levels according to a respective quantization level of each of one or more subnetwork models to generate one or more quantized subnetwork models.” Further, Applicant argues that Diao does not describe "quantizing, by a server, a global model, the global model being quantized at multiple different quantization levels according to a respective quantization level of each of one or more subnetwork models to generate one or more quantized subnetwork models," as recited in claim 1.
Examiner's response:
The examiner respectfully disagrees and finds the arguments unpersuasive because the rejection relies on the combined teachings of Abdelmoniem and Diao, rather than requiring either reference individually.
Abdelmoniem teaches the server-side quantization operation and the association of a respective quantization level with each client. In particular, Abdelmoniem discloses that the server maintains a single version of the model in floating-point format and quantizes the model during the configuration phase of each round. Abdelmoniem further discloses that the server uses multiple quantization levels, including 16, 8, 4, 3, 2-bits levels, from which it makes a pre-client decision and sends the model version that best matches the capabilities of the selected clients in each round. Algorithm 1 describes that the server determines the quantization level
Q
k
for each client k and sends a corresponding custom quantized
Q
k
(
w
t
)
to device
k
, where
Q
k
represents the pre-client (i.e., respective) quantization level selected for each generated quantized model. Thus, Abdelmoniem associates a respective quantization level with each client and provides the corresponding quantized model to that client. Specifically, Abdelmoniem teaches that the pre-client quantization precision
Q
k
is determined using the computational profile of the client, such that the quantized model provided to the client corresponds to the client’s computational capabilities. (See, page 4, Sections: 3 and 3.1; page 6, Section: 4.2).
Furthermore, Abdelmoniem also states that “the FL server could send several bit width versions of the quantized model to the client, and the client caches them locally.”
Diao teaches the subnetwork models aspect. Specifically, Diao explicitly teaches adaptively distributing subnetworks according to client capabilities and allocating subsets of global-model parameters according to the corresponding capabilities of local clients. Diao also describes multiple computation-complexity levels for the local models. Therefore, Diao teaches that different devices having capabilities may be provided by the server with different client-specific subnetworks corresponding to their respective computation and communication capabilities.
Accordingly, the combined teachings of Abdelmoniem and Diao provide the claimed limitations. Diao teaches the client-specific subnetwork models structure and the distributing of the subnetwork models to the corresponding local client, while Abdelmoniem teaches the server-side quantization of the global model at multiple quantization levels based on the selection of a respective quantization level for each client according to the client’s computational capabilities, and sending the quantized version to the respective client.
Applicant’s assertion that Diao does not disclose quantizing the subnetworks is not persuasive. The rejection relies on Abdelmoniem for the server-side quantization operation. It is noted that the rejection is based on the combined teachings of references, nonobviousness cannot be established by attacking each reference individually. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Accordingly, Applicant’s arguments are not persuasive and the rejection under U.S.C. § 103 35 based on is maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 3, 5-8, 10-11, 13, 15-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem et al., (NPL: "Towards mitigating device heterogeneity in federated learning via adaptive model quantization." (2021)) in view of Diao et al., (NPL: “Heterofl: Computation and communication efficient federated learning for heterogeneous clients.” (2021)), hereinafter, Abdelmoniem in view of Diao.
Regarding Currently Amended Claim 1,
Abdelmoniem discloses the following:
A processor-implemented method performed by one or more processors, the processor-implemented method comprising: (Abdelmoniem, [Abstract] “We propose AQFL, a simple and practical approach leveraging adaptive model quantization to homogenize the computing resources of the clients.” [p. 5, Section: 4] “Platform: We run experiments using the HeterFL simulation environment [5] on a GPU cluster. HeterFL is an extension of Flash [39], which simulates wall-clock accurate execution times of FL tasks... Implementation We implement quantization using the built-in quantization API of TensorFlow [35, 36].”)
quantizing, by a server, a global model, the global model being quantized at multiple different quantization levels according to a respective quantization level of each of one or more subnetwork models to generate one or more quantized subnetwork models, (Abdelmoniem, [P. 4, Section: 3] “To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round... the FL server may use Quantizable DNNs, which are a special type of quantized DNNs that can flexibly adjust the bit-width on the fly, i.e., turning on different bit mode by applying different quantization levels [13].” [Pp. 4-5, Section: 3, Algorithm 1] “3: K online clients check-in with the server. 4: Server selects, at random, a subset St of K devices 5: Server collects information from selected clients to decide the quantization level Qk for each client k...” [P. 5, Section: 3.1] “Dynamic client environment: how can the FL server adapt its quantization decision for each device when device workload and network conditions vary? … To resolve this, the server could try to learn and model the environment of the clients to predict the optimal quantization level. For instance, the server could employ online reinforcement-learning to build an adaptive decision model for the quantization level… For instance, the FL server could send several bit width versions of the quantized model to the client and the client caches them locally.” [p. 6, Section: 4.2] “We use 5 commonly used quantization levels (i.e., 16, 8, 4, 3, 2-bits) [13, 20].”) the one or more subnetwork models being assigned to one or more of multiple devices according to device processing capabilities; (Abdelmoniem, [p. 4, Section: 3] “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round... The server, which maintains computational profile(s) of the dummy model in different precision for reference device(s), can map the client profiles to a reference profile that will likely allow the device to finish training and send updates within the deadline.” [p. 6, Section: 4.3] “AQFL-enabled server picks the right-sized quantized model based on the device’s computational class, which reduces the training run time for slow devices.”) [Examiner’s Note: The proposed AQFL teaches quantizing a global model using different quantization levels and providing per-client quantized versions. Under the BRI, the quantized versions of the global model read on the quantized subnetwork models. The quantization level and the resulting quantized model version is identified to match the capabilities of each client device.]
distributing, by the server, to at least one device of the multiple devices, a quantized subnetwork model; (Abdelmoniem, [p. 3, Section: 2.1] “the server proceeds with sending an FL task and a checkpoint of the global model to each client. Then, the clients start executing the FL task on their local datasets.” [p. 4, Section: 3] “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round... [Algorithm 1] 6: Server sends a custom quantized Qk(wt) to device k parfor k = 0,...,K 1 do.” [p.4, Section: 3.1] “Then, during the configuration phase, the server sends the custom quantized model Qk(wt) to each client selected to participate in the training. Then, the clients train using the quantized model and upload the model updates to the server.”)
receiving, by the server, a model update from the at least one device based on local data; (Abdelmoniem, [p. 3, Section: 2.1] “the server proceeds with sending an FL task and a checkpoint of the global model to each client. Then, the clients start executing the FL task on their local datasets.” [Algorithm 1] 8: run SGD for E epochs on Fk with step-size a new version of the model is produced Qk(wt+1 k) send the updated model Qk(wt+1 k )) to the server 11: end for end parfor...” [p.4, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server.”) and
generating, by the server, an updated global model according to an aggregation function based on the model update from each of the at least one device. (Abdelmoniem, [p. 3, Section: 2.1] “Then, the server aggregates the updates received before the deadline using the configured aggregation mechanism. This stage and hence the round is deemed successful when the server updates the global model.” [Algorithm 1] 8: Server updates the model:
w
t
+
1
=
1
K
∑
k
∈
S
t
Q
k
'
(
w
k
t
+
1
)
.
” [Pp.4-5, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server. Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.”)
As outlined above, Abdelmoniem teaches the server-side quantization of the global model at multiple quantization levels based on the selection of a respective quantization level for each client according to the client’s computational capabilities, and sending the quantized version to the respective client. Abdelmoniem does not explicitly define the “one or more subnetwork models” generated by the server. However, it would have been obvious in view of Diao.
Hereinafter, Abdelmoniem in view of Diao teaches the limitations:
each of one or more subnetwork models to generate one or more quantized subnetwork models, the one or more subnetwork models being assigned to one or more of multiple devices according to device processing capabilities; (Diao, [Pp. 1-2, Section: 1] “In this work, we propose a new federated learning framework called HeteroFL to train heterogeneous local models with varying computation complexities and still produce a single global inference model... It is natural to adaptively distribute subnetworks according to clients’ capabilities.” [Pp. 3-4, Section: 3.1] “It is possible to have multiple computation complexity levels
W
l
-
p
⊂
W
l
p
-
1
·
·
·
⊂
W
l
-
1
as illustrated in Fig. 1... With this construction, we can adaptively allocate subsets of global model parameters according to the corresponding capabilities of local clients... The intuition is thus to perform global aggregation across all local models, at least on one subnetwork. To stabilize global model aggregation, we also allocate a fixed subnetwork for every computation complexity level. Our proposed inclusive subsets of global model parameters also guarantee that smaller local models will aggregate with more local models... We empirically found that this approach produces better results than uniformly sampled subnetworks for each client or computation complexity level.” [p. 4, Section: 3.3] “Because we need to optimize local models for multiple epochs, local model parameters at different computation complexity levels will digress to various scales.” [P. 6, Section: 4] “To study the effectiveness of our proposed HeteroFL framework, we construct five different computation complexity levels {a,b,c,d,e} with the hidden channel shrinkage ratio r = 0.5.”)
Abdelmoniem and Diao are from the same field of endeavor, and their disclosure generally relates to (Heterogeneous Federated Learning).
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skill in the art to modify the combination of Abdelmoniem and Diao to incorporate the proposed HeteroFL framework as taught by Diao. One would have been motivated to make such a combination in order to adaptively distribute subnetworks according to clients’ capabilities. Doing so would provide efficiency in both computation and communication (Diao [Abstract]).
Regarding Currently Amended Claim 3, Abdelmoniem in view of Diao teaches the elements of claim 1 as outlined above, and further teaches:
further comprising quantizing, by the server, the updated global model at the multiple different quantization levels according to the respective quantization level of each of the one or more subnetwork models to generate one or more quantized updated subnetwork models. (Abdelmoniem, [Pp. 4-5, Section: 3, Algorithm 1] “2: for t = 0,...,T 1 do 3: K online clients check-in with the server. 4: Server selects, at random, a subset St of K devices 5: Server collects information from selected clients to decide the quantization level Qk for each client k. 6: Server sends a custom quantized Qk(wt) to device k... 8: Server updates the model:
w
t
+
1
=
1
K
∑
k
∈
S
t
Q
k
'
(
w
k
t
+
1
)
.
” [Pp. 4-5, Section: 3] “the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.”) [Examiner’s Note: the AQFL framework provides iterative loop where the server re-quantize the updated global model at multiple different quantization levels (e.g., 16, 8, 4, 3, and 2 bits) during each round. This reads on the claimed quantizing the updated global model to generate quantized updated subnetwork models.]
Regarding Original Claim 5, Abdelmoniem in view of Diao teaches the elements of claim 1 as outlined above, and further teaches:
in which the model update from the at least one device is a quantized model update based on a quantization by the at least one device. (Abdelmoniem, [P. 4, Section: 3, Algorithm 1] “8: run SGD for E epochs on Fk with step-size a new version of the model is produced Qk(wt+1 k) send the updated model Qk(wt+1 k)) to the server 11: end for end parfor...” [Pp.4-5, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server. Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.”) [Examiner’s Note: The AQFL algorithm shows that each client device applying its quantization to produce a quantized update to be sent to the server.]
Regarding Currently Amended Claim 6, Abdelmoniem discloses the following:
receiving, by a device, from a server, a quantized subnetwork model, (Abdelmoniem, [p. 3, Section: 2.1] “the server proceeds with sending an FL task and a checkpoint of the global model to each client. Then, the clients start executing the FL task on their local datasets.” [Algorithm 1] 6: Server sends a custom quantized Qk(wt) to device k parfor k = 0,...,K 1 do.”) the quantized subnetwork model corresponding to a global model, (Abdelmoniem, [p. 2, Section: 2] “The FL server performs secure aggregation of the local models pushed by the clients. 5.The FL server updates the state of the global model.” [p. 4, Section: 3] “the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round.”) the global model being quantized at multiple different quantization levels according to a respective quantization level of each of one or more subnetwork models to generate one or more quantized subnetwork models, the one or more subnetwork models being assigned to one or more of multiple devices according to device processing capabilities; (Abdelmoniem, Abdelmoniem, [P. 4, Section: 3] “To this end, the server would maintain a single version of the model in floating-point format and quantize the model during the configuration phase of each round... the FL server may use Quantizable DNNs, which are a special type of quantized DNNs that can flexibly adjust the bit-width on the fly, i.e., turning on different bit mode by applying different quantization levels [13].” [Pp. 4-5, Section: 3, Algorithm 1] “3: K online clients check-in with the server. 4: Server selects, at random, a subset St of K devices 5: Server collects information from selected clients to decide the quantization level Qk for each client k...” [P. 5, Section: 3.1] “Dynamic client environment: how can the FL server adapt its quantization decision for each device when device workload and network conditions vary? … To resolve this, the server could try to learn and model the environment of the clients to predict the optimal quantization level. For instance, the server could employ online reinforcement-learning to build an adaptive decision model for the quantization level… For instance, the FL server could send several bit width versions of the quantized model to the client and the client caches them locally.” [p. 6, Section: 4.2] “We use 5 commonly used quantization levels (i.e., 16, 8, 4, 3, 2-bits) [13, 20].” [p. 4, Section: 3] “The server can then make a per-client decision and send the model version that best matches the capabilities of the selected clients in each round... The server, which maintains computational profile(s) of the dummy model in different precision for reference device(s), can map the client profiles to a reference profile that will likely allow the device to finish training and send updates within the deadline.” [p. 6, Section: 4.3] “AQFL-enabled server picks the right-sized quantized model based on the device’s computational class, which reduces the training run time for slow devices.”) [Examiner’s Note: The proposed AQFL teaches quantizing a global model using different quantization levels and providing per-client quantized versions. Under the BRI, the quantized versions of the global model read on the quantized subnetwork models. The quantization level and the resulting quantized model version is identified to match the capabilities of each client device.]
generating, by the device, a model update for the quantized subnetwork model based on local data; (Abdelmoniem, [p. 3, Section: 2.1] “the server proceeds with sending an FL task and a checkpoint of the global model to each client. Then, the clients start executing the FL task on their local datasets.” [Algorithm 1] 8: run SGD for E epochs on Fk with step-size a new version of the model is produced Qk(wt+1 k) send the updated model Qk(wt+1 k )) to the server 11: end for end parfor...” [p.4, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server.” [p. 2, Section: 2] “1.The clients check-in with the FL server, and then the server typically selects a sample of clients for training and pushes a copy of the up-to-date global model. 2. The clients perform an equal number of local optimization steps as determined by task designer.”) and
transmitting, by the device, the model update to the server, the server generating an updated global model based on the model update. (Abdelmoniem, [p. 3, Section: 2.1] “Then, the server aggregates the updates received before the deadline using the configured aggregation mechanism. This stage and hence the round is deemed successful when the server updates the global model.” [Algorithm 1] 8: Server updates the model:
w
t
+
1
=
1
K
∑
k
∈
S
t
Q
k
'
(
w
k
t
+
1
)
.
” [Pp.4-5, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server. Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” [p. 2, Section: 2] “3. Clients encrypt all or a subset of the updated local model parameters with encryption[32], differential privacy[3], or secret sharing [8] techniques and push the encrypted model to the FL server.”)
As outlined above, Abdelmoniem teaches the server-side quantization of the global model at multiple quantization levels based on the selection of a respective quantization level for each client according to the client’s computational capabilities, and sending the quantized version to the respective client. Abdelmoniem does not explicitly define the “one or more subnetwork models” generated by the server. However, it would have been obvious in view of Diao.
Hereinafter, Abdelmoniem in view of Diao teaches the limitations:
each of one or more subnetwork models to generate one or more quantized subnetwork models, the one or more subnetwork models being assigned to one or more of multiple devices according to device processing capabilities; (Diao, [Pp. 1-2, Section: 1] “In this work, we propose a new federated learning framework called HeteroFL to train heterogeneous local models with varying computation complexities and still produce a single global inference model... It is natural to adaptively distribute subnetworks according to clients’ capabilities.” [Pp. 3-4, Section: 3.1] “It is possible to have multiple computation complexity levels
W
l
-
p
⊂
W
l
p
-
1
·
·
·
⊂
W
l
-
1
as illustrated in Fig. 1... With this construction, we can adaptively allocate subsets of global model parameters according to the corresponding capabilities of local clients... The intuition is thus to perform global aggregation across all local models, at least on one subnetwork. To stabilize global model aggregation, we also allocate a fixed subnetwork for every computation complexity level. Our proposed inclusive subsets of global model parameters also guarantee that smaller local models will aggregate with more local models... We empirically found that this approach produces better results than uniformly sampled subnetworks for each client or computation complexity level.” [p. 4, Section: 3.3] “Because we need to optimize local models for multiple epochs, local model parameters at different computation complexity levels will digress to various scales.” [P. 6, Section: 4] “To study the effectiveness of our proposed HeteroFL framework, we construct five different computation complexity levels {a,b,c,d,e} with the hidden channel shrinkage ratio r = 0.5.”)
Abdelmoniem and Diao are from the same field of endeavor and their disclosure generally relates to (Heterogeneous Federated Learning).
Accordingly, at the effective filing date, it would have been prima facie obvious to one ordinarily skill in the art to modify the combination of Abdelmoniem and Diao to incorporate the proposed HeteroFL framework as taught by Diao. One would have been motivated to make such a combination in order to adaptively distribute subnetworks according to clients’ capabilities. Doing so would provide efficiency in both computation and communication (Diao [Abstract]).
Regarding Original Claim 7, Abdelmoniem in view of Diao teaches the elements of claim 6 as outlined above, and further teaches:
in which the updated global model is generated using an aggregation function based on the model update by the device. (Abdelmoniem, [P. 2, Section: 2] “4.The FL server performs secure aggregation of the local models pushed by the clients. 5. The FL server updates the state of the global model.” [p. 3, Section: 2.1] “Then, the server aggregates the updates received before the deadline using the configured aggregation mechanism. This stage and hence the round is deemed successful when the server updates the global model.” [Algorithm 1] 8: Server updates the model:
w
t
+
1
=
1
K
∑
k
∈
S
t
Q
k
'
(
w
k
t
+
1
)
.
” [Pp.4-5, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server. Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.” Further see Diao Section: 3.1, Equations 1-3.).
Regarding Original Claim 8, Abdelmoniem in view of Diao teaches the elements of claim 6 as outlined above, and further teaches:
further comprising quantizing, by the device, the model update to generate a quantized model update, the quantized model update being used to generate the updated global model. (Abdelmoniem, [Algorithm 1] 8: run SGD for E epochs on Fk with step-size a new version of the model is produced Qk(wt+1 k) send the updated model Qk(wt+1 k )) to the server 11: end for end parfor...” [Pp.4-5, Section: 3.1] “Then, the clients train using the quantized model and upload the model updates to the server. Finally, the server de-quantizes the per-client model updates and aggregates them to update the global model, which is typically stored in floating-point precision.”)
Regarding Original Claim 10, Abdelmoniem in view of Diao teaches the elements of claim 8 as outlined above, and further teaches:
further comprising repeating the receiving, the generating, and the transmitting for multiple training rounds; and in which the quantizing is performed in a subset of the training rounds based on the device processing capabilities. (Abdelmoniem, [p. 3, Section: 2] “The central server, which performs aggregation and hosts the global model, invokes the above process frequently and continues so until the model accuracy or loss function converges to a target value... FL tasks are broken into rounds and a round is concluded when certain stages complete successfully during which clients are expected to stay connected throughout their duration. In each round, the server drives the stages towards the completion of the round. The main stages in FL systems are selection, configuration, and reporting[7]... The server waits until a time limit or deadline, for the participating clients to upload their model updates... the round fails and the received updates are ignored, in which case the round is restarted from scratch.” [Pp. 4-5, Section: 3] “the system could rely on client-side mechanisms to make the decisions which provides higher quality decisions... the FL server could send several bit width versions of the quantized model to the client and the client caches them locally.” [Algorithm 1] “2: for t = 0,...,T 1 do 3: K online clients check-in with the server. 4: Server selects, at random, a subset St of K devices 5: Server collects information from selected clients to decide the quantization level Qk for each client k. 6: Server sends a custom quantized Qk(wt) to device k... 8: Server updates the model:
w
t
+
1
=
1
K
∑
k
∈
S
t
Q
k
'
(
w
k
t
+
1
)
.
”)
Regarding Currently Amended Claim 11,
The claim recites substantially similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 1 is directed to a processor-implemented method, and claim 11 is directed to an apparatus.
Regarding Currently Amended Claim 13,
The claim recites substantially similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale.
Regarding Original Claim 15,
The claim recites substantially similar limitations as corresponding claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale.
Regarding Currently Amended Claim 16,
The claim recites substantially similar limitations as corresponding claim 6 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 6 is directed to a processor-implemented method, and claim 16 is directed to an apparatus.
Regarding Original Claim 17,
The claim recites substantially similar limitations as corresponding claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale.
Regarding Currently Amended Claim 18,
The claim recites substantially similar limitations as corresponding claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale.
Regarding Original Claim 20,
The claim recites substantially similar limitations as corresponding claim 10 and is rejected for similar reasons as claim 10 using similar teachings and rationale.
Claim(s) 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem in view of Diao as outlined above and further in view of Wang et al., (Pub. No.: US 20240303506 A1) is entitled to U.S. Provisional Application filed Mar. 10, 2023.
Regarding Original Claim 2, Abdelmoniem in view of Diao teaches the elements of claim 1 as outlined above:
While Abdelmoniem in view of Diao teaches the iterative federated training process and global model update, Abdelmoniem in view of Diao does not appear to explicitly teach:
further comprising fine-tuning the updated global model using public data of a backup device.
However, Wang, in combination with Abdelmoniem and Diao, teaches the limitation:
further comprising fine-tuning the updated global model using public data of a backup device. (Wang, [0082] “model update engine 210 uses supervised training data 214 on server 204 to perform one or more rounds of supervised training updates 226 of that global model 222 after federated self-supervised training of global model 222 is complete... training data 214 could be stored in database 230 on server 204... Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214.” [0103] “After updating of the global version is complete, model update engine 210 performs step 412, in which model update engine 210 fine tunes one or more global versions of the machine learning model using a supervised training dataset.” [0013] “the ability to perform self-supervised training of the machine learning model using a relatively large quantity of unlabeled training data at the clients before performing fine-tuning of the machine learning model at the server using a relatively small quantity of labeled training data.”) [Examiner’s Note: the claimed “public data” broadly encompasses the server-side training data that is used to fine-tuned the global aggregated model.]
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Abdelmoniem and Diao to incorporate the federated self-supervised learning techniques as taught by Wang. One would have been motivated to make such a combination in order to improve the ability of the machine learning model to generalize to different types of environments or tasks at the clients without compromising the privacy of the data at the clients (Wang [0013]).
Regarding Original Claim 12,
The claim recites substantially similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale.
Claim(s) 4, 9, 14, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Abdelmoniem in view of Diao as outlined above and further in view of Kaur et al., (NPL: “RS-FAIRFRS: Communication Efficient Fair Federated Recommender System.” (February 2023)).
Regarding Original Claim 4, Abdelmoniem in view of Diao teaches the elements of claim 1 as outlined above:
Abdelmoniem in view of Diao does not appear to explicitly teach:
applying a regularization process to the updated global model to reduce a bias toward server data.
However, Kaur, in combination with Abdelmoniem and Diao, teaches the limitation:
applying a regularization process to the updated global model to reduce a bias toward server data. (Kaur, [Pp. 5-6, Section: 4.2] “We then propose fairness metric which when added as a constraint in optimization function at server helps achieve a fair global model. Furthermore, we discuss a two phase mechanism which helps in achieving global as well as local fairness... FairMF trains the data at server
D
s
e
r
v
e
r
for obtaining global fairness objective. The goal of FairMF is to optimize the loss function defined as the combination of regularized MF (equation 1) and fairness penalty (equation 2). The final loss function is
min
U
,
V
L
M
F
+
⋋
f
L
a
p
(3) The hyperparameter
⋋
f
acts as a fairness penalizer... Server runs FairMF for some iterations
t
s
and obtains final
U
f
a
i
r
and
V
f
a
i
r
. Finally,
V
f
a
i
r
and
V
i
i
=
1
m
(aggregated item vectors) are communicated to all the clients. We provide the exact procedure for FairMF in Algorithm 1.” [Pp. 6-7, Section: 4.3] “The assumption of
D
s
e
r
v
e
r
helps in obtaining fair item vectors
V
f
a
i
r
which are communicated to all the clients for local fairness. Alongwith these,
V
is also sent to each client to retain the federated properties of FedRec and allow benefits of the participation of other clients. For initial round, the initialized item vectors are communicated, however as the training proceeds,
V
gets updated by the aggregated item vectors. Further, the server aggregates item vectors sent by
C
τ
clients only. The fair item vectors
V
f
a
i
r
and aggregated item vectors communicated by the server are received by each client
u
and local training happen at all the clients.... The updated item gradients are uploaded to the server to train FairMF on the
D
s
e
r
v
e
r
and obtains
U
f
a
i
r
and
V
f
a
i
r
. This procedure is repeated till convergence.” [P. 9, Section: 5.2] “Fairness of RS-FAIRFRS: To analyze fairness of RS-FAIRFRS,.. Our results in Appendix show that even with only 25% and 50% of the items at the server, our model is able to reduce bias.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Abdelmoniem, Diao, and Kaur, to incorporate the RS-FAIRFRS techniques for federated learning dual-fair update as taught by Kaur. One would have been motivated to make such a combination in order to reduce communication cost and demographic bias with improved model accuracy (Kaur [Conclusion]).
Regarding Original Claim 9, Abdelmoniem in view of Diao teaches the elements of claim 8 as outlined above:
Abdelmoniem in view of Diao does not appear to explicitly teach:
applying a first regularization process to the updated global model to reduce a first bias toward server data or a second regularization process to the quantized model update to reduce a second bias toward device data.
However, Kaur, in combination with Abdelmoniem and Diao, teaches the limitation:
applying a first regularization process to the updated global model to reduce a first bias toward server data or a second regularization process to the quantized model update to reduce a second bias toward device data. (Kaur, [Pp. 5-6, Section: 4.2] “We then propose fairness metric which when added as a constraint in optimization function at server helps achieve a fair global model. Furthermore, we discuss a two phase mechanism which helps in achieving global as well as local fairness... FairMF trains the data at server
D
s
e
r
v
e
r
for obtaining global fairness objective. The goal of FairMF is to optimize the loss function defined as the combination of regularized MF (equation 1) and fairness penalty (equation 2). The final loss function is
min
U
,
V
L
M
F
+
⋋
f
L
a
p
(3) The hyperparameter
⋋
f
acts as a fairness penalizer... Server runs FairMF for some iterations
t
s
and obtains final
U
f
a
i
r
and
V
f
a
i
r
. Finally,
V
f
a
i
r
and
V
i
i
=
1
m
(aggregated item vectors) are communicated to all the clients. We provide the exact procedure for FairMF in Algorithm 1.” [Pp. 6-7, Section: 4.3] “The assumption of
D
s
e
r
v
e
r
helps in obtaining fair item vectors
V
f
a
i
r
which are communicated to all the clients for local fairness. Alongwith these,
V
is also sent to each client to retain the federated properties of FedRec and allow benefits of the participation of other clients. For initial round, the initialized item vectors are communicated, however as the training proceeds,
V
gets updated by the aggregated item vectors. Further, the server aggregates item vectors sent by
C
τ
clients only. The fair item vectors
V
f
a
i
r
and aggregated item vectors communicated by the server are received by each client
u
and local training happen at all the clients.... The updated item gradients are uploaded to the server to train FairMF on the
D
s
e
r
v
e
r
and obtains
U
f
a
i
r
and
V
f
a
i
r
. This procedure is repeated till convergence.” [P. 9, Section: 5.2] “Fairness of RS-FAIRFRS: To analyze fairness of RS-FAIRFRS,.. Our results in Appendix show that even with only 25% and 50% of the items at the server, our model is able to reduce bias.”)
Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, having the combination of Abdelmoniem, Diao, and Kaur, to incorporate the RS-FAIRFRS techniques for federated learning dual-fair update as taught by Kaur. One would have been motivated to make such a combination in order to reduce communication cost and demographic bias with improved model accuracy (Kaur [Conclusion]).
Regarding Original Claim 14,
The claim recites substantially similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale.
Regarding Original Claim 19,
The claim recites substantially similar limitations as corresponding claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
(Pub. No.: US 20250356176 A1) – “Tejas SUBRAMANYA” relates to “Quantized federated learning.”
[Abstract] “Method, comprising: receiving an indication of one or more supported bit-widths for local learning by a first node among plural nodes; generating a respective quantized version of a model for at least one of the supported bit-widths; providing the generated respective quantized versions of the model for the at least one of the supported bit-widths or a link to location from where the first node may download the at least one quantized version of the model for the at least one of the supported bit-widths to the first node.”
(Pub. No.: US 20220245527 A1) – “Chang-Sik Choi” relates to “Techniques for adaptive quantization level selection in federated learning.”
[0088]–[0112] “The process flow 300 may further be implemented by the server 305 and the worker 310 to potentially reduce latency associated with training a machine learning model (e.g., a global model shared by the server 305 and the worker 310) using federated learning techniques and ensuring global convergence of the machine learning model (e.g., based on selecting quantization levels based on channel conditions between the server 305 and the worker 310), among other benefits…. In some examples, the indication may identify a subset of quantization levels of the set of quantization levels, for example, if the server 305 determined a set of quantization levels at 334, where the subset of quantization levels corresponds to the determined set of quantization levels.”
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SADIK ALSHAHARI whose telephone number is (703)756-4749. The examiner can normally be reached Monday Friday, 9 A.M - 6 P.M. ET..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached on (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.A.A./Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121