Prosecution Insights
Last updated: October 01, 2026
Application No. 18/613,765

METHODS AND SYSTEMS FOR IMPROVED FEDERATED LEARNING AND IMPLEMENTATIONS THEREOF

Non-Final OA §103
Filed
Mar 22, 2024
Priority
Mar 24, 2023 — provisional 63/491,970
Examiner
MESFIN, MATTHEWOS
Art Unit
Tech Center
Assignee
Rutgers, The State University of New Jersey
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
8 currently pending
Career history
5
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 Claims 1, 4, 6-13 are rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022) and Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020). Regarding claim 1, Qi teaches a method for heterogeneous federated learning through two-way knowledge distillation (Page 189, Abstract, “we propose a heterogenous Federated learning framework based on Bidirectional Knowledge Distillation (FedBKD)”), comprising: configuring1 … a local variational autoencoder (VAE) (Pages 193-194, Equations 4-7) of each user device (Page 195, Figure 22) training … the local variational autoencoder of each user device using local data in each user device (Page 194, Column 2, Paragraph 2, “In this phase, each client trains a CVAE using its private data”) transmitting to a server device… trained local variational autoencoders from a plurality of user devices… (Page 194, Column 2, Paragraph 2, “…and uploads the trained CVAE to the cloud server”) performing backward knowledge distillation to distill knowledge in the merged… model to the local… model of each user device (Page 190, Column 2, Paragraph 1, “…the cloud-to-client distillation is the single-teacher-multiple students process, which distills the knowledge from the single global model back to multiple heterogeneous local networks”) using data generated by the… variational autoencoder (Page 194, Column 2, “To alleviate the problem of data heterogeneity, we used CVAE to generate a synthetic dataset in the cloud server”) to obtain an updated local… model for each user device.3 Qi fails to teach configuring a local complex model of each user device and initializing a local unified model, performing forward knowledge distillation to distill knowledge in the trained local complex model to the local unified model of each user device4, transmitting to a server device local unified models from a plurality of user devices after completion of the forward knowledge distillation, merging the local unified models and the local variational autoencoders at the server device, transmitting from the server device to each user device a merged variational autoencoder, and the use of a merged variational autoencoder. However, Xing teaches: … a local complex model of each user device and… a local unified model of each user device (Page 4, Figure 15) training the local complex model (Page 3, Column 2, Paragraph 4, “The overview of EFDLS is shown in Fig. 1. In the system, users train their models locally”) … using local data in each user device (Page 1, Column 2, Paragraph 2, “…invented federated learning (FL). FL allows users to collectively harvest the advantages of shared models trained from their local data”) performing forward knowledge distillation to distill knowledge in the trained local complex model to the local unified model of each user device (Page 4, Figure 16) transmitting to a server device local unified models… from a plurality of user devices after completion of the forward knowledge distillation (Page 3, Column 2, Paragraph 4, “For each connected user, its student model’s hidden layers’ weights are uploaded to the EFDLS server periodically”), and Zhao teaches: merging local…7 autoencoders (Page 43, Abstract, “In addition, we propose a multimodal FedAvg algorithm to aggregate local autoencoders”) at the server device (Page 45, Column 2, Paragraph 3, ‘Local models from both types of clients are sent to the server and are aggregated into a global model by using a multimodal version of the FedAvg algorithm”), transmitting from the server device to each user device a merged… 8autoencoder (Page 46, Figure 19) , and Marsan teaches: configuring a… local model of each user device (Page 38, Paragraph 2, “Regarding the configuration parameters of the federating averaging algorithm, we select ten clients randomly at every round to involve in the training phase, the configuration is as following: … the local batch size is 32…10”) and initializing a… local model11 of each user device (Page 12, Figure 2.2, “Central server transmits the initial model to several nodes”), and merging the…12 local models at the server device (Page 12, Figure 2.2, “Central server pools model results13 and generate one global model…”) transmitting from the server device to each user device a merged unified model (Page 15, Figure Algorithm 114, Page 14 Figure 2.4) Qi, Xing, Zhao and Marsan are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing, Zhao and Marsan, and have the forward distillation occur locally, as well as having a method of merging a local model and local variational autoencoders at the server. By having the forward distillation occur locally, one can reduce the network overhead, and the merging of variational autoencoders on differing client devices allows one to extract aligned information to better encode higher-level data (see Page 44, Column 1, Paragraph 1 of Zhao). Regarding claim 4, Qi teaches wherein the local variational autoencoder (VAE) comprises a conditional variational autoencoder (CVAE) (see claim 1 analysis). Regarding claim 6, Qi fails to teach wherein the step of merging the local unified models comprises averaging the local unified models. However, Xing teaches the use of a local unified model (Page 4, Figure 115), and Marsan teaches wherein the merging of models comprises averaging the local models (Page 15, Algorithm 1, “wt+1 ← K ∑ k=1 (nk/N * wk t+1)”)16 Qi, Xing and Marsan are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and Marsan and work with model weights when averaging. Doing so allows for a global model with the proper internal architecture to handle heterogenous local models. Regarding claim 7, Qi teaches the use of a local variational autoencoder (see claim 1 analysis). Qi fails to teach the merging of local variational autoencoders comprises averaging the local variational autoencoders. However, Zhao teaches the merging of the local… autoencoder (see claim 1 analysis) comprises averaging the local… autoencoders (Page 43, Abstract, “In addition, we propose a multimodal FedAvg17 algorithm to aggregate local autoencoders”). Qi and Zhao are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Zhao and have the averaging of local autoencoders to create a global one. Doing so allows for an autoencoder that learns from multiple clients instead of just the local one. Regarding claim 8, Qi fails to teach wherein a local complex model of a user device is different from a second local complex model of another user device. However, Xing teaches a local complex model of a user device is different from a second local complex model of another user device (Abstract, Page 1, “EFDLS consists of a central server and multiple mobile users, where different users may run different TSC tasks”, Page 4, Figure 118). Qi and Xing are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and make each local complex model of each device different from one another. The use of heterogenous models expands the system to incorporate a larger variety of potential devices. Regarding claim 9, Qi fails to teach the further limitations of the claim. However, Xing teaches wherein the local complex model comprises a model based on linear regression, logistic regression, decision trees, support vector machines (SVM), naive Bayes, k- nearest neighbors or K-nearest neighbors (k-NN), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, or neural networks (Page 4, Figure 1, “…Conv x 9128 represents a 1-D convolutional neural network”). Qi and Xing are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and have the local complex model be based on a neural network. Doing so allows the system to handle and compute much more complex data and patterns. Regarding claim 10, Qi fails to teach the further limitations of the claim. However, Xing teaches wherein the local complex model comprises one or more machine learning models (see claim 9 analysis). Qi and Xing are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and have the local complex model be based on a machine learning model. Doing so allows the system to handle and compute much more complex data and patterns. Regarding claim 11, Qi fails to teach the further limitations of the claim. However, Xing teaches wherein the local complex model comprises a neural network, a convolutional neural network (CNN), a deep convolutional neural network (DCNN), a cascaded deep convolutional neural network, a simplified CNN, a shallow CNN, or a combination thereof (see claim 9 analysis). Qi and Xing are considered analogous to the invention because all are directed towards federated learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and have the local complex model be based on a convolutional neural network. Doing so opens the system up to handling more complex input data. Regarding claim 12, Qi teaches wherein the local data of a user device is not shared with another user device or the server device (Page 189, Abstract, “A public dataset is generated by conditional variational autoencoder (CVAE) and stored in the cloud server for supporting the obtaining of heterogeneous knowledge without sharing the private data of IoT devices”). Regarding claim 13, Qi teaches wherein the local unified model, the local variational autoencoder, the merged unified model (Page 195, Figure 219), or the merged variational autoencoder does not contain personally identifiable information (Page 190, Column 2, Paragraph 1, “In order to tackle the issue of lacking public dataset in the cloud server, we adopt conditional variational autoencoder (CVAE) [27], [28], [29] to generate a synthetic dataset, which is stored in the cloud server and available for each client without uploading a portion of its private data.”20) Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022), Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020) and McMahan et al. (“Communication-Efficient Learning of Deep Networks from Decentralized Data”, 202321). Regarding claim 2, Qi fails to teach that steps (b) to step (g) of claim 1 are repeated. However, McMahan teaches a federated learning system, where the steps for federated learning are repeated at least ten times (Page 6, Column 2, Table 222). Qi and McMahan are considered analogous to the invention because all are directed towards machine learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of McMahan and repeat the steps for federated learning at least ten times. Doing so allows the models of the system to converge to a higher accuracy. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022), Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020) and Venaille et al. (“Application of Neural Networks to Image-Based Control of Robot Arms”, 1994). Regarding claim 3, Qi fails to teach the further limitations of the claim. However, Xing teaches the use of a local complex model, and said local complex model being updated (see claim 1 analysis), and Venaille teaches a method of neural network updating, where the steps are repeated until the difference between each iteration is under 2% (Page 267, Column 2, Table 2) Qi, Xing and Venaille are considered analogous to the invention because all are directed towards machine learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and Venaille and set a convergence threshold for how many times to iterate through the system. Doing so allows the system to reach a specified level of accuracy without any unnecessary computation spent on extra iterations. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022), Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020) and Sariyildiz et al. (“Gradient Matching Generative Networks for Zero-Shot Learning”, 2019). Regarding claim 5, Qi teaches the use of a conditional variational autoencoder (see claim 1 analysis). Qi fails to teach the use of a cosine similarity regularization term in a loss function. However, Sariyildiz teaches a neural network that uses a cosine similarity regularization term in a loss function (Page 2166, Column 2, Paragraph 2, “Here, LCLS(f,x,a) is the loss function used in training… we measure the discrepancy between gr and gs via the cosine similarity”). Qi and Sariyildiz are considered analogous to the invention because all are directed towards machine learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Xing and incorporate a loss function that uses cosine similarity. Doing so provides a less complex and thus more computationally efficient way for autoencoders to learn. Claims 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022), Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020) and Jain et al. (“US 11456080”, 2020). Regarding claim 14, Qi teaches the use of local data in each user device (see claim 1 analysis). Qi fails to teach that the local data comprises location data of a user. However, Jain teaches data that comprises location data of a user (Column 6, Lines 5-11, “The system can also use data about each individual's community to enhance estimates or predictions of disease exposure, likelihood of infection, and other items. For example, the system can collect information about where a user resides, works, and generally spends time, as well as the types of locations”). Qi and Jain are considered analogous to the invention because all are directed towards methods of manipulating stored data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Jain and incorporate location data locally within the federated learning system. Doing so allows for the processing of sensitive personal data without it being shared. Regarding claim 15, Qi fails to teach wherein the local data comprises exposure status of a user to a contagious disease. However, Jain teaches data that comprises exposure status of a user to a contagious disease (Column 8, Lines 9-24, “…data package for managing a disease to a user device of a user… acquire and report monitoring data… scores for the user based on the monitoring data, the one or more scores comprising at least one of (i) an exposure score indicating a level of exposure of the user to the disease”). Qi and Jain are considered analogous to the invention because all are directed towards methods of manipulating stored data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Jain and incorporate location data locally within the federated learning system. Doing so allows for the processing of sensitive personal data without it being shared. Regarding claim 16, Qi fails to teach wherein the contagious disease is COVID-19, influenza, or respiratory syncytial virus. However, Jain teaches wherein the contagious disease is COVID-19, influenza, or respiratory syncytial virus (Column 8, Line 38, “In some implementations, the disease is COVID-19.”). Qi and Jain are considered analogous to the invention because all are directed towards methods of manipulating stored data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Jain and incorporate location data locally within the federated learning system. Doing so allows for the processing of sensitive personal data without it being shared. Claims 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022), Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020) and Liu et al. (“Hierarchical Federated Learning With Quantization: Convergence Analysis and System Design”, 2022). Regarding claim 17, Qi fails to teach wherein the step of transmitting from the server device to each user device is performed through one or more intermediate layers. However, Liu teaches wherein the step of transmitting from the server device to each user device is performed through one or more intermediate layers (Page 6, Figure 1). Qi and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and incorporate intermediate layers and a hierarchical structure within the federated learning system. Doing so allows for faster convergence speed (see page 2, column 2, paragraph 2 of Liu). Regarding claim 18, Qi fails to teach wherein the one or more intermediate layers comprise at least one distributed unit layer and/or at least one edge user layer. However, Liu wherein the one or more intermediate layers comprise at least one distributed unit layer (Page 6, Figure 123). Qi and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and incorporate a distributed unit layer as part of intermediate layers within the federated learning system. Doing so improves the ability to scale the system without sacrificing latency speed. Regarding claim 19, Qi fails to teach wherein the at least one distributed unit layer or the at least one edge user layer comprises one or more edge server devices deployed in proximity to the plurality of user devices. However, Liu wherein the at least one distributed unit layer comprises one or more edge server devices deployed in proximity to the plurality of user devices (Page 6, Figure 1). Qi and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and incorporate edge server devices in proximity to user devices as part of the distributed unit layer within the federated learning system. Doing so further improves the ability to scale the system without sacrificing latency speed. Claims 20-23 are rejected under 35 U.S.C. 103 as being unpatentable over Qi et al. (“FedBKD: Heterogenous Federated Learning via Bidirectional Knowledge Distillation for Modulation Classification in IoT-Edge System”, 2022) in view of Xing et al. (“An Efficient Federated Distillation Learning System for Multitask Time Series Classification”, 2022) and in further view of Zhao et al. (“Multimodal Federated Learning on IoT Data”, 2022), Marsan et al. (“Federated learning for vehicle trajectory prediction”, 2020), Liu et al. (“Hierarchical Federated Learning With Quantization: Convergence Analysis and System Design”, 2022) and Kang et al. (“Content Caching based on Popularity and Priority of Content using seq2seq LSTM in ICN”, 202324) Regarding claim 20, Qi fails to teach wherein the one or more intermediate layers comprise an edge aggregation layer between the at least one distributed unit layer and the at least one edge user layer, and wherein the edge aggregation layer performs functions comprising caching frequently accessed content, processing data, and/or providing low-latency access to applications and services. However: Liu teaches wherein the one or more intermediate layers comprises an edge aggregation layer between …the at least on edge user layer (Page 6, Figure 1), and wherein the edge aggregation layer performs functions comprising caching frequently accessed content, processing data, and/or providing low-latency access to applications and services (Page 4, Column 2, Paragraph 4, “Edge aggregation enjoys a lower propagation latency compared with cloud aggregation. Hence, in Hier-Local-QSGD, each edge server efficiently aggregates the models within its local area for several times before the cloud aggregation”), and Kang teaches a four-layer hierarchical network architecture25, wherein a layer of edge servers exists between client devices and an additional distributed layer (Page 3, Figure 1). Qi and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and Kang and incorporate an edge aggregation layer that provides low latency access to applications and services. Reducing latency improves the synchronization between the client and server devices. Regarding claim 21, Qi teaches a trained local variational autoencoder (see claim 1 analysis). Qi fails to teach: transmitting to the one or more intermediate layers the local unified models and the trained local variational autoencoders from the plurality of user devices after completion of the forward knowledge distillation, wherein each of the one or more edge server devices of the one or more intermediate layers receives the local unified models and the trained local variational autoencoders of at least a subset of the plurality of user devices; merging the local unified models and the local variational autoencoders of at least the subset of the plurality of user devices at the one or more edge server devices; and transmitting from the one or more edge server devices to the server device merged unified models and merged variational autoencoders to perform further merging at step (e). However: Xing teaches the use of local unified models (see claim 1 analysis) and forward knowledge distillation (see claim 1 analysis). Zhao teaches the merging local…26 autoencoders (see claim 1 analysis). Liu teaches transmitting to the one or more intermediate layers (Page 6, Figure 127)…, wherein each of the one or more edge server devices of the one or more intermediate layers receives the… of at least a subset of the plurality of user devices (Page 6, Figure 128)…; and transmitting from the one or more edge server devices to the server device…to perform further merging at step (e) (Page 6, Figure 129); Marsan teaches the merging of local models (see claim 1 analysis). Qi, Xing, Zhao, Marsan and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and pass the variational autoencoder and unified model through intermediate layers before transmitting to the server. Doing so allows for faster convergence speed (see page 2, column 2, paragraph 2 of Liu). Regarding claim 22, Qi teaches a final merged unified model (see claim 1 analysis), and a variational autoencoder (see claim 1 analysis). Qi fails to teach transmitting from the server device, through the one or more intermediate layers, to each user device a final merged unified model and a final merged variational autoencoder. However: Zhao teaches a final merged…30autoencoder (see claim 1 analysis), and Liu teaches transmitting from the server device, through the one or more intermediate layers, to each user device (Page 6, Figure 131). Qi and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and pass the variational autoencoder and unified model through intermediate layers before transmitting back to each client device. Doing so allows for faster convergence speed (see page 2, column 2, paragraph 2 of Liu). Regarding claim 23, steps (a) through (c) are covered in the rejection of claim 1, and steps (d) through (f), as well as step (h) are covered in the rejection of claim 22. New limitations introduced in this claim will be covered below. Qi teaches a method for efficient implementation of a method of heterogeneous federated learning through two-way knowledge distillation according to claim 1 (Page 189, Abstract, “to mitigate the data and model heterogeneity, we propose a heterogenous Federated learning framework based on Bidirectional Knowledge Distillation (FedBKD)”) comprising: a variational autoencoder (see claim 1 analysis), performing backward knowledge distillation to distill knowledge in the merged… model to the local… model of each user device (Page 190, Column 2, Paragraph 1, “…the cloud-to-client distillation is the single-teacher-multiple students process, which distills the knowledge from the single global model back to multiple heterogeneous local networks”) using data generated by the… variational autoencoder (Page 194, Column 2, “To alleviate the problem of data heterogeneity, we used CVAE to generate a synthetic dataset in the cloud server”) to obtain an updated local… model for each user device. Qi fails to teach further merging the merged local unified models and the merged local variational autoencoders at the server device to obtain a final merged unified model and a final merged local variational autoencoder. However, Xing teaches the use of local complex and unified models (see claim 1 analysis), Zhao teaches the merging of an autoencoder (see claim 1 analysis), Marsan teaches the merging of local models (see claim 1 analysis), and Liu teaches further merging and a final model32 (Page 6, Figure 1). Qi, Xing, Zhao, Marsan and Liu are considered analogous to the invention because all are directed towards methods of federated learning. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of Liu and pass the variational autoencoder and unified model through intermediate layers before performing backwards distillation and transmitting back to each client device. Doing so allows for faster convergence speed (see page 2, column 2, paragraph 2 of Liu). Regarding claim 24, Qi fails to teach that steps (b) to step (g) of claim 1 are repeated. However, McMahan teaches a federated learning system, where the steps for federated learning are repeated at least ten times (see claim 2 analysis). Qi and McMahan are considered analogous to the invention because all are directed towards machine learning methods. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Qi to incorporate the teachings of McMahan and repeat the steps for federated learning at least ten times. Doing so allows the models of the system to converge to a higher accuracy. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEWOS MESFIN whose telephone number is (571)270-0782. The examiner can normally be reached Monday-Friday 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEWOS MESFIN/Examiner, Art Unit 2145 /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145 1 Interpreting configuring as the set-up of the variational autoencoder 2 Figure shows that the autoencoder (the CVAE, as that’s a type of autoencoder) is set up for each client 3 The use of complex and unified models to be taught in Xing 4 Knowledge distillation that occurs locally 5 The figure highlights that it is over multiple devices/users, and the local/complex models (the student/teacher models). 6 Represented by arrow showing distillation from teacher model to student 7 Zhao does not teach merging variational autoencoders or merging local united models, but in combination with the variational autoencoder taught in Qi, the merging taught in Qi, and the use of local united models taught in Xing, the whole limitation is covered. 8 See footnote 9 9 Figure shows communication in two channels (server-to-client and client-to-server) with the averaging (merging) occurring in the server 10 Reference contains further configuration. We define configuring as the set-up of model hyperparameters. 11 This reference teaches the actions of configuring and initializing a model. In combination with the structure of a local complex and local united model taught in Xing, the whole limitation is covered. 12 Through the teaching of the local unified model in Xing et al, it can be substituted and the whole limitation is covered. 13 It’s clear from Algorithm 1 (Page 15) that the model results being pooled (i.e. merged) is the model weights 14 The algorithm and figure shows that the merged models in the server get updated in the client model 15 Local unified model understood as student model 16 Shows the averaging of all the weights 17 FedAvg is a well-known term in the art. It is a federated learning algorithm that averages model weights of local devices to create a global server-side model. 18 The abstract teaches the devices doing distinct and different tasks, and the figure teaches a first and second complex model. 19 Understood as global model on server side 20 The merged unified model uses only the synthetic dataset, as shown in the figure, which means it doesn’t contain any personally identifiable information 21 Prior to EFD of application 22 Rounds for MNIST CNN represented in IID and Non-IID columns. All show at least 10 rounds. 23 The one distributed unit layer is represented by the group of edge server layers. A distributed unit layer is a layer in a type of decentralized networking architecture, which is showcased with the multiple edge servers. 24 Prior to EFD of application 25 Which can thus also be interpreted as a distributed network 26 Zhao does not teach merging variational autoencoders, but in combination with the variational autoencoder taught in Qi, the whole limitation is covered. 27 Transmitting shown by arrows in figure 28 Subset shown by circle enclosing phone devices 29 Further merging indicated by cloud aggregation 30 Zhao does not teach merging variational autoencoders, but in combination with the variational autoencoder taught in Qi, the whole limitation is covered. 31 Indicated by backwards arrows 32 “Further merging”, “merging the merged” and “final” are words that refer to the steps of a hierarchical learning system, where merging first occurs first at the edge serves before being done in the main server. The figure in Liu, in combination with the federated learning taught by Xing, Zhao, Marsan and Qi, covers the limitation.
Read full office action

Prosecution Timeline

Mar 22, 2024
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month