Prosecution Insights
Last updated: October 01, 2026
Application No. 18/615,830

SYSTEMS AND METHODS FOR CONTINUOUS MODEL TRAINING ON A PEER-TO-PEER NETWORK

Non-Final OA §103
Filed
Mar 25, 2024
Priority
Mar 28, 2023 — IN 202311022691
Examiner
ZENG, WENWEI
Art Unit
Tech Center
Assignee
JPMorgan Chase Bank, N.A.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
23 currently pending
Career history
18
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on July 19, 2024, was considered by the examiner. The submission is in compliance with the provisions of 37 CFR 1.97. Claim Objections Claim 1 is objected to because of the following informalities: the limitation “replicating, by the first node, the first version of the first version of the machine learning model on a second node of the peer-to-peer distributed network;” should read “replicating, by the first node, the first version of the machine learning model …”, Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. Claims 1, 3, 5, 6, 8, 10, 12, 14, 17, and 19 are rejected under 35 U.S.C. 103 over Lim, W. et al., in “Federated Learning in Mobile Edge Networks: A Comprehensive Survey,” published on April 8, 2020, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9060868 , (hereafter, Lim). in view of Ranathunga, T. et al., in "Blockchain-Based Decentralized Model Aggregation for Cross-Silo Federated Learning in Industry 4.0," published on November 1, 2022 , available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9933813 , (hereafter, Ranathunga), further in view of Hu, C., Jiang, J., & Wang, Z. (published on August 21, 2019). In “Decentralized federated learning: A segmented gossip approach,” available at https://arxiv.org/pdf/1908.07782 , (hereafter, Hu). Claim 1: Regarding claim 1, Lim teaches “1. A method comprising: training, at a first node of a peer-to-peer distributed network, a first version of a machine learning model on a first private data set;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Lim mentions that the data owners correlate to devices or client nodes. Here, Lim shows that N can be any number from 1st, 2nd, 3rd device or any subsequent device, which includes a first node. Each of these devices has their own private dataset Di, (i.e. first node trains model on a first private data set) to train its local model wi (where i = 1 for a first version of the machine learning model). Further, Lim teaches “sending, by the first node, the first version of the machine learning model to an aggregation server provided in the peer-to-peer distributed network;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Here, Lim continues the method for the first device or first node, where each data owner (i.e. node) uses its dataset to train a local model and send that local model information to the FL server. FL server here relates to an aggregation server. Further, Lim teaches “training, by the second node, the machine learning model on a second private data set, resulting in a second version of the machine learning model;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Lim mentions that the data owners correlate to devices or client nodes. Here, Lim shows that N can be any number from 1st, 2nd, 3rd device or any subsequent device, which includes a second node. Each of these devices has their own private dataset Di, (i.e. second node trains model on a second private data set) to train its local model wi (where i = 2 for a second version of the machine learning model). Further, Lim teaches “sending, by the second node, the second version of the machine learning model to the aggregation server;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Here, Lim continues the method for the second device or second node, where each data owner (i.e. node) uses its dataset to train a local model and send that local model information to the FL server. FL server here relates to an aggregation server. Further, Lim teaches “vi. aggregating, by the aggregation server, the first version of the machine learning model and the second version of the machine learning model resulting in an aggregated machine learning model;” See Lim in page 2032 of Introduction, paragraph 3, describe “To guarantee that training data remains on personal devices and to facilitate collaborative machine learning of complex models among distributed devices, a decentralized ML approach called Federated Learning (FL) is introduced in [21]. In FL, mobile devices use their local data to cooperatively train an ML model required by an FL server. They then send the model updates, i.e., the model’s weights, to the FL server for aggregation. The steps are repeated in multiple rounds until a desirable accuracy is achieved.” Here, Lim shows that the FL server relates to being an aggregation server, which aggregate models sent from each device node. Further, Lim teaches “vii. and sending, by the aggregation server, the aggregated machine learning model to the first node” See Lim describe in page 2036, section B. Federated Learning in step 3 "Step 3 (Global model aggregation and update): The server aggregates the local models from participants and then sends the updated global model parameters wt+1 G back to the data owners." Here, Lim shows that the server (i.e. aggregation server) in this case, sends the updated global model information (i.e. aggregated machine learning model) back to each individual data owner (i.e. send this model back to first node, second node, and subsequent nodes). However, Lim did not teach “ A method comprising: training, at a first node of a peer-to-peer distributed network, a first version of a machine learning model on a first private data set;” or “sending, by the first node … to an aggregation server provided in the peer-to-peer distributed network;” or “replicating, by the first node, the first version of the first version of the machine learning model on a second node of the peer-to-peer distributed network;” In an analogous method, Ranathunga teaches “ A method comprising: training, at a first node of a peer-to-peer distributed network, a first version of a machine learning model on a first private data set;” See Ranathunga in page 4451, section II. Related work, describe “Apart from these, a generic solution known as FPPDL is presented in [8] for decentralized FL for peer-to-peer interaction, rather then the traditional client-server FL architecture. Participants utilize a private Blockchain as an information-sharing platform and also employ HE for secure model aggregation.” Here, Ranathunga shows using a peer-to-peer as part of this method of federated distributed learning. Further, see Ranathunga in page 4453, section C. System overview, describe “as shown in Fig. 1, the proposed Bloackchain-enabled hierarchical aggregation process for FL begins with 1) local model training by edge devices which subsequently 2) register their local models on-chain along with their meta-data and 3) share them off-chain with validators for 4) model verification and validation… Each client aggregates the model it received from the client below in the hierarchy with it is own local model and shares the intermediate aggregated model (Fig. 1) with the client above in the hierarchy… Thus, no temporary control/access to the local model updates is given to a single client for model aggregation. Moreover, peer-to-peer model parameter sharing enables clients to join/leave the network without the perils of a single point of failure. The Blockchain consensus mechanism is used to achieve consensus with respect to the topology, where the position of a node in the hierarchy is a direct outcome of its current credibility score (Section V-A.1). The rationale behind using a credibility-based score is to reward/punish the clients according to their contribution in the global model building.” Here, Ranathunga incorporates using training models for devices with a peer-to-peer network and help devices interact and communicate in a distributed network. Further, Ranathunga teaches “sending, by the first node … to an aggregation server provided in the peer-to-peer distributed network;” See Ranathunga in page 4451, section II. Related work, describe “Apart from these, a generic solution known as FPPDL is presented in [8] for decentralized FL for peer-to-peer interaction, rather then the traditional client-server FL architecture. Participants utilize a private Blockchain as an information-sharing platform and also employ HE for secure model aggregation.” Here, Ranathunga shows using a peer-to-peer as part of this method of federated distributed learning. Further, see Ranathunga in page 4453, section C. System overview, describe “as shown in Fig. 1, the proposed Bloackchain-enabled hierarchical aggregation process for FL begins with 1) local model training by edge devices which subsequently 2) register their local models on-chain along with their meta-data and 3) share them off-chain with validators for 4) model verification and validation… Each client aggregates the model it received from the client below in the hierarchy with it is own local model and shares the intermediate aggregated model (Fig. 1) with the client above in the hierarchy… Thus, no temporary control/access to the local model updates is given to a single client for model aggregation. Moreover, peer-to-peer model parameter sharing enables clients to join/leave the network without the perils of a single point of failure. The Blockchain consensus mechanism is used to achieve consensus with respect to the topology, where the position of a node in the hierarchy is a direct outcome of its current credibility score (Section V-A.1). The rationale behind using a credibility-based score is to reward/punish the clients according to their contribution in the global model building.” Here, Ranathunga incorporates using training models for devices with a peer-to-peer network and help devices interact and communicate in a distributed network. Also, see Ranathunga in page 4452, Section III. Foundations, A. FedAvg for Model Aggregation describe “As introduced in Section I-A, FL enables distributed clients to train a global ML model in a collaborative fashion without exposing their raw data to each other. If dk represents the set of local training data samples for client Ck , the total size of data samples from n clients, D=∑nk=1dk . For an FL training round t , each client trains the global model (with the parameters represented as wt ) with its local data and calculates local parameters wkt+1 . After the completion of local training, every client uploads it is local model parameters (wkt+1) to the global server”. Here, Ranathunga describes sending a local model by each client to a global server (i.e. aggregation server), using peer to peer network. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Lim and incorporate into the teachings of Ranathunga because both references teach using two nodes to each train and send models to an aggregation server. One of ordinary skill in the art would be motivated to do so because “Extensive off-chain experiments were performed to test the framework’s ability to minimize model convergence time and maximize global model accuracy. An improved model aggregation algorithm was also proposed which results in 10% improvement in global model accuracy and faster model convergence time”, (see Ranathunga in page 4460, section VII. Conclusion). However, Lim in view of Ranathunga did not teach “replicating, by the first node, the first version of the first version of the machine learning model on a second node of the peer-to-peer distributed network;” See Hu in page 2, section 3. Segmented Gossip Aggregation describe “Now consider the network topology with n workers. An all reduce worker pushes n−1 local model replicates to the other workers through n − 1 links while a gossip worker is expected to push one local model replicate out through only one link. Within a datacenter where the workers are connected by the local area network, they can always communicate with each other at maximum bandwidth thus the gossip worker can achieve great speed up as the transmission size is drastically reduced.” Here, Hu describes a worker device or node that sends a copy of a model (Hu mentions this as ‘model replicates’) to other workers (another node) in a network. This model includes a first local model of that worker device since this process applies for all worker devices and their respective local model versions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim and Ranathunga with the teachings of Hu by using the teachings of Lim and Ranathunga using two nodes to each train and send models to an aggregation server, with Hu’s teaching of replicating a model on a peer to peer network. One of ordinary skill in the art would be motivated to do so because by integrating Hu’s framework into the methods of Lim and Ranathunga, one with ordinary skill in the art would achieve the goal of providing a method where “the analysis is validated by Fig.5(a) that with the increase of the number of model replicas, the accuracy of each iteration becomes better,” (see Hu in page 6, section 5.2 Experiment Result, in Impact of Model Replicas). Claim 3: Regarding claim 3, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. Further, Lim teaches “ training, by the third node, the first version of the machine learning model or the second version of the machine learning model using a third private data set, resulting in a third version of the machine learning model;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Lim mentions that the data owners correlate to devices or client nodes. Here, Lim shows that N can be any number from 1st, 2nd, 3rd device or any subsequent device, which includes a second node. Each of these devices has their own private dataset Di, (i.e. second node trains model on a second private data set) to train its local model wi (where i = 2 for a second version of the machine learning model). Further, Lim teaches “iii and sending, by the third node, the third version of the machine learning model to the aggregation server.” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Lim mentions that the data owners correlate to devices or client nodes. Here, Lim shows that N can be any number from 1st, 2nd, 3rd device or any subsequent device, which includes a third node. Each of these devices has their own private dataset Di, (i.e. third node trains model on a third private data set) to train its local model wi (where i = 3 for a third version of the machine learning model) and send its version to a FL server (i.e. aggregation server). This training process repeats. Further, see Lim in pages 2035-2036, step 3, B. Federated Learning, describe “Step 1 (Task initialization): The server decides the training task, i.e., the target application, and the corresponding data requirements. The server also specifies the hyper parameters of the global model and the training process, e.g., learning rate. Then, the server broadcasts the initialized global model w0G and task to selected participants. Step 2 (Local model training and update): Based on the global model wtG , where t denotes the current iteration index, each participant respectively uses its local data and device to update the local model parameters wti . The goal of participant i in iteration t is to find optimal parameters wti that minimize the loss function L(wti) , i.e., wt∗i=argminwtiL(wti).(3) PNG media_image1.png 90 910 media_image1.png Greyscale The updated local model parameters are subsequently sent to the server. Step 3 (Global model aggregation and update): The server aggregates the local models from participants and then sends the updated global model parameters wt+1G back to the data owners.” Here, Lim in step 3 mentions the data owner or each node receives an updated global model (i.e. aggregated model). Then, using the process from page 2035 of data owners and the FL server, a third node sends updated local model (i.e. third version of model) to the server. Further, Ranathunga teaches “training, by the third node, the first version of the machine learning model or the second version of the machine learning model using a third private data set, resulting in a third version of the machine learning model;” See Ranathunga in page 4453, section C. System overview, describe “Model aggregation begins with the client at the lowest level of the hierarchy, termed as the initiator, which shares its local model update with the client directly above it in the hierarchy. Client Mi+1 aggregates Mi ’s local model parameters with its own and shares the aggregated model with Mi+2 and so on. The final aggregated model for that round is created at client Mk and is shared with other clients prior to the next iteration. Until then, the clients are in possession of the versions of models (i.e., g′i,g′i+1,..,g′k−1 , etc.) aggregated by them. Unlike existing decentralized aggregation mechanisms, every client (except initiator) contributes to the model aggregation process and can possess a different version of the aggregated model till the next round of training.” Here, Ranathunga mentions that client Mi+1 aggregates Mi local model then shares this aggregated model with Mi+2, and each client has their own versions of models. The term client is interpreted to mean the same as a device or node. Here, Ranathunga shows that in the training process, the third node or client Mi+2 gets local model information from each first node Mi and second node Mi+1 to create a final aggregated model. Before the aggregation, each device has its own local model version. Further, see Ranathunga in page 4454 section V. Proposed Framework, A. Topology creation, part 1) credibility score in step 1 describe "Step 1:All clients share their training data size information for round t on the Blockchain. Peer nodes in the Blockchain network create a block containing the received information (|dt1|,|dt2|,..,|dtn|) , verifies the transactions and generates a new block. Once the block is verified, it is added to the Blockchain and subsequently downloaded by all clients." Here, Ranathunga shows that each client device or node uses a dataset of information, denoted as blocks from dt ... dtn, respectively for node 1 to node n. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Lim and incorporate into the teachings of Ranathunga because both references teach using 2 nodes to each train and send models to an aggregation server. One of ordinary skill in the art would be motivated to do so because “Extensive off-chain experiments were performed to test the framework’s ability to minimize model convergence time and maximize global model accuracy. An improved model aggregation algorithm was also proposed which results in 10% improvement in global model accuracy and faster model convergence time”, (see Ranathunga in page 4460, section VII. Conclusion). Further, Hu teaches “3. The method of claim 1, further comprising: replicating, by the first node or the second node, the first version of the machine learning model or the second version of the machine learning model on a third node in the peer-to-peer distributed network;” See Hu in page 2, section 3. Segmented Gossip Aggregation describe “Now consider the network topology with n workers. An all reduce worker pushes n−1 local model replicates to the other workers through n − 1 links while a gossip worker is expected to push one local model replicate out through only one link. Within a datacenter where the workers are connected by the local area network, they can always communicate with each other at maximum bandwidth thus the gossip worker can achieve great speed up as the transmission size is drastically reduced.” Here, Hu describes a worker device or node that sends a copy of a model (Hu mentions this as ‘model replicates’) to other workers (another node) in a network. Here, workers can be either a first node or a second node. Since this replication process occurs for any n workers, where n can equal 3 or a third node, this method can use either a first or second device and replicate either a first version or second version of a model onto a third device. Examiner construes a version of the model to mean local model of each respective device, such as a first version of a model relating to a first local model created by a first device. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim and Ranathunga with the teachings of Hu by using the teachings of Lim and Ranathunga using 2 nodes to each train and send models to an aggregation server, with Hu’s teaching of replicating a model on a peer to peer network. One of ordinary skill in the art would be motivated to do so because by integrating Hu’s framework into the methods of Lim and Ranathunga, one with ordinary skill in the art would achieve the goal of providing a method where “the analysis is validated by Fig.5(a) that with the increase of the number of model replicas, the accuracy of each iteration becomes better,” (see Hu in page 6, section 5.2 Experiment Result, in Impact of Model Replicas). Claim 5: Regarding claim 5, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. Further, Lim teaches “5. The method of claim 1, wherein the first node sends weights for the first version of the machine learning model to the aggregation server” See Lim in page 2, second paragraph in Introduction describe “To guarantee that training data remains on personal devices and to facilitate collaborative machine learning of complex models among distributed devices, a decentralized ML approach called Federated Learning (FL) is introduced in [21]. In FL, mobile devices use their local data to cooperatively train an ML model required by an FL server. They then send the model updates, i.e., the model’s weights, to the FL server for aggregation. The steps are repeated in multiple rounds until a desirable accuracy is achieved.” Here, Lim describes using a first device to send weights of the first local model, then mentioned steps are repeated in multiple rounds which also applies to multiple subsequent devices and their respective model weights. See Figure 3 in Lim where N participants refer to multiple devices, N can be 1, 2, 3, or any subsequent devices. PNG media_image2.png 612 846 media_image2.png Greyscale Claim 6: Regarding claim 6, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. Further, Lim teaches “6. The method of claim 1, wherein the second node sends weights for the second version of the machine learning model to the aggregation server” See Lim in page 2, second paragraph in Introduction describe “To guarantee that training data remains on personal devices and to facilitate collaborative machine learning of complex models among distributed devices, a decentralized ML approach called Federated Learning (FL) is introduced in [21]. In FL, mobile devices use their local data to cooperatively train an ML model required by an FL server. They then send the model updates, i.e., the model’s weights, to the FL server for aggregation. The steps are repeated in multiple rounds until a desirable accuracy is achieved.” Here, Lim describes using a device to send weights of its own local model, then mentioned steps are repeated in multiple rounds which also applies to multiple subsequent devices and their respective model weights. See Figure 3 in Lim where N participants refer to multiple devices, N can be 1, 2, 3, or any subsequent devices. Lim in Figure 3 shows that N = 2, which relates to a second node, which can also send weights of its local model (i.e. second version of the machine learning model) to the FL server (i.e. aggregation server). PNG media_image3.png 722 994 media_image3.png Greyscale Claim 8: Regarding claim 8, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. Further, Ranathunga teaches “The method of claim 1, wherein the peer-to-peer distributed network is a permissioned blockchain network” See Ranathunga in page 4450 Introduction, section C. Research contributions, mention in item 4 “The proposed framework is evaluated for two Industry 4.0 use cases: a) predictive maintenance and b) vision-based product quality inspection, in terms of its effect on accuracy, model convergence time, and overall system fairness. This work also provides a benchmark analysis of the implementation of the proposed framework using a state-of-the-art permission-ed Blockchain platform and provides the minimal hardware configurations required to achieve the expected throughput and latency requirement of the use case.” Here, Ranathunga describes using a permissioned blockchain platform for the network. Further, see Ranathunga in page 4452, section III. Foundations, part B. Blockchain for Decentralized Data Sharing describe “A Blockchain is a shared ledger of transactions that have been executed and shared among network participants.” Ranathunga defines a blockchain and how it works among participants in a network.’ It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Lim and incorporate into the teachings of Ranathunga because both references teach using 2 nodes to each train and send models to an aggregation server. One of ordinary skill in the art would be motivated to do so because “Extensive off-chain experiments were performed to test the framework’s ability to minimize model convergence time and maximize global model accuracy. An improved model aggregation algorithm was also proposed which results in 10% improvement in global model accuracy and faster model convergence time”, (see Ranathunga in page 4460, section VII. Conclusion). Claim 10: Regarding claim 10, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. Further, Lim teaches “10. The method of claim 1, wherein the first node trains the first version of the machine learning model until a desired accuracy is reached” See Lim in page 2, second paragraph in Introduction describe “To guarantee that training data remains on personal devices and to facilitate collaborative machine learning of complex models among distributed devices, a decentralized ML approach called Federated Learning (FL) is introduced in [21]. In FL, mobile devices use their local data to cooperatively train an ML model required by an FL server. They then send the model updates, i.e., the model’s weights, to the FL server for aggregation. The steps are repeated in multiple rounds until a desirable accuracy is achieved.” Here, Lim describes using a first device to send weights of the first local model, then mentioned steps are repeated in multiple rounds which also applies to multiple subsequent devices and their respective model weights. Lim also describes this training process is repeated until a desirable accuracy is achieved. See Figure 3 in Lim for details where N participants refer to multiple devices, N can be 1, 2, 3, or any subsequent devices. Claim 12: Regarding claim 12, Lim further teaches “12. A system, comprising: … a first node; a second node; … wherein: the first node is configured to train a first version of a machine learning model on a first private data set;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Lim mentions that the data owners correlate to devices or client nodes. Here, Lim shows that N can be any number from 1st, 2nd, 3rd device or any subsequent device, which includes a first node. Each of these devices has their own private dataset Di, (i.e. first node trains model on a first private data set) to train its local model wi (where i = 1 for a first version of the machine learning model). However, Lim did not teach “12. A system, comprising: a peer-to-peer distributed network comprising: … and an aggregation server; …” In an analogous art, Ranathunga teaches “12. A system, comprising: a peer-to-peer distributed network comprising: … and an aggregation server; …” See Ranathunga in page 4453, section C. System overview, describe “as shown in Fig. 1, the proposed Bloackchain-enabled hierarchical aggregation process for FL begins with 1) local model training by edge devices which subsequently 2) register their local models on-chain along with their meta-data and 3) share them off-chain with validators for 4) model verification and validation… Each client aggregates the model it received from the client below in the hierarchy with it is own local model and shares the intermediate aggregated model (Fig. 1) with the client above in the hierarchy… Thus, no temporary control/access to the local model updates is given to a single client for model aggregation. Moreover, peer-to-peer model parameter sharing enables clients to join/leave the network without the perils of a single point of failure. The Blockchain consensus mechanism is used to achieve consensus with respect to the topology, where the position of a node in the hierarchy is a direct outcome of its current credibility score (Section V-A.1). The rationale behind using a credibility-based score is to reward/punish the clients according to their contribution in the global model building.” Here, Ranathunga incorporates using training models for devices with a peer-to-peer network and help devices interact and communicate in a distributed network. Further, see Ranathunga in page 4452, Section III. Foundations, A. FedAvg for Model Aggregation describe “As introduced in Section I-A, FL enables distributed clients to train a global ML model in a collaborative fashion without exposing their raw data to each other. If dk represents the set of local training data samples for client Ck , the total size of data samples from n clients, D=∑nk=1dk . For an FL training round t , each client trains the global model (with the parameters represented as wt ) with its local data and calculates local parameters wkt+1 . After the completion of local training, every client uploads it is local model parameters (wkt+1) to the global server”. Here, Ranathunga describes sending a local model by each client to a global server (i.e. aggregation server), using a peer to peer network. It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the base reference of Lim and incorporate into the teachings of Ranathunga because both references teach using two nodes to each train and send models to an aggregation server. One of ordinary skill in the art would be motivated to do so because “Extensive off-chain experiments were performed to test the framework’s ability to minimize model convergence time and maximize global model accuracy. An improved model aggregation algorithm was also proposed which results in 10% improvement in global model accuracy and faster model convergence time”, (see Ranathunga in page 4460, section VII. Conclusion). Regarding claim 12, the claim recites similar additional limitations as corresponding independent claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Claim 14: Regarding claim 14, the claim recites similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Claim 17: Regarding claim 17, the claim recites similar limitations as corresponding claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale. Claim 19: Regarding claim 19, the claim recites similar limitations as corresponding claim 10 and is rejected for similar reasons as claim 10 using similar teachings and rationale. Claims 2 and 13 are rejected under 35 U.S.C. 103 over Lim in view of Ranathunga, further in view of Hu, further in view of Shayan, M. et al. in “Biscotti: A blockchain system for private and secure federated learning,” published on December 11, 2020, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9292450 , (hereafter, Shayan). Claim 2: Regarding claim 2, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. However, did not teach “The method of claim 1, wherein the first version is a genesis version of the machine learning model.” In an analogous art, Shayan teaches “The method of claim 1, wherein the first version is a genesis version of the machine learning model” See Shayan in page 1516, section 4.1 Initializing the Training describe “Biscotti peers initialize the training process using information in the first (genesis) block. We assume that a trusted authority facilitates and bootstraps the training process by publicly distributing the genesis block out of band to all peers in the system. The authority is only trusted for this step: they are not entrusted with the individual SGD updates of the peers, which could potentially leak private information of a peer's data. Each peer obtains the following information from the genesis block: (1) the initial model state w0 and expected number of iterations T, (2) the public key PK for creating commitments to SGD updates,” Here, Shayan describes using an initial model state w0 created from a genesis block or version of a model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim, Ranathunga, and Hu, and incorporate with the teachings of Shayan by using the teachings of Lim, Ranathunga, and Hu of using two nodes to train and send models to an aggregation server, with Shayan’s teachings of a genesis version of a machine learning model. One of ordinary skill in the art would be motivated to do so because by integrating Shayan’s framework into the methods of Lim, Ranathunga, and Hu, one with ordinary skill in the art would achieve “Inspired by prior work [32], Biscotti uses consistent hashing based on PoF in combination with verifiable random functions (VRFs) [41] to select key roles for peers who will help coordinate the privacy and security of model updates,” (See Shayan in page 1514, Introduction). Claim 13: Regarding claim 13, the claim recites similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Claims 4 and 15 are rejected under 35 U.S.C. 103 over Lim in view of Ranathunga, further in view of Hu, further in view of Llisterri Gimenez, N., et al. in “On-device training of machine learning models on microcontrollers with federated learning,” published on February 14, 2022, available at https://www.mdpi.com/2079-9292/11/4/573 , (hereafter, Llisterri-Gimenez). Claim 4: Regarding claim 4, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 3. However, Lim in view of Ranathunga, further in view of Hu, did not teach “4. The method of claim 3, wherein the aggregated machine learning model comprises an aggregation of the first version of the machine learning model, the second version of the machine learning model, and the third version of the machine learning model, and the aggregation server further replicates the aggregated machine learning model on the third node.” In an analogous art, Llisterri-Gimenez teaches “4. The method of claim 3, wherein the aggregated machine learning model comprises an aggregation of the first version of the machine learning model, the second version of the machine learning model, and the third version of the machine learning model, and the aggregation server further replicates the aggregated machine learning model on the third node.” See Llisterri-Gimenez in page 7, section 3.7. Federated Learning Training describe "Figure 4 shows the principal diagram of the federated learning workflow. After the configuration, the server creates a neural network model and initializes it with random parameters. This model is sent to all the clients. As soon as the clients receive the model, they start training it with their local data. The local data are generated by the user saying the keywords. In the independent and identically distributed (IID) data case, all the clients must use the same keywords. Meanwhile, the server starts a timer and sets a ten-second countdown for the first federated learning round. When the countdown ends, the server tries to connect with all the clients and asks them to send their trained model. The clients have five seconds to accept the request, or they will be discarded for this iteration. The models that are received from the connected clients are aggregated to create a new global model. This global model will be sent back to the connected clients to be trained again, and the server will restart the countdown for the next round." Here, Llisterri-Gimenez shows that the global aggregated model will be sent back to all clients, which includes a third node. The global model sent back is like providing a copy of that model to all nodes, including to a third node. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim, Ranathunga, and Hu, with the teachings of Llisterri-Gimenez, by using the teachings of Lim, Ranathunga, and Hu of using two nodes to each train and send models to an aggregation server, with Llisterri-Gimenez’s teaching of an aggregated machine learning model comprises an aggregation of the first version of the machine learning model, the second version of the machine learning model, and the third version of the machine learning model, and the aggregation server further replicates the aggregated machine learning model on the third node. One of ordinary skill in the art would be motivated to do so because by integrating Llisterri-Gimenez’s framework into the methods of Lim, Ranathunga, and Hu, one with ordinary skill in the art would achieve a method of “evaluated the loss function in the forward path of the training process, and we observed how the machine learning model trained on the device improves during the training after tens of spoken words,” (See Llisterri-Gimenez in page 14, section 6. Conclusions and Future Work). Claim 15: Regarding claim 15, the claim recites similar limitations as corresponding claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Claims 7 and 16 are rejected under 35 U.S.C. 103 over Lim in view of Ranathunga, further in view of Hu further in view of Jimenez Gutierrez, D., et al. in “Application of federated learning techniques for arrhythmia classification using 12-lead ecg signals,” published on August 23, 2022, available at https://arxiv.org/pdf/2208.10993v1 , (hereafter, Jimenez_Gutierrez), Claim 7: Regarding claim 7, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. However, Lim in view of Ranathunga, further in view of Hu, did not teach “7. The method of claim 1, wherein the aggregation server sends weights for the aggregated machine learning model to the first node and the second node.” In an analogous art, Jimenez_Gutierrez teaches “7. The method of claim 1, wherein the aggregation server sends weights for the aggregated machine learning model to the first node and the second node” See Jimenez_Gutierrez in page 4, section 3. Methodology, describe "When the training is concluded, the weights of the resulting model are transmitted to the global server. The global server examines all the individual weights using a weight aggregation method and transmits the resulting model to all organizations." Here, Jimenez_Gutierrez mentions that the global server (i.e. aggregation server) sends weights of the resulting model and the model itself (i.e. aggregated machine learning model) to all organizations (where organizations here are viewed as all nodes, including first and second nodes). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim, Ranathunga, and Hu, and incorporate with the teachings of Jimenez_Gutierrez by using the teachings of Lim, Ranathunga, and Hu of using two nodes to each train and send models to an aggregation server, with Jimenez_Gutierrez’s teaching of an aggregation server send weights of the aggregated model to the two nodes. One of ordinary skill in the art would be motivated to do so because by integrating Jimenez_Gutierrez’s framework into the methods of Lim, Ranathunga, and Hu, one with ordinary skill in the art would achieve “apart from improving the predictive performance of our model, it will also reduce the overall computational requirements of training the AI model,” (see Jimenez_Gutierrez in page 6, section 3.4 Feature selection, first paragraph on page). Claim 16: Regarding claim 16, the claim recites similar limitations as corresponding claims 5-7 and is rejected for similar reasons as claims 5-7 using similar teachings and rationale. Claims 9 and 18 are rejected under 35 U.S.C. 103 over Lim in view of Ranathunga, further in view of Hu, further in view of Aledhari, M. et al., in "Federated Learning: A Survey on Enabling Technologies, Protocols, and Applications," published on July 31, 2020, available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9153560&tag=1 , (hereafter, Aledhari). Claim 9: Regarding claim 9, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 1. Further, Lim teaches “9. The method of claim 1, further comprising: … training, by the first node, the first version of the machine learning model with an additional first private dataset, resulting in an updated first version of the aggregated machine learning model;” See Lim mention in page 2035 section B. Federated Learning "In general, there are two main entities in the FL system, i.e., the data owners (viz. participants) and the model owner (viz. FL server). Let N={1,…,N} denote the set of N data owners, each of which has a private dataset Di∈N . Each data owner i uses its dataset Di to train a local model wi and send only the local model parameters to the FL server. Then, all collected local models are aggregated w=∪i∈Nwi to generate a global model wG ." Lim mentions that the data owners correlate to devices or client nodes. Here, Lim shows that N can be any number from 1st, 2nd, 3rd device or any subsequent device, which includes a first node. Each of these devices has their own private dataset Di, (i.e. first node trains model on a first private data set) to train its local model wi (where i = 1 for a first version of the machine learning model). Further, Lim teaches “sending, by the first node, the updated first version of the aggregated machine learning model to the aggregation server;” See Lim in abstract, page describe "Recently, in light of increasingly stringent data privacy legislations and growing privacy concerns, the concept of Federated Learning (FL) has been introduced. In FL, end devices use their local data to train an ML model required by the server. The end devices then send the model updates rather than raw data to the server for aggregation." Also, see Lim in page 2036, in step 2, "Step 2 (Local model training and update): Based on the global model wtG , where t denotes the current iteration index, each participant respectively uses its local data and device to update the local model parameters wti . The goal of participant i in iteration t is to find optimal parameters wti that minimize the loss function L(wti) , i.e., wt∗i=argminwtiL(wti).(3) PNG media_image4.png 16 1008 media_image4.png Greyscale .The updated local model parameters are subsequently sent to the server." Here, Lim mentions that each participant i or each node including i = 1 or a first node, send any updated local model information (i.e. updated first version of the aggregated machine learning model) to the server (i.e. aggregation server). Further, Lim teaches “sending, by the second node, the updated second version of the aggregated machine learning model to the aggregation server;” See Lim in abstract, page describe “Recently, in light of increasingly stringent data privacy legislations and growing privacy concerns, the concept of Federated Learning (FL) has been introduced. In FL, end devices use their local data to train an ML model required by the server. The end devices then send the model updates rather than raw data to the server for aggregation.” Further, See Lim in page 2036, in step 2, "Step 2 (Local model training and update): Based on the global model wtG , where t denotes the current iteration index, each participant respectively uses its local data and device to update the local model parameters wti . The goal of participant i in iteration t is to find optimal parameters wti that minimize the loss function L(wti) , i.e., wt∗i=argminwtiL(wti).(3) PNG media_image4.png 16 1008 media_image4.png Greyscale .The updated local model parameters are subsequently sent to the server." Here, Lim mentions that each participant i or each node including i = 2 or a second node, send any updated local model information (i.e. updated second version of the aggregated machine learning model) to the server (i.e. aggregation server). Further, Ranathunga teaches “ … incorporating, by the first node, the aggregated machine learning model into the first version of the machine learning model;” See Ranathunga in page 4449, in Introduction, part A. Background mention "Traditional FL is centralized and works based on a client-server paradigm where multiple parties (clients) download a global ML model from a centralized server. Through successive rounds of FL, clients train the global model with their own data and only share the trained model parameters (known as local model updates) rather than their raw training data, with the server. Thus, protecting potential commercial intelligence. The local models received on the central server side are then aggregated. Subsequently, this aggregated model is downloaded again by all clients for the next round of training." Here, Ranatunga mentions that the aggregated model is then downloaded again by all clients or nodes. This downloading step refers to incorporating by a first device or node the aggregated machine learning model into its local model. Further, see Ranathunga in page 4452, Section IV. System Architecture, describe in part A. FL Clients "This work deals with a cross-silo FL scenario for Industry 4.0 applications, as such FL clients (C1,C2,…,Cn) are represented as a limited number (2–100) of manufacturing organizations. A global model (G) architecture is agreed upon before the commencement of a federated model training process. Clients utilize on-premise local edge devices (E1,E2,…,En) to train the global model using their local data sets (d1,d2,…,dn) . As shown in Fig. 1, edge devices are responsible for secure model training and orchestration in case of resource unavailability, and sharing their local updates (g1,g2,…,gn) with their respective edge servers." Also, see Ranathunga in page 4452, Section III. Foundations, A. FedAvg for Model Aggregation describe "As introduced in Section I-A, FL enables distributed clients to train a global ML model in a collaborative fashion without exposing their raw data to each other. If dk represents the set of local training data samples for client Ck , the total size of data samples from n clients, D=∑nk=1dk . For an FL training round t , each client trains the global model (with the parameters represented as wt ) with its local data and calculates local parameters wkt+1 . After the completion of local training, every client uploads it is local model parameters (wkt+1) to the global server wkt+1=wt−ηk∂k(1) PNG media_image5.png 3 1036 media_image5.png Greyscale where, ηk and ∂k represent learning rate and average gradient of Ck computed on its local data set with the current global model (wt) . FedAvg is the algorithm to aggregate such local model updates received from multiple client in order to create a more robust global model (wt+1) for the next round of training*"* Here, Ranathunga shows that each device or node E1, E2, and subsequent devices incorporate updated model updates into its local model. An example is a first node E1 incorporates updated model information into its version of a local model. Further, Ranathunga teaches “incorporating, by the second node, the aggregated machine learning model into the second version of the machine learning model;” See Ranathunga in page 4449, in Introduction, part A. Background mention "Traditional FL is centralized and works based on a client-server paradigm where multiple parties (clients) download a global ML model from a centralized server. Through successive rounds of FL, clients train the global model with their own data and only share the trained model parameters (known as local model updates) rather than their raw training data, with the server. Thus, protecting potential commercial intelligence. The local models received on the central server side are then aggregated. Subsequently, this aggregated model is downloaded again by all clients for the next round of training." Here, Ranathunga mentions that the aggregated model is then downloaded again by all clients or nodes. This downloading step refers to incorporating by a second device or node the aggregated machine learning model into its local model. Further, see Ranathunga in page 4452, Section IV. System Architecture, describe in part A. FL Clients "This work deals with a cross-silo FL scenario for Industry 4.0 applications, as such FL clients (C1,C2,…,Cn) are represented as a limited number (2–100) of manufacturing organizations. A global model (G) architecture is agreed upon before the commencement of a federated model training process. Clients utilize on-premise local edge devices (E1,E2,…,En) to train the global model using their local data sets (d1,d2,…,dn) . As shown in Fig. 1, edge devices are responsible for secure model training and orchestration in case of resource unavailability, and sharing their local updates (g1,g2,…,gn) with their respective edge servers." Also, see Ranathunga in page 4452, Section III. Foundations, A. FedAvg for Model Aggregation describe “As introduced in Section I-A, FL enables distributed clients to train a global ML model in a collaborative fashion without exposing their raw data to each other. If dk represents the set of local training data samples for client Ck , the total size of data samples from n clients, D=∑nk=1dk . For an FL training round t , each client trains the global model (with the parameters represented as wt ) with its local data and calculates local parameters wkt+1 . After the completion of local training, every client uploads it is local model parameters (wkt+1) to the global server wkt+1=wt−ηk∂k(1) PNG media_image5.png 3 1036 media_image5.png Greyscale where, ηk and ∂k represent learning rate and average gradient of Ck computed on its local data set with the current global model (wt) . FedAvg is the algorithm to aggregate such local model updates received from multiple client in order to create a more robust global model (wt+1) for the next round of training.” Here, Ranathunga shows that each device or node E1, E2, and subsequent devices incorporate updated model updates into its local model. An example is a second node E2 incorporates updated model information into its version of a local model. Further, Ranathunga teaches “aggregating, by the aggregation server, the updated first version of the aggregated machine learning model and the updated second version of the aggregated machine learning model, resulting in a second aggregated machine learning model;” See Ranathunga in page 4452, Section III. Foundations, A. FedAvg for Model Aggregation describe "As introduced in Section I-A, FL enables distributed clients to train a global ML model in a collaborative fashion without exposing their raw data to each other. If dk represents the set of local training data samples for client Ck , the total size of data samples from n clients, D=∑nk=1dk . For an FL training round t , each client trains the global model (with the parameters represented as wt ) with its local data and calculates local parameters wkt+1 . After the completion of local training, every client uploads it is local model parameters (wkt+1) to the global server wkt+1=wt−ηk∂k(1) PNG media_image5.png 3 1036 media_image5.png Greyscale where, ηk and ∂k represent learning rate and average gradient of Ck computed on its local data set with the current global model (wt) . FedAvg is the algorithm to aggregate such local model updates received from multiple client in order to create a more robust global model (wt+1) for the next round of training wt+1=∑k=1ndkDwkt+1.(2) PNG media_image5.png 3 1036 media_image5.png Greyscale " Here, Ranathunga mentions that the global server (i.e. aggregation server) uses the FedAvg algorithm to aggregate updates from local models received from multiple client devices or nodes (including updated first version and second version of local models) into creating a more robust global model (i.e. second aggregated machine learning model). Further, Ranathunga teaches “and sending, by the aggregation server, the second aggregated machine learning model to the first node and the second node. See Ranathunga in page 4449, in Introduction, part A. Background mention "Traditional FL is centralized and works based on a client-server paradigm where multiple parties (clients) download a global ML model from a centralized server. Through successive rounds of FL, clients train the global model with their own data and only share the trained model parameters (known as local model updates) rather than their raw training data, with the server. Thus, protecting potential commercial intelligence. The local models received on the central server side are then aggregated. Subsequently, this aggregated model is downloaded again by all clients for the next round of training." Here, Ranathunga mentions that the aggregated model is then downloaded again by all clients or nodes. This downloading step refers to sending the second aggregated machine learning model by all the devices including a first node and second node. Examiner construes second aggregated machine learning model to mean an updated model after it has been aggregated by local models of devices. However, Lim in view of Ranathunga, further in view of Hu, did not teach “training, by the first node, the first version of the machine learning model with an additional first private dataset, …” In an analogous art, Aledhari teaches “training, by the first node, the first version of the machine learning model with an additional first private dataset, …” See Aledhari in page 140702, Section II. Related Works, C. Federated Learning, describe “Most techniques for personalizing a model usually require two steps; In the first step, a global model is built. In the second step, the global model is modified for each client via private data from that client.” Aledhari here mentions each client device or node has its own private data. Also, see Aledhari in page 140704, Section III. Architectures and Platforms of Federated Learning describe “These parameters are the private information of each worker and they are kept unknown to the master as well as other workers.” Here, Aledhari mentions each worker device or node has its own private data. Further, see Aledhari in page 140705, Section III. Architectures and Platforms of Federated Learning, step 3, "3) The Personalization Stage - In this last stage, each device trains a personalized model in order to capture specific characteristics and requirements. This is based on the global model’s information and its own personal intel. The specific learning operations at this stage depend on the adopted federated learning mechanism." Each device trains its own local model. Further, see Aledhari in page 140706, Section III. Architectures and Platforms of Federated Learning, describe “According to the authors, the FADL framework trains some parts of a model using all data sources plus additional parts using data via certain data resources. To test their framework, the authors used ICU hospital data.” Here, Aledhari mentions that the additional private data includes additional parts, in addition to the private data already used by the device. The data sources here also include private data from each client. Aledhari describes that this method uses this type of data along with additional data parts for training. Further, Aledhari teaches “training, by the second node, the aggregated machine learning model with an additional second private dataset, resulting in an updated second version of the aggregated machine learning model;” See Aledhari in page 140702, Section II. Related Works, C. Federated Learning, describe “Most techniques for personalizing a model usually require two steps; In the first step, a global model is built. In the second step, the global model is modified for each client via private data from that client.” Aledhari here mentions each client device or node has its own private data. Also, see Aledhari in page 140704, Section III. Architectures and Platforms of Federated Learning describe “These parameters are the private information of each worker and they are kept unknown to the master as well as other workers.” Here, Aledhari mentions each worker device or node has its own set of private data. Further, see Aledhari in page 140705, Section III. Architectures and Platforms of Federated Learning, step 3, "3) The Personalization Stage - In this last stage, each device trains a personalized model in order to capture specific characteristics and requirements. This is based on the global model’s information and its own personal intel. The specific learning operations at this stage depend on the adopted federated learning mechanism." Each device trains its own local model. Later, see Aledhari in page 140706, Section III. Architectures and Platforms of Federated Learning, describe “According to the authors, the FADL framework trains some parts of a model using all data sources plus additional parts using data via certain data resources. To test their framework, the authors used ICU hospital data.” Here, Aledhari mentions that the additional private data includes additional parts, in addition to the private data already used by the device. The data sources here also include private data from each client. Aledhari describes that this method uses this type of data along with additional data parts for training. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim, Ranathunga, and Hu, and incorporate with the teachings of Aledhari by using the teachings of Lim, Ranathunga, and Hu of using two nodes to each train and send models to an aggregation server, with Aledhari’s teaching of an training with an additional private dataset. One of ordinary skill in the art would be motivated to do so because by integrating Aledhari’s framework into the methods of Lim, Ranathunga, and Hu, one with ordinary skill in the art would achieve “Upon experimentation, the authors were able to conclude their proposed framework was able to improve training speed nine times faster without sacrificing accuracy,” (see Aledhari in page 140704, third paragraph of page, part of section III. Architectures and Platforms of Federated Learning). Claim 18: Regarding claim 18, the claim recites similar limitations as corresponding claim 9 and is rejected for similar reasons as claim 9 using similar teachings and rationale. Claims 11 and 20 are rejected under 35 U.S.C. 103 over Lim in view of Ranathunga, further in view of Hu, further in view of Wink, T., et al., in “An approach for peer-to-peer federated learning,” published on August 4, 2021 , available at https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9502443 , (hereafter, Wink), and further in view of Renda, A., et al., in “Federated learning of explainable AI models in 6G systems: Towards secure and automated vehicle networking,” published on August 20, 2022, available at https://www.mdpi.com/2078-2489/13/8/395 , (hereafter, Renda). Claim 11: Regarding claim 11, Lim in view of Ranathunga, further in view of Hu, teach the limitations of claim 10. However, Lim in view of Ranathunga, further in view of Hu did not teach “The method of claim 10, further comprising: incorporating, by the first node, the aggregated machine learning model into the first version of the machine learning model; and deploying, by the first node, the first version of the machine learning model for inference generation to a production environment for an implementing organization, wherein the implementing organization is a participant of the peer-to-peer distributed network.” In an analogous art, Wink teach “The method of claim 10, further comprising: incorporating, by the first node, the aggregated machine learning model into the first version of the machine learning model; and deploying, by the first node, the first version of the machine learning model for inference generation to a production environment for an implementing organization, wherein the implementing organization is a participant of the peer-to-peer distributed network” See Wink in page 151, Section II. Related Work, part A. Federated Optimization describe “Federated optimization (FO) is a method to create and continuously improve a neural-network model based on distributed training data. It was developed at Google to train models with the help of a large number of mobile devices [2 , 11 , 12 , 13 , 14 , 15] and is used today, e.g ., for keyboard prediction [16] . A schematic illustration of a training round in FO is shown in Figure 3 . Upon starting a training round, the central server randomly determines a subset of available nodes. Each selected node trains its local copy of the previously deployed model using its own data and sends updated model weights to the server afterwards. The server computes average values from the weights it has actually received and updates his copy of the model accordingly.” Here, Wink shows a deployment of a selected device node’s local copy of a model onto a server that participates in the peer-to-peer network. Further, see Wink page 151, in section II. Related work, B. Further Distributed Model Training Approaches describe “BrainTorrent [25] is a system to train a common model in a peer-to-peer manner. To start a training round, a participant P first asks all other peers whether they meanwhile updated, i.e ., re-trained their local models.” Here, Wink shows a participant within a peer to peer learning network. Later, see Wink page 151, in section II. Related work, B, Further Distributed Model Training Approaches continue mentioning “In [27] a framework is outlined to share and validate DNN model parameters based on distributed ledger technology. An entity, e.g ., a self-driving car, first trains a locally deployed model using its own locally obtained data. After completing the training, model parameters are added to the Blockchain, thus other users can verify the parameters’ integrity before creating and offering new models.” Here, Wink shows deploying model using local data (relates to running and deploying the model), then added to the Blockchain network environment. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim, Ranathunga, and Hu, and incorporate with the teachings of Wink by using the teachings of Lim, Ranathunga, and Hu of using two nodes to each train and send models to an aggregation server, with Wink’s teaching of deploying a model into the peer to peer distributed network. One of ordinary skill in the art would be motivated to do so because by integrating Wink’s framework into the methods of Lim, Ranathunga, and Hu, one with ordinary skill in the art would achieve “it can help improve data secrecy and privacy protection especially in systems that serve smaller groups of collaborating data owners. Group members are able to collaboratively aggregate all corresponding ML model weights, and thus create a common ML model of a desired quality without involving a central server or third-party entity in the process,” (see Wink in page 150, Introduction, paragraph 7). However, Lim in view of Ranathunga, further in view of Hu, and further in view of Wink, did not teach “and deploying, by the first node, the first version of the machine learning model for inference generation to a production environment for an implementing organization, …” In an analogous field, Renda teaches “and deploying, by the first node, the first version of the machine learning model for inference generation to a production environment for an implementing organization, …” See Renda in page 1, in Introduction describe "Considering the above-mentioned challenges, in this article we envision the use of the federated learning (FL) concept applied jointly with XAI models and discuss its applicability to automated vehicle networking use cases to be encountered in B5G/6G setups." Also, see page 7, figure 2, where Renda illustrates "Figure 2. Example of video flow (red arrows) and related QoS/QoE metrics reporting (dashed black arrows) in a MEC-enabled FED-XAI architecture. Interaction among FM, CE and a real time XAI dashboard is also shown." Renda mentions real time which relates to making predictions in a real time production environment. Further, see Renda in page 8, section 3.2. Details of the Proposed FED-XAI Framework, describe "For building (or updating) the FED-XAI model, the involved FMs train (or update) the local model based on recent data (𝐐𝐎𝐒(𝑖) and 𝐐𝐎𝐄(𝑖) for each 𝑖=𝑛−𝑚,𝑛−𝑚+1, …,𝑛 , where m is a predefined time window), and share it with the CE. Once the CE produces the aggregated FED-XAI model, the latter is sent back to the FMs that will use it to perform the QoE prediction for their corresponding UE. The results of the prediction feed a dashboard that displays them in real time and explains how they were obtained. The above scenario will be evaluated in a real-time distributed testbed, which embodies both the communication and computation aspects of the system, as well as the application logic. The communication is realized by Simu5G, a modular simulator of 3GPP-compliant New Radio based on OMNeT++ [22], which also works in real time and interfaces with external applications [23]." Here, Renda mentions that this model was run in real time. Also, see Renda in page 11, section Standardization Impact of an Interoperable FED-XAI Implementation "An interoperable implementation of the FED-XAI concept with a focus on an automotive scenario is expected to stimulate discussion within Standards Development Organizations (SDOs) on specifying the involved architectural entities (e.g., FED-XAI CE and FMs)," Renda here mentions using this model for an organization. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the references of Lim, Ranathunga, and Hu, and incorporate with the teachings of Renda by using the teachings of Lim, Ranathunga, and Hu of using two nodes to each train and send models to an aggregation server, with Renda’s teaching of deploying a model for inference generation for a production environment. One of ordinary skill in the art would be motivated to do so because by integrating Renda’s framework into the methods of Lim, Ranathunga, and Hu, one with ordinary skill in the art would achieve “on the one side, XAI permits improving user experience of the offered communication services by helping end users trust (by design) that in-network AI functionality issues appropriate action recommendations. On the other side, FL ensures security and privacy of both vehicular and user data across the whole system,” (see Renda, in abstract, page 1). Claim 20: Regarding claim 20, the claim recites similar limitations as corresponding claim 11 and is rejected for similar reasons as claim 11 using similar teachings and rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WENWEI ZENG whose telephone number is (571)272-7111. The examiner can normally be reached Monday-Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WenWei Zeng/Examiner, Art Unit 2146 /SHAHID K KHAN/Primary Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Mar 25, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month