DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In view of the appeal brief filed on 5 June 2026, PROSECUTION IS HEREBY REOPENED. New grounds of rejection are set forth below.
To avoid abandonment of the application, appellant must exercise one of the following two options:
(1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or,
(2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid.
A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing below:
/OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 5, 8, 10-14, 16-18, 21, 23, and 31 is/are rejected under 35 U.S.C. 103 as being unpatentable over Szeto (US 11,461,690) in view of Sivaraman (US 2021/0397885).
As per claim 1, Szeto teaches a method for distributed learning at a local computing device [a distributed, online machine learning system including private data servers training and implementing machine learning algorithms (abstract, etc.); where the private servers are local computing devices], the method comprising:
training a local model of a first model type on local data, wherein the local data comprises a first set of labels [the modeling engine generates one or more private data distributions from the local private training data set, and uses the private data distribution(s) to generate a (local) set of proxy data (col. 4, lines 47-56; etc.) and the (local) proxy model is trained on the synthetic (local) proxy data constructed from the (local) private data (col. 4, line 64 to col. 5, line 1; col. 24, lines 36-39; etc.); where the training can include supervised training utilizing inputs associated with known outputs as the training data (col. 8, lines 39-44; etc.), which associated outputs can be class labels (col. 8, lines 59-65; etc.); where the locally created training data used to train the local proxy model is therefore local data comprising a first set of labels];
testing the local model on a portion of global data pertaining to the first set of labels [the modeling engine generates one or more private data distributions from the local private training data set, and uses the private data distribution(s) to generate a (local) set of proxy data (col. 4, lines 47-56; etc.), the proxy data can be partitioned into training and validations sets, where the training proxy data is used to generate the trained proxy model and then the validation proxy data would be provided to the trained proxy model for validation (testing) (col. 26, lines 19-27; etc.) and, if the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (col. 5, lines 12-17; etc.), including aggregating received proxy data into a global proxy set (col. 28, line 65 to col. 29, line 3); therefore the proxy data used for training and testing the local proxy model is a subset/portion of the aggregated global (proxy) data used pertaining to the labels/data used by the local machine for training], wherein the global data comprises a second set of labels and the first set of labels is a strict subset of the second set of labels [If the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (col. 5, lines 12-17; etc.), including aggregating received proxy data into a global proxy set (col. 28, line 65 to col. 29, line 3), where the proxy data can be partitioned into training and validations sets, which can then be used for cross-fold validation, where the training proxy data is used to generate the trained proxy model and then the validation proxy data would be provided to the trained proxy model for validation. Additionally/alternatively, the trained proxy model can be sent to other modeling engines in the ecosystem which can then attempt to validate the trained proxy model on their respective similar training data sets (col. 26, lines 19-33; etc.); therefore the subset of proxy data used for testing the local proxy model is a subset/portion of the aggregated global (proxy) data pertaining to the labels/data used by the local machine for training (which training proxy data is another subset of the local proxy data, which is a subset of the global proxy data)];
as a result of testing the local model on the portion of the global data pertaining to the first set of labels, producing [data] corresponding to the first set of labels [Once each of the validating devices completes their validation efforts, results are provided back to the original modeling engine for evaluation and derivation of model similarity score (col. 26, lines 33-37; etc.) where, if the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (col. 5, lines 12-17; etc.), which proxy data can include class labels (col. 8, lines 59-65; col. 34, lines 18-21; etc.); where sending the data based upon the outcome of the validation and similarity score is a result of testing the local model on the portion of the global data pertaining to the first set of labels (see above)]; and
sending [the data] corresponding to the first set of labels to a central computing device [Once each of the validating devices completes their validation efforts, results are provided back to the original modeling engine for evaluation and derivation of model similarity score (col. 26, lines 33-37; etc.) where, if the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (col. 5, lines 12-17; etc.), which proxy data can include class labels (col. 8, lines 59-65; col. 34, lines 18-21; etc.)].
While Szeto teaches producing a set of data corresponding to the first set of labels as a result of testing the local model and sending the data to a central computing device (see above) as well as sending proxy related data along with the results (see, e.g., Szeto: col. 10, lines 51-59), it has not been relied upon for teaching that the data includes a first set of probabilities corresponding to the first set of labels.
Sivaraman teaches producing a first set of probabilities corresponding to the first set of labels [the classifier can include a CNN (or DNN) and a softmax function, which generates confidence scores from the outputs of the CNN, turning the outputs into probabilities that add up to one (paras. 0035-36, 0039; figs. 1-2; etc.); (Examiner’s Note: fig. 1 also shows additional elements, including a class assignment component including merging classifications and other values that also produces output probabilities. However, this is not relevant to the combination relied upon)].
Szeto and Sivaraman are analogous art, as they are within the same field of endeavor, namely training classifiers for image data, including in distributed learning on multiple devices/systems.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize a softmax final layer on a classifier to output multiple confidence/probability values for multiple classes, as taught by Sivaraman, for the classifier outputs in the system taught by Szeto providing output data to a central device/server, to produce the set of probabilities and send them to the central device/server.
Because both Szeto and Sivaraman teach training classifiers for classifying images, it would have been obvious to one of ordinary skill in the art to utilize a softmax final layer on a classifier to output multiple confidence/probability values for multiple classes, as taught by Sivaraman, for the classifier in the system taught by Szeto providing output data to a central device/server, to produce the set of probabilities and send them to the central device/server, to achieve the predictable result of providing a multi-class classifier output for the model(s) when multiple classes are useful. Sivaraman also provides motivation as [providing the softmax and/or normalization allows multiclass classification for a variety of desired inputs and classes (abstract; para. 0038, etc.)]. Additionally, it would provide more proxy related information, which can be used by the central server to determine properties of the models (see, e.g., Szeto: col. 10, lines 51-59).
As per claim 2, Szeto/Sivaraman teaches receiving a second set of probabilities from the central computing device; and updating the local model based on the second set of probabilities [As proxy data is generated and relayed to the global model server, the global model server aggregates the data and generates an updated global model. Once the global model is updated, it can be determined whether the updated global model is an improvement over the previous version of the global model. If the updated global model is an improvement (e.g., the predictive accuracy is improved), new parameters may be provided to the private data servers via the updated model instructions 230. At the private data server 124, the performance of the trained actual model (e.g., whether the model improves or worsens) can be evaluated to determine whether the models instructions provided by the updated global model result in an improved trained actual model. Parameters associated with various machine learning model versions may be stored so that earlier machine learning models may be later retrieved, if needed (Szeto: col. 11, lines 44-59; etc.), which includes updating the trained actual model from these updates, and/or upon a schedule, determining certain shifts, etc. (Szeto: col. 18, lines 5-25; etc.) or when the local private data server determines that the model accuracy is low, in which case it may request additional updates from the global model server (Szeto: col. 16, lines 10-14; etc.); where the model output data can include a softmax function, which generates confidence scores from the outputs of the CNN, turning the outputs into probabilities that add up to one (Sivaraman: paras. 0035-36, 0039; figs. 1-2; etc.); thus, providing the probabilities output by the global model as part of the proxy related information associated with the classifier labels from the global model, and updating the local model based upon these outputs].
As per claim 3, Szeto/Sivaraman teaches, after training the local model of a first model type on local data, distilling the local model to create a distilled local model of a second model type [the modelling engine creates a trained actual model using local private training data, then generates one or more private data distributions from the local private training data set, and uses the private data distribution(s) to generate a (local) set of proxy data, which it uses to train a local proxy model (Szeto: col. 4, line 37 to col. 5, line 1; col. 24, lines 36-39; etc.); in this case the local model trained on the local private data is training the local model of a first model type on local data, and creating the proxy data and training the proxy model is distilling the local model to create a distilled local model of a second model type (Examiner’s Note: while this changes the rejection and relies on a different local model as the local model of the first type, it is consistent with the claim language changing what testing the local model refers to (testing the distilled model instead))],
wherein testing the local model on a portion of the global data pertaining to the first set of labels comprises testing the distilled local model of the second model type [the modeling engine generates one or more private data distributions from the local private training data set, and uses the private data distribution(s) to generate a (local) set of proxy data (Szeto: col. 4, lines 47-56; etc.), the proxy data can be partitioned into training and validations sets, where the training proxy data is used to generate the trained proxy model and then the validation proxy data would be provided to the trained proxy model for validation (testing) (Szeto: col. 26, lines 19-27; etc.) and, if the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (Szeto: col. 5, lines 12-17; etc.), including aggregating received proxy data into a global proxy set (Szeto: col. 28, line 65 to col. 29, line 3); where in this case, testing the local proxy model on portion of the aggregated global proxy data pertaining to the first set of labels is testing the distilled model of the second model type (the local proxy model). (Examiner’s Note: while this changes the rejection and relies on a different local model as the local model of the first type, it is consistent with the claim language changing what testing the local model refers to (testing the distilled model instead))].
As per claim 5, Szeto/Sivaraman teaches wherein the first set of probabilities correspond to softmax probabilities computed by the local model [where the model output data can include a softmax function, which generates confidence scores from the outputs of the CNN, turning the outputs into probabilities that add up to one (Sivaraman: paras. 0035-36, 0039; figs. 1-2; etc.)].
As per claim 8, Szeto teaches a method for distributed learning at a central computing device, the method comprising:
providing a central model of a first model type [The global model server using the global modeling engine aggregates data from different locations into a global model (col. 16, lines 12-14; etc.)];
receiving a first set of [data] corresponding to a first set of labels from a first local computing device [Once each of the validating devices completes their validation efforts, results are provided back to the original modeling engine for evaluation and derivation of model similarity score (col. 26, lines 33-37; etc.) where, if the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (col. 5, lines 12-17; etc.), which proxy data can include class labels (col. 8, lines 59-65; col. 34, lines 18-21; etc.); from multiple devices, one of which would provide the first set of data/labels from a first local computing device, while another provides the second set of labels from a second computing device, etc. (see, e.g., fig. 1)];
receiving a second set of [data] corresponding to a second set of labels from a second local computing device, wherein the second set of labels is different than the first set of labels [Once each of the validating devices completes their validation efforts, results are provided back to the original modeling engine for evaluation and derivation of model similarity score (col. 26, lines 33-37; etc.) where, if the model similarity requirement is satisfied during validation, the modeling engine can transmit the set of proxy data to a non-private computing device to integrate the proxy data into an aggregated model (col. 5, lines 12-17; etc.), which proxy data can include class labels (col. 8, lines 59-65; col. 34, lines 18-21; etc.); from multiple devices, one of which would provide the first set of labels from a first local computing device, while another provides the second set of labels from a second computing device, etc. (see, e.g., fig. 1, which shows different devices with different data and different models)];
updating the central model by combining the first and second set of [data] based on the first and second set of labels [The global model server using the global modeling engine aggregates data from different locations into the global model (col. 16, lines 12-14; etc.), where, as proxy data 260 is generated and relayed to the global model server 130, the global model server aggregates the data and generates an updated global model (col. 11, lines 44-46; etc.)]; and
sending model parameters for the updated central model to one or more of the first and second local computing devices [As proxy data is generated and relayed to the global model server, the global model server aggregates the data and generates an updated global model. Once the global model is updated, it can be determined whether the updated global model is an improvement over the previous version of the global model. If the updated global model is an improvement (e.g., the predictive accuracy is improved), new parameters may be provided to the private data servers via the updated model instructions 230. At the private data server 124, the performance of the trained actual model (e.g., whether the model improves or worsens) can be evaluated to determine whether the models instructions provided by the updated global model result in an improved trained actual model. Parameters associated with various machine learning model versions may be stored so that earlier machine learning models may be later retrieved, if needed (col. 11, lines 44-59; etc.), which includes updating the trained actual model from these updates, and/or upon a schedule, determining certain shifts, etc. (col. 18, lines 5-25; etc.)].
While Szeto teaches producing sets of data corresponding to the different sets of labels as a result of testing the local models on each device, and sending the data to a central computing device (see above) as well as sending proxy related data along with the results (see, e.g., Szeto: col. 10, lines 51-59), it has not been relied upon for teaching that the first and second sets of data include a first set of probabilities corresponding to the first set of labels and a second set of probabilities corresponding to the second set of labels.
Sivaraman teaches a first set of probabilities corresponding to the first set of labels and a second set of probabilities corresponding to the second set of labels [the classifier can include a CNN (or DNN) and a softmax function, which generates confidence scores from the outputs of the CNN, turning the outputs into probabilities that add up to one (paras. 0035-36, 0039; figs. 1-2; etc.); for each of the classifiers of each of the devices in Szeto, above].
Szeto and Sivaraman are analogous art, as they are within the same field of endeavor, namely training classifiers for image data, including in distributed learning on multiple devices/systems.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize a softmax final layer on a classifier to output multiple confidence/probability values for multiple classes, as taught by Sivaraman, for the outputs of the classifiers in the multiple devices taught by Szeto providing output data to a central device/server, to produce the sets of probabilities and send them to the central device/server.
Because both Szeto and Sivaraman teach training classifiers for classifying images, it would have been obvious to one of ordinary skill in the art to utilize a softmax final layer on a classifier to output multiple confidence/probability values for multiple classes, as taught by Sivaraman, for the outputs of the classifiers in the multiple devices taught by Szeto providing output data to a central device/server, to produce the sets of probabilities and send them to the central device/server, to achieve the predictable result of providing a multi-class classifier output for the model(s) when multiple classes are useful. Sivaraman also provides motivation as [providing the softmax and/or normalization allows multiclass classification for a variety of desired inputs and classes (abstract; para. 0038, etc.)]. Additionally, it would provide more proxy related information, which can be used by the central server to determine properties of the models (see, e.g., Szeto: col. 10, lines 51-59).
As per claim 10, Szeto/Sivaraman teaches wherein updating the central model by combining the first and second set of probabilities based on the first and second set of labels comprises averaging probabilities of the first and second sets of probabilities corresponding to labels belonging to both the first and second set of labels [The global model server using the global modeling engine aggregates data from different locations into a global model (Szeto: col. 16, lines 12-14; etc.) where each classifier can include a CNN (or DNN) and a softmax function, which generates confidence scores from the outputs of the CNN, turning the outputs into probabilities that add up to one (Sivaraman: paras. 0035-36, 0039; figs. 1-2; etc.) and where the outputs can be merged by averaging (Sivaraman: para. 0041; Sivaraman: col. 21, lines 10-14; col. 25, lines 54-60; etc.)].
As per claim 11, Szeto/Sivaraman teaches wherein updating the central model by combining the first and second set of probabilities based on the first and second set of labels further comprises normalizing the combined first and second sets of probabilities [The global model server using the global modeling engine aggregates data from different locations into a global model (Szeto: col. 16, lines 12-14; etc.) where a normalization process may be applied to adjust the classification confidence scores (Sivaraman: paras. 0006, 0022-23, etc.)].
As per claim 12, Szeto/Sivaraman teaches wherein sending model parameters for the updated central model to one or more of the first and second local computing devices comprises sending model parameters for the updated central model to both of the first and second local computing devices [As proxy data is generated and relayed to the global model server, the global model server aggregates the data and generates an updated global model. Once the global model is updated, it can be determined whether the updated global model is an improvement over the previous version of the global model. If the updated global model is an improvement (e.g., the predictive accuracy is improved), new parameters may be provided to the private data servers via the updated model instructions 230. At the private data server 124, the performance of the trained actual model (e.g., whether the model improves or worsens) can be evaluated to determine whether the models instructions provided by the updated global model result in an improved trained actual model. Parameters associated with various machine learning model versions may be stored so that earlier machine learning models may be later retrieved, if needed (Szeto: col. 11, lines 44-59; etc.), which includes updating the trained actual model from these updates, and/or upon a schedule, determining certain shifts, etc. (Szeto: col. 18, lines 5-25; etc.); which includes sending the global model parameters for the updated central model to multiple (both the first and second) local devices].
As per claim 13, Szeto/Sivaraman teaches sending to both of the first and second local computing devices information about a common model type, and wherein the first and second sets of probabilities are model parameters based on the common model type [As proxy data is generated and relayed to the global model server, the global model server aggregates the data and generates an updated global model. Once the global model is updated, it can be determined whether the updated global model is an improvement over the previous version of the global model. If the updated global model is an improvement (e.g., the predictive accuracy is improved), new parameters may be provided to the private data servers via the updated model instructions 230. At the private data server 124, the performance of the trained actual model (e.g., whether the model improves or worsens) can be evaluated to determine whether the models instructions provided by the updated global model result in an improved trained actual model. Parameters associated with various machine learning model versions may be stored so that earlier machine learning models may be later retrieved, if needed (Szeto: col. 11, lines 44-59; etc.), which includes updating the trained actual model from these updates, and/or upon a schedule, determining certain shifts, etc. (Szeto: col. 18, lines 5-25; etc.); which includes parameters for classifiers (common model type) in both devices; and where the models can be the same type or implementation of machine learning algorithm (Szeto: col. 28, lines 5-8; etc.)].
As per claim 14, Szeto/Sivaraman teaches wherein the central model is a classifier-type model [the aggregated global model can be a classifier (Szeto: col. 21, lines 45-48; etc.)].
As per claim 16, see the rejection of claim 1, above, wherein Szeto/Sivaraman also teaches a user computing device comprising:
a memory;
a processor coupled to the memory, wherein the processor is configured to: [perform the method] [the system can be implemented across multiple devices via instructions executed from memories by one or more processors (Szeto: col. 6, lines 28-60; etc.)].
As per claim 17, see the rejection of claim 2, above.
As per claim 18, see the rejection of claim 3, above.
As per claim 21, Szeto/Sivaraman teaches wherein the local model is a classifier-type model [the local model can operate as a classifier (Szeto: col. 21, lines 45-48) producing class labels (Szeto: col. 8, lines 59-65; etc.)].
As per claim 23, see the rejection of claim 8, above, wherein Szeto/Sivaraman also teaches a central computing device or server comprising:
a memory; and
a processor coupled to the memory, wherein the processor is configured to: [perform the method] [the system can be implemented across multiple devices via instructions executed from memories by one or more processors (Szeto: col. 6, lines 28-60; etc.)].
As per claim 31, Szeto/Sivaraman teaches a non-transitory computer readable storage medium storing a computer program comprising instructions which when executed by processing circuitry causes the processing circuitry to perform the method of claim 1 [the system can be implemented across multiple devices via instructions executed from memories (storage medium) by one or more processors (Szeto: col. 6, lines 28-60; etc.)].
Claim(s) 4 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Szeto (US 11,461,690), in view of Sivaraman (US 2021/0397885), and further in view of Jeong et al. (Communication-Efficient On-Device Machine Learning: Federated Distillation and Augmentation under Non-IID Private Data, Nov 2018, pgs. 1-6)
As per claim 4, Szeto/Sivaraman teaches the method of claim 2, as described above.
While Szeto/Sivaraman teaches updating the local model comprising averaging (see above), it has not been relied upon for teaching wherein updating the local model based on the second set of probabilities comprises a weighted average of the local model with a version of the local model from a previous iteration.
Jeong teaches wherein updating the local model based on the second set of probabilities comprises a weighted average of the local model with a version of the local model from a previous iteration [the server (central computing device) produces a global-average logit vector per label, from the per-label local-average logit vectors received from each device, which are then downloaded to each device and used to update its model (pg. 2, section 2; pg. 3, Algorithm 1; etc.); where the logit averaging process (used for updating the local model based on the second set of probabilities) can use a weighted average (pg. 5, section 5; etc.); for updating the local model based on the second set of probabilities of Szeto/Sivaraman, above].
Szeto/Sivaraman and Jeong are analogous art, as they are within the same field of endeavor, namely distributed/federated training of models.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize weighted averaging of local models for updating the local model, as taught by Jeong, in the updating of local models using sets of probabilities in the system taught by Szeto/Sivaraman.
Jeong provides motivation as [utilizing the weighted average in a way that the weight increases with local computation time improves the performance of the model(s) (Jeong: pg. 5, section 5; etc.)].
As per claim 19, see the rejection of claim 4, above.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Szeto (US 11,461,690), in view of Sivaraman (US 2021/0397885), and further in view of Vasseur (US 2015/0193694).
As per claim 6, Szeto/Sivaraman teaches wherein the local model is a classifier-type model [the local model can operate as a classifier (Szeto: col. 21, lines 45-48) producing class labels (Szeto: col. 8, lines 59-65; etc.)],
While Szeto/Sivaraman teaches utilizing local classifiers (see above) for multiple types of data (see, e.g., Szeto: col. 7, lines 62-67; etc.), it has not been relied upon for teaching that the local data corresponds to an alarm dataset for a telecommunications operator.
Vasseur teaches wherein the local model is a classifier-type model, and the local data corresponds to an alarm dataset for a telecommunications operator [a system of classifiers is trained to detect attacks on a network and generate alarms (for the telecommunications operator of the network) (para. 0077, etc.), which can include distributed training of local models and training of a central global model (para. 0110, etc.)].
Szeto/Sivaraman and Vasseur are analogous art, as they are within the same field of endeavor, namely distributed learning.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to use distributed learning to train multiple local and global classifier models to classify attack/alarm data, as taught by Vasseur, using the distributed learning system/method for training models taught by Szeto/Sivaraman.
Vasseur provides motivation as [using a distributed learning method to train classifiers to detect attacks and other network conditions in large-scale networks provide protection while allowing the models to work in a very constrained environment with limited overhead (para. 0076, etc.)].
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Szeto (US 11,461,690), in view of Sivaraman (US 2021/0397885), and further in view of Gou (US 2020/0110982).
As per claim 9, Szeto/Sivaraman teaches the method of claim 8, as described above.
While Szeto/Sivaraman teaches a system performing distributed learning, including updating model parameters and creating proxy data/models (see above), it has not been relied upon for teaching distilling the updated central model to create a distilled central model of a second model type, and wherein the model parameters for the updated central model correspond to the distilled central model of the second model type.
Gou teaches distilling the updated central model to create a distilled central model of a second model type, and wherein the model parameters for the updated central model correspond to the distilled central model of the second model type [knowledge distillation may be used to compress cumbersome models (e.g, first predictive model 410) into light-weight models (e.g., second predictive model 420) for various purposes (e.g., simplifying the deployment of a model and/or the like). For example, a small second (e.g., student) predictive model 420 may be trained using knowledge distilled from a cumbersome first (e.g., teacher) predictive model 410 (which may be pre-trained, trained before training the second model, and/or the like) (para. 0160, etc.); for the updating the central model of Szeto/Sivaraman, above].
Szeto/Sivaraman and Gou are analogous art, as they are within the same field of endeavor, namely training ML model using different kinds of knowledge transfer, as well as distributed learning.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to include performing knowledge distillation to compress a model into a model of a different type, as taught by Gou, for updating the central model in the system/method taught by Szeto/Sivaraman.
Gou provides motivation as [distilling the model to a compressed model makes deployment simpler and makes the model less cumbersome (para. 0160, etc.)].
Response to Arguments
Applicant’s arguments, see the appeal brief, filed 5 June 2026, with respect to the rejection(s) of claim(s) 1-6, 8-14, 16-19, 21, 23, and 31 under 35 U.S.C. 102/13 have been fully considered and are persuasive, regarding the teachings of Jeong and producing the first set of probabilities as a result of testing the local model on the portion of the global data. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Szeto and Sivaraman, which have been relied upon for teaching the training, testing, and production of the probabilities as described in the rejections above.
Conclusion
The following is a summary of the treatment and status of all claims in the application as recommended by M.P.E.P. 707.07(i): claims 7, 15, 20, 22, 24-30, and 32 are cancelled; claims 1-6, 8-14, 16-19, 21, 23, and 31 are rejected.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Deng (US 2020/0401886) and Deng (US 11,868,884) – discloses a federated learning service, including a central model and passing model output probability scores.
Han (US 11,681,913) – discloses a distributed learning system including global and local models, where each local model/terminal device receives a subset of the global data set.
Hassanzadeh (US 2021/0352093) – discloses a distributed learning system including using a subset of global data to test local models.
Blanchard (US 2020/0380340) – discloses a federated learning system including passing estimate vectors and computing an average with a given probability p.
Anil et al. (Large Scale Distributed Neural Network Training Through Online Distillation, April 2018, pgs. 1-12) – discloses a distributed learning method/system including model distillation.
Kairouz et al. (Advances and Open Problems in Federated Learning, Dec 2019, pgs. 1-105 – cited in an IDS) – discloses various systems/methods for federated learning.
Bonawitz et al. (Towards Federated Learning at Scale: System Design, March 2019, pgs. 1-15 – cited in an IDS) – discloses a federated learning system including training on private local data.
Yang et al. (Federated Machine Learning: Concept and Applications, Feb 2019, pgs. 1-19 – cited in an IDS) – discloses various systems/methods for federated learning.
Manamohan (US 11,436,692) – discloses a distributed training system including evaluating each local model to calculating validation metrics and sharing the validation metrics with a leader/central node as validation completes.
Karame (US 2021/0051169) – discloses a system/method for defending against model poisoning, including a centralized defense performing validation/testing with global data and local test datasets.
Das Gupta (US 2020/0351344) – discloses a distributed learning system using edge computers and central hub, including validating a global model on various local edge systems that were not used during training.
Vaidyanathan (US 6,941,287) – discloses a distributed learning system including dividing data into training/validation/testing subsets.
Zhu (US 11,699,080) – discloses distributed learning including using predicted class label probabilities to determine which data should be shared with the central device/server.
Duerig (US 2020/0401929) – discloses a distributed learning system performing knowledge distillation including selecting outputs from different models for inclusion in the distillation data based upon performance evaluations.
The examiner requests, in response to this Office action, that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line number(s) in the specification and/or drawing figure(s). This will assist the examiner in prosecuting the application.
When responding to this office action, Applicant is advised to clearly point out the patentable novelty which he or she thinks the claims present, in view of the state of the art disclosed by the references cited or the objections made. He or she must also show how the amendments avoid such references or objections. See 37 CFR 1.111(c).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GEORGE GIROUX whose telephone number is (571)272-9769. The examiner can normally be reached M-F 10am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571-272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GEORGE GIROUX/Primary Examiner, Art Unit 2128