DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to pending claims 1-20 filed 1/19/2024.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 6-7, 10-19 are rejected under 35 U.S.C. 103 as being unpatentable over Li-cross ("Incentive and knowledge distillation based federated learning for cross-silo applications", published 2022) in view of Rachakonda ("Privacy enhancing and scalable federated learning to accelerate AI implementation in cross-silo and IoMT environments", published 2022).
For claim 1, Li discloses: a computer-implemented method comprising:
accessing, by a computing system, a set of server models that includes at least a first server model and a second server model (fig.1, §III.A ¶1-2 discloses various “clients” relative to a central federation server, however, these local clients are serves as well based on their interaction with local workers (fig.1 bottom), such as for data purchases and reward contracts, and hosting of local machine learning models), wherein each server model of the set of server model has a one-to-one correspondence with a separate client-device subset of a plurality of client-device subsets that includes at least a first client-device subset corresponding to the first server model and a second client-device subset corresponding to the second server model (§I¶3 contemplates cross-silo scenarios comprising various isolated data islands, hence, each local server model having its own isolated client device subset, e.g., IoT devices (§III.A ¶2)), and wherein each client-device subset of the plurality of client-device subsets is a subset of a set of client devices and is disjoint from each other client device subset of the plurality of client-device subsets (ibid: data islands or separated data silos contemplates disjoint subsets of client devices);
accessing, by the computing system, a set of distillation training data (§III.C Step 2 contemplates a 5000 element public distillation training dataset);
updating, by the computing system, the second server model via a first knowledge distillation process based on the set of distillation training data (§III.C, Alg.2, particularly lines 6, 13-14 discloses training process, i.e., Train(…) discloses updating server model weights w_i via training based on distillation set X_p), wherein the first server model is employed as a teacher model of the first knowledge distillation process and the second server model is employed as a student model of the first knowledge distillation process (ibid: as the consensus data is based on an average of the various models (see Alg.2: line 8, §III.C eq.5-6), the first model function as a teacher model via the consensus model, and the second functions as a student).
Li does not disclose: causing, by the computing system, a transmission of the updated second server model to at least a first portion of client devices of the second client-device subset.
Rachakonda discloses: causing, by the computing system, a transmission of the updated second server model to at least a first portion of client devices of the second client-device subset (§III.E, fig.5: models are transmitted to worker devices for further inferencing). Hence, combination with the worker-client system of Li §III.A yielding determining heterogenous server models via the distillation process followed by transmission of the local client models to worker via the disclosed federated averaging technique.
It would have been obvious before the effective filing date to one of ordinary skill in the art to modify the method of Li by incorporating the federated averaging technique of Rachakonda. Both concern the art of federated learning in sensitive environments such as healthcare (Li Abs) and the incorporation would have, according to Rachakonda, allow practical implementation of federated learning in the healthcare domain (p.745 col.1 ¶2), allow for federated aggregation for a large number of workers (p.745 col.2 last ¶).
For claim 2, Li modified by Rachakonda discloses the method of claim 1, as described above. Li further discloses: wherein the first knowledge distillation process comprises:
generating, by the computing system, a first set of predictions based on the set of distillation training data and the first server model (§III.C Step 2, Alg.2:6, 13-14: predictions are generated during training in order to bring predictions closer in line with consensus, with first predictions are mapped to averaged consensus data (§III.C eq.5-6), the first predictions being based on the training dataset and the first model, as part of its averaged elements);
generating, by the computing system, a second set of predictions based on the set of distillation training data and the second server model (ibid: additional predictions are generated for other silos based on the distillation training data and the local model);
determining, by the computing system, a first performance metric for the second server model based on a comparison of the first set of predictions and the second set of predictions (ibid: training is performed by the client model to adjust its weights w_i to approach the consensus f_k over the course of the 5000 samples of the public dataset, hence, performance metrics are determine based on the difference between the two sets of predictions for adjusting the weights of the second server model; see also §IV.B disclosing implementation as 2-3 layer neural networks, hence, determining performance metrics for training via backpropagation in the neural networks); and
updating, by the computing system, the second server model based on the first performance metric for the second server model (ibid: weights are adjusted during training based on performance metric (loss function, etc.)).
For claim 6, Li modified by Rachakonda discloses the method of claim 1, as described above. Li modified by Rachakonda further discloses: updating, by the computing system, the first server model via a second knowledge distillation process based on the set of distillation training data, wherein the second server model is employed as a teacher model of the second knowledge distillation process and the first server model is employed as a student model of the second knowledge distillation process (Li §III.C Step 2, Alg.2:6, 13-14: as the distillation process is symmetric, the first server model is updated via a second instantiation of the distillation process, based on the public training dataset, with the second server model employed as a teacher model via the averaging process (§III.C eq.5-6)); and
causing, by the computing system, a transmission of the updated first server model to at least a first portion of client devices of the first client-device subset (Rachakonda §III.E, fig.5).
For claim 7, Li modified by Rachakonda discloses the method of claim 6, as described above. Li further discloses: wherein the second knowledge distillation process comprises:
generating, by the computing system, a first set of predictions based on the set of distillation training data and the second server model (§III.C Step 2, Alg.2:6, 13-14: predictions are generated during training in order to bring predictions closer in line with consensus, with first predictions are mapped to averaged consensus data (§III.C eq.5-6), the first predictions being based on the training dataset and the first model, as part of its averaged elements);
generating, by the computing system, a second set of predictions based on the set of distillation training data and the first server model (ibid: additional predictions are generated for other silos based on the distillation training data and the local model);
determining, by the computing system, a first performance metric for the first server model based on a comparison of the first set of predictions and the second set of predictions (ibid: training is performed by the client model to adjust its weights w_i to approach the consensus f_k over the course of the 5000 samples of the public dataset, hence, performance metrics are determine based on the difference between the two sets of predictions for adjusting the weights of the second server model; see also §IV.B disclosing implementation as 2-3 layer neural networks, hence, determining performance metrics for training via backpropagation in the neural networks); and
updating, by the computing system, the first server model based on the first performance metric for the first server model (ibid: weights are adjusted during training based on performance metric (loss function, etc.)).
For claim 10, Li modified by Rachakonda discloses the method of claim 1, as described above. Li further discloses: wherein the first server model is encoded via a first set of parameters (Li fig.1 contemplates various server models with respective weights) and accessing the set of server models comprises:
accessing, by the computing system, a set of client-level values for the first set of parameters for each client device of the first client-device subset (Rachakonda §III.E, fig.5 shows accessing client level values for workers, such as worker trained model weights); and
determining, by the computing system, a first set of server-level values for the first set of parameters based on a federated averaging process applied to the set of server-level values for the first set of parameters for each client device of the first client-device subset (Rachakonda fig.5, §III.D ¶2: federated averaging is performed on the worker weights to generate server-level values). Hence, combination with the worker-client system of Li §III.A yielding determining heterogenous server models via the disclosed federated averaging.
For claim 11, Li modified by Rachakonda discloses the method of claim 10, as described above. Li modified by Rachakonda further discloses: wherein the second server model is encoded via a second set of parameters (fig.1 contemplates various server models with respective weights) and accessing the set of server models further comprises:
accessing, by the computing system, a set of client-level values for the second set of parameters for each client device of the second client-device subset (§III.E, fig.5 shows accessing client level values for workers, such as worker trained model weights); and
determining, by the computing system, a first set of server-level values for the second set of parameters based on a federated averaging process applied to the set of server-level values for the second set of parameters for each client device of the second client-device subset (fig.5, §III.D ¶2: federated averaging is performed on the worker weights to generate server-level values).
For claim 12, Li modified by Rachakonda discloses the method of claim 11, as described above. Li modified by Rachakonda further discloses: wherein the first knowledge distillation process comprises:
implementing, by the computing system, the teacher model based on the first set of server-level values for the first set of parameters (Li’s disclosure of distillation (fig.1) combined with Rachakonda’s determining parameters based on federated averaging of workers (fig.5, §III.D) yields a technique where the teacher model as the aggregate (§III.C eq.5-6) is base on the first set of server-level values); and
implementing, by the computing system, the student model based on the first set of server-level values for the second set of parameters (Li fig.1, §III.C eq.5-6, Rachakonda fig.5, §III.D, as above: likewise, the student model would be the aggregate of the workers).
For claim 13, Li modified by Rachakonda discloses the method of claim 12, as described above. Li modified by Rachakonda further discloses: wherein the first knowledge distillation process further comprises:
generating, by the computing system, a first set of labeled training data based on the set of distillation training data (Li §III.C eq.5-6 discloses generating labeled training data based on the distillation training dataset), wherein labels for the first set of labeled training data are determined based on logits calculated by the teacher model (ibid: labels for data are based on weighted averaging, hence, based on logits of the teacher model); and
determining, by the computing system, a second set of server-level values for the second set of parameters based on a supervised learning process and the first set of labeled training data (Li §III.C Alg.2:6, 14: second server-value parameters and weights are determined or calibrated based on supervised training on the labeled training data via supervised training); and
updating, by the computing system, the second server model such that the second set of parameters are encoded by the second set of server-level values for the second set of parameters (ibid: parameters are updated according to values during training).
For claim 14, Li modified by Rachakonda discloses the method of claim 11, as described above. Li modified by Rachakonda further discloses: updating, by the computing system, the first server model via a second knowledge distillation process based on the set of distillation training data, wherein the second server model is employed as a teacher model of the second knowledge distillation process and the first server model is employed as a student model of the second knowledge distillation process (Li §III.C eq.5-6, Alg.2:6,14: likewise, the first server model is updated via distillation based on distillation training data via a weighted average based on the second server model); and
causing, by the computing system, a transmission of the updated first server model to at least a first portion of client devices of the first client-device subset (Rachakonda fig.5 discloses provision of models to workers for further inferencing and training).
For claim 15, Li modified by Rachakonda discloses the method of claim 14, as described above. Li modified by Rachakonda further discloses: wherein the second knowledge distillation process comprises:
implementing, by the computing system, the teacher model based on the first set of server-level values for the second set of parameters (Li §III.C eq.5-6: teacher model is weighted average including first set of server-level values for the second set of parameters); and
implementing, by the computing system, the student model based on the first set of server-level values for the first set of parameters (Li §III.C alg.2:6, 14: student model is implemented based on the first set of server-level values for the first set of parameters).
For claim 16, Li modified by Rachakonda discloses the method of claim 15, as described above. Li modified by Rachakonda further discloses: wherein the second knowledge distillation process further comprises:
generating, by the computing system, a first set of labeled training data based on the set of distillation training data (Li §III.C eq.5-6 discloses generating labeled training data based on the distillation training dataset), wherein labels for the first set of labeled training data are determined based on logits calculated by the teacher model (ibid: labels for data are based on weighted averaging, hence, based on logits of the teacher model); and
determining, by the computing system, a second set of server-level values for the first set of parameters based on a supervised learning process and the first set of labeled training data (Li §III.C Alg.2:6, 14: second server-value parameters and weights are determined or calibrated based on supervised training on the labeled training data via supervised training); and
updating, by the computing system, the first server model such that the first set of parameters are encoded by the second set of server-level values for the first set of parameters (ibid: parameters are updated according to values during training).
For claim 17, Li modified by Rachakonda discloses the method of claim 1, as described above. Li modified by Rachakonda further discloses: updating the second server model comprises:
computing, by the computing system, a gradient of the second server model (Rachakonda §III.D ¶2).
Claim 18 discloses a system corresponding to the method of claim 1 and hence is rejected for the same reasons. Rachakonda further discloses: one or more processors (p.753 col.1 ¶1 contemplates application of typical computer hardware including processor to machine learning tasks); and
one or more non-transitory computer-readable media that store instructions that when executed by the one or more processors, cause the computer system to perform operations comprising (ibid: computer memory including RAM for storing dataset, processing instructions).
Claim(s) 19 recite systems analogous to the above methods and are hence rejected for the same reasons.
Claim(s) 3-5, 8-9, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li-cross ("Incentive and knowledge distillation based federated learning for cross-silo applications", published 2022) in view of Rachakonda ("Privacy enhancing and scalable federated learning to accelerate AI implementation in cross-silo and IoMT environments", published 2022) in view of Thakur ("Knowledge Distillation — Make your neural networks smaller", published 6/30/2022).
For claim 3, Li modified by Rachakonda discloses the method of claim 2, as described above. Li further discloses: wherein generating the first set of predictions comprises:
generating, by the computing system, a first distribution for a set of labels based on the first server model and the set of distillation training data (§III.C Step 2, Alg.2:6, 13-14: first distributions for labels (see §IV.A) are generated based on the first server model and the distillation dataset), wherein the set of labels is associated with each of the first server model and the second server model (§III.C Step 2, §IV.A: labels and outputs are shared among models).
Li modified by Rachakonda does not disclose: wherein the first distribution is a probability distribution.
Thakur discloses: wherein the first distribution is a probability distribution (p.2-3: “Knowledge” contemplates comparing models based on probability distribution).
It would have been obvious before the effective filing date to one of ordinary skill in the art to modify Li modified by Rachakonda by incorporating the knowledge distillation framework of Thakur. Both concern the art of knowledge distillation, and the incorporation would have, according to Thakur, allow predictions of probabilities for generating knowledge inferences (p.2-3).
For claim 4, Li modified by Rachakonda discloses the method of claim 3, as described above. Li further discloses: wherein generating the second set of predictions comprises:
generating, by the computing system, a second probability (Thakur p.2-3) distribution for the set of labels based on the second server model and the set of distillation training data (§III.C Step 2, Alg.2:6,13-14: second probability distributions are generated for the set of labels based on training on the public data set).
For claim 5, Li modified by Rachakonda discloses the method of claim 4, as described above. Thakur further discloses: wherein the first performance metric for the second server model includes a relative entropy loss between the first probability distribution and the second probability distribution (p.6 contemplates an overall loss including KL divergence term, which is a relative entropy loss between the two distributions).
For claim 8, Li modified by Rachakonda discloses the method of claim 7, as described above. Li further discloses: wherein generating the first set of predictions comprises:
generating, by the computing system, a first distribution for a set of labels based on the second server model and the set of distillation training data (§III.C Step 2, Alg.2:6, 13-14: first distributions for labels (see §IV.A) are generated based on the first server model and the distillation dataset), wherein the set of labels is associated with each of the first server model and the second server model (§III.C Step 2, §IV.A: labels and outputs are shared among models).
Li modified by Rachakonda does not disclose: wherein the first distribution is a probability distribution.
Thakur discloses: wherein the first distribution is a probability distribution (p.2-3: “Knowledge” contemplates comparing models based on probability distribution).
It would have been obvious before the effective filing date to one of ordinary skill in the art to modify Li modified by Rachakonda by incorporating the knowledge distillation framework of Thakur. Both concern the art of knowledge distillation, and the incorporation would have, according to Thakur, allow predictions of probabilities for generating knowledge inferences (p.2-3).
For claim 9, Li modified by Rachakonda modified by Thakur discloses the method of claim 8, as described above. Li further discloses: wherein generating the second set of predictions comprises:
generating, by the computing system, a second probability (Thakur p.2-3) distribution for the set of labels based on the first server model and the set of distillation training data (§III.C Step 2, Alg.2:6,13-14: second probability distributions are generated for the set of labels based on training on the public data set).
Claim(s) 20 recite systems analogous to the above methods and are hence rejected for the same reasons.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Shae (US 20240054352 A1) discloses multi-level distillation clustering based on data distribution, see fig.3B.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LIANG LI whose telephone number is (303)297-4263. The examiner can normally be reached Mon-Fri 9-12p, 3-11p MT (11-2p, 5-1a ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. The examiner is available for interviews Mon-Fri 6-11a, 2-7p MT (8-1p, 4-9p ET).
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Jennifer Welch can be reached on (571)272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center and the Private Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from Patent Center or Private PAIR. Status information for unpublished applications is available through Patent Center or Private PAIR to authorized users only. Should you have questions about access to Patent Center or the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/LIANG LI/
Primary examiner AU 2143