NON-FINAL REJECTION, FIRST DETAILED ACTION
Status of Prosecution
The present application, 18/404,069 filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The application was filed in the Office on Dec. 4, 2024 and is a continuation of PCT/CN2022/100555 filed on June 22, 2022 which in turn claims priority to Chinese application CN202110770808.8 filed July 7, 2021.
Applicant’s preliminary amendment received Jan 23, 2024 is acknowledged and entered.
Claims 1-20 are pending and all are rejected. Claims 1 and 11 are independent.
Status of Claims
Claims 1, 10-11 and 20 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom, United States Patent 10,153,676 B1, published on Dec. 11, 2018.
Claims 2, 4-5, 12 and 14-15 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom and in further view of Saxton et al., (“Saxton”), United States Patent Application Publication US 2021/0383222 published on Dec. 9, 2021.
Claims 3 and 13 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom in view of Saxton in view of non-patent literature Dinh et al. (“Dinh”), “Federated Learning Over Wireless Networks: Convergence Analysis and Resource Allocation,” published in February 2021.
Claims 6-7, 9, 16-17 and 19 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom in further view of non-patent literature Zhou et al. (“Zhou”), CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes,” published in October 2020.
Claims 8 and 18 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom in view of Zhou and in further view of non-patent literature Su et al. (“Su”), “Data and Channel-Adaptive Sensor Scheduling for Federated Edge Learning via Over-the-Air Gradient Aggregation,” published in July 2021 (manuscript received Jan. 8, 2021).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
A.
Claims 1, 10-11 and 20 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018.
As to Claim 1, Goetz teaches: A method, wherein a plurality of participating nodes jointly train an intelligent model (Goetz: p. 2, clients (i.e. nodes) of a cohort; “Once the client has generated the synthetic data, the client uploads this synthetic data to the server…”), and the method is performed by one of the plurality of participating nodes, the method comprising:
performing a Kth time of model training on the intelligent model, to obtain first gradient information (Goetz: p. 2 “during each training round” (of which one may be a Kth time), data is uploaded; p. 3, synthetic data tries to approximate the gradient update the standard federated learning would have transmitted back to the server, (θ) a “true update” (i.e. first gradient information)); and
sending first synthetic gradient information to a central node, wherein the first synthetic gradient information comprises synthetic information of the first gradient information(Goetz: p. 2, clients (i.e. nodes) of a cohort; “Once the client has generated the synthetic data, the client uploads this synthetic data to the server…”).
Goetz may not explicitly teach: sending first synthetic gradient information to a central node, wherein the first synthetic gradient information comprises synthetic information of the first gradient information and residual gradient information; and
the residual gradient information represents a residual estimate of synthetic gradient information that is not transmitted to the central node before the Kth time of model training, wherein K is a positive integer.
Strom teaches in general concepts related to distributing the training of models over multiple computing nodes (Strom: Abstract). Specifically, Strom teaches that a residual gradient representing the residual estimate of the gradient is retained (Strom: col. 3, lines 49 to 52, “Rather than discarding the smaller update values of the gradient altogether, they can be saved in a "residual gradient" along with elements from previous and subsequent iterations.”).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have modified the Goetz disclosures and teachings by retaining and sending the residual gradient information with the rest of the synthetic information as taught and suggested by Strom. Such a person would have been motivated to do so with a reasonable expectation of success to allow for updates that may be in aggregate be retained and then sent together in a manner when ready to be used to update (Strom: col. 3, lines 66 to col. 4, line 5).
As to Claim 10, Goetz and Strom teaches the limitations of claim 1.
Goetz further teaches: receiving model parameter information from the central node; and performing the Kth time of model training on the intelligent model, wherein the intelligent model is a model configured based on the model parameter information (Goetz: Algorithms 1 and 2 deal with server and client updates, with transmission of model parameters).
As to Claim 11, it is rejected for similar reasons as claim 1. Strom further teaches circuity and transceivers (col. 13, lines 1-3).
As to Claim 20, it is rejected for similar reasons as claim 10.
B.
Claims 2, 4-5, 12 and 14-15 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of Saxton et al., (“Saxton”), United States Patent Application Publication US 2021/0383222 published on Dec. 9, 2021.
As to Claim 2, Goetz and Strom teaches the limitations of claim 1.
Goetz and Strom may not explicitly teach: wherein the residual gradient information comprises a residual estimate of second synthetic gradient information weighted by a weighting coefficient, and the second synthetic gradient information comprises synthetic gradient information that is last sent to the central node before the Kth time of model training.
Saxton teaches in general concepts related to training a neural network by estimating objective function curvature based on current and previous gradients(Saxton: Abstract). Of relevance here is that Saxton teaches that for each iteration of training, the previous gradients’ weight coefficients are updated and adjusted accordingly (Saxton: cls. 7-8).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom by weighting the previous gradient information that hadn’t been sent as taught and suggested by Saxton. Such a person would have done so with an expectation for success, because of updating the curvature of the gradient with better performance (Saxton: par. 0027, the improved performance may allow for reduction in computational resources).
As to Claim 4, Goetz, Strom and Saxton teach the limitations of claim 2.
Goetz, Strom and Saxton as combined further teaches: wherein the second synthetic gradient information comprises synthetic gradient information that is sent to the central node after a Qth time of model training, Q is a positive integer less than K (Examiner asserts that the second synthetic information may be sent at a time Q as it is not restricted by the combination); and
the weighting coefficient is associated with one or more of:
a learning rate of the Kth time of model training, or
a learning rate of the Qth time of model training (Goetz: p. 2, the learning rate is used for the optimization problem of the synthetic gradient problem, which Examiner asserts could be applicable at either of the iterations).
As to Claim 5, Goetz, Strom and Saxton teach the limitations of claim 2.
Strom further teaches: wherein the second synthetic gradient information comprises the synthetic gradient information that is sent to the central node after the Qth time of model training, the residual gradient information further comprises synthetic information of N pieces of gradient information, and the N pieces of gradient information are gradient information that is obtained through N times of model training after the Qth time of model training and before the Kth time of model training and that is not sent to the central node before the Kth time of model training, wherein K is greater than Q, N=K-Q-1, and Q is a positive integer (Strom: col. 3, lines 49 to 52, “Rather than discarding the smaller update values of the gradient altogether, they can be saved in a "residual gradient" along with elements from previous and subsequent iterations.”; Examiner asserts that he different information may be accumulated at different times including at Nth time per the constraints).
As to Claim 12, it is rejected for similar reasons as claim 2.
As to Claim 14, it is rejected for similar reasons as claim 4.
As to Claim 15, it is rejected for similar reasons as claim 5.
C.
Claims 3 and 13 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of Saxton et al., (“Saxton”), United States Patent Application Publication US 2021/0383222 published on Dec. 9, 2021 in view of non-patent literature Dinh et al. (“Dinh”), “Federated Learning Over Wireless Networks: Convergence Analysis and Resource Allocation,” published in February 2021.
As to Claim 3, Goetz, Strom and Saxton teach the limitations of claim 2.
Goetz, Strom and Saxton may not explicitly teach: wherein the residual estimate of the second synthetic gradient information is associated with the second synthetic gradient information, a transmission power corresponding to the second synthetic gradient information, and channel information corresponding to the second synthetic gradient information.
Dinh teaches in general concepts related FEDL, a FL algorithm which can handle heterogeneous UE data without further assumptions except strongly convex and smooth loss functions (Dinh: Abstract). Specifically, Dinh teaches that transmission and channel information is considered in evaluating and optimizing a federated learning system (Dinh: p. 399, “Specifically, the total wall-clock training time of FL includes not only the UE computation time (which depend on UEs’ CPU types and local data sizes) but also the communication time of all UEs (which depends on UEs’ channel gains, transmission power, and local data sizes”). A rate of achievable transmission rate takes into account transmission power and the average channel gain during a training time of the federated learning (Dinh: eq. 14).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom-Saxton by utilizing the transmission power and channel information as taught and suggested by Dinh. Such a person would have done so with an expectation for success, to improve communication efficiency.
As to Claim 13, it is rejected for similar reasons as claim 3.
D.
Claims 6-7, 9, 16-17 and 19 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of non-patent literature Zhou et al. (“Zhou”), CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes,” published in October 2020.
As to Claim 6, Goetz and Strom teaches the limitations of claim 1.
Goetz further teaches: where a transmission power corresponding to the first synthetic gradient information is greater than a power threshold and the method further comprises:
sending the first synthetic gradient information to the central node.
Zhou teaches in general concepts related to dealing with federated learning in a manner to have computation efficiency using Lyapunov optimization theory (Zhou: Abstract). Specifically, Zhou teaches that the transmission power is considered and modeled for a cost formulation (Zhou: p. 9346, “To model this cost, we use SP to denote the data size of both the model parameters and the gradients interacted between the cloud and each edge node and each global iteration, and PBi to denote the price of the bandwidth of the WAN connecting the cloud and the edge node i ∈ I.”).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom by modeling the transmission power cost and to consider it for a threshold amount as taught and suggested by Zhou. Such a person would have done so with an expectation for success, to allow for optimizing the three-way tradeoff among communication, computation and throughput (Zhou: p. 9342, “In response, CEFL leverages Lyapunov optimization theory to rigorously design an online control algorithm, which is able to independently and concurrently make the aforementioned four control decisions in an online fashion, without requiring any future information as a priori.”).
As to Claim 7, Goetz, Strom and Zhou teaches the limitations of claim 6.
Zhou further teaches: where the transmission power of the first synthetic gradient information is associated with communication price metric information (Zhou: p. 9345, resource usage price), channel information corresponding to the first synthetic gradient information (Examiner asserts this is related to the data size and other aspects of the gradient information), and the first synthetic gradient information(Zhou: p. 9346, “to model this cost we use SP to denote the data size of both the model parameters and gradients interacted between the cloud and each node”), wherein the communication price metric information represents a cost volume of communication between one participating node and the central node (Zhou: p. 9344, the number of newly arrived data samples at edge node I, is related to the cost) .
As to Claim 9, Goetz, Strom and Zhou teaches the limitations of claim 7.
Zhou further teaches: receiving the communication price metric information from the central node (Zhou: p. 9350, “Second, for the cloud that solves the problem, it only needs to collect the queue backlog and resource price information from each edge node at each corresponding time slot t.”).
As to Claim 16, it is rejected for similar reasons as claim 6.
As to Claim 17, it is rejected for similar reasons as claim 7.
As to Claim 19, it is rejected for similar reasons as claim 9.
E.
Claims 8 and 18 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of non-patent literature Zhou et al. (“Zhou”), CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes,” published in October 2020 and in further view of non-patent literature Su et al. (“Su”), “Data and Channel-Adaptive Sensor Scheduling for Federated Edge Learning via Over-the-Air Gradient Aggregation,” published in July 2021 (manuscript received Jan. 8, 2021).
As to Claim 8, Goetz, Strom and Zhou teaches the limitations of claim 6.
Goetz, Strom and Zhou may not explicitly teach: wherein the power threshold is in direct proportion to the communication price metric information and/or the power threshold is in direct proportion to an activation power of the participating node, and the communication price metric information represents the cost volume of communication between the participating node and the central node.
Su teaches in general concepts related to over-the-air gradient aggregation and improving communication efficiency for federated edge learning applications (Su: Abstract). Specifically, Su teaches that the activation power of an edge device (i.e. a participating node) is considered in the power threshold (Su: p. 1645, Pon is the activation power for each sensor).
It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom-Zhou by modeling the transmission power cost with activation power proportionally as taught and suggested by Su. Such a person would have done so with an expectation for success, to allow for optimizing the cost metrics for federated learning.
As to Claim 18, it is rejected for similar reasons as claim 8.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES T TSAI whose telephone number is (571)270-3916. The examiner can normally be reached M-F 8-5 Eastern.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES T TSAI/ Primary Examiner, Art Unit 2177