Prosecution Insights
Last updated: August 17, 2026
Application No. 18/404,069

INTELLIGENT MODEL TRAINING METHOD AND APPARATUS

Non-Final OA §103
Filed
Jan 04, 2024
Priority
Jul 07, 2021 — CN 202110770808.8 +1 more
Examiner
TSAI, JAMES T
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
62%
Grant Probability
Moderate
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
192 granted / 307 resolved
+2.5% vs TC avg
Strong +57% interview lift
Without
With
+56.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
35 currently pending
Career history
331
Total Applications
across all art units

Statute-Specific Performance

§101
11.6%
-28.4% vs TC avg
§103
63.2%
+23.2% vs TC avg
§102
10.1%
-29.9% vs TC avg
§112
9.9%
-30.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 307 resolved cases

Office Action

§103
NON-FINAL REJECTION, FIRST DETAILED ACTION Status of Prosecution The present application, 18/404,069 filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . The application was filed in the Office on Dec. 4, 2024 and is a continuation of PCT/CN2022/100555 filed on June 22, 2022 which in turn claims priority to Chinese application CN202110770808.8 filed July 7, 2021. Applicant’s preliminary amendment received Jan 23, 2024 is acknowledged and entered. Claims 1-20 are pending and all are rejected. Claims 1 and 11 are independent. Status of Claims Claims 1, 10-11 and 20 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom, United States Patent 10,153,676 B1, published on Dec. 11, 2018. Claims 2, 4-5, 12 and 14-15 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom and in further view of Saxton et al., (“Saxton”), United States Patent Application Publication US 2021/0383222 published on Dec. 9, 2021. Claims 3 and 13 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom in view of Saxton in view of non-patent literature Dinh et al. (“Dinh”), “Federated Learning Over Wireless Networks: Convergence Analysis and Resource Allocation,” published in February 2021. Claims 6-7, 9, 16-17 and 19 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom in further view of non-patent literature Zhou et al. (“Zhou”), CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes,” published in October 2020. Claims 8 and 18 are rejected under 35 USC. § 103 as being unpatentable over Goetz in view of Strom in view of Zhou and in further view of non-patent literature Su et al. (“Su”), “Data and Channel-Adaptive Sensor Scheduling for Federated Edge Learning via Over-the-Air Gradient Aggregation,” published in July 2021 (manuscript received Jan. 8, 2021). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. A. Claims 1, 10-11 and 20 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018. As to Claim 1, Goetz teaches: A method, wherein a plurality of participating nodes jointly train an intelligent model (Goetz: p. 2, clients (i.e. nodes) of a cohort; “Once the client has generated the synthetic data, the client uploads this synthetic data to the server…”), and the method is performed by one of the plurality of participating nodes, the method comprising: performing a Kth time of model training on the intelligent model, to obtain first gradient information (Goetz: p. 2 “during each training round” (of which one may be a Kth time), data is uploaded; p. 3, synthetic data tries to approximate the gradient update the standard federated learning would have transmitted back to the server, (θ) a “true update” (i.e. first gradient information)); and sending first synthetic gradient information to a central node, wherein the first synthetic gradient information comprises synthetic information of the first gradient information(Goetz: p. 2, clients (i.e. nodes) of a cohort; “Once the client has generated the synthetic data, the client uploads this synthetic data to the server…”). Goetz may not explicitly teach: sending first synthetic gradient information to a central node, wherein the first synthetic gradient information comprises synthetic information of the first gradient information and residual gradient information; and the residual gradient information represents a residual estimate of synthetic gradient information that is not transmitted to the central node before the Kth time of model training, wherein K is a positive integer. Strom teaches in general concepts related to distributing the training of models over multiple computing nodes (Strom: Abstract). Specifically, Strom teaches that a residual gradient representing the residual estimate of the gradient is retained (Strom: col. 3, lines 49 to 52, “Rather than discarding the smaller update values of the gradient altogether, they can be saved in a "residual gradient" along with elements from previous and subsequent iterations.”). It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the application to have modified the Goetz disclosures and teachings by retaining and sending the residual gradient information with the rest of the synthetic information as taught and suggested by Strom. Such a person would have been motivated to do so with a reasonable expectation of success to allow for updates that may be in aggregate be retained and then sent together in a manner when ready to be used to update (Strom: col. 3, lines 66 to col. 4, line 5). As to Claim 10, Goetz and Strom teaches the limitations of claim 1. Goetz further teaches: receiving model parameter information from the central node; and performing the Kth time of model training on the intelligent model, wherein the intelligent model is a model configured based on the model parameter information (Goetz: Algorithms 1 and 2 deal with server and client updates, with transmission of model parameters). As to Claim 11, it is rejected for similar reasons as claim 1. Strom further teaches circuity and transceivers (col. 13, lines 1-3). As to Claim 20, it is rejected for similar reasons as claim 10. B. Claims 2, 4-5, 12 and 14-15 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of Saxton et al., (“Saxton”), United States Patent Application Publication US 2021/0383222 published on Dec. 9, 2021. As to Claim 2, Goetz and Strom teaches the limitations of claim 1. Goetz and Strom may not explicitly teach: wherein the residual gradient information comprises a residual estimate of second synthetic gradient information weighted by a weighting coefficient, and the second synthetic gradient information comprises synthetic gradient information that is last sent to the central node before the Kth time of model training. Saxton teaches in general concepts related to training a neural network by estimating objective function curvature based on current and previous gradients(Saxton: Abstract). Of relevance here is that Saxton teaches that for each iteration of training, the previous gradients’ weight coefficients are updated and adjusted accordingly (Saxton: cls. 7-8). It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom by weighting the previous gradient information that hadn’t been sent as taught and suggested by Saxton. Such a person would have done so with an expectation for success, because of updating the curvature of the gradient with better performance (Saxton: par. 0027, the improved performance may allow for reduction in computational resources). As to Claim 4, Goetz, Strom and Saxton teach the limitations of claim 2. Goetz, Strom and Saxton as combined further teaches: wherein the second synthetic gradient information comprises synthetic gradient information that is sent to the central node after a Qth time of model training, Q is a positive integer less than K (Examiner asserts that the second synthetic information may be sent at a time Q as it is not restricted by the combination); and the weighting coefficient is associated with one or more of: a learning rate of the Kth time of model training, or a learning rate of the Qth time of model training (Goetz: p. 2, the learning rate is used for the optimization problem of the synthetic gradient problem, which Examiner asserts could be applicable at either of the iterations). As to Claim 5, Goetz, Strom and Saxton teach the limitations of claim 2. Strom further teaches: wherein the second synthetic gradient information comprises the synthetic gradient information that is sent to the central node after the Qth time of model training, the residual gradient information further comprises synthetic information of N pieces of gradient information, and the N pieces of gradient information are gradient information that is obtained through N times of model training after the Qth time of model training and before the Kth time of model training and that is not sent to the central node before the Kth time of model training, wherein K is greater than Q, N=K-Q-1, and Q is a positive integer (Strom: col. 3, lines 49 to 52, “Rather than discarding the smaller update values of the gradient altogether, they can be saved in a "residual gradient" along with elements from previous and subsequent iterations.”; Examiner asserts that he different information may be accumulated at different times including at Nth time per the constraints). As to Claim 12, it is rejected for similar reasons as claim 2. As to Claim 14, it is rejected for similar reasons as claim 4. As to Claim 15, it is rejected for similar reasons as claim 5. C. Claims 3 and 13 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of Saxton et al., (“Saxton”), United States Patent Application Publication US 2021/0383222 published on Dec. 9, 2021 in view of non-patent literature Dinh et al. (“Dinh”), “Federated Learning Over Wireless Networks: Convergence Analysis and Resource Allocation,” published in February 2021. As to Claim 3, Goetz, Strom and Saxton teach the limitations of claim 2. Goetz, Strom and Saxton may not explicitly teach: wherein the residual estimate of the second synthetic gradient information is associated with the second synthetic gradient information, a transmission power corresponding to the second synthetic gradient information, and channel information corresponding to the second synthetic gradient information. Dinh teaches in general concepts related FEDL, a FL algorithm which can handle heterogeneous UE data without further assumptions except strongly convex and smooth loss functions (Dinh: Abstract). Specifically, Dinh teaches that transmission and channel information is considered in evaluating and optimizing a federated learning system (Dinh: p. 399, “Specifically, the total wall-clock training time of FL includes not only the UE computation time (which depend on UEs’ CPU types and local data sizes) but also the communication time of all UEs (which depends on UEs’ channel gains, transmission power, and local data sizes”). A rate of achievable transmission rate takes into account transmission power and the average channel gain during a training time of the federated learning (Dinh: eq. 14). It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom-Saxton by utilizing the transmission power and channel information as taught and suggested by Dinh. Such a person would have done so with an expectation for success, to improve communication efficiency. As to Claim 13, it is rejected for similar reasons as claim 3. D. Claims 6-7, 9, 16-17 and 19 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of non-patent literature Zhou et al. (“Zhou”), CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes,” published in October 2020. As to Claim 6, Goetz and Strom teaches the limitations of claim 1. Goetz further teaches: where a transmission power corresponding to the first synthetic gradient information is greater than a power threshold and the method further comprises: sending the first synthetic gradient information to the central node. Zhou teaches in general concepts related to dealing with federated learning in a manner to have computation efficiency using Lyapunov optimization theory (Zhou: Abstract). Specifically, Zhou teaches that the transmission power is considered and modeled for a cost formulation (Zhou: p. 9346, “To model this cost, we use SP to denote the data size of both the model parameters and the gradients interacted between the cloud and each edge node and each global iteration, and PBi to denote the price of the bandwidth of the WAN connecting the cloud and the edge node i ∈ I.”). It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom by modeling the transmission power cost and to consider it for a threshold amount as taught and suggested by Zhou. Such a person would have done so with an expectation for success, to allow for optimizing the three-way tradeoff among communication, computation and throughput (Zhou: p. 9342, “In response, CEFL leverages Lyapunov optimization theory to rigorously design an online control algorithm, which is able to independently and concurrently make the aforementioned four control decisions in an online fashion, without requiring any future information as a priori.”). As to Claim 7, Goetz, Strom and Zhou teaches the limitations of claim 6. Zhou further teaches: where the transmission power of the first synthetic gradient information is associated with communication price metric information (Zhou: p. 9345, resource usage price), channel information corresponding to the first synthetic gradient information (Examiner asserts this is related to the data size and other aspects of the gradient information), and the first synthetic gradient information(Zhou: p. 9346, “to model this cost we use SP to denote the data size of both the model parameters and gradients interacted between the cloud and each node”), wherein the communication price metric information represents a cost volume of communication between one participating node and the central node (Zhou: p. 9344, the number of newly arrived data samples at edge node I, is related to the cost) . As to Claim 9, Goetz, Strom and Zhou teaches the limitations of claim 7. Zhou further teaches: receiving the communication price metric information from the central node (Zhou: p. 9350, “Second, for the cloud that solves the problem, it only needs to collect the queue backlog and resource price information from each edge node at each corresponding time slot t.”). As to Claim 16, it is rejected for similar reasons as claim 6. As to Claim 17, it is rejected for similar reasons as claim 7. As to Claim 19, it is rejected for similar reasons as claim 9. E. Claims 8 and 18 are rejected under 35 USC. § 103 as being unpatentable over non-patent literature, Goetz et al. (“Goetz”), “Federated Learning via Synthetic Data, published September 29, 2020 in view of Strom United States Patent 10,153,676 B1, published on Dec. 11, 2018 and in further view of non-patent literature Zhou et al. (“Zhou”), CEFL: Online Admission Control, Data Scheduling, and Accuracy Tuning for Cost-Efficient Federated Learning Across Edge Nodes,” published in October 2020 and in further view of non-patent literature Su et al. (“Su”), “Data and Channel-Adaptive Sensor Scheduling for Federated Edge Learning via Over-the-Air Gradient Aggregation,” published in July 2021 (manuscript received Jan. 8, 2021). As to Claim 8, Goetz, Strom and Zhou teaches the limitations of claim 6. Goetz, Strom and Zhou may not explicitly teach: wherein the power threshold is in direct proportion to the communication price metric information and/or the power threshold is in direct proportion to an activation power of the participating node, and the communication price metric information represents the cost volume of communication between the participating node and the central node. Su teaches in general concepts related to over-the-air gradient aggregation and improving communication efficiency for federated edge learning applications (Su: Abstract). Specifically, Su teaches that the activation power of an edge device (i.e. a participating node) is considered in the power threshold (Su: p. 1645, Pon is the activation power for each sensor). It would have been obvious to a person having ordinary skill in the art at a time before the effective filing date of the invention to have modified Goetz-Strom-Zhou by modeling the transmission power cost with activation power proportionally as taught and suggested by Su. Such a person would have done so with an expectation for success, to allow for optimizing the cost metrics for federated learning. As to Claim 18, it is rejected for similar reasons as claim 8. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES T TSAI whose telephone number is (571)270-3916. The examiner can normally be reached M-F 8-5 Eastern. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAMES T TSAI/ Primary Examiner, Art Unit 2177
Read full office action

Prosecution Timeline

Jan 04, 2024
Application Filed
Jan 23, 2024
Response after Non-Final Action
Jul 28, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699916
MACHINE LEARNING MODEL TRAINING CHECKPOINTS
5y 2m to grant Granted Aug 04, 2026
Patent 12694327
AUTOMATED FEW-SHOT LEARNING TECHNIQUES FOR ARTIFICIAL INTELLIGENCE-BASED QUERY ANSWERING SYSTEMS
4y 7m to grant Granted Jul 28, 2026
Patent 12682206
QUANTIZED NEURAL NETWORK TRAINING AND INFERENCE
4y 5m to grant Granted Jul 14, 2026
Patent 12682207
ENHANCING SILENT FEATURES WITH ADVERSARIAL NETWORKS FOR IMPROVED MODEL VERSIONS
3y 9m to grant Granted Jul 14, 2026
Patent 12675718
AHEAD-OF-TIME GATE-FUSION TRANSPILATION FOR SIMULATION
4y 6m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+56.9%)
3y 3m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 307 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month