Prosecution Insights
Last updated: October 02, 2026
Application No. 18/067,503

COMPRESSION OF MODEL WEIGHTS FOR DISTRIBUTED AND FEDERATED LEARNING

Final Rejection §103
Filed
Dec 16, 2022
Examiner
MILLER, ALEXANDRIA JOSEPHINE
Art Unit
2142
Tech Center
2100 — Computer Architecture & Software
Assignee
VMware, Inc.
OA Round
2 (Final)
24%
Grant Probability
At Risk
3-4
OA Rounds
2m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants only 24% of cases
24%
Career Allowance Rate
8 granted / 34 resolved
-31.5% vs TC avg
Strong +73% interview lift
Without
With
+72.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
14 currently pending
Career history
70
Total Applications
across all art units

Statute-Specific Performance

§101
29.4%
-10.6% vs TC avg
§103
56.0%
+16.0% vs TC avg
§102
4.4%
-35.6% vs TC avg
§112
6.8%
-33.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 34 resolved cases

Office Action

§103
DETAILED ACTION Claims 1, 3-8, 10-15, and 17-21 are presented for examination. This office action is in response to submission of application on 10-FEBRUARY-2026. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 16-DECEMBER-2022 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Amendment The amendment filed 10-FEBRUARY-2026 in response to the non-final office action mailed 13-NOVEMBER-2025 has been entered. Claims 1, 3-8, 10-15, 17-21 remain pending in the application. With regards to the non-final office action’s rejection under 101, the amendments to the claims have overcome the original rejection with regards to the claims being directed towards an abstract idea. With regards to the non-final office action’s rejections under 103, the amendments to the claims necessitated a new consideration of the art. After this consideration, the examiner respectfully disagrees with the applicant’s arguments that the art referenced in the previous office action does not teach the amendment claim limitations. A new 103 rejection over the prior art has been provided. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-8, 10-15, and 17-21 are rejected under 35 U.S.C. 103 as being unpatentable over Zhu et al. (Pub. No. US 20220114475 A1, filed October 9th 2020, hereinafter Zhu) in view of Agrawal et al. (Pub. No. US 20190171935 A1, filed December 4th 2017, hereinafter Agrawal) further in view of Fan et al. (Pub. No. CN 109672891 A, published April 23rd 2019, hereinafter Fan). Regarding claim 1: Claim 1 recites: A method performed by a parameter server directing a distributed learning (DL) or federated learning (FL) procedure for training an artificial neural network (ANN) across a plurality of client computing devices over a plurality of rounds, the method comprising: receiving from client computing devices that are participating in a current round of the DL or FL procedure, model updates generated by local ANN copies executing on the participating client computing devices; aggregating the received model updates to determine one or more global gradients for updating the ANN; applying the determined global gradients to update the ANN comprising: determining whether the current round of the DL or FL procedure is an anchor round; upon determining that the current round is an anchor round: compressing a model weight vector comprising model weights of the ANN using a first compression scheme, the compressing resulting in an anchor point; transmitting the anchor point to of each of the participating client computing devices; and saving the anchor point for use in subsequent rounds of the DL or FL procedure; and upon determining that the current round is not an anchor round: computing a sum of global gradients determined in one or more rounds since the last anchor round; compressing the sum using a second compression scheme having lower encoding complexity than the first compression scheme, the compressing of the sum resulting in a correction; and for each of the participating client devices; and for each participating client devices: transmitting the correction to the participating client computing device upon determining that the participating client computing device has already received the saved anchor point; and transmitting the saved anchor point and the correction to the participating client computing device upon determining that the participating client computing device has not already received the saved anchor point Zhu discloses performed by a parameter server directing a distributed learning (DL) or federated learning (FL) procedure for training an artificial neural network (ANN) across a plurality of client computing devices over a plurality of rounds, the method comprising: receiving from client computing devices that are participating in a current round of the DL or FL procedure, model updates generated by local ANN copies executing on the participating client computing devices; aggregating the received model updates to determine one or more global gradients for updating the ANN; applying the determined global gradients to update the ANN comprising: Zhu teaches a method of federated learning for training a machine learning model across a plurality of client computing devices over a plurality of rounds (Paragraph 9), which may include a neural network (Paragraph 41). While Zhu’s own methodology avoids a central server, the use of a centralized parameter server in federated learning is common in the art as reference by Zhu (Paragraph 8). The federated learning method as taught by Zhu comprises receiving from participating client computing devices model updates generated by the respective clients’ models (Paragraph 9), then updating original parameters with a weighted aggregation of the received model updates e.g. the second set of local model parameters to determine updated gradients, the training the machine learning model using the new parameters thereby applying the determined gradients to update the model (Paragraph 9). Zhu discloses determining whether the current round of the DL or FL procedure is an anchor round: Zhu teaches that within a distributed learning system, certain steps may be performed once every certain number of rounds (Paragraph 78). An anchor round is defined within the application as being determined by being every certain number of rounds, as can be seen in the present application’s specification (Paragraph 21), and therefore performing particular steps every number of rounds would demonstrate determining whether a current round of the distributed learning procedure is an anchor round. Zhu discloses upon determining that the current round is an anchor round: compressing a model weight vector comprising model weights of the ANN using a [first compression scheme], the compressing resulting in an anchor point: Zhu teaches that every certain number of rounds, i.e., upon determining that the current round is an anchor round, aggregating a set of weight coefficients (for example, by computing a weighted sum), which would be the model weight vector comprising current model weights of the ANN (Paragraph 88) wherein the aggregation and averaging fulfill a similar purpose to the compression scheme as it simplifies the weight coefficients. This aggregation could be considered an anchor point as it is used to update the machine learning model for future use, making it the new baseline for the model (Paragraph 91). However, Zhu by itself only provides the framework for a compression scheme, but does not teach the compression scheme itself. This aspect is taught further below by Agrawal. Zhu discloses transmitting the anchor point to of each of the participating client computing devices: Zhu teaches that a client device identifies neighboring client devices to transmit its weighting coefficient, i.e. the anchor point, to (Paragraph 78). The neighboring devices would be the participating clients in this scenario as they are participating in the training. Zhu discloses transmitting the retrieved anchor point and the correction to the participating client upon determining that the participating client has not already received the saved anchor point Zhu teaches that upon receiving consent from a client system to participate in a training task, transmitting a set of initial model parameters (Paragraph 21). This would be an indication that the participating client has not already received the saved anchor point as the client is not yet part of training. The initial model parameters would include the first retrieved anchor point as the anchor point is part of the model parameters. However, Zhu does not fully disclose a compression scheme. This limitation is disclosed below by Agrawal in the same field of endeavor of machine learning: Agrawal teaches a compression scheme for the residual gradient weights of a neural network (Paragraph 3), which may be integrated with the teachings of Zhu in order to compress the weights. Agrawal is analogous art to the present application because it is in the same field of endeavor of machine learning. Furthermore, Zhu does not fully disclose saving the anchor point for use in subsequent round. Instead, this limitation is in part disclosed by Agrawal in the same field of endeavor of machine learning. Agrawal discloses saving [the anchor point] for use in subsequent rounds. Previously, Zhu has disclosed the anchor point: Agrawal teaches a recursive process wherein a residue is calculated using a previously stored residue from a previous round (Paragraph 55). This would be an example of a saved value for use in subsequent rounds. Furthermore, Zhu teaches an anchor point as seen above. In combination, the two disclose saving by the parameter server the anchor point for use in subsequent rounds. Furthermore, Agrawal discloses upon determining that the current round is not an anchor round: computing a sum of global gradients determined in one or more rounds since the last anchor round: Agrawal teaches computing a residue by summing a previous residue along with a latest gradient value, which over time would be a sum of global gradients determined in one or more rounds (Paragraph 55), wherein Zhu has previous taught determining is a current round is or is not an anchor round. Agrawal discloses retrieving a saved anchor point for a prior anchor round; and for each participating client computing device in the current round: Agrawal teaches that a maximum of the absolute value of the residue is identified and transmitting to a plurality of other learner system, i.e. participating clients in the current round (Paragraph 68). This would be retrieving a saved anchor point for a previous anchor round as the maximum of the absolute value of the residue is a saved anchor point for a prior anchor round since it would the total sum of the gradients since the last anchor round. Agrawal discloses transmitting the correction to the participating client computing device upon determining that the participating client computing device has already received the saved anchor point: Agrawal teaches exchanging values of a residual with another client system (Paragraph 68) which would indicate that the participating client has already received a saved anchor point as a saved anchor point is already present to be exchanged. The exchanged value would be the transmitted value. However, neither Zhu nor Agrawal discloses compressing the sum using a second compression scheme that has lower encoding complexity than the first compression scheme, the compressing of the sum resulting in a correction. Instead, Fan in the same field of endeavor of compression and encoding discloses this limitation: Fan recites: “However, due to design's earlier JPEG image coding standard, the compression rate of the JPEG image is not high, which also offers the double compressed JPEG image as possible”. “The invention the JPEG image compression for the second time, the compression rate is high, compared with normal ZIP compression scheme such as JPEG image compression has greatly improved performance, and the quality of JPEG image encoding complexity is low” Fan teaches using a compression scheme with a lower encoding complexity for a second image compression, wherein the greatly improved performance would be an example of a correction. Fan further demonstrates a first and second compression as further shown by “however, due to design's earlier JPEG image coding standard, the compression rate of the JPEG image is not high, which also offers the double compressed JPEG image as possible”. This does not teach the incorporation of an anchor round, but the higher- (i.e. second) and lower- (i.e. first) compressions of Fan may be used in combination with the anchor round and non-anchor rounds compression of an accumulated sum of global gradients to generate a correction, as Zhu in view of Agrawal already teach the concept of compression, wherein Fan expands upon that usage. Fan is analogous art to the present application because it is in the same field of endeavor of compression and encoding. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that utilized the teachings of Zhu, the teachings of Agrawal and the teachings of Fan. This would have provided the advantage of accelerating the training of neural networks via compression (Agrawal, Paragraph 54) as well as improving the performance of compression (“the compression rate is high, compared with normal […] compression scheme […] has greatly improved performance”, Fan). Regarding claim 3, which depends upon claim 1: Claim 3 recites: The method of claim 1 wherein the parameter server assigns a bandwidth allocation for transmitting anchor points and corrections, and wherein the bandwidth allocation is dynamically changed as the DL or FL procedure progresses Zhu in view of Agrawal further in view of Fan disclose the method of claim 1 upon which claim 3 relies. Furthermore, Zhu discloses the limitations of claim 3: Zhu teaches a management of data traffic and resource usage by a scheduler (Paragraph 116), which would include assigning a bandwidth allocation for transmitting the previously taught anchor points and corrections, wherein the scheduler uses information such as progress in training in order to manage traffic and resource usage (Paragraph 116). Regarding claim 4, which depends upon claim 1: Claim 4 recites: The method of claim 1 wherein the current round is an anchor round if a round number of the current round is a multiple of a value k Zhu in view of Agrawal further in view of Fan disclose the method of claim 1 upon which claim 4 relies. Furthermore, Zhu discloses the limitations of claim 4: Zhu teaches performing certain steps every certain number of rounds (Paragraph 78). This would be wherein the current round is an anchor round if a round number of the current round is a multiple of a value k, as the value k would be the “certain number” required by Zhu. Regarding claim 5, which depends upon claim 4: Claim 5 recites: The method of claim 4 wherein the value k is adaptively modified by the parameter server as the DL or FL procedure progresses, based on a status of the DL or FL procedure or a resource load on the parameter server Zhu in view of Agrawal further in view of Fan disclose the method of claim 4 upon which claim 5 relies. Furthermore, Agrawal discloses the limitations of claim 5: Agrawal teaches that the time at which certain steps are triggered may rely on fulfilling a particular condition, such as a maximum value being exceeded during rounds of training (Paragraph 68). A maximum value being exceeded during training would be an example of a status of the distributed learning procedure. In combination with Zhu, which teaches the value k as used for an anchor round as can be seen above, a method that uses the value k of Zhu that is adaptively modified by the parameter server at the distributed learning procedure progresses, based on a status of the distributed learning procedure would be obvious in light of the advantages provided below. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that utilized the teachings of Zhu and the teachings of Agrawal. This would have provided the advantage of accelerating the training of neural networks via compression (Agrawal, Paragraph 54). Regarding claim 6, which depends upon claim 1: Claim 6 recites: The method of claim 1 further comprising: selecting a client computing device to participate in a future round of the DL or FL procedure; retrieving a saved anchor point for a prior anchor round; transmitting the retrieved anchor point to the selected client computing device before the future round; and upon reaching the future round: computing another sum of global gradients determined in one or more rounds since the prior anchor round, and compressing said another sum using the second compression scheme, the compressing of said another sum resulting in another correction; and transmitting said another correction to the selected client computing device. Zhu in view of Agrawal further in view of Fan disclose the method of claim 1 upon which claim 6 relies. Furthermore, Agrawal discloses selecting a client computing device to participate in a future round of the DL or FL procedure; retrieving a saved anchor point for a prior anchor round; transmitting the retrieved anchor point to the selected client computing device before the future round; and upon reaching the future round: Agrawal teaches that a maximum of the absolute value of the residue is identified and transmitting to a plurality of other learner system, i.e. selected clients in the current round (Paragraph 68). This would be retrieving a saved anchor point for a previous anchor round as the maximum of the absolute value of the residue is a saved anchor point for a prior anchor round since it would the total sum of the gradients since the last anchor round. This would be for a future round as the updates of the residue is required to begin a future round. Agrawal discloses computing another sum of global gradients determined in one or more rounds since the prior anchor round: Agrawal teaches computing a residue by summing a previous residue along with a latest gradient value, which over time would be a sum of global gradients determined in one or more rounds (Paragraph 55), wherein Zhu has previously taught an anchor round. Agrawal discloses transmitting said [another correction] to the selected client computing device: Agrawal teaches transmitting the resulting values to the clients as has been previously described (Paragraph 68). However, Agrawal does not describe another correction, which is taught below by Fan. However, Agrawal does not disclose every limitation of claim 6. Instead, Fan discloses compressing said another sum using the second compression scheme, the compressing of said another sum resulting in another correction: Fan recites: “The invention the JPEG image compression for the second time, the compression rate is high, compared with normal ZIP compression scheme such as JPEG image compression has greatly improved performance, and the quality of JPEG image encoding complexity is low” Fan teaches using a compression scheme with a lower encoding complexity for a second image compression (i.e. another sum as opposed to the first compression), wherein the greatly improved performance would be an example of a correction. This could be applied to the method of Zhu in view of Agrawal for purposes of the improvement described below. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that utilized the teachings of Zhu in view of Agrawal and the teachings of Fan. This would have provided the advantage of accelerating the training of neural networks via compression (Agrawal, Paragraph 54) as well as improving the performance of compression (“the compression rate is high, compared with normal […] compression scheme […] has greatly improved performance”, Fan). Regarding claim 7, which depends upon claim 1: Claim 7 recites: The method of claim 1 wherein a regularization term is added to a loss function used by the participating client computing devices to generate a gradient for the ANN, the regularization term being designed to minimize differences between model weights of the ANN and the compressed model weights in the anchor point Zhu in view of Agrawal further in view of Fan disclose the method of claim 1 upon which claim 7 relies. Furthermore, Zhu discloses the limitations of claim 7: Zhu teaches a loss function used by the client device to generate a gradient for a set of weights, wherein backpropagation is used to minimize the error and in doing so update the model (Paragraph 43). This would be a regularization term designed to minimize differences between model weights of the ANN and the compressed model weights in anchor point as the anchor point has previously been used to update the model of Zhu. Claims 8 and 10-14 recite a non-transitory computer readable storage medium that parallels the method of claims 1 and 3-7 respectively. Therefore, the analysis discussed above with respect to claims 1 and 3-7 also applies to claims 8 and 10-14 respectively. Accordingly, claims 8 and 10-14 are rejected based on substantially the same rationale as set forth above with respect to claims 1 and 3-7 respectively. Claims 15 and 17-21 recite a system that parallels the method of claims 1 and 3-7 respectively. Therefore, the analysis discussed above with respect to claims 1 and 3-7 also applies to claims 15 and 17-21 respectively. Accordingly, claims 15 and 17-21 are rejected based on substantially the same rationale as set forth above with respect to claims 1 and 3-7 respectively. Response to Arguments Applicant’s arguments filed 10-FEBRUARY-2026 have been fully considered, but the examiner believes that not all are fully persuasive. Regarding the applicant’s remarks on the non-final office action’s 103 rejection of the claims, the applicant argues that neither Zhu nor Agrawal nor Fan fully teach the amended limitations of these claims. As such, the applicant argues that all claims dependent on the above would additionally not be obvious under 103. However, the examiner believes that Zhu in view of Agrawal in view of Fan does teach the amended limitations and respectfully requests applicant’s consideration of the following: The applicant argues that Zhu does not teach compression, as aggregation or averaging does not satisfy the compression limitations recited in the independent claims: The examiner acknowledges this argument and clarifies the incorporation of Agrawal for the teaching of a compression scheme. Agrawal teaches a compression scheme for the residual gradient weights of a neural network (Paragraph 3), which may be integrated with the teachings of Zhu in order to compress the weights. The applicant argues that Fan does not teach the use of a higher-complexity compression scheme with an anchor round for compressing a model weight of an ANN, and associate the use of a lower-complexity compression scheme with a non-anchor round for compressing an accumulated sum of global gradients to generate a correction: While Fan does not directly teach the use of a higher- and lower- complexity compression scheme for these specific purposes, the examiner believes that these teachings of Fan may be incorporated into the teachings of Zhu and Agrawal. However, Fan does teach two compression schemes to be used at separate stages. Fan references “the invention the JPEG image compression for the second time, the compression rate is high, compared with normal ZIP compression scheme such as JPEG image compression has greatly improved performance”, demonstrating a first and second compression as further shown by “however, due to design's earlier JPEG image coding standard, the compression rate of the JPEG image is not high, which also offers the double compressed JPEG image as possible”. This does not teach the incorporation of an anchor round, but the higher- (i.e. second) and lower- (i.e. first) compressions of Fan may be used in combination with the anchor round and non-anchor rounds compression of an accumulated sum of global gradients to generate a correction, as Zhu in view of Agrawal already teach the concept of compression and Fan expands upon that usage. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDRIA JOSEPHINE MILLER whose telephone number is (703)756-5684. The examiner can normally be reached Monday-Thursday: 7:30 - 5:00 pm, every other Friday 7:30 - 4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /A.J.M./Examiner, Art Unit 2142 /Mariela Reyes/Supervisory Patent Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

Dec 16, 2022
Application Filed
Nov 13, 2025
Non-Final Rejection mailed — §103
Feb 03, 2026
Interview Requested
Feb 10, 2026
Applicant Interview (Telephonic)
Feb 10, 2026
Examiner Interview Summary
Feb 10, 2026
Response Filed
Sep 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699886
METHOD FOR NEURAL NETWORK WITH WEIGHT QUANTIZATION
4y 3m to grant Granted Aug 04, 2026
Patent 12614069
METHOD OF OPTIMIZING NEURAL NETWORK MODEL THAT IS PRE-TRAINED, METHOD OF PROVIDING A GRAPHICAL USER INTERFACE RELATED TO OPTIMIZING NEURAL NETWORK MODEL, AND NEURAL NETWORK MODEL PROCESSING SYSTEM PERFORMING THE SAME
4y 2m to grant Granted Apr 28, 2026
Patent 12608625
AUTOMATICALLY TRAINING AND IMPLEMENTING ARTIFICIAL INTELLIGENCE-BASED ANOMALY DETECTION MODELS
4y 1m to grant Granted Apr 21, 2026
Patent 12566943
METHOD AND APPARATUS WITH NEURAL NETWORK QUANTIZATION
4y 5m to grant Granted Mar 03, 2026
Patent 12481890
SYSTEMS AND METHODS FOR APPLYING SEMI-DISCRETE CALCULUS TO META MACHINE LEARNING
4y 4m to grant Granted Nov 25, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
24%
Grant Probability
96%
With Interview (+72.7%)
3y 11m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 34 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month