Prosecution Insights
Last updated: October 02, 2026
Application No. 17/954,824

MANAGEMENT OF FEDERATED LEARNING

Final Rejection §101§102§103
Filed
Sep 28, 2022
Examiner
GONZALES, VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
4 (Final)
79%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
424 granted / 539 resolved
+23.7% vs TC avg
Moderate +11% lift
Without
With
+11.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
14 currently pending
Career history
557
Total Applications
across all art units

Statute-Specific Performance

§101
21.0%
-19.0% vs TC avg
§103
41.7%
+1.7% vs TC avg
§102
13.8%
-26.2% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 539 resolved cases

Office Action

§101 §102 §103
Detailed Action This action is written in response to the remarks and amendments dated 25 June 2026. This action is made final. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments The Examiner is persuaded by the Applicant’s arguments with respect to §101 for independent claims 1/9/13/22/26; accordingly, in view of the Applicant’s amendments, these rejections are withdrawn. However, the examiner maintains the §101 rejection for dependent claims 15-19. The Applicants argue that the previous art of record does not anticipate or render obvious the claims as currently amended. The Examiner provides updated prior art rejections below necessitated by the current amendments. Claim Rejections - 35 USC § 101 Claims 15-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. In determining whether the claims are subject matter eligible, the Examiner applies the 2019 USPTO Patent Eligibility Guidelines.1 Step 1: Is the claim to a process, machine, manufacture, or composition of matter? Yes—claim 15 recites a method, which is a process. Step 2A, prong one: Does the claim recite an abstract idea, law of nature or natural phenomenon? Yes—the claim recites one or more limitations which—under their broadest reasonable interpretation—covers performance of the limitation in the mind (see table below). Claim limitation Examiner analysis 15. The method of claim 13, further comprising: estimating a quantity of local epochs for the training procedure based at least in part on a computational capability of the UE or a link capacity associated with the UE; and This is a mental process akin to a human evaluation/judgment. comparing the estimated quantity of local epochs to the minimum quantity of epochs, wherein the second message is transmitted based at least in part on the comparing. This is a mental process akin to a human evaluation/judgment. Because the claim recites limitations which can practically be implemented as mental processes, the claim recites a mental process. Step 2A, prong two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No practical application is recited, and the recited method does not seem to be clearly and substantially linked to any particular real-world technological problem. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No—the additional limitations are addressed below: [From parent claim 13] receiving a first message indicating a training configuration for a training procedure for training of a predictive model, the training configuration comprising a set of training parameters, wherein the set of training parameters comprises a minimum quantity of epochs for the training of the predictive model to be performed at the UE; and ‘Receiving’ is insignificant pre-solution activity: gathering information which appears to preexist the recited method; alternately, information which is to be used in subsequent steps. transmitting a second message indicating whether the UE has implemented the training configuration prior to the training procedure based at least in part on one or more constraints of the UE and on the set of training parameters. This is insignificant extra-solution activity: transmitting results from a preceding step or transmitting information that pre-existed the preceding step. The Examiner notes that generating, choosing, configuring or optimizing the recited training configuration are not recited in the claim. Likewise, actually performing the prescribed training is not recited in the claim. For the reasons above, claim 15 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to dependent claims 16-19. The additional limitations of the dependent claims are addressed briefly below. Taken alone, the additional elements of the dependent claims above do not amount to significantly more than the above-identified judicial exception (the abstract idea). Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements taken individually. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide conventional computer implementation. Claim limitation Examiner analysis 16. The method of claim 15, wherein the second message indicates that the UE refrains from implementing the training configuration based at least in part on the estimated quantity of local epochs being less than the minimum quantity of epochs. This is merely further details about the information received in the parent claim. 17. The method of claim 15, wherein the second message indicates that the UE implements the training configuration based at least in part on the estimated quantity of local epochs being equal to or greater than the minimum quantity of epochs. This is merely further details about the information received in the parent claim.‘Transmitting’ is insignificant pre-solution activity: transmitting information which appears to preexist the recited method. 18. The method of claim 17, further comprising: performing the training procedure for the predictive model in accordance with the training configuration and based at least in part on the estimated quantity of local epochs; and transmitting a report indicating a set of model parameters for the predictive model based at least in part on performing the training procedure. This is well-understood/routine/conventional. See eg Russell textbook, p. 742, discussing neural network training according to specified number of epochs. (S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach, 2nd Ed., 2003, chapt 18-21, pp. 649-789.) ‘Transmitting’ is insignificant pre-solution activity: transmitting information which appears to preexist the recited method. 19. The method of claim 18, wherein the second message further comprises an indication that the predictive model is ready for activation at the UE, the method further comprising: receiving a third message comprising an indication to activate the training procedure based at least in part on transmitting the second message, wherein performing the training procedure is based at least in part on receiving the third message. ‘Receiving’ is insignificant pre-solution activity: gathering information which appears to preexist the recited method; alternately, information which is to be used in subsequent steps. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. The following are the references relied upon in the rejections below: Jiang (Jiang, Zhifeng, Wei Wang, Bo Li, and Qiang Yang. "Towards efficient synchronous federated training: A survey on system optimization strategies." IEEE Transactions on Big Data 9, no. 2 (2022): 437-454. Published 23 May 2022.) Claims 1 and 9 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Nishio. Regarding claim 1, (alternate rejection) Nishio discloses a method for wireless communications at a network node, comprising: transmitting an indication of a training configuration for training of a predictive model to one or more user equipments (UEs), the training configuration comprising a set of training parameters based at least in part on one or more constraints of the one or more UEs, … P. 3, protocol 2 (reproduced below). PNG media_image1.png 418 434 media_image1.png Greyscale P. 3, “First, the new Resource Request step asks random clients to inform the MEC operator of their resource information such as wireless channel states, computational capacities (e.g., if they can spare CPUs or GPUs for updating models), and the size of data resources relevant to the current training task (e.g., if the server is going to train a 'dog vs-cat' classifier, the number of images containing dogs or cats). Then, the operator refers to this information in the subsequent Client Selection step to estimate the time required for the Distribution and Scheduled Update and Upload steps and to determine which clients go to these steps (the specific algorithms for scheduling clients are explained later). In the Distribution step, a global model is distributed to the selected clients via multicast from the BS because it is bandwidth effective for transmitting the same content (i.e., the global model) to client populations.” Protocol 2, step 3: “Client Selection: Using the information, the MEC operator determines which of the clients go to the subsequent steps to complete the steps within a certain deadline. “. Protocol 2, step 4: “Distribution: The server distributes the parameters of the global model to the selected clients.” ‘training parameters’ :: ‘parameters’ from step 4 (reproduced above). wherein the set of training parameters comprises a minimum quantity of epochs for the training of the predictive model to be performed at a UE of the one or more UEs; The Examiner interprets “minimum quantity of epochs” according to its broadest reasonable interpretation as encompassing a specification of an exact number of epochs, which specifies both a minimum and a maximum required for local training. P. 5, second col., “number of epochs in each round”. receiving, from at least one UE of the one or more UEs, a message indicating whether the at least one UE has implemented the training configuration. Fig. 2 (reproduced below): “Updated model”. PNG media_image2.png 348 350 media_image2.png Greyscale The Examiner notes that the clients each respond with an updated model only after the training configuration has been received and implemented (ie after training has occurred). Regarding claim 9, the rejection of claim 1 supra applies equally here. Nishio further discloses the application of the disclosed method at a server. Abstract: “Specifically, FedCS solves a client selection problem with resource constraints, which allows the server to aggregate as many client updates as possible and to accelerate performance improvement in ML models.” (Emphasis added.) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. The following are the references relied upon in the rejections below: Ahmadi (Sassan Ahmadi (ed.), Chapter 5 - Radio Resource Control Functions, LTE-Advanced, Academic Press, 2014. PP. 227-287.) Jiang (Jiang, Zhifeng, Wei Wang, Bo Li, and Qiang Yang. "Towards efficient synchronous federated training: A survey on system optimization strategies." IEEE Transactions on Big Data 9, no. 2 (2022): 437-454. Published 23 May 2022.) Lian (Lian, Zhuotao, and Chunhua Su. "Decentralized federated learning for internet of things anomaly detection." In Proceedings of the 2022 ACM on Asia conference on computer and communications security, pp. 1249-1251. Published 30 May 2022.) Liu (Liu J, Huang J, Zhou Y, Li X, Ji S, Xiong H, Dou D. From distributed machine learning to federated learning: A survey. Knowledge and information systems. 2022 Apr; 64(4):885-917.) Postel (Postel J. Rfc0821: Simple mail transfer protocol. August 1982. 72 pages.) Toumi (Toumi, Abdelmalek, Jean-Christophe Cexus, and Ali Khenchaf. "A proposal learning strategy on CNN architectures for targets classification." 2022 6th International Conference on Advanced Technologies for Signal and Image Processing (ATSIP). IEEE, conference date 24-27 May 2022.) Wang (WO 2022/010685 A1, cited by Applicant on IDS dated 11/29/23) Claims 1, 3, 5, 8-9 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Wang and Toumi. Regarding claim 1, Wang discloses a method for wireless communications at a network node, comprising: transmitting an indication of a training configuration for training of a predictive model to one or more user equipments (UEs), the training configuration comprising a set of training parameters based at least in part on one or more constraints of the one or more UEs, … [0058] "the core network server 302 (not illustrated) selects the initial ML configuration and communicates the ML configuration UEs through base station 120.” [0056] “the base station receives a UE capability information message (not illustrated) from each UE and selects the initial ML configuration based on a common UE capability between the UEs 111, 112, and 113." receiving, from at least one UE of the one or more UEs, a message indicating whether the at least one UE has implemented the training configuration. Fig. 6 “Request each UE to report updated ML information using a training procedure 620” [Wingdings font/0xE0] “Transmit updated ML information 635”. Toumi discloses the following further limitation which Wang does not disclose wherein the set of training parameters comprises a minimum quantity of epochs for the training of the predictive model to be performed at a UE of the one or more UEs; P. 2, “A. Early stopping techniques The ”early stopping” technique makes it possible to adjust the training process (number of epochs) according to the performance training evolution on the validation dataset. The objective is to train the network long enough to reach the best performance. To do this, we follow the evolution of the classification accuracy at each learning epoch, measured on the validation base. When it stops increasing, we stop the learning. We consider that the accuracy stops increasing when it does not increase during a predetermined number of epochs (Early Stopping number). Moreover, to ensure minimum learning, a minimum number of learning epochs is defined. When the training process is stopped, the training parameters which give us the best performance (accuracy) with the ”best epoch” in the training process are retained and stored.” (Emphasis added.) At the time of filing, it would have been obvious to a skilled machine learning engineer to include the minimum number of epochs as a training hyperparameter (as identified by Toumi) among the hyperparameters chosen and communicated to UEs in the Wang system. Machine learning (ML) models invariably comprise hyperparameters which define the model. These must be chosen by engineers to optimize performance for the particular task at hand before the model’s internal parameters can be learned. Wang discusses these “architecture configurations” in their specification: [0007] “Dynamic reconfiguration of a DNN, such as by modifying various architecture configurations ( e.g., number of layers, layer processing algorithms, down-sampling configurations) and parameter configurations (e.g., coefficients or weights, layer connections, kernel sizes), also provides an ability to adapt how the DNNs process the wireless communications based on changing operating conditions.” Because every hyperparameter of a ML model must be specified before a model can be trained, and “number of epochs” (or “minimum number of epochs”) is one such hyperparameter for neural networks, the inclusion of this hyperparameter would have been obvious. Regarding claim 3, Wang discloses the following further limitation wherein selecting the set of training parameters for the training configuration based at least in part on an estimated link capacity associated with the one or more UEs, a computational capability associated with the one or more UEs, or a combination thereof. [0057] “As another example, the base station 120 receives signal and/or link quality measurements through RRC messages and/or Media Access Control (MAC) layer messages and selects the set of DEs based on the DEs having commensurate (e.g., within a threshold value or range to one another) signal and/or link quality parameters.” (Emphasis added.) Regarding claim 5, Wang discloses the further limitation wherein the set of training parameters further comprises a model structure identifier associated with the predictive model, a baseline parameter set identifier associated with the predictive model, a training validity area, a maximum quantity of epochs for the training of the predictive model, a training deadline, a set of weights, a periodicity, a server address, or a combination thereof. [The Examiner notes that this is a Markush group.] ‘baseline parameter set’ :: [0009] “baseline ML configuration” ‘set of weights’ :: [0007] “Dynamic reconfiguration of a DNN, such as by modifying various architecture configurations ( e.g., number of layers, layer processing algorithms, down-sampling configurations) and parameter configurations (e.g., coefficients or weights, layer connections, kernel sizes), also provides an ability to adapt how the DNNs process the wireless communications based on changing operating conditions.” (Emphasis added.) Regarding claims 8, Wang discloses the further limitation comprising: transmitting, to the one or more UEs, a message comprising an indication to activate the training of the predictive model at the one or more UEs. Fig. 6, “Direct each UE to form a DNN using the initial ML configuration”. [0036] “The core network federated learning manager 314 indicates, to the UE 110 and through the base station 120, when to initiate a training procedure and/or when to report updated ML information learned from the training procedure (e.g., offline training) and/or from processing wireless communications ( e.g., online training).” Regarding claim 9, the rejection of claim 1 supra applies equally here. Wang further discloses the application of the disclosed method at a server. [0002] "This document describes techniques and apparatuses for federated learning for deep neural networks (DNNs) in a wireless communication system") [0032] “In FIG. 3, the core network server 302 may provide all or part of a function, entity, service, and/or gateway in the core network 150”. Regarding claim 11, Wang discloses the further limitation wherein the set of training parameters for the training configuration is based at least in part on an estimated link capacity associated with the one or more UEs, a computational capability associated with the one or more UEs, or a combination thereof. [0056] “the base station receives a UE capability information message (not illustrated) from each UE and selects the initial ML configuration based on a common UE capability between the UEs 111, 112, and 113." [0058] "the core network server 302 (not illustrated) selects the initial ML configuration and communicates the ML configuration UEs through base station 120.” [0114] “For example, the network entity 105-a or the server 203 may select the set of training parameters based on an estimated link capacity associated with one or more of the UEs 115,” Regarding claim 12, Wang discloses the further limitation wherein the set of training parameters further comprises a model structure identifier associated with the predictive model, a baseline parameter set identifier associated with the predictive model, a training validity area, a maximum quantity of epochs for the training of the predictive model, a training deadline, a set of weights, a periodicity, a server address, or a combination thereof. The Examiner notes that this is a Markush group. ‘baseline parameter set’ :: [0009] “baseline ML configuration” ‘set of weights’ :: [0007] “Dynamic reconfiguration of a DNN, such as by modifying various architecture configurations ( e.g., number of layers, layer processing algorithms, down-sampling configurations) and parameter configurations (e.g., coefficients or weights, layer connections, kernel sizes), also provides an ability to adapt how the DNNs process the wireless communications based on changing operating conditions.” (Emphasis added.) Claims 2, 10 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wang, Toumi and Ahmadi. Regarding claims 2 and 10, Wang discloses the following further limitation comprising: transmitting, to the one or more UEs, an indication of … configured for downloading a model structure and a baseline parameter set associated with the predictive model. [0058] "the core network server 302 (not illustrated) selects the initial ML configuration and communicates the ML configuration UEs through base station 120." Ahmadi discloses the following further limitation which Wang does not disclose comprising an indication of a data radio bearer configured for … P. 227, “signaling radio bearer (SRB)”. P. 233, “The RRC connection establishment involves SRB1 establishment. The procedure is also used to transfer the initial NAS dedicated information/message from the UE to the E-UTRAN. The UE initiates the RRC connection establishment procedure when the upper layers request establishment of an RRC connection while the UE is in the RRC_IDLE state.” See also p. 235-36. At the time of filing, it would have been obvious to a person of ordinary skill to apply radio bearer services (as taught by Ahmadi) to the federated learning system of Wang/Toumi because the former a standard component of 5g wireless communication protocols. Regarding claim 14, Wang discloses the further limitation comprising: … downloading the model structure and the baseline parameter set …. Fig. 6, “Direct each UE to form a DNN using the initial ML configuration”. Ahmadi discloses the following further limitation which Wang/Toumi does not disclose: receiving an indication of a data radio bearer configured for downloading a model structure and a baseline parameter set associated with the predictive model; and P. 227, “signaling radio bearer (SRB)”. P. 233, “The RRC connection establishment involves SRB1 establishment. The procedure is also used to transfer the initial NAS dedicated information/message from the UE to the E-UTRAN. The UE initiates the RRC connection establishment procedure when the upper layers request establishment of an RRC connection while the UE is in the RRC_IDLE state.” See also p. 235-36. The obviousness analysis of claims 2/10 applies equally here. Claim 7, 20, 23 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Wang, Toumi and Postel. Regarding claim 7, Postel discloses the following further limitation which Wang/Toumi does not disclose comprising: receiving, from at least one UE of the one or more UEs, a message indicating that the [node] is ready for activation at the at least one UE. P. 41, sec. 4.3, “The communication between the sender and receiver is intended to be an alternating dialogue, controlled by the sender. As such, the sender issues a command and the receiver responds with a reply. The sender must wait for this response before sending further commands. One important reply is the connection greeting. Normally, a receiver will send a 220 "Service ready" reply when the connection is completed. The sender should wait for this greeting message before sending any commands.” At the time of filing, it would have been obvious to a person of ordinary skill to apply the message transfer protocol “service ready” (as taught by Postel) to the Wang/Toumi system because this will ensure that model data or training data is not sent before the receiving node is ready to receive it. Regarding claim 20, Postel discloses the following further limitation which Wang/Liu do not disclose comprising: configuring the training procedure in accordance with the set of training parameters, wherein the second message further comprises an indication that configuration of the training procedure is complete. P. 36 “250 Requested mail action okay, completed”. The obviousness analysis of claim 7 applies equally here. Regarding claim 23, Postel discloses the following further limitation which Wang/Lian do not disclose comprising: receiving, from at least one UE of the one or more UEs, a message indicating that the [node] is ready for activation at the at least one UE. P. 41, sec. 4.3, “The communication between the sender and receiver is intended to be an alternating dialogue, controlled by the sender. As such, the sender issues a command and the receiver responds with a reply. The sender must wait for this response before sending further commands. One important reply is the connection greeting. Normally, a receiver will send a 220 "Service ready" reply when the connection is completed. The sender should wait for this greeting message before sending any commands.” At the time of filing, it would have been obvious to a person of ordinary skill to apply the message transfer protocol “service ready” (as taught by Postel) to the Wang/Lian system because this will ensure that model data or training data is not sent before the receiving node is ready to receive it. Regarding claim 27, Postel discloses the following further limitation which Wang/Lian do not disclose comprising: transmitting a message indicating that the [node] is ready … based at least in part on receiving the indication. P. 41, sec. 4.3, “The communication between the sender and receiver is intended to be an alternating dialogue, controlled by the sender. As such, the sender issues a command and the receiver responds with a reply. The sender must wait for this response before sending further commands. One important reply is the connection greeting. Normally, a receiver will send a 220 "Service ready" reply when the connection is completed. The sender should wait for this greeting message before sending any commands.” The obviousness analysis of claim 23 applies equally here. Claims 13 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Wang, Liu and Toumi. Regarding claim 13, Wang discloses a method for wireless communications at a user equipment (UE), comprising: receiving a first message indicating a training configuration for a training procedure for training of a predictive model, the training configuration comprising a set of training parameters, [0056] “In aspects, the base station receives a UE capability information message (not illustrated) from each UE and selects the initial ML configuration based on a common UE capability between the UEs 111, 112, and 113.” [0058] “Alternatively, or additionally, the core network server 302 (not illustrated) selects the initial ML configuration and communicates the ML configuration UEs through base station 120.” transmitting a second message indicating whether the UE has implemented the training configuration … based at least in part on one or more constraints of the UE and on the set of training parameters. [0067] “In response to detecting the condition and/or in response to performing a training procedure, at 635, 636, and/or 637, the UEs 111, 112, and 113 transmit a message that indicates updated ML information to the base station 120.” (Emphasis added.) The Examiner notes that neural networks are trained iteratively. See eg discussion at Wang [0072]: “This allows the base station 120 to analyze updated ML information and adapt the common DNN to optimize (and re-optimize and/or iteratively optimize) the processing as the operating environment changes …”. (Emphasis added.) See also fig. 6, items 635-640. [0114] “For example, the network entity 105-a or the server 203 may select the set of training parameters based on an estimated link capacity associated with one or more of the UEs 115,” Liu discloses the following further limitation which Wang does not disclose: transmitting a second message indicating whether the UE has implemented the training configuration … based at least in part on the set of training parameters. P. 891, “The monitoring enables the users to get the real-time status of the distributed training process. As the training process of FL models can be very long, e.g., from several hours to days [79], it is of much importance to track the execution status, which allows the user to verify whether the training proceeds normally. The log service is generally supported by major FL systems, which can be used to analyze the training process. In addition, the log generated during the training process can be used to debug the system or adjust the FL model.” (Emphasis added.) At the time of filing, it would have been obvious to a person of ordinary skill to transmit a confirmation message from the UE to the global server (as taught by Liu) in combination with the Wang system. There are only three possibilities for when such a message could be transmitted: prior to training, during training, or after training. All three are valid options, but prior to training confers the advantage of the earliest notice to the global server, so that system engineers could be immediately informed as to the status of the UE. T Toumi discloses the following further limitation which Wang/Liu do not disclose wherein the set of training parameters comprises a minimum quantity of epochs for the training of the predictive model to be performed at the UE; and P. 2, “A. Early stopping techniques The ”early stopping” technique makes it possible to adjust the training process (number of epochs) according to the performance training evolution on the validation dataset. The objective is to train the network long enough to reach the best performance. To do this, we follow the evolution of the classification accuracy at each learning epoch, measured on the validation base. When it stops increasing, we stop the learning. We consider that the accuracy stops increasing when it does not increase during a predetermined number of epochs (Early Stopping number). Moreover, to ensure minimum learning, a minimum number of learning epochs is defined. When the training process is stopped, the training parameters which give us the best performance (accuracy) with the ”best epoch” in the training process are retained and stored.” (Emphasis added.) At the time of filing, it would have been obvious to a skilled machine learning engineer to include the minimum number of epochs as a training hyperparameter (as identified by Toumi) among the hyperparameters chosen and communicated to UEs in the Wang/Liu system. Machine learning (ML) models invariably comprise hyperparameters which define the model. These must be chosen by engineers to optimize performance for the particular task at hand before the model’s internal parameters can be learned. Wang discusses these “architecture configurations” in their specification: [0007] “Dynamic reconfiguration of a DNN, such as by modifying various architecture configurations ( e.g., number of layers, layer processing algorithms, down-sampling configurations) and parameter configurations (e.g., coefficients or weights, layer connections, kernel sizes), also provides an ability to adapt how the DNNs process the wireless communications based on changing operating conditions.” Because every hyperparameter of a ML model must be specified before a model can be trained, and “number of epochs” (or “minimum number of epochs”) is one such hyperparameter for neural networks, the inclusion of this hyperparameter would have been obvious. Regarding claim 21, Wang discloses the further limitation wherein the set of training parameters comprises a model structure identifier associated with the predictive model, a baseline parameter set identifier associated with the predictive model, a training validity area, a maximum quantity of epochs for the training procedure, a minimum quantity of epochs for the training procedure, a training deadline, a set of weights, a periodicity, a server address, or a combination thereof. The Examiner notes that this is a Markush group. ‘baseline parameter set’ :: [0009] “baseline ML configuration” ‘set of weights’ :: [0007] “Dynamic reconfiguration of a DNN, such as by modifying various architecture configurations ( e.g., number of layers, layer processing algorithms, down-sampling configurations) and parameter configurations (e.g., coefficients or weights, layer connections, kernel sizes), also provides an ability to adapt how the DNNs process the wireless communications based on changing operating conditions.” (Emphasis added.) Claims 22, 24-26 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Nishio and Lian. Regarding claim 22, Nishio discloses a method for wireless communications at a server, comprising: transmitting, to a set of user equipments (UEs), a training configuration for a training procedure associated with training of a predictive model, the training configuration comprising a first set of model parameters … based at least in part on one or more constraints of the set of UEs; P. 3, “First, the new Resource Request step asks random clients to inform the MEC operator of their resource information such as wireless channel states, computational capacities (e.g., if they can spare CPUs or GPUs for updating models), and the size of data resources relevant to the current training task (e.g., if the server is going to train a 'dog vs-cat' classifier, the number of images containing dogs or cats). Then, the operator refers to this information in the subsequent Client Selection step to estimate the time required for the Distribution and Scheduled Update and Upload steps and to determine which clients go to these steps (the specific algorithms for scheduling clients are explained later). In the Distribution step, a global model is distributed to the selected clients via multicast from the BS because it is bandwidth effective for transmitting the same content (i.e., the global model) to client populations.” Protocol 2, step 3: “Client Selection: Using the information, the MEC operator determines which of the clients go to the subsequent steps to complete the steps within a certain deadline. “. Protocol 2, step 4: “Distribution: The server distributes the parameters of the global model to the selected clients.” ‘training parameters’ :: ‘parameters’ from step 4 (reproduced above). receiving, from one or more UEs of the set of UEs, one or more reports indicating one or more subsets of model parameters output from the training procedure for the predictive model at the one or more UEs; P. 3, protocol 2, step 5: “Scheduled Update and Upload: The clients update global models and upload the new parameters using the RBs allocated by the MEC operator.” Lian discloses the following further limitations which Nishio does not disclose: … a first set of model parameters associated with a first parameter set identifier … P. 1250, second col. “Each client will get the initial model and other participant in formation. And get the initial version number, where the version starts from 0.” transmitting an indication of a second parameter set identifier of a second set of model parameters, the second parameter set identifier different from the first parameter set identifier, wherein the second parameter set identifier indicates an aggregation of the one or more subsets of model parameters received via the one or more reports from the one or more UEs of the set of UEs. P. 1250, second col., “After the model is obtained, the customers in customer 𝑖 and𝑈𝑖 are weighted averaged. where |𝐷𝑗| represents the local data volume of user 𝑗. After completion, the client 𝑖 updates the local model via stochastic gradient descent, and updates the version number, initializing the set 𝑈𝑖. The training ends after 𝑅 rounds.” The Examiner notes that the passage above describes an iterative process. See p. 3, algorithm 2, step 7: “7: All steps but Initialization are iterated for multiple rounds until the global model achieves a desired performance or the final deadline arrives.” At the time of filing, it would have been obvious to a person of ordinary skill to apply the message transfer protocol (including model version numbering) disclosed by Lian with the federated learning system of Nishio because this will ensure that clients are updating the most current global models, and not wasting resources updating old models. Regarding claim 24, Nishio discloses the further limitation comprising: transmitting, to the set of UEs, a message indicating the second set of model parameters for the predictive model, the second set of model parameters comprising an updated set of model parameters for the predictive model. P. 3, fig. 2 (reproduced supra), “global model & schedule”, “ACK”. Regarding claim 25, Nishio discloses the further limitation wherein the second parameter set identifier comprises a temporary parameter set … ; The Examiner notes that this is a Markush group. P. 3, protocol 2 (reproduced supra), step 4: “Distribution: The server distributes the parameters of the global model to the selected clients.” The Examiner interprets “temporary parameter set identifier” according to its broadest reasonable interpretation as encompassing the parameters of the global model which are transmitted to the clients. These parameters are temporary because they are subsequently updated by the clients. Lian discloses the following further limitation which Nishio does not disclose: wherein the second parameter set identifier comprises a temporary parameter set identifier or a combination of the first parameter set identifier and a version tag. P. 1250, second col., “After the model is obtained, the customers in customer 𝑖 and 𝑈𝑖 are weighted averaged. where |𝐷𝑗| represents the local data volume of user 𝑗. After completion, the client 𝑖 updates the local model via stochastic gradient descent, and updates the version number, initializing the set 𝑈𝑖. The training ends after 𝑅 rounds.” Regarding claim 26, Nishio discloses a method for wireless communications at a user equipment (UE), comprising: receiving a training configuration for a training procedure associated with a predictive model, the training configuration comprising a first set of model parameters … the first set of model parameters based at least in part on one or more constraints of the UE; P. 3, “First, the new Resource Request step asks random clients to inform the MEC operator of their resource information such as wireless channel states, computational capacities (e.g., if they can spare CPUs or GPUs for updating models), and the size of data resources relevant to the current training task (e.g., if the server is going to train a 'dog vs-cat' classifier, the number of images containing dogs or cats). Then, the operator refers to this information in the subsequent Client Selection step to estimate the time required for the Distribution and Scheduled Update and Upload steps and to determine which clients go to these steps (the specific algorithms for scheduling clients are explained later). In the Distribution step, a global model is distributed to the selected clients via multicast from the BS because it is bandwidth effective for transmitting the same content (i.e., the global model) to client populations.” Protocol 2, step 3: “Client Selection: Using the information, the MEC operator determines which of the clients go to the subsequent steps to complete the steps within a certain deadline. “. Protocol 2, step 4: “Distribution: The server distributes the parameters of the global model to the selected clients.” ‘training parameters’ :: ‘parameters’ from step 4 (reproduced above). transmitting a report indicating a subset of model parameters output from the training procedure for the predictive model at the UE; and P. 3, protocol 2, step 5: “Scheduled Update and Upload: The clients update global models and upload the new parameters using the RBs allocated by the MEC operator.” receiving an indication of a second parameter set … associated with a second set of model parameters based at least in part on transmitting the report, the second parameter set identifier different from the first parameter set identifier. The Examiner notes that the training procedure described throughout Nishio is in iterative process performed over many epochs (training cycles). Lian discloses the following further limitations which Nishio does not disclose: … a first set of model parameters associated with a first parameter set identifier … P. 1250, second col. “Each client will get the initial model and other participant in formation. And get the initial version number, where the version starts from 0.” receiving an indication of a second parameter set identifier associated with a second set of model parameters based at least in part on transmitting the report, the second parameter set identifier different from the first parameter set identifier, wherein the second parameter set identifier indicates an aggregation of the one or more subsets of model parameters received via the one or more reports from the one or more UEs of the set of UEs. P. 1250, second col., “After the model is obtained, the customers in customer 𝑖 and𝑈𝑖 are weighted averaged. where |𝐷𝑗| represents the local data volume of user 𝑗. After completion, the client 𝑖 updates the local model via stochastic gradient descent, and updates the version number, initializing the set 𝑈𝑖. The training ends after 𝑅 rounds.” The Examiner notes that the passage above describes an iterative process. See p. 3, algorithm 2, step 7: “7: All steps but Initialization are iterated for multiple rounds until the global model achieves a desired performance or the final deadline arrives.” The obviousness analysis of claim 22 applies equally here. Regarding claim 30, Nishio discloses the further limitation wherein the second parameter set identifier comprises a temporary parameter set … The Examiner notes that this is a Markush group. P. 3, protocol 2 (reproduced supra), step 4: “Distribution: The server distributes the parameters of the global model to the selected clients.” The Examiner interprets “temporary parameter set identifier” according to its broadest reasonable interpretation as encompassing the parameters of the global model which are transmitted to the clients. These parameters are temporary because they are subsequently updated by the clients. Lian discloses the following further limitation which Nishio does not disclose: wherein the second parameter set identifier comprises a temporary parameter set identifier or a combination of the first parameter set identifier and a version tag. P. 1250, second col., “After the model is obtained, the customers in customer 𝑖 and 𝑈𝑖 are weighted averaged. where |𝐷𝑗| represents the local data volume of user 𝑗. After completion, the client 𝑖 updates the local model via stochastic gradient descent, and updates the version number, initializing the set 𝑈𝑖. The training ends after 𝑅 rounds.” Claims 28-29 are rejected under 35 U.S.C. 103 as being unpatentable over Nishio, Lian and Wang. Regarding claim 28, Wang discloses the following further limitation comprising: receiving a message indicating the second set of model parameters for the predictive model, the second set of model parameters comprising an updated set of model parameters for the predictive model. Fig. 6 “Request each UE to report updated ML information using a training procedure 620” [Wingdings font/0xE0] “Transmit updated ML information 635”. The examiner notes that this is an iterative process. At the time of filing, it would have been obvious to apply an iterative training / retraining scheme (as taught by Wang) in combination with the Nishio/Lian system because this would provide for improved predictive performance, as well as provide for updating of models using new data. Regarding claim 29, Wang discloses the following further limitation comprising: performing the training procedure for the predictive model using the second set of model parameters and based at least in part on the second parameter set identifier. Fig. 6 “Request each UE to report updated ML information using a training procedure 620” [Wingdings font/0xE0] “Perform training 630” [Wingdings font/0xE0] “Transmit updated ML information 635”. The examiner notes that this is an iterative process. Claims 1, 9 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang. Regarding claim 1, (alternate rejection) Jiang discloses a method for wireless communications at a network node, comprising: transmitting an indication of a training configuration for training of a predictive model to one or more user equipments (UEs), the training configuration comprising a set of training parameters based at least in part on one or more constraints of the one or more UEs, … P. 438, fig. 1 (reproduced below). PNG media_image3.png 244 698 media_image3.png Greyscale P. 441, second col., “Besides heuristic methods, some researchers use reinforcement learning (RL) algorithms to learn which clients to select in the presence of data heterogeneity. For example, FAVOR [19] seeks to reduce the number of rounds to reach a target accuracy with a deep Q-learning network (DQN) [105]. To capture each client’s statistical characteristics, it takes the low-dimension representations of local models as the RL states. Compared to random selection, FAVOR can reduce the communication rounds by up to 49% in three image classification tasks.” P. 440, first col., “Resource Heterogeneity: due to the variability in hardware specifications and system-level constraints, clients in federated training typically possess different capabilities in computation (CPU/GPU/NPU, memory, and storage), communication (connectivity and bandwidth) and power (battery level and lifespan) [103].” P. 443, second col., “Load Balancing. Given the variations in computing power and data volume, clients may not finish the training process at the same time. To mitigate the resulting straggler effects, [42], [43] suggest balancing the amount of training data across clients. Specifically, they turn to RL techniques for determining the optimal number of data units used in a communication round for each participant, intending to minimize the time and energy consumption and maximize the volume of involved data. …. HeteroFL [46] assigns sub-models with different widths of hidden channels to clients so that clients with fewer capabilities can train smaller sub-models.” (Emphasis added.) P. 444, first col., “Apart from the computational load, it is sometimes also beneficial to balance the communication load across clients, especially when the network conditions are complicated as in wireless connections. For example, targeting a mobile edge computing (MEC) scenario where Time Division Multiple Access (TDMA) is implemented, [47] optimizes both the data batch size and uplink/downlink frame time slots for each client to achieve the maximum learning efficiency. In addition to coping with CPU computing, the authors further extend the optimization problem to the scenario where devices are equipped with GPUs for training.” (Emphasis added.) wherein the set of training parameters comprises a minimum quantity of epochs for the training of the predictive model to be performed at a UE of the one or more UEs; The Examiner interprets “minimum quantity of epochs” according to its broadest reasonable interpretation as encompassing a specification of an exact number of epochs, which specifies both a minimum and a maximum required for local training. P. 448, sec. 5.2.2, “Although there exist some heterogeneity-aware efforts like load balancing where the number of batches, batch size, and number of local epochs can vary across clients (Section 3.2.3), we anticipate that the design space for heterogeneity-aware client configurations could be larger, e.g., using different compression ratios or synchronization frequencies.” receiving, from at least one UE of the one or more UEs, a message indicating whether the at least one UE has implemented the training configuration. Fig. 1 (reproduced above): ‘Reporting’ phase, including “Drop out or miss deadline? Y/N”. Fig. 2 (reproduced below): “Updated model”. PNG media_image2.png 348 350 media_image2.png Greyscale The Examiner notes that the clients each respond with an updated model only after the training configuration has been received and implemented (ie after training has occurred). Although Jiang discloses each of the limitations of claim 1, as illustrated above, it does so in the context of a survey paper. At the time of filing, it would have been obvious to a person of ordinary skill to combine the client selection techniques taught by Jiang with the training load configuration techniques taught by Jiang because each of these approaches can help ensure an efficient use of processing or communication resources to accomplish a particular computing task. Regarding claim 9, (alternate rejection) the rejection of claim 1 supra applies equally here. Jiang further discloses the application of the disclosed method at a server. P. 438, second col., “Configuration. The server next sends the global model status and configuration profiles (e.g., the number of local epochs or the reporting deadline) to each of the selected clients. Based on the instructed configuration, the clients perform local model training independently with their private data.” Regarding claim 13, (alternate rejection) Jiang discloses a method for wireless communications at a user equipment (UE), comprising: receiving a first message indicating a training configuration for a training procedure for training of a predictive model, the training configuration comprising a set of training parameters, P. 438, fig. 1 (reproduced below). PNG media_image3.png 244 698 media_image3.png Greyscale P. 441, second col., “Besides heuristic methods, some researchers use reinforcement learning (RL) algorithms to learn which clients to select in the presence of data heterogeneity. For example, FAVOR [19] seeks to reduce the number of rounds to reach a target accuracy with a deep Q-learning network (DQN) [105]. To capture each client’s statistical characteristics, it takes the low-dimension representations of local models as the RL states. Compared to random selection, FAVOR can reduce the communication rounds by up to 49% in three image classification tasks.” P. 440, first col., “Resource Heterogeneity: due to the variability in hardware specifications and system-level constraints, clients in federated training typically possess different capabilities in computation (CPU/GPU/NPU, memory, and storage), communication (connectivity and bandwidth) and power (battery level and lifespan) [103].” P. 443, second col., “Load Balancing. Given the variations in computing power and data volume, clients may not finish the training process at the same time. To mitigate the resulting straggler effects, [42], [43] suggest balancing the amount of training data across clients. Specifically, they turn to RL techniques for determining the optimal number of data units used in a communication round for each participant, intending to minimize the time and energy consumption and maximize the volume of involved data. …. HeteroFL [46] assigns sub-models with different widths of hidden channels to clients so that clients with fewer capabilities can train smaller sub-models.” (Emphasis added.) P. 444, first col., “Apart from the computational load, it is sometimes also beneficial to balance the communication load across clients, especially when the network conditions are complicated as in wireless connections. For example, targeting a mobile edge computing (MEC) scenario where Time Division Multiple Access (TDMA) is implemented, [47] optimizes both the data batch size and uplink/downlink frame time slots for each client to achieve the maximum learning efficiency. In addition to coping with CPU computing, the authors further extend the optimization problem to the scenario where devices are equipped with GPUs for training.” (Emphasis added.) wherein the set of training parameters comprises a minimum quantity of epochs for the training of the predictive model to be performed at the UE; and The Examiner interprets “minimum quantity of epochs” according to its broadest reasonable interpretation as encompassing a specification of an exact number of epochs, which specifies both a minimum and a maximum required for local training. P. 448, sec. 5.2.2, “Although there exist some heterogeneity-aware efforts like load balancing where the number of batches, batch size, and number of local epochs can vary across clients (Section 3.2.3), we anticipate that the design space for heterogeneity-aware client configurations could be larger, e.g., using different compression ratios or synchronization frequencies.” transmitting a second message indicating whether the UE has implemented the training configuration prior to the training procedure based at least in part on one or more constraints of the UE and on the set of training parameters. Fig. 1 (reproduced above): ‘Reporting’ phase, including “Drop out or miss deadline? Y/N”. Fig. 2 (reproduced below): “Updated model”. PNG media_image2.png 348 350 media_image2.png Greyscale The Examiner notes that the clients each respond with an updated model only after the training configuration has been received and implemented (ie after training has occurred). The obviousness analysis of claim 1 applies equally here. Allowable Subject Matter Claims 15-19 are allowable over the prior art, but are objected to as depending upon a rejected parent claim. (Claims 15-19 are also rejected under §101.) Additional Relevant Prior Art The Examiner identified the following reference as being relevant to the Applicant’s claimed invention, although it is not relied upon in any rejection: Russell discloses neural network training according to specified number of epochs. (S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach, 2nd Ed., 2003, chapt 18-21, pp. 649-789.) Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Vincent Gonzales whose telephone number is (571) 270-3837. The examiner can normally be reached on Monday-Friday 7 a.m. to 4 p.m. MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092. Information regarding the status of an application may be obtained from the USPTO Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. /Vincent Gonzales/Primary Examiner, Art Unit 2124 1 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50, Jan. 7, 2019.
Read full office action

Prosecution Timeline

Show 5 earlier events
Feb 05, 2026
Applicant Interview (Telephonic)
Feb 05, 2026
Examiner Interview Summary
Feb 17, 2026
Response after Non-Final Action
Mar 05, 2026
Request for Continued Examination
Mar 13, 2026
Response after Non-Final Action
Apr 07, 2026
Non-Final Rejection mailed — §101, §102, §103
Jun 25, 2026
Response Filed
Aug 21, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737626
JOINT INPUT PERTUBATION AND TEMPERATURE SCALING FOR NEURAL NETWORK CALIBRATION
3y 3m to grant Granted Sep 15, 2026
Patent 12705472
Prefetching Weights For Use In A Neural Network Processor
2y 9m to grant Granted Aug 11, 2026
Patent 12675989
FUSION MODEL TRAINING USING DISTANCE METRICS
2y 3m to grant Granted Jul 07, 2026
Patent 12651182
IDENTIFYING TRAITS OF PARTITIONED GROUP FROM IMBALANCED DATASET
4y 11m to grant Granted Jun 09, 2026
Patent 12639623
FAIR SELECTIVE CLASSIFICATION VIA A VARIATIONAL MUTUAL INFORMATION UPPER BOUND FOR IMPOSING SUFFICIENCY
4y 4m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
79%
Grant Probability
90%
With Interview (+11.2%)
3y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 539 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month