Prosecution Insights
Last updated: August 17, 2026
Application No. 18/623,569

FEDERATED LEARNING

Non-Final OA §101§103§112
Filed
Apr 01, 2024
Priority
Apr 07, 2023 — IN 202311026165 +1 more
Examiner
LEE, MICHAEL CHRISTOPHER
Art Unit
Tech Center
Assignee
Nokia Corporation
OA Round
1 (Non-Final)
62%
Grant Probability
Moderate
1-2
OA Rounds
11m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
95 granted / 153 resolved
+2.1% vs TC avg
Strong +26% interview lift
Without
With
+26.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
53 currently pending
Career history
197
Total Applications
across all art units

Statute-Specific Performance

§101
30.1%
-9.9% vs TC avg
§103
45.2%
+5.2% vs TC avg
§102
10.5%
-29.5% vs TC avg
§112
12.8%
-27.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 153 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Regarding Indian Patent App. No. IN202311026165 (filed 4/7/2023) and Republic of Finland App. No. FI20236024 (filed 9/14/2023), receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statements submitted on 6/5/2024, 10/31/2024, and 10/6/2025 have been considered. Preliminary Amendment The 4/1/2024 Preliminary Amendment has been considered. Claims 1-15 are cancelled and claims 16-35 have been added. Drawings The drawings are objected to because Fig. 4 should be resubmitted using India Ink to comply with the applicable sections of 37 CFR 1.84 set forth below. In particular, such figures should be drawings using India ink or its equivalent. (a) Drawings. There are two acceptable categories for presenting drawings in utility and design patent applications. (1) Black ink. Black and white drawings are normally required. India ink, or its equivalent that secures solid black lines, must be used for drawings; or Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Objections Claims 16, 22, and 29 are objected to because of the following informalities: In claim 16, the examiner suggests amending this to recite “An Apparatus” in line 1. In claim 22, the examiner suggests amending “multiply-and-accumulate, MAC,” to read “multiply-and-accumulate (MAC)” in line 3. In claim 29, line 6, “cause the apparatus to” should read “cause the second apparatus to”). Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claim 31 is rejected under 35 U.S.C. 112(d) as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 31 purports to depend from claim 31 (itself), which is improper. For purposes of compact prosecution, claim 31 will be interpreted as depending from claim 30. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 16-35 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Step 1 of the Alice/Mayo framework, Claims 16-28 are directed to an apparatus (a machine), Claim 29 is directed to a system (a machine), Claims 30-34 are directed to a method (a process), and Claim 35 is directed to a non-transitory computer readable medium (an article of manufacture), which each fall within one of the four statutory categories of inventions. Regarding Claim 1 Step 2A, prong 1 (Is the claim directed to a law of nature, a natural phenomenon or an abstract idea). Claim 1 recites the following mental processes, that in each case under the broadest reasonable interpretation, covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components (e.g., “processor”, “memory”, and “computational model”). determine, based on one or more resources of a client device, whether a first computational model architecture can be trained locally by the client device within a target training time; (under the broadest reasonable interpretation, a human such as a ML engineer, can review the sources of a client device and mentally determine whether a first computational model can be trained locally within a target training time, for example, if the model is extremely complex, the client device is an older generation Blackberry device, and the target training time is only 5 minutes, the human can mentally determine that the first computational model cannot be trained locally within the target training time) select, if the first computational model architecture cannot be trained locally by the client device within the target training time, a modified version of the first computational model architecture that can be trained by the client device within the target training time; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally determine that the first computational architecture cannot be trained locally (as explained above), and can view a list of related and modified versions, and can mentally select a model that can be trained within the target training time, such as selecting a neural network with a single input node, a single hidden layer (with 1 node), and a single output node) Step 2A, prong 2 (Does the claim recite additional elements that integrate the judicial exception into a practical application?). The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements (e.g., “processor”, “memory”, and “computational model”) which are recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Regarding the “Apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:” limitation, such limitations are recited at a high-level of generality and amount to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional elements of a generic apparatus, processor, and memory. These additional elements are recited at a high-level of generality and amount to no more than mere instructions to apply the exception using generic computer components (generic apparatus, processor, and memory). Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “provide the selected modified version of the first computational model architecture for local training by the client device” limitation, such additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of transmitting data, i.e. post-solution activity of transmitting data from the claimed process (see MPEP 2106.05(g)). Accordingly, at Step 2A, prong two, after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application. Step 2B (Does the claim recite additional elements that amount to significantly more than the judicial exception?) In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements (e.g., “processor”, “memory”, and “computational model”) are recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Regarding the “Apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “provide the selected modified version of the first computational model architecture for local training by the client device” limitation, as discussed above, the additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. post-solution activity of transmitting data from the claimed process. The courts have found limitations directed to transmitting information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Accordingly, at Step 2B after considering all claim elements individually and as an ordered combination, it is determined that the claims do not integrate the judicial exception into a practical application. Regarding Claim 17 Step 2A, Prong 1 estimate a total training time for the client device to train locally the first computational model architecture based on the one or more resources of the client device, wherein determining whether the first computation model architecture can be trained locally by the client device comprises determining if the total training time for the client device is within the target training time. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate the training time for the first client device to train locally the first computational model architecture based on the client device’s resources, for example by using his/her experience training the model on similar or identical devices, and determining an average based on such similar or identical devices, and then comparing that estimated time to the target training time mentally) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 18 Step 2A, Prong 1 identify a plurality of client devices; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally identify 2 or more client devices, such as by mentally looking at a list of client devices) estimate, for each client device in the plurality of client devices, a respective total training time; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate a total training time for each device, such as by basing estimations on training times for similar or identical client devices) determine, for each client device, whether the first computational model architecture can be trained locally by said client device within the target training time; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally compare the estimated total training time to the target training time to determine if the first computational model architecture can be trained locally by the client device within the target training time) select, for each client device that cannot be trained locally within the target training time, a respective modified version of the first computational model architecture that can be trained by the client device within the target training time; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally select a respective modified version that can be trained within the target training time, such as by selecting a simple model with a single input node, a single hidden layer (with 1 node), and a single output node) Step 2A, Prong 2 Regarding the “provide, to each client device that cannot be trained locally within the target training time, the respective modified version of the first computational model architecture” limitation, such additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of transmitting data, i.e. post-solution activity of transmitting data from the claimed process (see MPEP 2106.05(g)). Step 2B Regarding the “provide, to each client device that cannot be trained locally within the target training time, the respective modified version of the first computational model architecture” limitation, as discussed above, the additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. post-solution activity of transmitting data from the claimed process. The courts have found limitations directed to transmitting information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Regarding Claim 19 Step 2A, Prong 2 Regarding the “provide, to each client device that can train locally the first computational model architecture within the target training time, the first computational model architecture” limitation, such additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of transmitting data, i.e. post-solution activity of transmitting data from the claimed process (see MPEP 2106.05(g)). Step 2B Regarding the “provide, to each client device that can train locally the first computational model architecture within the target training time, the first computational model architecture” limitation, as discussed above, the additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. post-solution activity of transmitting data from the claimed process. The courts have found limitations directed to transmitting information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Regarding Claim 20 Step 2A, Prong 1 wherein the total training time for a respective client device is estimated based at least partly on characteristics of one or more hardware resources of the respective client device. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate the training time based on characteristics of the hardware resources of the device, such as by estimating the training time based on known training times for similar devices having similar characteristics of hardware resources, e.g., same processor and same amount of memory) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 21 Step 2A, Prong 1 wherein the one or more hardware resources comprise at least one of the following: the respective client device's processing, memory or additional hardware unit resources. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate the training time based on characteristics of the hardware resources of the device, such as by estimating the training time based on known training times for similar devices having similar characteristics of hardware resources, e.g., same processor and same amount of memory) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 22 Step 2A, Prong 1 wherein the total training time for the respective client device is estimated based at least partly on: a number of multiply-and-accumulate, MAC, operations required to train the first computational model architecture; and an estimated time taken to perform the number of MAC operations using the one or more hardware resources of the respective client device. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate both the number of MAC operations and the amount of time per MAC operation, and mentally multiply those numbers to determine the total training time) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 23 Step 2A, Prong 1 wherein the total training time for the respective client device is estimated further based on one or more characteristics of the first computational model architecture. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate the total training time by looking at known training times for the same first computational model architecture using similar client device hardware profiles) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 24 Step 2A, Prong 1 wherein the total training time for the respective client device is estimated further based on data indicative of a current utilization of the one or more hardware resources of the respective client device. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally estimate the total training time based on the current utilization of hardware resources of the device, e.g., the device is currently utilizing a particular processor, and X GB of RAM are available after the operating system has used its share of memory, etc.) Regarding Step 2A, Prong 2, the claim does not include any additional elements that integrate the judicial exception into a practical application and regarding Step 2B, there are no additional elements recited that amount to significantly more than the judicial exception. Regarding Claim 25 Step 2A, Prong 2 Regarding the “wherein the total training time for the respective client device is estimated based on use of an empirical model, trained based on resource profiles for a plurality of different client device types, ... wherein the empirical model is further caused to provide as output the estimated total training time for the respective client device” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of an empirical model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (an empirical model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Regarding the “wherein the empirical model is further caused to receive as input: a time to train the identified computational model architecture using the one or more hardware resources of the respective client device; and the data indicative of current utilization of the one or more hardware resources of the respective client device” limitation, such additional element of a data gathering step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. pre-solution activity of gathering data for use in the claimed process (see MPEP 2106.05(g)). Step 2B Regarding the “wherein the total training time for the respective client device is estimated based on use of an empirical model, trained based on resource profiles for a plurality of different client device types, ... wherein the empirical model is further caused to provide as output the estimated total training time for the respective client device” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding the “wherein the empirical model is further caused to receive as input: a time to train the identified computational model architecture using the one or more hardware resources of the respective client device; and the data indicative of current utilization of the one or more hardware resources of the respective client device” limitation, as discussed above, the additional element of a data transmitting step is recited at a high level of generality and amounts to extra-solution activity of receiving data, i.e. post-solution activity of transmitting data from the claimed process. The courts have found limitations directed to transmitting information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). Regarding Claim 26 Step 2A, Prong 2 Regarding the “wherein the modified version of the first computational model architecture comprises at least one of the following: fewer hidden layers than the first computational model architecture; one or more convolutional layers with a reduced filter or kernel size than corresponding convolutional layers of the first computational model architecture; or fewer nodes in one or more layers than in corresponding layer(s) of the first computational model architecture” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of a generic neural network model having different layer structures. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (a neural network). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Step 2B Regarding the “wherein the modified version of the first computational model architecture comprises at least one of the following: fewer hidden layers than the first computational model architecture; one or more convolutional layers with a reduced filter or kernel size than corresponding convolutional layers of the first computational model architecture; or fewer nodes in one or more layers than in corresponding layer(s) of the first computational model architecture” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding Claim 27 Step 2A, Prong 1 wherein the selecting of the modified version of the first computational model architecture further comprises: accessing one or more candidate modified versions of the first computational model, each having an associated training complexity; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally access candidate modified versions, such as looking at a table on paper that has candidate modified versions and an associated training variable written down next to each candidate modified version) selecting the particular candidate modified version as the modified version of the first computational model architecture. (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally select a particular candidate modified version based on criteria) Step 2A, Prong 2 Regarding the “iteratively testing the candidate modified versions in descending order of complexity until it is determined that the client device can train locally a particular candidate modified version within the target training time” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception. In particular, the claim only recites the additional element of training and testing a neural network model. This additional element is recited at a high-level of generality and amounts to no more than mere instructions to apply the exception using a generic computer component (training and testing a neural network model). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)). Step 2B Regarding the “iteratively testing the candidate modified versions in descending order of complexity until it is determined that the client device can train locally a particular candidate modified version within the target training time” limitation, such limitation is recited at a high-level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, because the limitation merely provides instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not add significantly more than the judicial exception. (See MPEP 2106.05(f)). Regarding Claim 28 Step 2A, Prong 1 identify, a client device of the plurality of client devices with a smallest capacity locally trained computational model; (under the broadest reasonable interpretation, a human such as a ML engineer, can mentally identify a client device that has the smallest capacity locally trained computational model, e.g., by reviewing a list of client devices and their associated locally trained computational model file names) Step 2A, Prong 2 transmit, to an other client device of the plurality of client devices not having the smallest capacity computational model, an indication of the smallest capacity computational model for local re-training based on their respective locally trained computational model; receive, from each said other client device, a respective second set of updated parameters representing the re-trained smallest capacity computational model; average or aggregate the first set of parameters from the client device having the smallest capacity computational model and the second sets of updated parameters; and transmit, to each client device, the averaged or aggregated updated parameters. The examiner finds that the above limitations reflect an improvement to federated learning of machine models. As discussed at page 16, line 26 – page 17, line 3 of the specification, client devices with less resources can become a “bottleneck” to the federated training process, and one of ordinary skill would understand that the limitations of this claim 28 reflect an improvement to federated learning technologies in order to ameliorate this “bottleneck”. Therefore, claim 28 is considered to be subject matter-eligible under 35 U.S.C. 101. Regarding Claim 29 Step 2A, Prong 2 receive, from the first apparatus, the indication of the smallest capacity computational model; re-train the smallest capacity computational model based on a locally trained computational model; transmit, to the first apparatus, the respective second set of updated parameters representing the re-trained smallest capacity computational model; receive, from the first apparatus, averaged or aggregated updated parameters representing a common computational model; and re-train the locally trained computational model based on the common computational model. The examiner finds that the above limitations reflect an improvement to federated learning of machine models. As discussed at page 16, line 26 – page 17, line 3 of the specification, client devices with less resources can become a “bottleneck” to the federated training process, and one of ordinary skill would understand that the limitations of this claim 29 reflect an improvement to federated learning technologies in order to ameliorate this “bottleneck”. Therefore, claim 29 is considered to be subject matter-eligible under 35 U.S.C. 101. Claim 30 recites a method that corresponds to the apparatus of claim 16 and is therefore rejected for the same reasons explained above with respect to claim 16. Claim 31 depends from claim 30 and recites a method that corresponds to the apparatus of claim 17and is therefore rejected for the same reasons explained above with respect to claims 17 and 30. Claim 32 depends from claim 31 and recites a method that corresponds to the apparatus of claim 18 and is therefore rejected for the same reasons explained above with respect to claims 18 and 31. Claim 33 depends from claim 32 and recites a method that corresponds to the apparatus of claim 19 and is therefore rejected for the same reasons explained above with respect to claims 19 and 32. Claim 34 depends from claim 30 and recites a method that corresponds to the apparatus of claim 20 and is therefore rejected for the same reasons explained above with respect to claims 20 and 30. Claim 35 recites a non-transitory computer readable medium that corresponds to the apparatus of claim 16 and is therefore rejected for the same reasons explained above with respect to claim 16. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 16-19, 26, 30-33, and 35 are rejected under 35 U.S.C. 103 as being unpatentable over US 20250103904 A1, hereinafter referenced as YUE, in view of US 20210019599 A1, hereinafter referenced as MAZZAWI. Regarding Claim 16 YUE teaches: Apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: (YUE, para. 0204: “The computer-readable storage medium 1113, having stored thereon the computer program 1112, may comprise instructions which, when executed on at least one processor 1108, cause the at least one processor 1108 to carry out the actions described herein, as performed by the first node 111. In some embodiments, the computer-readable storage medium 1113 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, a memory stick, or stored in the cloud space.”) determine, based on one or more resources of a client device, whether a first computational model architecture can be trained locally by the client device within a target training time; (YUE, para. 0014: “The main idea of Federated Learning may be understood to be to build machine-learning models based on data sets that may be distributed in different network functions. A client NWDAF, e.g., deployed in a domain or network function, may locally train the local ML model with its own data, and share it to the server NWDAF. With local ML models from different client NWDAFs, the server NWDAF may aggregate them into a global or optimal ML model or ML model parameters and send them back to the client NWDAFs for inference.” YUE, para. 0088: “According to a second option, the respective information may indicate a respective first capability to complete one or more training tasks of the ongoing distributed machine-learning or federated learning process, e.g., available computation resource, achievable speed/required time for completing tasks, etc.”; YUE, para. 0104: “In this Action 705, the first node 111 may determine, based on, that is, considering or using, the obtained respective information for the one or more selected nodes 120, at least one of the following. According to a first option, the first node 111 may determine a time needed to complete a service to a consumer of the ongoing distributed machine-learning or federated learning process using the one or more selected nodes 120. The first node 111 may estimate the time for completing training if model provision is required. According to a first option, the first node 111 may determine a level of accuracy for providing the service to the consumer using the one or more selected nodes 120. The first node 111 may estimate the time for completing training and inference if analytic results, e.g., statistics or predictions, are required.”: YUE, para. 0107: “By in this Action 705 determining the time or the level of accuracy for providing the required service to the consumer with the current one or more selected nodes 120, e.g., Client NWDAF(s), the first node 111 be enabled to judge whether the requirement from the consumer, e.g., NWDAF service consumer, may be satisfied with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), and whether to continue the ongoing distributed machine-learning or federated learning process with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), or whether to terminate it.”; YUE, para. 0113: “By determining whether the ongoing distributed machine-learning or federated learning process is to be continued with any of the one or more selected nodes 120 in this Action 706, the first node 111 may be enabled to know whether or not to terminate the procedure to avoid unnecessary resource usage and time consumption.” Examiner’s Note: YUE discloses determining the time for training a model on a consumer’s Network Data Analytics Function (NWDAF) and available computational resources, where the training time is within a “required time for completing tasks” related to model training and provisioning) if the first computational model architecture cannot be trained locally by the client device within the target training time (YUE, para. 0104: “In this Action 705, the first node 111 may determine, based on, that is, considering or using, the obtained respective information for the one or more selected nodes 120, at least one of the following. According to a first option, the first node 111 may determine a time needed to complete a service to a consumer of the ongoing distributed machine-learning or federated learning process using the one or more selected nodes 120. The first node 111 may estimate the time for completing training if model provision is required. According to a first option, the first node 111 may determine a level of accuracy for providing the service to the consumer using the one or more selected nodes 120. The first node 111 may estimate the time for completing training and inference if analytic results, e.g., statistics or predictions, are required.”: YUE, para. 0107: “By in this Action 705 determining the time or the level of accuracy for providing the required service to the consumer with the current one or more selected nodes 120, e.g., Client NWDAF(s), the first node 111 be enabled to judge whether the requirement from the consumer, e.g., NWDAF service consumer, may be satisfied with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), and whether to continue the ongoing distributed machine-learning or federated learning process with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), or whether to terminate it.”; YUE, para. 0113: “By determining whether the ongoing distributed machine-learning or federated learning process is to be continued with any of the one or more selected nodes 120 in this Action 706, the first node 111 may be enabled to know whether or not to terminate the procedure to avoid unnecessary resource usage and time consumption.” Examiner’s Note: YUE discloses determining if a model should be trained, or if the training process should be aborted) provide the ... first computational model architecture for local training by the client device. (YUE, para. 0119: “The first node 111 may then send the updated machine-learning model, or one or more parameters, e.g., weights associated with it, to the one or more selected nodes 120. In other words, the first node 111 may distribute the updated, aggregated machine-learning model.”) However, YUE fails to explicitly teach: select, ... a modified version of the first computational model architecture that can be trained by the client device within the target training time; and the selected modified version However, in a related field of endeavor (determining architectures for neural networks, see para. 0002), MAZZAWI teaches and makes obvious: select, ... a modified version of the first computational model architecture that can be trained by the client device within the target training time; and (MAZZAWI, para. 0006: “The described techniques greatly reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures. As a particular example, the system incrementally and greedily constructs candidate networks that will be trained (networks having “mutated architectures”), so that “full size” candidate neural networks are only trained once the space of smaller candidate neural networks has been sufficiently explored. Additionally, the system dynamically selects the number of training steps that a candidate architecture will be trained for based on the size of the candidate, reducing the time and resources consumed by the training even further, as smaller candidate neural networks can be trained for fewer training steps without adversely impacting the quality of the architecture search. Moreover, the system employs parameter value transfer when generating a mutated architecture, reducing the amount of training required for training the mutated architecture.”; MAZZAWI, para. 0048: “the system 100 can select a final architecture for the neural network using the architectures and performance measures in the maintained data 130.” MAZZAWI, para. 0056: “The system selects, based on the performance measures in the population data, a candidate architecture from the set of candidate architectures (step 202).” MAZZAWI, para. 0073: “Thus, during the early stages of the search, shallow architectures having relatively few blocks will train for a shorter time, increasing the computational efficiency of the overall framework.”; Examiner’s Note: MAZZAWI teaches neural architecture search techniques, where simpler models are created during the earlier stages of the search, and more complex models are provided later, so there is a range of neural architectures to select from based on certain performance measures; the YUE-MAZZAWI combination now determines if a model can be trained within a certain period of time as in YUE, and if it cannot, a simpler model can be chosen from the options provided by MAZZAWI according to performance measures, such as the training time requirements of YUE) provide the selected modified version of the first computational model architecture for local training by the client device. (MAZZAWI, para. 0006: “The described techniques greatly reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures. As a particular example, the system incrementally and greedily constructs candidate networks that will be trained (networks having “mutated architectures”), so that “full size” candidate neural networks are only trained once the space of smaller candidate neural networks has been sufficiently explored. Additionally, the system dynamically selects the number of training steps that a candidate architecture will be trained for based on the size of the candidate, reducing the time and resources consumed by the training even further, as smaller candidate neural networks can be trained for fewer training steps without adversely impacting the quality of the architecture search. Moreover, the system employs parameter value transfer when generating a mutated architecture, reducing the amount of training required for training the mutated architecture.”; MAZZAWI, para. 0048: “the system 100 can select a final architecture for the neural network using the architectures and performance measures in the maintained data 130.” MAZZAWI, para. 0056: “The system selects, based on the performance measures in the population data, a candidate architecture from the set of candidate architectures (step 202).” MAZZAWI, para. 0073: “Thus, during the early stages of the search, shallow architectures having relatively few blocks will train for a shorter time, increasing the computational efficiency of the overall framework.”; Examiner’s Note: MAZZAWI teaches neural architecture search techniques, where simpler models are created during the earlier stages of the search, and more complex models are provided later, so there is a range of neural architectures to select from based on certain performance measures; the YUE-MAZZAWI combination now selects a model according to the training time requirements of YUE and provides the selected model to the consumer NWDAF of YUE) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI as explained above. As disclosed by MAZZAWI, one of ordinary skill would have been motivated to do so in order to “reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures.” (para. 0006). Regarding Claim 17 YUE and MAZZAWI disclose the apparatus of claim 16 as explained above. YUE further teaches: estimate a total training time for the client device to train locally the first computational model architecture based on the one or more resources of the client device, (YUE, para. 0088: “According to a second option, the respective information may indicate a respective first capability to complete one or more training tasks of the ongoing distributed machine-learning or federated learning process, e.g., available computation resource, achievable speed/required time for completing tasks, etc.”; YUE, para. 0104: “In this Action 705, the first node 111 may determine, based on, that is, considering or using, the obtained respective information for the one or more selected nodes 120, at least one of the following. According to a first option, the first node 111 may determine a time needed to complete a service to a consumer of the ongoing distributed machine-learning or federated learning process using the one or more selected nodes 120. The first node 111 may estimate the time for completing training if model provision is required. According to a first option, the first node 111 may determine a level of accuracy for providing the service to the consumer using the one or more selected nodes 120. The first node 111 may estimate the time for completing training and inference if analytic results, e.g., statistics or predictions, are required.”) wherein determining whether the first computation model architecture can be trained locally by the client device comprises determining if the total training time for the client device is within the target training time. (YUE, para. 0107: “By in this Action 705 determining the time or the level of accuracy for providing the required service to the consumer with the current one or more selected nodes 120, e.g., Client NWDAF(s), the first node 111 be enabled to judge whether the requirement from the consumer, e.g., NWDAF service consumer, may be satisfied with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), and whether to continue the ongoing distributed machine-learning or federated learning process with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), or whether to terminate it.”; YUE, para. 0113: “By determining whether the ongoing distributed machine-learning or federated learning process is to be continued with any of the one or more selected nodes 120 in this Action 706, the first node 111 may be enabled to know whether or not to terminate the procedure to avoid unnecessary resource usage and time consumption.”) Regarding Claim 18 YUE and MAZZAWI disclose the apparatus of claim 17 as explained above. YUE further teaches: identify a plurality of client devices; (YUE, para. 0022: “The NWDAF service consumer may use the discovery mechanism from NRF as defined in clause 6.3.13 of TS 23.501, v. 17.3.0 to identify NWDAFs with certain capabilities, e.g., analytics aggregation, covering certain area of interest, e.g., providing data/analytics for specific TAI(s).”) estimate, for each client device in the plurality of client devices, a respective total training time; (YUE, para. 0104: “In this Action 705, the first node 111 may determine, based on, that is, considering or using, the obtained respective information for the one or more selected nodes 120, at least one of the following. According to a first option, the first node 111 may determine a time needed to complete a service to a consumer of the ongoing distributed machine-learning or federated learning process using the one or more selected nodes 120. The first node 111 may estimate the time for completing training if model provision is required. According to a first option, the first node 111 may determine a level of accuracy for providing the service to the consumer using the one or more selected nodes 120. The first node 111 may estimate the time for completing training and inference if analytic results, e.g., statistics or predictions, are required.”; Examiner’s Note: As set forth in MPEP 2144.04 VI.B, duplication of parts or steps “has no patentable significance unless a new and unexpected result is produced” and therefore estimating a training time for more than one device is merely a duplication of this estimating step) determine, for each client device, whether the first computational model architecture can be trained locally by said client device within the target training time; (YUE, para. 0104: “In this Action 705, the first node 111 may determine, based on, that is, considering or using, the obtained respective information for the one or more selected nodes 120, at least one of the following. According to a first option, the first node 111 may determine a time needed to complete a service to a consumer of the ongoing distributed machine-learning or federated learning process using the one or more selected nodes 120. The first node 111 may estimate the time for completing training if model provision is required. According to a first option, the first node 111 may determine a level of accuracy for providing the service to the consumer using the one or more selected nodes 120. The first node 111 may estimate the time for completing training and inference if analytic results, e.g., statistics or predictions, are required.”: YUE, para. 0107: “By in this Action 705 determining the time or the level of accuracy for providing the required service to the consumer with the current one or more selected nodes 120, e.g., Client NWDAF(s), the first node 111 be enabled to judge whether the requirement from the consumer, e.g., NWDAF service consumer, may be satisfied with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), and whether to continue the ongoing distributed machine-learning or federated learning process with the current one or more selected nodes 120, e.g., selected Client NWDAF(s), or whether to terminate it.”; YUE, para. 0113: “By determining whether the ongoing distributed machine-learning or federated learning process is to be continued with any of the one or more selected nodes 120 in this Action 706, the first node 111 may be enabled to know whether or not to terminate the procedure to avoid unnecessary resource usage and time consumption.” Examiner’s Note: YUE discloses determining if a model should be trained, or if the training process should be aborted) provide, to each client device that cannot be trained locally within the target training time, ...the first computational model architecture. (YUE, para. 0119: “The first node 111 may then send the updated machine-learning model, or one or more parameters, e.g., weights associated with it, to the one or more selected nodes 120. In other words, the first node 111 may distribute the updated, aggregated machine-learning model.”) However, YUE fails to explicitly teach: select, for each client device that cannot be trained locally within the target training time, a respective modified version of the first computational model architecture that can be trained by the client device within the target training time; and respective modified version of However, in a related field of endeavor (determining architectures for neural networks, see para. 0002), MAZZAWI teaches and makes obvious: select, for each client device that cannot be trained locally within the target training time, a respective modified version of the first computational model architecture that can be trained by the client device within the target training time; and (MAZZAWI, para. 0006: “The described techniques greatly reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures. As a particular example, the system incrementally and greedily constructs candidate networks that will be trained (networks having “mutated architectures”), so that “full size” candidate neural networks are only trained once the space of smaller candidate neural networks has been sufficiently explored. Additionally, the system dynamically selects the number of training steps that a candidate architecture will be trained for based on the size of the candidate, reducing the time and resources consumed by the training even further, as smaller candidate neural networks can be trained for fewer training steps without adversely impacting the quality of the architecture search. Moreover, the system employs parameter value transfer when generating a mutated architecture, reducing the amount of training required for training the mutated architecture.”; MAZZAWI, para. 0048: “the system 100 can select a final architecture for the neural network using the architectures and performance measures in the maintained data 130.” MAZZAWI, para. 0056: “The system selects, based on the performance measures in the population data, a candidate architecture from the set of candidate architectures (step 202).” MAZZAWI, para. 0073: “Thus, during the early stages of the search, shallow architectures having relatively few blocks will train for a shorter time, increasing the computational efficiency of the overall framework.”; Examiner’s Note: MAZZAWI teaches neural architecture search techniques, where simpler models are created during the earlier stages of the search, and more complex models are provided later, so there is a range of neural architectures to select from based on certain performance measures; the YUE-MAZZAWI combination now determines if a model can be trained within a certain period of time as in YUE, and if it cannot, a simpler model can be chosen from the options provided by MAZZAWI according to performance measures, such as the training time requirements of YUE) provide, to each client device that cannot be trained locally within the target training time, the respective modified version of the first computational model architecture (MAZZAWI, para. 0006: “The described techniques greatly reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures. As a particular example, the system incrementally and greedily constructs candidate networks that will be trained (networks having “mutated architectures”), so that “full size” candidate neural networks are only trained once the space of smaller candidate neural networks has been sufficiently explored. Additionally, the system dynamically selects the number of training steps that a candidate architecture will be trained for based on the size of the candidate, reducing the time and resources consumed by the training even further, as smaller candidate neural networks can be trained for fewer training steps without adversely impacting the quality of the architecture search. Moreover, the system employs parameter value transfer when generating a mutated architecture, reducing the amount of training required for training the mutated architecture.”; MAZZAWI, para. 0048: “the system 100 can select a final architecture for the neural network using the architectures and performance measures in the maintained data 130.” MAZZAWI, para. 0056: “The system selects, based on the performance measures in the population data, a candidate architecture from the set of candidate architectures (step 202).” MAZZAWI, para. 0073: “Thus, during the early stages of the search, shallow architectures having relatively few blocks will train for a shorter time, increasing the computational efficiency of the overall framework.”; Examiner’s Note: MAZZAWI teaches neural architecture search techniques, where simpler models are created during the earlier stages of the search, and more complex models are provided later, so there is a range of neural architectures to select from based on certain performance measures; the YUE-MAZZAWI combination now selects a model according to the training time requirements of YUE and provides the selected model to the consumer NWDAF of YUE) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI as explained above. As disclosed by MAZZAWI, one of ordinary skill would have been motivated to do so in order to “reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures.” (para. 0006). Regarding Claim 19 YUE and MAZZAWI disclose the apparatus of claim 18 as explained above. YUE further teaches: provide, to each client device that can train locally the first computational model architecture within the target training time, the first computational model architecture. (YUE, para. 0119: “The first node 111 may then send the updated machine-learning model, or one or more parameters, e.g., weights associated with it, to the one or more selected nodes 120. In other words, the first node 111 may distribute the updated, aggregated machine-learning model.”) Regarding Claim 26 YUE and MAZZAWI disclose the apparatus of claim 18 as explained above. However, YUE fails to explicitly teach: wherein the modified version of the first computational model architecture comprises at least one of the following: fewer hidden layers than the first computational model architecture; one or more convolutional layers with a reduced filter or kernel size than corresponding convolutional layers of the first computational model architecture; or fewer nodes in one or more layers than in corresponding layer(s) of the first computational model architecture. However, in a related field of endeavor (determining architectures for neural networks, see para. 0002), MAZZAWI teaches and makes obvious: wherein the modified version of the first computational model architecture comprises at least one of the following: fewer hidden layers than the first computational model architecture; one or more convolutional layers with a reduced filter or kernel size than corresponding convolutional layers of the first computational model architecture; or fewer nodes in one or more layers than in corresponding layer(s) of the first computational model architecture. (MAZZAWI, para. 0034: “Each neural network block in each candidate architecture is selected from a set of possible neural network blocks. Thus, the search space for the final architecture is the set of possible combinations of neural network blocks in the set that include at most the maximum number of blocks. A neural network block is a combination of one or more neural network layers that receives one or more input tensors and generates as output one or more output tensors.” MAZZAWI, para. 0043: “The architecture generation engine 120 then generates, based on the results of the determining, a mutated architecture 112 by either (i) adding the selected neural network block as a new neural network block in the selected candidate architecture or (ii) replacing one of the neural network blocks in the selected candidate architecture with the selected neural network block.”; MAZZAWI, para. 0044: “By generating the mutated architectures in this manner, the engine 120 grows architectures adaptively and incrementally via greedy mutations to reduce the sample complexity of the search process.” Examiner’s Note: MAZZAWI teaches mutating different neural architectures by adding blocks, which each include one or more layers (corresponding to recited “fewer hidden layers than the first computational model architecture” limitation); the YUE-MAZZAWI combination now utilizes MAZZAWI to determine a selection of neural network models of varying complexity, such that a modified version of a model can include more or fewer neutral network layers) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI as explained above. As disclosed by MAZZAWI, one of ordinary skill would have been motivated to do so in order to “reduce the time and resource consumption of this training by using a number of techniques that also result in improved performance in discovering new architectures.” (para. 0006). Claim 30 recites a method that corresponds to the apparatus of claim 16 and is therefore rejected for the same reasons explained above with respect to claim 16. Claim 31 depends from claim 30 and recites a method that corresponds to the apparatus of claim 17and is therefore rejected for the same reasons explained above with respect to claims 17 and 30. Claim 32 depends from claim 31 and recites a method that corresponds to the apparatus of claim 18 and is therefore rejected for the same reasons explained above with respect to claims 18 and 31. Claim 33 depends from claim 32 and recites a method that corresponds to the apparatus of claim 19 and is therefore rejected for the same reasons explained above with respect to claims 19 and 32. Claim 35 recites a non-transitory computer readable medium that corresponds to the apparatus of claim 16 and is therefore rejected for the same reasons explained above with respect to claim 16. Claims 20-21, 23-24, and 34 are rejected under 35 U.S.C. 103 as being unpatentable over YUE in view of MAZZAWI and further in view of Justus, Daniel, et al. "Predicting the computational cost of deep learning models." 2018 IEEE international conference on big data (Big Data). IEEE, 2018, hereinafter referenced as JUSTUS. Regarding Claim 20 YUE and MAZZAWI disclose the apparatus of claim 17 as explained above. However, YUE and MAZZAWI fail to explicitly teach: wherein the total training time for a respective client device is estimated based at least partly on characteristics of one or more hardware resources of the respective client device. However, in a related field of endeavor (determining cost envelopes for neural network training, see p. 3873, section I), JUSTUS teaches and makes obvious: wherein the total training time for a respective client device is estimated based at least partly on characteristics of one or more hardware resources of the respective client device. (JUSTUS, p. 3876, section IV.C: PNG media_image1.png 546 400 media_image1.png Greyscale Examiner’s Note: JUSTUS discloses that the predictive model uses several different hardware features; the YUE-MAZZAWI-JUSTUS combination now modifies YUE to estimate the training time using the hardware factors of JUSTUS) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI and JUSTUS as explained above. As disclosed by JUSTUS, one of ordinary skill would have been motivated to do so in order to determine “a priori if the training can be performed within the required cost envelope” which will conserve computing resources. (p. 3873, section I). Regarding Claim 21 YUE, MAZZAWI, and JUSTUS disclose the apparatus of claim 20 as explained above. However, YUE and MAZZAWI fail to explicitly teach: wherein the one or more hardware resources comprise at least one of the following: the respective client device's processing, memory or additional hardware unit resources. However, in a related field of endeavor (determining cost envelopes for neural network training, see p. 3873, section I), JUSTUS teaches and makes obvious: wherein the one or more hardware resources comprise at least one of the following: the respective client device's processing, memory or additional hardware unit resources. (JUSTUS, p. 3876, section IV.C: PNG media_image1.png 546 400 media_image1.png Greyscale Examiner’s Note: JUSTUS discloses that the predictive model uses several different hardware features, including GPU clock speed and memory bandwidth; the YUE-MAZZAWI-JUSTUS combination now modifies YUE to estimate the training time using the hardware factors of JUSTUS) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI and JUSTUS as explained above. As disclosed by JUSTUS, one of ordinary skill would have been motivated to do so in order to determine “a priori if the training can be performed within the required cost envelope” which will conserve computing resources. (p. 3873, section I). Regarding Claim 23 YUE, MAZZAWI, and JUSTUS disclose the apparatus of claim 20 as explained above. However, YUE and MAZZAWI fail to explicitly teach: wherein the total training time for the respective client device is estimated further based on one or more characteristics of the first computational model architecture. However, in a related field of endeavor (determining cost envelopes for neural network training, see p. 3873, section I), JUSTUS teaches and makes obvious: wherein the total training time for the respective client device is estimated further based on one or more characteristics of the first computational model architecture. (JUSTUS, p. 3875, section IV.A and B lists the different layer and layer-specific features taken into account when estimating the training time; Examiner’s Note: JUSTUS discloses that the predictive model uses several different layer and layer-specific features, including the number of layers, types of activation functions, etc., that all pertain to the characteristics of the model architecture; the YUE-MAZZAWI-JUSTUS combination now modifies YUE to estimate the training time using the layer and layer-specific factors of JUSTUS) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI and JUSTUS as explained above. As disclosed by JUSTUS, one of ordinary skill would have been motivated to do so in order to determine “a priori if the training can be performed within the required cost envelope” which will conserve computing resources. (p. 3873, section I). Regarding Claim 24 YUE, MAZZAWI, and JUSTUS disclose the apparatus of claim 20 as explained above. However, YUE and MAZZAWI fail to explicitly teach: wherein the total training time for the respective client device is estimated further based on data indicative of a current utilization of the one or more hardware resources of the respective client device. However, in a related field of endeavor (determining cost envelopes for neural network training, see p. 3873, section I), JUSTUS teaches and makes obvious: wherein the total training time for the respective client device is estimated further based on data indicative of a current utilization of the one or more hardware resources of the respective client device. (JUSTUS, p. 3876, section IV.C: PNG media_image1.png 546 400 media_image1.png Greyscale Examiner’s Note: JUSTUS discloses that their approach takes into account particular counts of GPUs and particular amounts of cores and memories; the YUE-MAZZAWI-JUSTUS combination now modifies YUE to estimate the training time using the actual hardware utilized for the training) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI and JUSTUS as explained above. As disclosed by JUSTUS, one of ordinary skill would have been motivated to do so in order to determine “a priori if the training can be performed within the required cost envelope” which will conserve computing resources. (p. 3873, section I). Claim 34 depends from claim 30 and recites a method that corresponds to the apparatus of claim 20 and is therefore rejected for the same reasons explained above with respect to claims 20 and 30. Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over YUE in view of MAZZAWI and JUSTUS and further in view of US 20100076915 A1, hereinafter referenced as XU. Regarding Claim 22 YUE, MAZZAWI, and JUSTUS disclose the apparatus of claim 20 as explained above. However, YUE, MAZZAWI, and JUSTUS fail to explicitly teach: wherein the total training time for the respective client device is estimated based at least partly on: a number of multiply-and-accumulate, MAC, operations required to train the first computational model architecture; and an estimated time taken to perform the number of MAC operations using the one or more hardware resources of the respective client device. However, in a related field of endeavor (neural network training, see para. 0005), XU teaches and makes obvious: wherein the total training time for the respective client device is estimated based at least partly on: a number of multiply-and-accumulate, MAC, operations required to train the first computational model architecture; and an estimated time taken to perform the number of MAC operations using the one or more hardware resources of the respective client device. (XU, para. 0177: “For the executing time of the hardware and software implementation, in every round, the time of the training process could be calculated using the time used for training one query multiplied by the number of the queries directly. The processing time for a query can be divided into three parts as described herein, forward process (forward propagation), lambda calculation and backward propagation. Here, the forward process refers to the typical multiply-and-accumulate calculations for every feature in the neural networks.”; Examiner’s Note: XU teaches calculating training time by determining the number of multiply-and-accumulate calculations and the time spent doing such calculations; the YUE-MAZZAWI-JUSTUS-XU combination now modifies YUE to estimate the training time using the estimated number and time related to multiply-and-accumulate operations as in XU) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI, JUSTUS, and XU as explained above. One of ordinary skill would have been motivated to do so in order to simulate the exact type of operations that the neural network will most likely have to compute. Claim 25 is rejected under 35 U.S.C. 103 as being unpatentable over YUE in view of MAZZAWI and JUSTUS and further in view of US 20220253748 A1, hereinafter referenced as RAND. Regarding Claim 25 YUE, MAZZAWI, and JUSTUS disclose the apparatus of claim 20 as explained above. However, YUE and MAZZAWI fail to explicitly teach: wherein the total training time for the respective client device is estimated based on use of an empirical model, trained based on resource profiles for a plurality of different client device types, wherein the empirical model is further caused to receive as input: a time to train the identified computational model architecture using the one or more hardware resources of the respective client device; the data indicative of current utilization of the one or more hardware resources of the respective client device, wherein the empirical model is further caused to provide as output the estimated total training time for the respective client device. However, in a related field of endeavor (determining cost envelopes for neural network training, see p. 3873, section I), JUSTUS teaches and makes obvious: wherein the total training time for the respective client device is estimated based on use of an empirical model, trained based on resource profiles for a plurality of different client device types, (JUSTUS, p. 3874, section I: “For our approach we build a generally applicable, data driven model for predicting the execution time for commonly used layers in deep neural networks. From this we deduce the execution time required for one training step (forward and backward pass) for processing an individual batch of data.”; JUSTUS, p. 3876, section IV.C: PNG media_image1.png 546 400 media_image1.png Greyscale Examiner’s Note: JUSTUS discloses a data driven model for estimating the time it will take to train a neural network, where hardware features of the training equipment are taken into account; the YUE-MAZZAWI-JUSTUS combination now uses the model of JUSTUS to predict the time to train devices as in YUE) the data indicative of current utilization of the one or more hardware resources of the respective client device, wherein the empirical model is further caused to provide as output the estimated total training time for the respective client device. (JUSTUS, p. 3874, section I: “For our approach we build a generally applicable, data driven model for predicting the execution time for commonly used layers in deep neural networks. From this we deduce the execution time required for one training step (forward and backward pass) for processing an individual batch of data.”; JUSTUS, p. 3876, section IV.C: PNG media_image1.png 546 400 media_image1.png Greyscale Examiner’s Note: JUSTUS discloses that their approach takes into account particular counts of GPUs and particular amounts of cores and memories; the YUE-MAZZAWI-JUSTUS combination now modifies YUE to estimate the training time using the actual hardware utilized for the training) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI and JUSTUS as explained above. As disclosed by JUSTUS, one of ordinary skill would have been motivated to do so in order to determine “a priori if the training can be performed within the required cost envelope” which will conserve computing resources. (p. 3873, section I). However, YUE, MAZZAWI, and JUSTUS fail to explicitly teach: wherein the empirical model is further caused to receive as input: a time to train the identified computational model architecture using the one or more hardware resources of the respective client device; However, in a related field of endeavor (deep neural networks, see para. 0029), RAND teaches and makes obvious: wherein the empirical model is further caused to receive as input: a time to train the identified computational model architecture using the one or more hardware resources of the respective client device; (RAND, para. 0038: “The conventional manner of achieving model training includes downloading all of the data from storage to the training instance via the Elastic Block Store (EBS). This needs to be done each time it is desired to spin up a new training session, and may cause a significant delay to the training start time. If a large data set is present, this may also incur significant storage costs, again, for each training instance. The use of data pipes, therefore, such as those used in accordance with Sagemaker pipe mode, for example, avoids this by essentially feeding the data directly to the algorithm as it is needed.”; Examiner’s Note: RAND discloses delaying the time to train a model in order to provide enough time to download all necessary training data; the YUE-MAZZAWI-JUSTUS-RAND combination now modifies the predictive model of YUE to take into account a specific starting time for training) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of YUE with MAZZAWI, JUSTUS, and RAND as explained above. As disclosed by RAND, one of ordinary skill would have been motivated to do so in order to provide sufficient time for all training data to be downloaded before commencing training. (para. 0038). Allowable Subject Matter Claims 27-29 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims (and with respect to claim 27, provided that the rejections under 35 U.S.C. 101 are overcome). The following is a statement of reasons for the indication of allowable subject matter: Claim 27 would be considered allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, and provided that the rejections under 35 U.S.C. 101 are overcome, because none of the references of record either alone or in combination fairly disclose or suggest the combination of limitations specified in the claim, including at least: wherein the selecting of the modified version of the first computational model architecture further comprises: accessing one or more candidate modified versions of the first computational model, each having an associated training complexity; The closest prior art of record teaches: US 20250103904 A1, hereinafter referenced as YUE, teaches estimating the amount of time it will take to train a machine learning model. (para. 0104). US 20210019599 A1, hereinafter referenced as MAZZAWI, teaches incrementally training neural networks having similar, but mutated architectures, to create a set of neural network architectures having varying complexities. (paras. 0006, 0048, 0056, and 0073). Justus, Daniel, et al. "Predicting the computational cost of deep learning models." 2018 IEEE international conference on big data (Big Data). IEEE, 2018, hereinafter referenced as JUSTUS, teaches a predictive model for estimating an amount of neural network training time using a large number of different features. (p. 3874, section I and p. 3876, section IV). However, the examiner has found that the distinct feature of the Applicant's claimed invention over the prior art is the explicit claiming of the aforementioned limitations in combination with all the other limitations as specified in the claim. Moreover, the examiner has found that one of ordinary skill would not have been motivated to create a particular indicator of “training complexity” without the hindsight aid of Applicant’s disclosure, as JUSTUS teaches just how complicated and how many features are required to predict neural network training times. To the extent that these features are not found in the prior art, claim 27 would be allowed if rewritten in independent form including all of the limitations of the base claim and any intervening claims, and provided that the rejections under 35 U.S.C. 101 are overcome. Claim 28 would be considered allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims because none of the references of record either alone or in combination fairly disclose or suggest the combination of limitations specified in the claim, including at least: identify, a client device of the plurality of client devices with a smallest capacity locally trained computational model; transmit, to an other client device of the plurality of client devices not having the smallest capacity computational model, an indication of the smallest capacity computational model for local re-training based on their respective locally trained computational model; receive, from each said other client device, a respective second set of updated parameters representing the re-trained smallest capacity computational model; average or aggregate the first set of parameters from the client device having the smallest capacity computational model and the second sets of updated parameters; and transmit, to each client device, the averaged or aggregated updated parameters. The closest prior art of record teaches: US 20250103904 A1, hereinafter referenced as YUE, teaches estimating the amount of time it will take to train a machine learning model. (para. 0104). US 20210019599 A1, hereinafter referenced as MAZZAWI, teaches incrementally training neural networks having similar, but mutated architectures, to create a set of neural network architectures having varying complexities. (paras. 0006, 0048, 0056, and 0073). Justus, Daniel, et al. "Predicting the computational cost of deep learning models." 2018 IEEE international conference on big data (Big Data). IEEE, 2018, hereinafter referenced as JUSTUS, teaches a predictive model for estimating an amount of neural network training time using a large number of different features. (p. 3874, section I and p. 3876, section IV). However, the examiner has found that the distinct feature of the Applicant's claimed invention over the prior art is the explicit claiming of the aforementioned limitations in combination with all the other limitations as specified in the claim. Moreover, the examiner has found that one of ordinary skill would not have been motivated to perform the recited series of limitations without the hindsight aid of Applicant’s disclosure. To the extent that these features are not found in the prior art, claim 28 would be allowed if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim 29 would be considered allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims because none of the references of record either alone or in combination fairly disclose or suggest the combination of limitations specified in the claim, including at least: a second apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive, from the first apparatus, the indication of the smallest capacity computational model; re-train the smallest capacity computational model based on a locally trained computational model; transmit, to the first apparatus, the respective second set of updated parameters representing the re-trained smallest capacity computational model; receive, from the first apparatus, averaged or aggregated updated parameters representing a common computational model; and re-train the locally trained computational model based on the common computational model. The closest prior art of record teaches: US 20250103904 A1, hereinafter referenced as YUE, teaches estimating the amount of time it will take to train a machine learning model. (para. 0104). US 20210019599 A1, hereinafter referenced as MAZZAWI, teaches incrementally training neural networks having similar, but mutated architectures, to create a set of neural network architectures having varying complexities. (paras. 0006, 0048, 0056, and 0073). Justus, Daniel, et al. "Predicting the computational cost of deep learning models." 2018 IEEE international conference on big data (Big Data). IEEE, 2018, hereinafter referenced as JUSTUS, teaches a predictive model for estimating an amount of neural network training time using a large number of different features. (p. 3874, section I and p. 3876, section IV). However, the examiner has found that the distinct feature of the Applicant's claimed invention over the prior art is the explicit claiming of the aforementioned limitations in combination with all the other limitations as specified in the claim. Moreover, the examiner has found that one of ordinary skill would not have been motivated to perform the recited series of limitations without the hindsight aid of Applicant’s disclosure. To the extent that these features are not found in the prior art, claim 29 would be allowed if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20250119913 A1 (Kheirkhah). “In step 11, if the WTRU has determined that it cannot complete the local training on time or that the 3GPP system delay and/or end-to-end latency will be too long for the AI/ML application server to receive the WTRU local trained model on time, the WTRU may decide to release the PDU Session resources by initiating the PDU Session Release procedure.” (para. 0139). Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C LEE whose telephone number is (571)272-4933. The examiner can normally be reached M-F 12:00 pm - 8:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571-272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL C. LEE/Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Apr 01, 2024
Application Filed
Jul 17, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12645972
Performing Property Estimation Using Quantum Gradient Operation on Quantum Computing System
3y 7m to grant Granted Jun 02, 2026
Patent 12603081
METHOD AND SERVER FOR A TEXT-TO-SPEECH PROCESSING
4y 7m to grant Granted Apr 14, 2026
Patent 12602605
QUANTUM COMPUTER ARCHITECTURE BASED ON MULTI-QUBIT GATES
3y 11m to grant Granted Apr 14, 2026
Patent 12591915
METHODS AND SYSTEMS FOR DETERMINING RECOMMENDATIONS BASED ON REAL-TIME OPTIMIZATION OF MACHINE LEARNING MODELS
5y 0m to grant Granted Mar 31, 2026
Patent 12585743
INTERFACE ACCESS PROCESSING METHOD, COMPUTER DEVICE AND STORAGE MEDIUM
1y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
62%
Grant Probability
88%
With Interview (+26.4%)
3y 3m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 153 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month