Prosecution Insights
Last updated: October 02, 2026
Application No. 18/216,982

FEDERATED LEARNING WITH MODEL DIVERSITY

Final Rejection §101§103§112
Filed
Jun 30, 2023
Examiner
SIPPEL, MOLLY CLARKE
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Robert Bosch GmbH
OA Round
2 (Final)
50%
Grant Probability
Moderate
3-4
OA Rounds
6m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
14 granted / 28 resolved
-5.0% vs TC avg
Strong +26% interview lift
Without
With
+25.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
18 currently pending
Career history
42
Total Applications
across all art units

Statute-Specific Performance

§101
34.4%
-5.6% vs TC avg
§103
31.6%
-8.4% vs TC avg
§102
10.0%
-30.0% vs TC avg
§112
22.8%
-17.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 28 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This action is responsive to the amendment filed on 06/23/2026. Claims 1-20 are pending in the case. Claims 1, 5, 11, and 16 are currently amended. Claims 1, 11, and 16 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/03/2026 is being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 5-7 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the enablement requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention. Regarding claim 5, the claim recites “wherein the loss is determined by: L i = l o s s ( ∑ j = 1 l i f W j D i ,   y i ) + λ i R ( g W 1 i D i ,   g W 2 i D i , … , g W l i i D i ) wherein D i represents the locally-stored data at a particular one of the clients i, and y i represents a corresponding label for the locally stored data; wherein λ is an adjustable hyperparameter corresponding to the regularization term R, and wherein g is the one or more latent features of the local machine models W 1 i ,   W 2 i ,   … ,   W l i i ”. This includes any reasonable interpretation of the loss equation, “ L i = l o s s ( ∑ j = 1 l i f W j D i ,   y i ) + λ i R ( g W 1 i D i ,   g W 2 i D i , … , g W l i i D i )”, including interpretations in which the second term of the equation is included in the summation and interpretations in which the second term of the equation is excluded in the summation; because it appears the equation is missing parentheses and is thus incomplete. There is an opening parentheses for the loss function and an opening parentheses for the function f W j and then it becomes unclear. If the closing parentheses following D i closes the function, then it is unclear how y i is meant to be evaluated. Further, there are no parentheses to designate the beginning and end of the terms of the summation. Regarding the breadth of the claims, the claims are specific to a method that trains neural networks with federated learning by sending at least portions of the server-maintained machine learning models from the server to the clients, training the models at the clients using a loss function, which is specified in claim 4 to include a regularization term and one or more latent features of the local machine learning models, and further specified to be determined by “ L i = l o s s ( ∑ j = 1 l i f W j D i ,   y i ) + λ i R ( g W 1 i D i ,   g W 2 i D i , … , g W l i i D i )” in claim 5. Regarding the nature of the invention, the invention is directed to a method described above. Regarding the state of the art, the state of the prior art clearly shows that using a loss function for training local models on client devices in a federated learning system is well known. However, the prior art fails to provide any evidence that the person of ordinary skill in the art would have been able to develop a method that performs training of neural networks with federated learning with the provided loss function (including interpretations in which the second term of the equation is included in the summation and/or interpretations in which the second term of the equation is excluded in the summation). Regarding the level of one of ordinary skill in the art, it is fairly high and the level of predictability in the art is fairly low. Regarding the amount of direction provided by the inventor, applicant’s specification provides little or no guidance in how the loss is to be calculated using the loss function (including interpretations in which the second term of the equation is included in the summation and/or interpretations in which the second term of the equation is excluded in the summation). The current claims cover a loss function in which the second term of the equation is included in the summation and/or interpretations in which the second term of the equation is excluded in the summation, which is not enabled by the applicant’s specification. No working examples are provided in the specification for the invention of claim 5. With no guidance from the specification or disclosures in the prior art, it would require an undue amount of experimentation to make and use the invention of claim 5. Claims 6-7 are rejected as being dependent upon a rejected base claim without curing any of the deficiencies. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 1, the claim recites: “ ( G w i ) T G w i ” in line 11. This limitation renders the claim indefinite for failing to particularly point out and distinctly claim the subject matter because it is unclear what this variable is meant to represent. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes the limitation has been interpreted as the concatenated latent features (in matrix form) multiplied by the concatenated latent features in a transposed matrix, reciting a new claim element. Claims 2-10 are rejected as being dependent upon a rejected base claim without curing any of the deficiencies. Regarding claim 7, the claim recites “the concatenated weights” in line 5. There is insufficient antecedent basis for this limitation in the claim. It is unclear if applicant is attempting to refer to a previously recited claim element or if applicant is attempting to recite a new claim element. For examination purposes, the limitation is considered to be “concatenated weights”, reciting a new claim element. Regarding claim 11, the claim recites: “ ( G w i ) T G w i ” in line 15. This limitation renders the claim indefinite for failing to particularly point out and distinctly claim the subject matter because it is unclear what this variable is meant to represent. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes the limitation has been interpreted as the concatenated latent features (in matrix form) multiplied by the concatenated latent features in a transposed matrix, reciting a new claim element. Claims 12-15 are rejected as being dependent upon a rejected base claim without curing any of the deficiencies. Regarding claim 16, the claim recites: “ ( G w i ) T G w i ” in line 13. This limitation renders the claim indefinite for failing to particularly point out and distinctly claim the subject matter because it is unclear what this variable is meant to represent. It is unclear if applicant is attempting to recite a new claim element or if applicant is attempting to refer to a previously recited claim element. For examination purposes the limitation has been interpreted as the concatenated latent features (in matrix form) multiplied by the concatenated latent features in a transposed matrix, reciting a new claim element. Claims 17-20 are rejected as being dependent upon a rejected base claim without curing any of the deficiencies. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1 Statutory Category: Claim 1 is directed to a method, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 1 recites, in part, “determining a respective loss for each of the plurality of local machine learning models by (i) computing, for the plurality of local machine learning models at that client, one or more latent features from selected intermediate layers using locally-stored data yielding concatenated latent features G w i , and (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical calculations, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular the claim recites: “training neural networks with federated learning”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes…updating respective weights for each of the plurality of local machine learning models”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Finally, the claim recites: “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Alternatively, “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “outputting, from the server and based on the training, a plurality of trained server-maintained machine learning models configured to provide model diversity and robustness to client data distribution shifts”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element: “training neural networks with federated learning” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional elements: “sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models” and “transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients” amount to adding insignificant extra-solution activity to the judicial exception and further, are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the additional element “at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes…updating respective weights for each of the plurality of local machine learning models” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Ryden, U.S. Patent Application Publication No. 20240155714, Paragraphs 0036-0037, “The FL steps performed in a wireless communication system involving access nodes and wireless devices typically comprise [0037] 1) Local Model Training: After receiving the global/initial model, all UEs perform local training based on the global model to update their local weights”. Further, the additional element “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Alternatively, the additional element “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Ryden, U.S. Patent Application Publication No. 20240155714, Paragraphs 0036-0038, “The FL steps performed in a wireless communication system involving access nodes and wireless devices typically comprise … 2) Global Model Aggregation: The gNB receives the local weights from each UE and updates the global model weights and then send back the updated global model weights to all the participants in the FL process”. Further, the additional element “outputting, from the server and based on the training, a plurality of trained server-maintained machine learning models configured to provide model diversity and robustness to client data distribution shifts” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 2, the rejection of claim 1 is incorporated, and further, the claim recites: “selecting the plurality of server-maintained machine learning models from a pool of machine learning models”. This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III). Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 3, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein the plurality of server-maintained machine learning models are selected from the pool based on resource limits associated with the plurality of clients”. This limitation is a continuation of the “selecting the plurality of server-maintained machine learning models from a pool of machine learning models” limitation identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 4, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the training of the plurality of local machine learning models at each client includes determining a loss based on a regularization term and one or more latent features of the local machine learning models”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 5, the rejection of claim 4 is incorporated, and further, the claim recites: “wherein the loss is determined by: L i = l o s s ( ∑ j = 1 l i f W j D i ,   y i ) + λ i R ( g W 1 i D i ,   g W 2 i D i , … , g W l i i D i ) wherein D i represents the locally-stored data at a particular one of the clients i, and y i represents a corresponding label for the locally stored data; wherein λ is an adjustable hyperparameter corresponding to the regularization term R, and wherein g is the one or more latent features of the local machine models W 1 i ,   W 2 i ,   … ,   W l i i ”. This limitation recites mathematical concepts in addition to those identified in the parent claim, in this case a mathematical formula or equation, as directed to “a claim that recites a numerical formula or equation will be considered as falling within the "mathematical concepts" grouping. In addition, there are instances where a formula or equation is written in text format that should also be considered as falling within this grouping”. See MPEP § 2106.04(a)(2)(I)(B). The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 6, the rejection of claim 5 is incorporated, and further, the claim recites: “at the server, aggregating information received from the clients to perform the training of the plurality of server-maintained machine learning models with the updated weights”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim and thus, recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 7, the rejection of claim 6 is incorporated, and further, the claim recites: “wherein the aggregating includes performing: min W 1 ,   W 2 ,   … , W L ⁡ ∑ i = 1 L 1 Z i ∑ i ∈ Z l W l ( i ) - W l F 2 + λ W c o n T W c o n - I F 2 wherein Z represents an index of the plurality of server-maintained machine learning models, and wherein W c o n denotes the concatenated weights of the plurality of server-maintained machine learning models W 1 ,   W 2 ,   … , W L ”. This limitation recites mathematical concepts in addition to those identified in the parent claim, in this case a mathematical formula or equation, as directed to “a claim that recites a numerical formula or equation will be considered as falling within the "mathematical concepts" grouping. In addition, there are instances where a formula or equation is written in text format that should also be considered as falling within this grouping”. See MPEP § 2106.04(a)(2)(I)(B). The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 8, the rejection of claim 1 is incorporated, and further, the claim recites: “determining that a first client of the plurality of clients is disconnected or otherwise unable to receive the at least portions of the plurality of server-maintained machine learning models from the server”. This limitation is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an observation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “connecting the first client to a neighboring client that is able to communicate with the server”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the limitation is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Iyer et al., U.S. Patent Application Publication No. 20250036959, Paragraph 0024, Lines 31-34, “ Typical electronic devices also include a set of one or more physical network interface(s) (NI(s)) to establish network connections (to transmit and/or receive code and/or data using propagating signals) with other electronic devices”. Further, the claim recites: “sending the portions of the plurality of server-maintained machine learning models from the neighboring client to the first client”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, this limitation is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 9, the rejection of claim 8 is incorporated, and further, the claim recites: “performing an interpolation of the portions of the plurality of server-maintained machine learning models received from the plurality of neighboring clients”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “wherein the connecting includes connecting the first client to a plurality of neighboring clients”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the limitation is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Iyer et al., U.S. Patent Application Publication No. 20250036959, Paragraph 0024, Lines 31-34, “ Typical electronic devices also include a set of one or more physical network interface(s) (NI(s)) to establish network connections (to transmit and/or receive code and/or data using propagating signals) with other electronic devices”. The claim is not patent eligible. Regarding claim 10, the rejection of claim 9 is incorporated, and further, the claim recites: “wherein the interpolation is W + = W + ∑ { i ∈ C b } A i ∙ ( W i - W ) , wherein W + is an interpolated model for models W received by the plurality of neighboring clients C b , and wherein A i is a linear combination weight for model W i ”. This limitation recites mathematical concepts in addition to those identified in the parent claim, in this case a mathematical formula or equation, as directed to “a claim that recites a numerical formula or equation will be considered as falling within the "mathematical concepts" grouping. In addition, there are instances where a formula or equation is written in text format that should also be considered as falling within this grouping”. See MPEP § 2106.04(a)(2)(I)(B). The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 11: Step 1 Statutory Category: Claim 11 is directed to a system, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 11 recites, in part, “determining a respective loss for each of the plurality of local machine learning models by (i) computing, for the plurality of local machine learning models at that client, one or more latent features from selected intermediate layers using locally-stored data yielding concatenated latent features G w i , and (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical calculations, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular the claim recites: “a system”, “memory storing instructions”, and “a plurality of processors”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “training neural networks with federated learning”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes…updating respective weights for each of the plurality of local machine learning models”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Finally, the claim recites: “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Alternatively, “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “outputting, from the server and based on the training, a plurality of trained server-maintained machine learning models configured to provide model diversity and robustness to client data distribution shifts”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a system”, “memory storing instructions”, and “a plurality of processors” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Further, the additional element “training neural networks with federated learning” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional elements: “sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models” and “transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients” amount to adding insignificant extra-solution activity to the judicial exception and further, are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the additional element “at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes…updating respective weights for each of the plurality of local machine learning models” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Ryden, U.S. Patent Application Publication No. 20240155714, Paragraphs 0036-0037, “The FL steps performed in a wireless communication system involving access nodes and wireless devices typically comprise [0037] 1) Local Model Training: After receiving the global/initial model, all UEs perform local training based on the global model to update their local weights”. Further, the additional element “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Alternatively, the additional element “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Ryden, U.S. Patent Application Publication No. 20240155714, Paragraphs 0036-0038, “The FL steps performed in a wireless communication system involving access nodes and wireless devices typically comprise … 2) Global Model Aggregation: The gNB receives the local weights from each UE and updates the global model weights and then send back the updated global model weights to all the participants in the FL process”. Further, the additional element “outputting, from the server and based on the training, a plurality of trained server-maintained machine learning models configured to provide model diversity and robustness to client data distribution shifts” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 12, the rejection of claim 11 is incorporated, and further, claim 12 is substantially similar to claim 2 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 13, the rejection of claim 12 is incorporated, and further, the claim recites: “wherein the plurality of server-maintained machine learning models are selected from the pool and sent to the plurality of clients based on resource limits of the plurality of clients”. This limitation is a continuation of the mental process limitation of claim 12, and thus the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 14, the rejection of claim 11 is incorporated, and further, claim 14 is substantially similar to claim 4 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 15, the rejection of claim 14 is incorporated, and further, claim 15 is substantially similar to claim 6 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 16: Step 1 Statutory Category: Claim 16 is directed to a machine, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 16 recites, in part, “determining a respective loss for each of the plurality of local machine learning models by (i) computing, for the plurality of local machine learning models at that client, one or more latent features from selected intermediate layers using locally-stored data yielding concatenated latent features G w i , and (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity”. This limitation, under the broadest reasonable interpretation, covers the recitation of mathematical calculations, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular the claim recites: “A non-transitory computer readable storage medium”, “a computer readable program code”, “computer readable instructions”, and “a computing system”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “train neural networks with federated learning”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes…updating respective weights for each of the plurality of local machine learning models”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Finally, the claim recites: “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Alternatively, “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “outputting, from the server and based on the training, a plurality of trained server-maintained machine learning models configured to provide model diversity and robustness to client data distribution shifts”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “A non-transitory computer readable storage medium”, “a computer readable program code”, “computer readable instructions”, and “a computing system” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Further, the additional element “train neural networks with federated learning” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional elements: “sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models” and “transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients” amount to adding insignificant extra-solution activity to the judicial exception and further, are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the additional element “at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes…updating respective weights for each of the plurality of local machine learning models” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Ryden, U.S. Patent Application Publication No. 20240155714, Paragraphs 0036-0037, “The FL steps performed in a wireless communication system involving access nodes and wireless devices typically comprise [0037] 1) Local Model Training: After receiving the global/initial model, all UEs perform local training based on the global model to update their local weights”. Further, the additional element “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Alternatively, the additional element “at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is well-understood, routine, and conventional as taught by activity supported under Berkheimer Option 2, Ryden, U.S. Patent Application Publication No. 20240155714, Paragraphs 0036-0038, “The FL steps performed in a wireless communication system involving access nodes and wireless devices typically comprise … 2) Global Model Aggregation: The gNB receives the local weights from each UE and updates the global model weights and then send back the updated global model weights to all the participants in the FL process”. Further, the additional element “outputting, from the server and based on the training, a plurality of trained server-maintained machine learning models configured to provide model diversity and robustness to client data distribution shifts” amounts to adding insignificant extra-solution activity to the judicial exception, and further, is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible. Regarding claim 17, the rejection of claim 16 is incorporated, and further, claim 17 is substantially similar to claim 2 and claim 12 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 18, the rejection of claim 17 is incorporated, and further, claim 18 is substantially similar to claim 13 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 19, the rejection of claim 16 is incorporated, and further, claim 19 is substantially similar to claim 4 and claim 14 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 20, the rejection of claim 19 is incorporated, and further, claim 20 is substantially similar to claim 6 and 15 respectively, and is rejected in the same manner and reasoning applying. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4, 8-12, 14-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Bhuyan et al., Multi-Model Federated Learning, 01/07/2022, https://arxiv.org/pdf/2201.02582, hereinafter referred to as "Bhuyan" in view of Wang et al., Deep Multimodal Hashing with Orthogonal Regularization, 07/25/2015, In Proceedings of the 24th International Conference on Artificial Intelligence (IJCAI'15), AAAI Press, 2291–2297, https://dl.acm.org/doi/10.5555/2832415.2832567, hereinafter referred to as “Wang” in further view of Yemini et al., Semi-Decentralized Federated Learning with Collaborative Relaying, 05/23/2022, https://arxiv.org/pdf/2205.10998, hereinafter referred to as “Yemini”. Regarding claim 1, Bhuyan teaches A method of training neural networks with federated learning (Bhuyan, Page 1, Abstract, Lines 3-5, “In this paper, we extend federated learning to the setting where multiple unrelated models are trained simultaneously”), the method comprising: sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models (Bhuyan, Page 1, Section II, Lines 1-4, “We consider the setting where the server trains M unrelated models in a distributed manner using a pool for clients. Each client has a separate dataset for each model. The server maintains a global version of each of the M models”; Bhuyan, Page 1, Section II, Paragraph 2, Lines 1-4, “At the start of each round, the server selects up to K clients to be used for training. K is a fixed parameter given as an input to our system. Each selected client receives global weights of the model it needs to train”); at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes determining a respective loss for each of the plurality of local machine learning models by (i) computing, for the plurality of local machine learning models at that client, one or more latent features from selected intermediate layers using the local-stored data yielding concatenated latent features G w i , and …; updating respective weights for each of the plurality of local machine learning models (Bhuyan, Page 1, Section II, Paragraph 2, Lines 4-7, “The selected clients train the global version of the model allotted to them on the corresponding local training dataset and send the updated model weights to the server”; Bhuyan, Page 2, Section III, Subsection B, Paragraph 3, Line 3, “Local loss, for model i of client k, is denoted by l t ( k , i ) ”); Bhuyan, Page 2, Algorithm 1, Step 10, “Update localModel[m,c] by local training”; A person of ordinary skill in the art would recognize that training the model suing a local training dataset would require computing latent features from intermediate layers); transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients (Bhuyan, Page 1, Section II, Paragraph 2, Lines 4-7, “The selected clients train the global version of the model allotted to them on the corresponding local training dataset and send the updated model weights to the server”; Bhuyan, Page 1, Section I, Lines 4-6, “The key feature in federated learning is that the local client datasets are never shared with each other or with the server”); at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients (Bhuyan, Page 1, Section II, Lines 7-10, “The server weights are an average of the corresponding client weights with importance given to each client equal to the proportion of total training samples in it”; see also: Bhuyan, Page 2, Algorithm 1, Steps 13 and 14). Bhuyan does not explicitly teach determining the loss by (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity nor outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts. Wang teaches determining the loss by (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity (Wang, Page 2294, Section 3.3, Paragraph 3, “Under the assumption that input Xv and Xt are orthogonal [Wang et al., 2010], if we impose the orthogonal constraints on each layer’s weight matrix as shown in Eq. 5, the representations of the layer h(mv) v and h(mt) t are approximately orthogonal … we can impose the orthogonal regularization on the weighting matrices”; Wang, Page 2294, Section 3.3, Paragraph 4, “instead of imposing hard orthogonality constraints, we add penalty terms on the objective function and propose the following final overall objective function”; see also Wang, Page 2294, Section 3.3, Equation 7 Final Term). It would be obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of Bhuyan to include applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I as taught by Wang. The motivation to do so would have been reducing redundancy which allows for better performance (Wang, Page 2294, Section 3.3, Paragraph 1). While Bhuyan does teach a plurality of trained server-maintained machine learning models (Bhuyan, Page 1, Abstract, Lines 5-8, “every client is able to train any one of M models at a time and the server maintains a model for each of the M models which is typically a suitably averaged version of the model computed by the clients”), the proposed combination thus far does not explicitly teach outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts. Yemini teaches outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts (Yemini, Page 3, Algorithm 2, Line 2, “Output: Global model x ( R ) ”). It would be obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include outputting the trained models as taught by Yemini. The motivation to do so would have been the ability to use the trained model by other devices aside from the server (Yemini, Page 2, Figure 1). Regarding claim 2, the rejection of claim 1 is incorporated, and further, Bhuyan teaches selecting the plurality of server-maintained machine learning models from a pool of machine learning models (Bhuyan, Page 2, Algorithm 1, Step 1, “globalModel[m] ← 0 Ɐ m ∈ {1, 2, …, M}”; “{1, 2, …, M}” is considered to be the “pool of machine learning models”). Regarding claim 4, the rejection of claim 1 is incorporated, and further the proposed combination teaches wherein the training of the plurality of local machine learning models at each client includes determining a loss based on a regularization term and one or more latent features of the local machine learning models (Wang, Page 2294, Section 3.3, Paragraph 3, “Under the assumption that input Xv and Xt are orthogonal [Wang et al., 2010], if we impose the orthogonal constraints on each layer’s weight matrix as shown in Eq. 5, the representations of the layer h(mv) v and h(mt) t are approximately orthogonal … we can impose the orthogonal regularization on the weighting matrices”; Wang, Page 2294, Section 3.3, Paragraph 4, “instead of imposing hard orthogonality constraints, we add penalty terms on the objective function and propose the following final overall objective function”; see also Wang, Page 2294, Section 3.3, Equation 7 Final Term). Regarding claim 8, the rejection of claim 1 is incorporated. The proposed combination thus far does not explicitly teach determining that a first client of the plurality of clients is disconnected or otherwise unable to receive the at least portions of the plurality of server-maintained machine learning models from the server; connecting the first client to a neighboring client that is able to communicate with the server; and sending the portions of the plurality of server-maintained machine learning models from the neighboring client to the first client. Yemini teaches determining that a first client of the plurality of clients is disconnected or otherwise unable to receive the at least portions of the plurality of server-maintained machine learning models from the server (Yemini, Page 2, Section II, Subsection B, Lines 1-7, “We consider a setting where the uplink connections between the clients and the PS are intermittent. As shown in Fig. 1, we model the connectivity of client i to the PS at round r by the Bernoulli random variable τ i (r) ∼ Bern( p i ), where τ i = 1 indicates the presence of an uplink communication opportunity, whereas τ i (r) = 0 indicates a blocked uplink”); connecting the first client to a neighboring client that is able to communicate with the server (Yemini, Page 3, Algorithm 1 COLREL-CLIENT: Collaborative Relaying, “Neighborhood of client I N i ”; see also Yemini, Page 2, Figure 1, “System model with intermittent uplink communication between clients and PS (dotted lines) and reliable communication between neighboring clients (solid lines)”); and sending the portions of the plurality of server-maintained machine learning models from the neighboring client to the first client (Yemini, Page 3, Algorithm 1, Steps 6 and 7, “Send ∆ x i to every j ∈ N i . Receive ∆ x j from every j ∈ N i ”). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include connecting disconnected clients to neighboring clients to receive the server-maintained machine learning models as taught by the proposed combination. The motivation to do so would have been to prevent client models from becoming stale, and thus improving convergence (Yemini, Page 1, Section I, Lines 13-14, “Stragglers deteriorate the convergence of FL as the computed local updates become stale”; Yemini, Page 1, Section I, Paragraph 3, Lines 8-10, “Using this approach, the PS receives new updates from disconnected clients, which would otherwise become stale and be discarded”). Regarding claim 9, the rejection of claim 8 is incorporated, and further, the proposed combination teaches wherein the connecting includes connecting the first client to a plurality of neighboring clients (Yemini, Page 3, Algorithm 1 COLREL-CLIENT: Collaborative Relaying, “Neighborhood of client I N i ”; see also Yemini, Page 2, Figure 1, “System model with intermittent uplink communication between clients and PS (dotted lines) and reliable communication between neighboring clients (solid lines)”). The proposed combination thus far does not explicitly teach performing an interpolation of the portions of the plurality of server-maintained machine learning models received from the plurality of neighboring clients. Yemini teaches performing an interpolation of the portions of the plurality of server-maintained machine learning models received from the plurality of neighboring clients (Yemini, Page 3, Algorithm 1, Step 8; Yemini, Page 3, Section C, Lines 4-7, “Then client i computes a weighted average of its own update and those of its neighbors in N i , i.e., ∆ x ~ i ( r + 1 ) = ∑ j ∈ N i ∪ { i } α i j ∆ x j ( r + 1 ) = ∑ j ∈ N i ∪ { i } α i j ( x j r , T - x r ) , where α i j is a non-negative importance weight assigned by client i while relaying the client j’s update”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include performing an interpolation of the portions of the plurality of server-maintained machine learning models received from the neighboring clients as taught by Yemini. The motivation to do so would have been improved convergence rate and accuracy of the global models (Yemini, Page 1, Abstract, Lines 11-14, “Numerical simulations substantiate our theoretical claims and demonstrate settings with intermittent connectivity between the clients and the PS, where our proposed algorithm shows an improved convergence rate and accuracy in comparison with the federated averaging algorithm”). Regarding claim 10, the rejection of claim 9 is incorporated, and further, the proposed combination teaches wherein the interpolation is W + = W + ∑ { i ∈ C b } A i ∙ ( W i - W ) , wherein W + is an interpolated model for models W received by the plurality of neighboring clients C b , and wherein A i is a linear combination weight for model W i (Yemini, Page 3, Section C, Lines 4-7, “Then client i computes a weighted average of its own update and those of its neighbors in N i , i.e., ∆ x ~ i ( r + 1 ) = ∑ j ∈ N i ∪ { i } α i j ∆ x j ( r + 1 ) = ∑ j ∈ N i ∪ { i } α i j ( x j r , T - x r ) , where α i j is a non-negative importance weight assigned by client i while relaying the client j’s update”). Regarding claim 11, Bhuyan teaches A system of training neural networks with federated learning (Bhuyan, Page 1, Abstract, Lines 3-5, “In this paper, we extend federated learning to the setting where multiple unrelated models are trained simultaneously”), the system comprising: memory storing instructions; and a plurality of processors that, when executing the instructions stored in the memory (Bhuyan, Page 3, Section IV, Subsection A, Lines 1-3, “The policies were tested on Synthetic(1,1) [15], [16] and Synthetic-IID [15], [16] datasets with logistic regression as the classification model”; Bhuyan, Page 4, Figure 1; A person of ordinary skill in the art would recognize that performing these experiments and determining this collection of results would require the use of a generic computer for each client, providing evidence for a “memory storing instructions” and “a plurality of processors”), collectively perform: sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models (Bhuyan, Page 1, Section II, Lines 1-4, “We consider the setting where the server trains M unrelated models in a distributed manner using a pool for clients. Each client has a separate dataset for each model. The server maintains a global version of each of the M models”; Bhuyan, Page 1, Section II, Paragraph 2, Lines 1-4, “At the start of each round, the server selects up to K clients to be used for training. K is a fixed parameter given as an input to our system. Each selected client receives global weights of the model it needs to train”); at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes determining a respective loss for each of the plurality of local machine learning models by (i) computing, for the plurality of local machine learning models at that client, one or more latent features from selected intermediate layers using the local-stored data yielding concatenated latent features G w i , and …; updating respective weights for each of the plurality of local machine learning models (Bhuyan, Page 1, Section II, Paragraph 2, Lines 4-7, “The selected clients train the global version of the model allotted to them on the corresponding local training dataset and send the updated model weights to the server”; Bhuyan, Page 2, Section III, Subsection B, Paragraph 3, Line 3, “Local loss, for model i of client k, is denoted by l t ( k , i ) ”); Bhuyan, Page 2, Algorithm 1, Step 10, “Update localModel[m,c] by local training”; A person of ordinary skill in the art would recognize that training the model suing a local training dataset would require computing latent features from intermediate layers); transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients (Bhuyan, Page 1, Section II, Paragraph 2, Lines 4-7, “The selected clients train the global version of the model allotted to them on the corresponding local training dataset and send the updated model weights to the server”; Bhuyan, Page 1, Section I, Lines 4-6, “The key feature in federated learning is that the local client datasets are never shared with each other or with the server”); at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients (Bhuyan, Page 1, Section II, Lines 7-10, “The server weights are an average of the corresponding client weights with importance given to each client equal to the proportion of total training samples in it”; see also: Bhuyan, Page 2, Algorithm 1, Steps 13 and 14). Bhuyan does not explicitly teach determining the loss by (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity nor outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts. Wang teaches determining the loss by (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity (Wang, Page 2294, Section 3.3, Paragraph 3, “Under the assumption that input Xv and Xt are orthogonal [Wang et al., 2010], if we impose the orthogonal constraints on each layer’s weight matrix as shown in Eq. 5, the representations of the layer h(mv) v and h(mt) t are approximately orthogonal … we can impose the orthogonal regularization on the weighting matrices”; Wang, Page 2294, Section 3.3, Paragraph 4, “instead of imposing hard orthogonality constraints, we add penalty terms on the objective function and propose the following final overall objective function”; see also Wang, Page 2294, Section 3.3, Equation 7 Final Term). It would be obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of Bhuyan to include applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I as taught by Wang. The motivation to do so would have been reducing redundancy which allows for better performance (Wang, Page 2294, Section 3.3, Paragraph 1). While Bhuyan does teach a plurality of trained server-maintained machine learning models (Bhuyan, Page 1, Abstract, Lines 5-8, “every client is able to train any one of M models at a time and the server maintains a model for each of the M models which is typically a suitably averaged version of the model computed by the clients”), the proposed combination thus far does not explicitly teach outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts. Yemini teaches outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts (Yemini, Page 3, Algorithm 2, Line 2, “Output: Global model x ( R ) ”). It would be obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include outputting the trained models as taught by Yemini. The motivation to do so would have been the ability to use the trained model by other devices aside from the server (Yemini, Page 2, Figure 1). Regarding claim 12, the rejection of claim 11 is incorporated, and further, Bhuyan teaches selecting the plurality of server-maintained machine learning models from a pool of machine learning models (Bhuyan, Page 2, Algorithm 1, Step 1, “globalModel[m] ← 0 Ɐ m ∈ {1, 2, …, M}”; “{1, 2, …, M}” is considered to be the “pool of machine learning models”). Regarding claim 14, the rejection of claim 11 is incorporated, and further, the proposed combination teaches wherein the training of the plurality of local machine learning models at each client includes determining a loss based on a regularization term and one or more latent features of the local machine learning models (Wang, Page 2294, Section 3.3, Paragraph 3, “Under the assumption that input Xv and Xt are orthogonal [Wang et al., 2010], if we impose the orthogonal constraints on each layer’s weight matrix as shown in Eq. 5, the representations of the layer h(mv) v and h(mt) t are approximately orthogonal … we can impose the orthogonal regularization on the weighting matrices”; Wang, Page 2294, Section 3.3, Paragraph 4, “instead of imposing hard orthogonality constraints, we add penalty terms on the objective function and propose the following final overall objective function”; see also Wang, Page 2294, Section 3.3, Equation 7 Final Term). Regarding claim 15, the rejection of claim 14 is incorporated, and further, the proposed combination teaches at the server, aggregating information received from the clients to perform the training of the plurality of server-maintained machine learning models with the updated weights (Bhuyan, Page 1, Section II, Lines 7-10, “The server weights are an average of the corresponding client weights with importance given to each client equal to the proportion of total training samples in it”; see also: Bhuyan, Page 2, Algorithm 1, Steps 13 and 14). Regarding claim 16, Bhuyan teaches A non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, cause a computing system to (Bhuyan, Page 3, Section IV, Subsection A, Lines 1-3, “The policies were tested on Synthetic(1,1) [15], [16] and Synthetic-IID [15], [16] datasets with logistic regression as the classification model”; Bhuyan, Page 4, Figure 1; A person of ordinary skill in the art would recognize that performing these experiments and determining this collection of results would require the use of a generic computer for each client, providing evidence for a “non-transitory computer readable storage medium”, “a computer readable program code”, and “a computing system”) train neural networks with federated learning (Bhuyan, Page 1, Abstract, Lines 3-5, “In this paper, we extend federated learning to the setting where multiple unrelated models are trained simultaneously”) by: sending at least portions of a plurality of server-maintained machine learning models from a server to a plurality of clients, yielding a plurality of local machine learning models (Bhuyan, Page 1, Section II, Lines 1-4, “We consider the setting where the server trains M unrelated models in a distributed manner using a pool for clients. Each client has a separate dataset for each model. The server maintains a global version of each of the M models”; Bhuyan, Page 1, Section II, Paragraph 2, Lines 1-4, “At the start of each round, the server selects up to K clients to be used for training. K is a fixed parameter given as an input to our system. Each selected client receives global weights of the model it needs to train”); at each client, training the plurality of local machine learning models with locally-stored data that is stored locally at that respective client, wherein the training at each client includes determining a respective loss for each of the plurality of local machine learning models by (i) computing, for the plurality of local machine learning models at that client, one or more latent features from selected intermediate layers using the local-stored data yielding concatenated latent features G w i , and …; updating respective weights for each of the plurality of local machine learning models (Bhuyan, Page 1, Section II, Paragraph 2, Lines 4-7, “The selected clients train the global version of the model allotted to them on the corresponding local training dataset and send the updated model weights to the server”; Bhuyan, Page 2, Section III, Subsection B, Paragraph 3, Line 3, “Local loss, for model i of client k, is denoted by l t ( k , i ) ”); Bhuyan, Page 2, Algorithm 1, Step 10, “Update localModel[m,c] by local training”; A person of ordinary skill in the art would recognize that training the model suing a local training dataset would require computing latent features from intermediate layers); transferring the respective updated weights from each client to the server without transferring the locally-stored data of the clients (Bhuyan, Page 1, Section II, Paragraph 2, Lines 4-7, “The selected clients train the global version of the model allotted to them on the corresponding local training dataset and send the updated model weights to the server”; Bhuyan, Page 1, Section I, Lines 4-6, “The key feature in federated learning is that the local client datasets are never shared with each other or with the server”); and at the server, training the plurality of server-maintained machine learning models with the updated weights sent from each of the clients (Bhuyan, Page 1, Section II, Lines 7-10, “The server weights are an average of the corresponding client weights with importance given to each client equal to the proportion of total training samples in it”; see also: Bhuyan, Page 2, Algorithm 1, Steps 13 and 14). Bhuyan does not explicitly teach determining the loss by (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity nor outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts. Wang teaches determining the loss by (ii) applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I to enforce model diversity (Wang, Page 2294, Section 3.3, Paragraph 3, “Under the assumption that input Xv and Xt are orthogonal [Wang et al., 2010], if we impose the orthogonal constraints on each layer’s weight matrix as shown in Eq. 5, the representations of the layer h(mv) v and h(mt) t are approximately orthogonal … we can impose the orthogonal regularization on the weighting matrices”; Wang, Page 2294, Section 3.3, Paragraph 4, “instead of imposing hard orthogonality constraints, we add penalty terms on the objective function and propose the following final overall objective function”; see also Wang, Page 2294, Section 3.3, Equation 7 Final Term). It would be obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of Bhuyan to include applying a regularization term R based on a distance between ( G w i ) T G w i and an identity matrix I as taught by Wang. The motivation to do so would have been reducing redundancy which allows for better performance (Wang, Page 2294, Section 3.3, Paragraph 1). While Bhuyan does teach a plurality of trained server-maintained machine learning models (Bhuyan, Page 1, Abstract, Lines 5-8, “every client is able to train any one of M models at a time and the server maintains a model for each of the M models which is typically a suitably averaged version of the model computed by the clients”), the proposed combination thus far does not explicitly teach outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts. Yemini teaches outputting, from the server and based on the training, a … trained server-maintained machine learning model… configured to provide model diversity and robustness to client data distribution shifts (Yemini, Page 3, Algorithm 2, Line 2, “Output: Global model x ( R ) ”). It would be obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include outputting the trained models as taught by Yemini. The motivation to do so would have been the ability to use the trained model by other devices aside from the server (Yemini, Page 2, Figure 1). Regarding claim 17, the rejection of claim 16 is incorporated, and further, Bhuyan teaches selecting the plurality of server-maintained machine learning models from a pool of machine learning models (Bhuyan, Page 2, Algorithm 1, Step 1, “globalModel[m] ← 0 Ɐ m ∈ {1, 2, …, M}”; “{1, 2, …, M}” is considered to be the “pool of machine learning models”). Regarding claim 19, the rejection of claim 16 is incorporated, and further, the proposed combination teaches wherein the training of the plurality of local machine learning models at each client includes determining a loss based on a regularization term and one or more latent features of the local machine learning models (Wang, Page 2294, Section 3.3, Paragraph 3, “Under the assumption that input Xv and Xt are orthogonal [Wang et al., 2010], if we impose the orthogonal constraints on each layer’s weight matrix as shown in Eq. 5, the representations of the layer h(mv) v and h(mt) t are approximately orthogonal … we can impose the orthogonal regularization on the weighting matrices”; Wang, Page 2294, Section 3.3, Paragraph 4, “instead of imposing hard orthogonality constraints, we add penalty terms on the objective function and propose the following final overall objective function”; see also Wang, Page 2294, Section 3.3, Equation 7 Final Term). Regarding claim 20, the rejection of claim 19 is incorporated, and further, the proposed combination teaches at the server, aggregating information received from the clients to perform the training of the plurality of server-maintained machine learning models with the updated weights (Bhuyan, Page 1, Section II, Lines 7-10, “The server weights are an average of the corresponding client weights with importance given to each client equal to the proportion of total training samples in it”; see also: Bhuyan, Page 2, Algorithm 1, Steps 13 and 14). Claims 3, 13, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Bhuyan in view of Wang in further view of Yemini in further view of Li et al., An Efficient Multi-Model Training Algorithm for Federated Learning, 2021 IEEE Global Communications Conference (GLOBECOM), Madrid, Spain, 2021, pp. 1-6, doi: 10.1109/GLOBECOM46510.2021.9685230, hereinafter referred to as "Li". Regarding claim 3, the rejection of claim 2 is incorporated. The proposed combination does not explicitly teach wherein the plurality of server-maintained machine learning models are selected from the pool based on resource limits associated with the plurality of clients. Li teaches wherein the plurality of server-maintained machine learning models are selected from the pool based on resource limits associated with the plurality of clients (Li, Page 4, Section IV, Paragraph 3, Lines 1-4, “Algorithm 1 gives the detailed algorithm design of LFMB. In Algorithm 1, lines 1-5 perform initialization, where lines 1-4 randomly assign models to all clients iteratively until each client has no spare resource to train any more model”; Li, Page 1, Abstract, Lines 7-9, “The objective is to effectively utilize the heterogeneous resources at clients for parallel multi-model training”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include selecting models based on resource limits of the clients as taught by Li. The motivation to do so would have been to increase training efficiency (Li, Page 1, Abstract, Lines 7-11, “The objective is to effectively utilize the heterogeneous resources at clients for parallel multi-model training and therefore maximize the overall training efficiency while ensuring a certain fairness among individual models”). Regarding claim 13, the rejection of claim 12 is incorporated. The proposed combination does not explicitly teach wherein the plurality of server-maintained machine learning models are selected from the pool and sent to the plurality of clients based on resource limits of the plurality of clients. Li teaches wherein the plurality of server-maintained machine learning models are selected from the pool and sent to the plurality of clients based on resource limits of the plurality of clients (Li, Page 4, Section IV, Paragraph 3, Lines 1-4, “Algorithm 1 gives the detailed algorithm design of LFMB. In Algorithm 1, lines 1-5 perform initialization, where lines 1-4 randomly assign models to all clients iteratively until each client has no spare resource to train any more model”; Li, Page 1, Abstract, Lines 7-9, “The objective is to effectively utilize the heterogeneous resources at clients for parallel multi-model training”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include selecting models based on resource limits of the clients as taught by Li. The motivation to do so would have been to increase training efficiency (Li, Page 1, Abstract, Lines 7-11, “The objective is to effectively utilize the heterogeneous resources at clients for parallel multi-model training and therefore maximize the overall training efficiency while ensuring a certain fairness among individual models”). Regarding claim 18, the rejection of claim 17 is incorporated. The proposed combination does not explicitly teach wherein the plurality of server-maintained machine learning models are selected from the pool and sent to the plurality of clients based on resource limits of the plurality of clients. Li teaches wherein the plurality of server-maintained machine learning models are selected from the pool and sent to the plurality of clients based on resource limits of the plurality of clients (Li, Page 4, Section IV, Paragraph 3, Lines 1-4, “Algorithm 1 gives the detailed algorithm design of LFMB. In Algorithm 1, lines 1-5 perform initialization, where lines 1-4 randomly assign models to all clients iteratively until each client has no spare resource to train any more model”; Li, Page 1, Abstract, Lines 7-9, “The objective is to effectively utilize the heterogeneous resources at clients for parallel multi-model training”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the federated learning method of the proposed combination to include selecting models based on resource limits of the clients as taught by Li. The motivation to do so would have been to increase training efficiency (Li, Page 1, Abstract, Lines 7-11, “The objective is to effectively utilize the heterogeneous resources at clients for parallel multi-model training and therefore maximize the overall training efficiency while ensuring a certain fairness among individual models”). Response to Arguments Applicant asserts, on pages 9-10 of the response, that applicant has submitted a terminal disclaimer to overcome the double patenting rejection of claims 1, 8, and 9-11. While no terminal disclaimer appears to have been filed, the amendments to the claims have differentiated the claims of the instant application from the copending 18/217,006 application, and thus the double patenting rejection has been withdrawn. Applicant’s arguments regarding the 35 U.S.C. 112(a) rejections of the claims have been fully considered but are unpersuasive. Applicant argues the regularization term is a separate additive term outside the summation, and that the specification provides direct guidance for computing the claimed loss. Examiner respectfully disagrees. It appears the equation is missing parentheses and is thus incomplete. There is an opening parentheses for the loss function and an opening parentheses for the function f W j and then it becomes unclear. If the closing parentheses following D i closes the function, then it is unclear how y i is meant to be evaluated. Further, there are no parentheses to designate the beginning and end of the terms of the summation, and thus the interpretation that the summation includes the regularization term is not unreasonable and as the claim nor specification provides guidance on where the summation begins and ends. Please refer to the updated 35 U.S.C. 112(a) rejection above. Applicant’s arguments regarding the 35 U.S.C. 101 rejections of the claims have been fully considered but are unpersuasive. Argument 1: Applicant argues, in page 10, final paragraph of the response, that the amended claim recite the technical mechanism that provides an improvement and points to “selected intermediate-layer latent features are computed and constrained using an orthogonality-based regularization term during client-side training”. Examiner’s Response: Examiner respectfully disagrees. The limitations relied on by the applicant that provide the alleged improvement have been identified as abstract ideas in the updated 35 U.S.C. 101 rejection above. It is important to note the judicial exception alone cannot provide the improvement, see MPEP 2106.05(a). Argument 2: Applicant argues, on page 11, paragraphs 1-2 of the response, that the amended claims integrate any abstract idea into a practical application because the amended claims recite how the federated-learning training is technically constrained. Applicant specifically pointed out computing latent features, and applying orthogonality-enforcing regularization. Examiner’s Response: Examiner respectfully disagrees. relied on by the applicant that provide the alleged improvement have been identified as abstract ideas in the updated 35 U.S.C. 101 rejection above. It is important to note the judicial exception alone cannot provide the improvement, see MPEP 2106.05(a). An improvement to regularization may be an improvement in an abstract idea, but not an improvement in the functioning of a computer, as a computer. Applicant’s arguments regarding the 35 U.S.C. 102 rejections of the claims have been fully considered but are unpersuasive. Applicant argues, on page 12 of the response, that the claimed invention “requires that the client train a plurality of local machine learning models at the client, determine respective losses for those models, and update respective weights for those models”, and thus, Bhuyan does not anticipate the claimed arrangement. Examiner respectfully disagrees. Applicant agrees that Bhuyan teaches a client that is capable of handling any job over time, and in a given round it receives and trains only the model assigned to it. Thus, Bhuyan does teach a federated learning method where a client trains a plurality of local machine learning models and determines respect losses for those models, and updates respective weights for those models, it simply does this over a series of rounds. The claims do not require the client train multiple models simultaneously in the same round, only that the client trains multiple models, which is taught by Bhuyan. With regard to applicant’s arguments toward the newly added limitations, Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant's arguments regarding the 35 U.S.C. 103 rejection rely upon the arguments asserted with respect to the 35 U.S.C. 102, and are thus unpersuasive. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOLLY CLARKE SIPPEL whose telephone number is (571)272-3270. The examiner can normally be reached Monday - Friday, 7:30 a.m. - 4:30 p.m. ET.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571)272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.C.S./ Examiner, Art Unit 2122 /MICHAEL H HOANG/ PRIMARY EXAMINER, Art Unit 2122
Read full office action

Prosecution Timeline

Jun 30, 2023
Application Filed
Mar 31, 2026
Non-Final Rejection mailed — §101, §103, §112
Jun 23, 2026
Response Filed
Sep 17, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748955
REINFORCEMENT LEARNING WITH ADAPTIVE RETURN COMPUTATION SCHEMES
4y 1m to grant Granted Sep 29, 2026
Patent 12670387
SYSTEM, METHOD, AND COMPUTER-READABLE MEDIA FOR LEAKAGE CORRECTION IN GRAPH NEURAL NETWORK BASED RECOMMENDER SYSTEMS
4y 1m to grant Granted Jun 30, 2026
Patent 12664398
SYSTEM, METHOD AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 9m to grant Granted Jun 23, 2026
Patent 12657427
Systems, Methods, and Computer Program Products for Determining Uncertainty from a Deep Learning Classification Model
4y 1m to grant Granted Jun 16, 2026
Patent 12632779
HYPERPARAMETER SELECTION USING BUDGET-AWARE BAYESIAN OPTIMIZATION
4y 5m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
50%
Grant Probability
76%
With Interview (+25.7%)
3y 10m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 28 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month