DETAILED ACTION
This Non-Final Rejection is responsive to the RCE amendment filed 2/17/2026. Claims 23-28 are newly added. Claims 1, 3-7, 9, 11-14, and 20-28 are pending. Claims 1, 12 and 20 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
The rejection of claims 1-9, 11 and 20 remain and 21-22 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more, have been withdrawn as necessitated by the amendment and arguments as indicated in the response to arguments.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-7, 9, 11-14, and 25-28, are rejected under 35 U.S.C. 103 as being unpatentable over of Amad-Ud-Din et al. US PG PUB 2022/0012601, in view of McCourt et al. US PG PUB 2021/0034924, and further in view Basu et al. US PG PUB 2020/0226496.
Regarding Claim 1, Amad-Ud-Din et al teaches a server aggregating a plurality of updated received from clients in a federated learning system (32, 35, and Fig.1),-- an aggregator in a federated learning environment, …computing devices, wherein the plurality of computing devices include respective parties in the federated learning environment and include respective local machine learning models;
Moreover, Amad-Ud-Din et al teaches that the server 202 collects or aggregates updates, having hyper parameter values, and validation set performances, received from clients (38 , fig.2)--receiving HPO results from each of the plurality of computing devices; generating a unified data which incorporates the HPO results from the plurality of computing devices by: combining the HPO results from the plurality of computing devices into a single set of HP/performance metric value pairs,
Additionally, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2)—and using the set of HP/performance metric value pairs to train a global machine learning model in the aggregator; determining optimal global hyperparameters, utilizing the unified performance metric surface; The updated master model is then distributed to the clients(40, fig.2)-- and sending the optimal global hyperparameters to the respective plurality of computing devices.
Amad-Ud-Din et al. fails to explicitly teach issuing a hyperparameter optimization (HPO) query. However, McCourt et al. teach "The API 105 may be specifically designed to include a limited number of API endpoints that reduce of complexity in creating an optimization work request.. The optimization work request, as referred to herein, generally relates to an API request that includes one or more hyperparameters that a user is seeking to optimize and one or more constraints that the user desires for the optimization trials performed by the intelligent optimization platform 110"(37).
to a plurality of computing devices; McCourt et al., teach "The API 105 may additionally be configured with logic that enables the API 105 to intelligently parse optimization work request data and/or augment the optimization work request data with metadata prior to passing the optimization work request to the shared work queue 135" and Par. 48. "the plurality of intelligent queue worker machines 120 include multiple, distinct intelligent queue worker machines 120 that coordinate optimization work request from the shared work queue 135 received via the API 105", signifies a hyperparameter optimization request from the API sent to the queue then to the computing devices being the worker machines(Fig. 1 and Par. 39). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al show improving an effectiveness including one or more objective performance metrics of a machine learning model (6-7).
In addition, Amad-Ud-Din et al. fails to explicitly teach generating a unified performance metric surface. However Basu et al. discloses “These operations provide an initial idea of the surface f (x)”, and Par. 23, “Let x denote the vector of hyperparameters that lie in the domain X•IRD. After training a selected machine-learning model using x, let f(x) denote a predetermined quality metric. As examples, f(x) may be an Area under the Curve (AUC), negative Root Mean Square Error (RMSE), a normalized discounted cumulative gain (NDCG), or other such quality or performance metrics”-- (Basu et al., Par. 22)-- a performance metric surface in f(x) and it being unified as a function with multiple inputs).
Amad-Ud-Din et al. and Basu et al. all teach analogous art to the present invention because they are all in the same field of endeavor to optimize hyperparameters. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Amad-Ud-Din et al. and Basu et al. The motivation to do so is to be able to reduce the processing cost, and time involved in the tuning of the hyperparameters (Basu et al., Par.9-11).
Regarding Claim 3, McCourt et al. teaches the following limitation.
The computer-implemented method of Claim 1, wherein the HPO query (McCourt et al., Par. 22 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model”, signifies a request to optimize hyperparameters);
includes a request to perform a plurality of HPO operations at the plurality of computing devices (McCourt et al., Par. 22 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model… execute a tuning operation based on a tuning of the joint tuning function; identify a Pareto efficient frontier curve defined by a plurality of distinct hyperparameter values based on the execution of the tuning operation”, signifies plurality of computing devices with request to complete multiple hyperparameter optimization operations.).
McCourt et al., Khodak et al. nor Amad-Ud-Din et al. teaches the following limitations however Basu et al. does.
and a performance metric to be optimized. (Basu et al., Fig. 3D and Par. 89, “operation 326, the master server 112 adjusts the hyperparameter values, the performance metric values, and/or quality metric values by the determined bias value. This adjusted set of values is then merged with the prior set of hyperparameter values and performance metric values and/or quality metric values (Operation 328)”, and Par. 90, “At Operation 330, the master server 112 identifies the merged set of hyperparameter values and corresponding performance metric values as the next set of values to use in performing the next iteration of Operations 320-330.”, signifies operations that find the best possible performance metrics to use for each iteration of hyperparameter optimization).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., the FedEx algorithm of Khodak et al., the Federated learning hyperparameter tuning server holding a global model and performance metrics of Amad-Ud-Din et al. with the optimizing the performance metrics of Basu et al. The motivation to do so is to be able to measure model accuracy (Basu et al., Par. 5 “f is a black-box function that corresponds to model accuracy;”) and improve the measure of accuracy during the optimization process.
Regarding Claim 4, McCourt et al. does not teach the following limitations however Khodak et al. does.
using all of the HPO results from the plurality of computing devices; (Khodak et al., EQ. 8 Sec. 3, Par. 5--“FedEx tunes client (on-device training) hyperparameters”, Sec. 3, Par. 5, “To obtain FedEx from weight-sharing we restrict to the case of tuning only the hyperparameters c of local training Locc.”and Algorithm 2, shows that clients with their devices optimizing their hyperparameters with Locc and sending their resulting model, hyperparameter and loss to be received by the server, and Sec. 4, Par. 1, “Note we use "client" and "task" interchangeably, as the goal is a meta-learning (personalization) result. The performance measure here is task-averaged regret, which takes the average over T clients of the regret they incur on its loss… Here it,i is the ith loss of client t, Wt,i the parameter chosen on its ith round from a compact parameter space W, and w; E argmillwEW LZ:,1 it,i(w) the task optimum.” And Algorithm 2, signifies a performance equation made from local results of each client from algorithm 2.).
McCourt et al., Khodak et al. nor Amad-Ud-Din et al. teaches the following limitations however Basu et al. does.
McCourt et al., Khodak et al. nor Amad-Ud-Din et al. teaches the following limitations however Basu et al. does.
wherein the unified performance metric surface is generated. (Basu et al., Par. 22., “These operations provide an initial idea of the surface f (x)”, and Par. 23, “Let x denote the vector of hyperparameters that lie in the domain X•IRD. After training a selected machine-learning model using x, let f(x) denote a predetermined quality metric. As examples, f(x) may be an Area under the Curve (AUC), negative Root Mean Square Error (RMSE), a normalized discounted cumulative gain (NDCG), or other such quality or performance metrics”, signifies a performance metric surface in f(x) and it being unified as a function with multiple inputs).
McCourt et al., Khodak et al., Amad-Ud-Din et al. and Basu et al. all teach analogous art to the present invention because they are all in the same field of endeavor to optimize hyperparameters. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al. with the FedEx algorithm of Khodak et al. The motivation to do so is to optimize local models on a personalized client device and receive the information their device for further updating of a server model.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., the FedEx algorithm of Khodak et al. and the Federated learning hyperparameter tuning server holding a global model and performance metrics of Amad-Ud-Din et al. The motivation to so is to have more personalized results for a specific user device (Amad-Ud-Din et al., Par. 10 “Hyper-parameter optimization in a Federated Learning mode enables providing more accurate personalized recommendations.”) and improve accuracy, reduce complexity and not relying on repeated offline testing of different configurations (Amad-Ud-Din et al., Par. 12 “the processor is configured to cause the server apparatus to maintain a pairwise history of received hyper-parameter values and corresponding validation set performance metrics obtained from the master machine learning model on the federated learning server. The aspects of the disclosed embodiments allows for adaptive tuning of the hyper-parameters while the Federated Learning model is being trained and does not rely on the repeated off-line testing of the hyper-parameter configurations. This continuous online tuning not only improves the accuracy of recommendations but also helps to achieve faster convergence thereby reducing the computational complexity.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., the FedEx algorithm of Khodak et al., the Federated learning hyperparameter tuning server
holding a global model and performance metrics of Amad-Ud-Din et al. with the performance metrics of Basu et al. The motivation to do so is be able measure model accuracy (Basu et al., Par. 5 “f is a black-box function that corresponds to model accuracy;”)
Regarding Claim 5, Amad-Ud-Din et al teaches that the server 202 collects or aggregates updates, having hyper parameter values, and validation set performances, received from clients (38 , 11, fig.2)-- wherein using the set of HP/performance metric value pairs to train the global machine learning model in the aggregator includes: mapping hyperparameters in the set of HP/performance metric value pairs to single scalar values in the global machine learning model,
Additionally, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2)— and assigning the set of HP/performance metric value pairs as targets of the global machine learning model.
Regarding Claim 6, Amad-Ud-Din et al. does. wherein the HP/performance metric value pairs include: the hyperparameters and corresponding loss values generated in response to the HPO operations utilizing the respective hyperparameters in the local machine learning models of the respective computing devices, and wherein the unified performance metric is a loss surface. (Amad-Ud-Din et al., Fig. 2-3 and Par. 59, 57-58, 39-40, “In one embodiment, the optimizer 204 is also configured to maintain 508 a pairwise history of hyperparameter values. The optimizer 204 can use this pairwise history of hyper-parameter values to train 510 an optimization model. The training of the optimization model can be used to generate the new set of hyper-parameters that will be used to update the master model of the server 202.”, signifies using performance metrics and hyperparameters, including log loss values received from the clients to train for performance optimization”). Basu teaches the “the unified performance metric surface” as indicated in claim 1 above.
Regarding Claim 7, Amad-Ud-Din et al teach that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances pair to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2)-- for the respective hyperparameters in the set of HP/performance metric value pairs. Amad-Ud-Din et al. fails to explicitly teach generating predictions utilizing the unified performance metric, and using the predictions to determine a performance metric value for the respective hyperparameters; and selecting ones of the hyperparameters that produce an optimal performance metric, when compared to other ones of the hyperparameters, as optimal global hyperparameters. However, McCourt et al teach “setting or defining a metric threshold and/or a metric constraint as a maximum metric value or a minimum metric value for feasible metric values that may be returned for given optimization or tuning trials…” (63, 58). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the Prior art, and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al teach optimizing hyperparameters by excluding hyperparameter points that exceed a maximum threshold(63). Basu teaches the “the unified performance metric surface” as indicated in claim 1 above.
Regarding Claim 9, Amad-Ud-Din et al. teach utilizing the optimal global hyperparameters to determine a structure of the global machine learning model, train the global machine learning model, or determine the global machine learning model structure and train the global machine learning model. (Amad-Ud-Din et al., Fig. 2 and Par. 39, “The hyper-parameter optimizer 204 uses the historical data (hyper-parameter values-validation set performances) to train an optimization model, such as for example a Bayesian optimization model. Given the current hyperparameters and performance values as a new input query, the optimization model infers the next set of potentially optimal hyper-parameters for the Federated Learning master model”, and Par. 40 “The Federated Learning master model 202 is configured to update the current hyper-parameters in the master model with the new values”, signifies updating hyperparameters and training a master model which is a global model that shares it’s hyperparameters to client devices).
Claim 12 is a computer program product that recites limitations similar to claims 1, and 8, and is likewise rejected.
Regarding Claim 13, Amad-Ud-Din et al teaches that the server 202 collects or aggregates updates, having hyper parameter values, and validation set performances, received from clients (38 , fig.2)-- wherein the aggregator is used to combine the HPO results from the plurality of computing devices into the single set of HP/performance metric value pairs.
Claim 14 is directed towards a computer program product that recites limitations similar to claim 3, and is likewise rejected.
Claim 25, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2). Amad-Ud-Din et al. fails to explicitly teach wherein the determining of the optimal global hyperparameters further includes an uncertainty quantification of a prediction of the global machine learning model. However, McCourt et al teach “generate a surrogate (or proxy) model that can be used to test the uncertainty and/or the likelihood that a candidate data point would perform well in an external model….”(51). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the Prior art, and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al ensuring good performance of the hyperparameter data point(50).
Regarding Claim 26, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2). Amad-Ud-Din et al. fails to explicitly teach wherein the uncertainty quantification penalizes a respective hyperparameter if the prediction includes uncertainty around the hyperparameter. However, McCourt et al teach “generate a surrogate (or proxy) model that can be used to test the uncertainty and/or the likelihood that a candidate data point would perform well in an external model….rank multiple suggestions for a given optimization work request (or across multiple optimization work requests for a given user) such that the suggestions having hyperparameter values most likely to perform the best can be passed or pulled via the API 105…” (51-52). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the Prior art, and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al teach using random forest algorithm to optimize hyperparameters and using hyperparameters that will perform best(52).
Regarding Claim 27, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances pair to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2)-- wherein the using of the set of HP/performance metric value pairs to train the global machine learning model in the aggregator comprises training a respective global machine learning model per each HP/performance metric value pair. Amad-Ud-Din et al. fails to explicitly teach and wherein the unified performance metric outputs a maximum value of the global machine learning models. However, McCourt et al teach “setting or defining a metric threshold and/or a metric constraint as a maximum metric value or a minimum metric value for feasible metric values that may be returned for given optimization or tuning trials…” (63, 58). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the Prior art, and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al teach optimizing hyperparameters by excluding hyperparameter points that exceed a maximum threshold(63). Basu teaches the “the unified performance metric surface” as indicated in claim 1 above.
Regarding Claim 28, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2). Amad-Ud-Din et al. fails to explicitly teach wherein the using of the set of HP/performance metric value pairs to train the global machine learning model in the aggregator comprises training a respective global machine learning model per each HP/performance metric value pair, and wherein the unified performance metric outputs an average of the global machine learning models. However, McCourt et al teach using averaged one-dependence estimators algorithm for the machine learning models (50). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to output the average of the global machine learning models, by combining the Prior art, and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al teach using random forest algorithm to optimize hyperparameters (50). Basu teaches the “the unified performance metric surface” as indicated in claim 1 above.
Claim 11, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Amad-Ud-Din et al. US PG PUB 2022/0012601, in view of McCourt et al. US PG PUB 2021/0034924, in view of Basu et al. US PG PUB 2020/0226496, and further view of Agrawal et al., US PG PUB 2020/0125961.
Regarding Claim 11, Amad-Ud-Din et al fail to teach, but McCourt teaches the following limitations.
The computer-implemented method of Claim 1, wherein the HPO query (McCourt et al., Par. 22 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model”, signifies a request to optimize hyperparameters).
includes a request to perform a plurality of HPO operations at the plurality of computing devices (McCourt et al., Par. 22 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model… execute a tuning operation based on a tuning of the joint tuning function; identify a Pareto efficient frontier curve defined by a plurality of distinct hyperparameter values based on the execution of the tuning operation”, signifies plurality of computing devices with request to complete multiple hyperparameter optimization operations.).
Amad-Ud-Din et al. do not teach the following limitations. However, Basu et al. does.
and a performance metric to be optimized, (Basu et al., Fig. 3D and Par. 89, “operation 326, the master server 112 adjusts the hyperparameter values, the performance metric values, and/or quality metric values by the determined bias value. This adjusted set of values is then merged with the prior set of hyperparameter values and performance metric values and/or quality metric values (Operation 328)”, and Par. 90, “At Operation 330, the master server 112 identifies the merged set of hyperparameter values and corresponding performance metric values as the next
set of values to use in performing the next iteration of Operations 320-330.”, signifies operations that find the best possible performance metrics to use for each iteration of hyperparameter optimization).
where the performance metric of the HPO query is selected (Basu et al., Par. 90 “the master server 112 identifies the merged set of hyperparameter values and corresponding performance metric values as the next set of values to use”, signifies identifying then selecting performance metrics to use along with the hyperparameter value.).
Amad-Ud-Din et al. do not teach the following limitations. However, Agrawal et al. does.
from the group consisting of: predictive machine learning metrics including absolute or relative accuracy or loss, (Agrawal et al., Par. 36 “Furthermore, since the performance scores are generated to select the most accurate algorithm variant, the more precise the performing score itself is, the more precise the generated model's prediction relative accuracy is compared”, signifies using relative accuracy for evaluating predictions of a machine learning model).
and resource metrics including runtime (Agrawal et al., Par. 47, “the computational cost metric may be generated by measuring the time duration for generating a machine learning model by training the reference variant on a training data set. The time duration may be measured based on performing training using computational resources that are pre-defined.”, signifies a metric for runtime of a computational task using resources).
and memory utilization. (Agrawal et al., Par. 48, “Other embodiments of generating computation cost metric include, in addition or alternative to the time duration measurement, measurements based on one or more of memory consumption, processor utilization, or I/0 activity measurement.”).
Amad-Ud-Din et al., McCourt et al., Basu et al., and Agrawal et al. all teach analogous art to the present invention because they are all in the same field of endeavor to optimize hyperparameters. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., the FedEx algorithm of Khodak et al., the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., the performance metrics of Basu et al. and model performance metrics of Agrawal et al. The motivation to do so is to be able to select the most accurate model algorithm (Agrawal et al., Par. 36, “the performance scores are generated to select the most accurate algorithm variant”).
Regarding Claim 22, Amad-Ud-Din et al fails to teach, but McCourt teaches the following limitations.
wherein the HPO query includes a request to perform a plurality of HPO operations at the plurality of computing devices and a performance metric to be optimized (McCourt et al., Par. 22, 42 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model… execute a tuning operation based on a tuning of the joint tuning function; identify a Pareto efficient frontier curve defined by a plurality of distinct hyperparameter values based on the execution of the tuning operation”. This signifies a plurality of computing devices are provided with a request to optimize hyperparameters, and to complete multiple hyperparameter optimization operations.
where the performance metric of the HPO query is selected from the group consisting of: predictive machine learning metrics including absolute or relative accuracy or loss; hyperparameters and corresponding loss values generated in response to the HPO operations utilizing the respective hyperparameters in the local machine learning models of the respective computing devices--McCourt et al teach in Par. 58, 31 identifying of a first predictive accuracy metric to increase the prediction accuracy of the machine learning model.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al show improving an effectiveness including one or more objective performance metrics of a machine learning model (6-7).
Amad-Ud-Din et al. do not teach the following limitations however Agrawal et al. does.
and resource metrics including runtime (Agrawal et al., Par. 47, “the computational cost metric may be generated by measuring the time duration for generating a machine learning model by training the reference variant on a training data set. The time duration may be measured based on performing training using computational resources that are pre-defined.”, signifies a metric for runtime of a computational task using resources).
and memory utilization. (Agrawal et al., Par. 48, “Other embodiments of generating computation cost metric include, in addition or alternative to the time duration measurement, measurements based on one or more of memory consumption, processor utilization, or I/0 activity measurement.”).
McCourt et al., Amad-Ud-Din et al., Basu et al., and Agrawal et al. all teach analogous art to the present invention because they are all in the same field of endeavor to optimize hyperparameters. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., the FedEx algorithm of Khodak et al., the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., the performance metrics of Basu et al. and model performance metrics of Agrawal et al. The motivation to do so is to be able to select the most accurate model algorithm (Agrawal et al., Par. 36, “the performance scores are generated to select the most accurate algorithm variant”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., the FedEx algorithm of Khodak et al., the Federated learning hyperparameter tuning server
holding a global model and performance metrics of Amad-Ud-Din et al. with the performance metrics of Basu et al. The motivation to do so is to be able to measure model accuracy (Basu et al., Par. 5 “f is a black-box function that corresponds to model accuracy;”) and improve the measure of accuracy during the optimization process.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Khodak et al., "Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-Sharing" (listed in the IDS filed 12/9/2021) in view of Amad-Ud-Din et al. US PG PUB 2022/0012601, and in further view of McCourt et al. US PG PUB 2021/0034924.
Regarding Claim 20, Khodak et al. disclose “Note we use "client" and "task" interchangeably, as the goal is a meta-learning (personalization) result. The performance measure here is task-averaged regret, which takes the average over T clients of the regret they incur on its loss… Here it,i is the ith loss of client t, Wt,i the parameter chosen on its ith round from a compact parameter space W, and w; E argmillwEW LZ:,1 it,i(w) the task optimum.” And Algorithm 2.( EQ. 8 and Sec. 4, Par. 1,), signifies a performance equation made from local results of each client from algorithm 2 -- causing the computing device to generate a local performance metric surface which incorporates the local results of performing the HPO operations. “We tune several hyperparameters of both aggregation and local training; for the former we tune the server learning rate schedule and momentum, found to be helpful for personalization [19]; for the latter we tune the learning rate, momentum, weight-decay, the number of local epochs, the batch-size, dropout, and proximal regularization.”, signifies tuning local hyperparameters to get the most optimal results (Sec. 5, Par. 2) , which signifies using the loss to set or tune the hyperparameters to certain values which are used to control the model and determining optimal values-- by mapping hyperparameters in the local results of performing the HPO operations to single scalar values in a local machine learning model at the computing device, and assigning the local results of performing the HPO operations as targets of the local machine learning model; determining optimal local hyperparameters
Khodak et al. discloses “Note we use "client" and "task" interchangeably, as the goal is a meta-learning (personalization) result. The performance measure here is task-averaged regret, which takes the average over T clients of the regret they incur on its loss… Here it,i is the ith loss of client t, Wt,i the parameter chosen on its ith round from a compact parameter space W, and w; E argmillwEW LZ:,1 it,i(w) the task optimum.”, and Sec. 5 Par. 3 “Because our baseline is running Algorithm 1, a standard hyperparameter tuning algorithm, to tune all hyperparameters…We run Algorithm 1 on the personalized objective”, signifies using an objective function which can act as a performance metric that is local to clients to tune both local and global hyperparameters(EQ. 8 and Sec. 4, Par. 1)-- by applying the optimal global hyperparameters to the local performance metric surface. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the references. The motivation as indicated by Khodak et al to do so is to be able to accelerate hyperparameter tuning (abst).
Khodak et al fail to explicitly teach, but Amad-Ud-Din et al. teach receiving, at a computing device from an aggregator in a federated learning environment, optimal global hyperparameters; wherein the aggregator includes a global machine learning model. “The Federated Learning master model 202 is configured to update the current hyper-parameters in the master model with the new values and distribute the updated copy of the master model across one or more of clients 200-a to 200-n.”, and Par. 57 ““The server 202 is also configured to aggregate the predictive model updates obtained from each client 200 and update the master model of the predictive model 314 of the server 202.”,” signifies global hyperparameters as the figure shows a global model from the server/aggregator sharing the hyperparameters with local models in client devices.( Fig. 2 and Par. 40).
In addition, Amad-Ud-Din et al. show further comprising: using the optimal local hyperparameters and the optimal global hyperparameters to re-train the local machine learning model. “The client side 200 can also able to update and train the local master model on the local or user's device, such as device 200n of FIG. 2. Based on the user's viewing of different videos in the Huawei video service, the different videos are randomly divided into training…” signifies global hyperparameters as the figure shows a global model from the server/aggregator sharing the hyperparameters with local models in client devices.( Fig. 2 and Par. 50). The master model stored at the client is updated using local training or hyperparameter data.
Moreover, Amad-Ud-Din et al teaches a server aggregating a plurality of updated received from clients in a federated learning system (32, 35, and Fig.1)-- teach sending local results of performing the HPO operations.
Additionally, Khodak et al fail to explicitly teach, but McCourt et al teach receiving,…, a hyperparameter optimization (HPO) query (McCourt et al., Par. 22 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model”, signifies a receiving a request to optimize hyperparameters).
causing the computing device to perform HPO operations in response to receiving the query; (McCourt et al., Par. 22 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model… execute a tuning operation based on a tuning of the joint tuning function; identify a Pareto efficient frontier curve defined by a plurality of distinct hyperparameter values based on the execution of the tuning operation”. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the performance metrics of Amad-Ud-Din et al, with the hyperparameter tuning requests of McCourt et al. The motivation to do so is be able to define parameters and objectives to getting the optimal hyperparameters (McCourt et al., Par. 16, “the multi-criteria tuning request are defined including: defining the first objective function of the model, defining the second objective function of the model, defining the one or more metric thresholds for constraining a hyperparameter search space during the tuning operation.”, and Par. 34, “metric thresholds and/or metric constraints may be defined for improved searching of the Pareto-efficient frontier in view of one or more real-world limitations of a subject model. In particular, metric thresholds and/or metric constraints may be applied to a hyperparameter search space that include the Pareto-efficient frontier for a given model with competing metrics. In such cases, the metric thresholds and/or metric constraint reduce and/or constrain the available hyperparameters within the search space forcing a directed search along preferred sections of the Pareto-efficient frontier which may include hyperparameter values that potentially address one or more real-world limitations of a subject model.”).
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Khodak et al., "Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-Sharing" (listed in the IDS filed 12/9/2021) in view of Amad-Ud-Din et al. US PG PUB 2022/0012601, in view of McCourt et al. US PG PUB 2021/0034924, as applied to claim 20 above, and further in view of Agrawal et al., US PG PUB 2020/0125961.
Regarding Claim 21, Khodak et al. disclose “Note we use "client" and "task" interchangeably, as the goal is a meta-learning (personalization) result. The performance measure here is task-averaged regret, which takes the average over T clients of the regret they incur on its loss… Here it,i is the ith loss of client t, Wt,i the parameter chosen on its ith round from a compact parameter space W, and w; E argmillwEW LZ:,1 it,i(w) the task optimum.”, and Sec. 5 Par. 3 “Because our baseline is running Algorithm 1, a standard hyperparameter tuning algorithm, to tune all hyperparameters…We run Algorithm 1 on the personalized objective”, signifies using an objective function which can act as a performance metric that is local to clients to tune both local and global hyperparameters(EQ. 8 and Sec. 4, Par. 1, fig.1)-- further comprising: using the optimal local hyperparameters and the optimal global hyperparameters to re-train the local machine learning model.
where the performance metric of the HPO query is selected from the group consisting of: predictive machine learning metrics including absolute or relative accuracy or loss-- (Khodak et al., EQ. 8 and Sec. 4, Par. 1, “Note we use "client" and "task" interchangeably, as the goal is a meta-learning (personalization) result. The performance measure here is task-averaged regret, which takes the average over T clients of the regret they incur on its loss… Here it,i is the ith loss of client t, Wt,i the parameter chosen on its ith round from a compact parameter space W, and w; E argmillwEW LZ:,1 it,i(w) the task optimum.” And Algorithm 2, signifies a performance equation made from local results, including local losses, of each client from algorithm 2.).
Moreover, Khodak et al do not teach the following limitation, wherein the HPO query includes a request to: perform a plurality of HPO operations at the computing device and optimize a performance metric. However, McCourt et al. teach "The API 105 may be specifically designed to include a limited number of API endpoints that reduce of complexity in creating an optimization work request.. The optimization work request, as referred to herein, generally relates to an API request that includes one or more hyperparameters that a user is seeking to optimize and one or more constraints that the user desires for the optimization trials performed by the intelligent optimization platform 110"(37, 39, fig.1). (McCourt et al. teach in Par. 22, 37 “a system for tuning hyperparameters for improving an effectiveness including one or more objective performance metrics of a machine learning model includes: a hyperparameter tuning service that tunes hyperparameters of the machine learning model of a subscriber to the hyperparameter tuning service, wherein the remote tuning service is implemented by a distributed network of computers that: receive a multi-criteria tuning request for tuning hyperparameters of a machine learning model… execute a tuning operation based on a tuning of the joint tuning function; identify a Pareto efficient frontier curve defined by a plurality of distinct hyperparameter values based on the execution of the tuning operation”, signifies plurality of computing devices with request to complete multiple hyperparameter optimization operations.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the hyperparameter tuning requests of McCourt et al., and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al show improving an effectiveness including one or more objective performance metrics of a machine learning model (6-7).
Khodak et al do not teach the following limitation, however Agrawal et al. do. teach and resource metrics including runtime (Agrawal et al., Par. 47, “the computational cost metric may be generated by measuring the time duration for generating a machine learning model by training the reference variant on a training data set. The time duration may be measured based on performing training using computational resources that are pre-defined.”, signifies a metric for runtime of a computational task using resources).
and memory utilization. (Agrawal et al., Par. 48, “Other embodiments of generating computation cost metric include, in addition or alternative to the time duration measurement, measurements based on one or more of memory consumption, processor utilization, or I/0 activity measurement.”).
Khodak et al., Amad-Ud-Din et al., McCourt et al., and Agrawal et al. all teach analogous art to the present invention because they are all in the same field of endeavor to optimize hyperparameters. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of the references. The motivation to do so is to be able to select the most accurate model algorithm (Agrawal et al., Par. 36, “the performance scores are generated to select the most accurate algorithm variant”).
Claims 23-24 are rejected under 35 U.S.C. 103 as being unpatentable over over of Amad-Ud-Din et al. US PG PUB 2022/0012601, in view of McCourt et al. US PG PUB 2021/0034924, in view Basu et al. US PG PUB 2020/0226496, further in view Paliwal et al, hereinafter Paliwal -US PG PUB 20230133880 A1, filed 10/29/2021.
Regarding Claim 23, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2). Amad-Ud-Din et al. fails to explicitly teach wherein the global machine learning model comprises a regressor which considers the hyperparameters as covariates and a corresponding performance metric as a dependent variable. However, Paliwal teaches using a random forest model where the mean or average prediction is returned and estimating relationships between a dependent variable and independent variables called covariates (76).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the Prior art, the model of Paliwal., and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because Paliwal teaches improving an effectiveness including one or more objective performance metrics of a machine learning model (6-7). Also McCourt et al teach using random forest algorithm to optimize hyperparameters (50).
Regarding Claim 24, Amad-Ud-Din et al teaches that a hyper-parameter optimizer 204 uses the hyper parameter values-validation set performances to train an optimization model by generating optimal hyperparameter configuration, which is used to update a master model (39-40, fig.2). Amad-Ud-Din et al. fails to explicitly teach wherein the regressor comprises a random forest regressor. However, McCourt et al teach using random forest algorithm (50). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the Prior art, the model of Paliwal., and the Federated learning hyperparameter tuning server holding a global model of Amad-Ud-Din et al., because McCourt et al teach using random forest algorithm to optimize hyperparameters (50).
Response to Arguments
Applicant's arguments filed 2/17/2026 regarding the 35 USC 101 rejection have been fully considered and are persuasive. Applicant indicates that “As in Ex parte Desjardins, the embodiments claimed in the present application do in fact improve the functioning of a computer. Reference is first made to 0003 of the present application, which states:
"Machine learning is a popular means of decision making and is used in many different fields today. Federated learning is a popular means of training machine learning models that provides data privacy for collaborative training entities. However, hyperparameters used within local and global models during the implementation of federated learning are currently manually assigned, and often require many manual iterations to fine-tune. There is therefore a need to dynamically implement hyperparameters for models used within a federated learning
environment." Moreover, 0074 of the present application proceeds to clarify that:
"In this way, optimal hyperparameters may be dynamically determined for the global model. This may reduce an amount of processing necessary to fine-tune the global model during federated learning, which may improve a performance of computing hardware performing the federated learning(e.g., the aggregator, etc.)."…”(pages 12-16). The examiner agrees, because the claims disclose an improved method to fine tune federated learning devices.
Applicant's arguments filed 2/17/2026 have been fully considered but they are not persuasive.
Regarding claims 1, and 12, the Applicant states that “For instance, looking again to the paragraphs cited in the rejection, it is clear that Basu is motivated to actively remove the number of samples that are evaluated. Basu's paragraph 0022 outlines that "[i]nstead of searching in the overall domain of X, the disclosed approach shrinks the search space to a "ball" (e.g., one or more points having a predetermined distance) around the maximum of the mean function of the posterior GP." Additionally, Basu's paragraph 0021 describes that "[b]ecause some function evaluations are computationally expensive, the disclosed embodiments begin with a subsampling approach that drastically reduces the sample size. This allows parallel executions to run on a sub-sampled dataset using a relatively large number of quasi-random points from X (e.g., the domain of hyperparameters)" (emphasis added).
It follows that the cited sections of Badu are unable to meet the claimed:
"generating a unified performance metric surface which incorporates the
HPO results from the plurality of computing devices by:
combining the HPO results from the plurality of computing devices into a single set of HP/performance metric value pairs, and using the set of HP/performance metric value pairs to train a global machine learning model in the aggregator..."
” (page 19). Applicant is directed towards the new grounds of rejection above, as necessitated by the amendment.
Additionally, Applicant is directed towards the rejection of claims 6, 11, and 22 above, as necessitated by the amendment.
Also, Applicant argues that “The art of record is also unable to meet the claimed "sending the optimal global hyperparameters to the respective plurality of computing devices." (page19). The Applicant is directed towards the new grounds of rejection above, showing the teaching for this limitation.
Claims 6,11, and 22 are rejected above in accordance with the submitted amendment.
New grounds of rejection have been made for claim 20 above in accordance with the submitted amendment.
Moreover, Applicant is directed towards the rejection of claim 21 above, as necessitated by the amendment.
Newly added claims 23-28 have been rejected as shown above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CESAR PAULA whose telephone number is (571)272-4128. The examiner can normally be reached Monday - Thursday 7am-4:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Wiley can be reached at (571)272-3923. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145