DETAILED ACTION
This action is responsive to Applicant’s reply filed July 17, 2026. This action is made final.
Status of the Claims
Claims 1, 10 and 19 are amended.
Claim status is currently pending and under examination for claims 1-20 of which independent claims are 1, 10 and 19.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Applicant’s amendments to the Claims have overcome each and every 35 USC § 101 rejections previously set forth in the Non-Final Office Action mailed April 20, 2026.
Applicant’s arguments regarding the art rejections are moot in view of the new grounds of rejection necessitated by Applicant’s amendment.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Fu, Jie. "Neural optimizers with hypergradients for tuning parameter-wise learning rates." AutoML, International Conference on Machine Learning (ICML). 2017.
Priyanshu, Aman, et al. "Efficient hyperparameter optimization for differentially private deep learning." arXiv preprint arXiv:2108.03888 (2021).
Claims 1-4, 6-7, 10-13, 15-16 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Fu / Priyanshu.
Regarding Claim 1, Fu teaches:
A method comprising ((P. 1, Abstract) “we train a neural optimizer to only control the learning rates of another optimizer using gradients of the training loss with respect to the learning rates”):
extracting, in a processor-based machine learning system comprising a first machine learning model arranged at least in part in parallel with a second machine learning model different than the first machine learning model, multiple parameters from a target machine learning model, different than the first and second machine learning models ((P. 2, Sec. 2.1, ¶1) “training a deep model with n parameters can be formulated as the problem of minimizing a function l”
(P. 6, Sec. 4) “an LSTM-based neural optimizer for learning parameter-wise learning rates using hypergradients directly from the optimizee”
Fu discloses Figure 1 on P. 3 (reproduced below) depicting an LSTM optimizer comprised of two LSTMs. The first LSTM is a ‘first machine learning model’ and the second LSTM is a ‘second machine learning model’. The LSTMs are arranged in parallel since they share the same parameters and are part of the same LSTM optimizer that determines learning rates. The LSTMs are different from each other since they have different hidden states (see Figure 1 Caption). The LSTM optimizer uses hypergradients (‘multiple parameters’) from an optimizee model (‘target machine learning model’) to determine learning rates (therefore extracting parameters from a target model). The optimizee model is a deep learning model and therefore is different from the LSTMs.
A computer with a processor is implied by using an LSTM optimizer to predict learning rates, therefore using an LSTM optimizer to predict learning rates is a ‘processor-based machine learning system’.
PNG
media_image1.png
384
1042
media_image1.png
Greyscale
),
wherein the multiple parameters comprise a learning rate, state information, a loss value, a gradient, and a weight ((P. 2, Sec. 1, Last Paragraph) “an LSTM-based optimizer learns to propose parameter wise learning rates for the hand-designed optimizer and is trained with hypergradients (the gradients with respect to learning rates).”
(P. 3, Sec. 2.2, ¶2) “we adopt the following rule:
w
t
+
1
=
w
t
-
g
t
h
∇
l
w
t
;
ϕ
⋅
∇
l
w
t
, where h(·), the input to the LSTM optimizer, is defined as the state description vector of the optimizee gradients at iteration t, and
g
t
h
∇
l
w
t
;
ϕ
=
α
t
.”
(P. 2, Sec. 2.1¶1-2) “l is optimized by iteratively adjusting
w
t
(the weight vector at time step t) using gradient information …
α
t
is the SGD learning rate at time t. … g(·) is the optimizer”
An LSTM optimizer
g
t
·
uses hypergradients (‘a gradient’), weights
w
t
, and state description vector h(·) (‘state information’) from an optimizee model as input to generate learning rate
α
t
. Weights
w
t
are input into LSTM optimizer
g
t
·
and weights
w
t
were updated by using a previous iteration’s learning rate (obtained from
g
t
-
1
·
=
α
t
-
1
), therefore weights
w
t
comprise a learning rate. To compute a gradient
∇
l
using weights
w
t
, loss values must be calculated for the optimizee model, therefore optimizer g(·) uses loss values of the optimizee model.),
and the target machine learning model is configured to execute tasks related to at least one of images, videos, voice, and text ((P. 1, Abstract) “The optimizee is trained by Adam on MNIST, and our neural optimizer learns to tune the learning rates for the Adam.”
An optimizee model (‘target machine learning model’) is trained on the MNIST dataset (image dataset), therefore the optimizee model executes tasks related to images.);
predicting, by the first machine learning model based on at least a first subset of the multiple parameters, a first learning rate associated with the target machine learning model ((P. 3, Sec. 2.2, ¶2) “To solve the above problems, our proposed one is trained to predict the learning rates. … we adopt the following rule:
w
t
+
1
=
w
t
-
g
t
h
∇
l
w
t
;
ϕ
⋅
∇
l
w
t
, where h(·), the input to the LSTM optimizer, is defined as the state description vector of the optimizee gradients at iteration t, and
g
t
h
∇
l
w
t
;
ϕ
=
α
t
.”
(P. 3, Sec. 2.2, ¶3) “we let all the LSTMs share the same parameters but have separate hidden states as shown in Fig. 1, and use the following pre-processing approaches:
h
k
⋅
=
(
log
∇
k
l
w
t
c
,
s
g
n
(
∇
k
l
w
t
)
… where c > 0 is a constant,
∇
k
l
w
t
is the gradient for the k-th parameter, and sgn(·) is the sign function.”
The LSTM optimizer
g
t
h
∇
l
w
t
;
ϕ
=
α
t
is used to predict learning rates. The LSTM optimizer is comprised of two LSTMs with different hidden states
h
k
⋅
. Since the LSTM optimizer
g
t
·
requires a hidden state h(·) as input to predict a learning rate, each LSTM with a different hidden state
h
k
⋅
therefore predicts its own learning rate (and therefore a first LSTM (‘first machine learning model’) predicts a first learning rate for the optimizee model (‘target machine learning model’) using a first subset of multiple parameters (hidden state, weights, gradients)).
See Figure 1 depicting each LSTM produces an output.),
predicting, by the second machine learning model based on at least a second subset of the multiple parameters, a second learning rate associated with the target machine learning model (The LSTM optimizer
g
t
h
∇
l
w
t
;
ϕ
=
α
t
is used to predict learning rates (see P. 3, Sec. 2.2, ¶2). The LSTM optimizer is comprised of two LSTMs with different hidden states
h
k
⋅
, (see P. 3, Sec. 2.2, ¶3). Since the LSTM optimizer
g
t
·
requires a hidden state h(·) as input to predict a learning rate, each LSTM with a different hidden state
h
k
⋅
therefore predicts its own learning rate (and therefore a second LSTM (‘second machine learning model’) predicts a second learning rate for the optimizee model (‘target machine learning model’) using a second subset of multiple parameters (hidden state, weights, gradients)).),
the second machine learning model comprising a neural network machine learning model ((P. 1, Abstract) “LSTM-based neural optimizers are competitive with state-of-the- art hand-designed optimization methods for short horizons”),
the neural network machine learning model being configured to receive as inputs at least the second subset of the multiple parameters extracted from the target machine learning model, and to generate as an output the predicted second learning rate ((P. 3, Sec. 2.2, ¶2) “To solve the above problems, our proposed one is trained to predict the learning rates. … we adopt the following rule:
w
t
+
1
=
w
t
-
g
t
h
∇
l
w
t
;
ϕ
⋅
∇
l
w
t
, where h(·), the input to the LSTM optimizer, is defined as the state description vector of the optimizee gradients at iteration t, and
g
t
h
∇
l
w
t
;
ϕ
=
α
t
.”
The second LSTM is a neural network that predicts a second learning rate for the optimizee model (‘target machine learning model’) using a second subset of multiple parameters (hidden state, weights, gradients) obtained from the optimizee model.);
choosing, in the processor-based machine learning system and based on the first learning rate and the second learning rate, a learning rate having a minimum loss value in the first learning rate and the second learning rate ((P. 2, Sec. 2.1, ¶1)“training a deep model with n parameters can be formulated as the problem of minimizing a function l … Usually, l is optimized by iteratively adjusting
w
t
(the weight vector at time step t) using gradient information”
(P. 2, Sec. 2.1, ¶2) “an LSTM optimizer g with its own set of parameters φ, is used to minimize the loss of optimizee l”
(P. 4, Sec. 2.3, ¶1-2) “SGD or its variants, whose learning rates are controlled by the LSTM optimizer, is used to train the optimizee till convergence. … we only allow the LSTM optimizer to propose learning rates every S time-steps”
To train the optimizee model, the loss function l for the model is minimized. LSTM optimizer g is used to predict (propose) learning rates each S time-steps to minimize the loss function l. Therefore, when the LSTM optimizer has its two LSTMs with different hidden states predict learning rates, the two learning rates (first and second learning rates) are compared to determine (choose) which learning rate minimizes the loss function l.);
adjusting, in the processor-based machine learning system, the target machine learning model based on the learning rate having the minimum loss value ((P. 2, Sec. 2.1, ¶1) “training a deep model with n parameters can be formulated as the problem of minimizing a function l”
(P. 2, Sec. 2.1, ¶2) “an LSTM optimizer g with its own set of parameters φ, is used to minimize the loss of optimizee l”
The LSTM optimizer predicts learning rates every S time-steps to train the optimizee model (‘target machine learning model’) until convergence (see P. 4, Sec. 2.3, ¶1-2). The optimizee model is trained until loss function l is minimized. The LSTM optimizer predicts learning rates that minimizes the loss function l, therefore the optimizee model is trained (adjusted) based on learning rates having a minimum loss value.);
and deploying, in the processor-based machine learning system, the adjusted target machine learning model for execution of one or more of the tasks related to the at least one of the images, videos, voice, and text ((P. 4, Sec. 3.1, ¶1) “the learning rate schedules proposed by the frozen LSTM on MNIST. We can see that our LSTM optimizer (NOH) improves the optimizee performance significantly in terms of convergence rate and final accuracy after training for 5 meta-iterations”
An optimizee model (‘adjusted target machine learning model’) is trained on the MNIST dataset (image dataset) and has its performance evaluated, therefore the optimizee model executes tasks related to images. A computer with a processor is implied by using an LSTM optimizer to predict learning rates, therefore using an LSTM optimizer to predict learning rates is a ‘processor-based machine learning system’.).
However, Fu does not teach using a reinforcement learning model to predict a first learning rate, which is taught by Priyanshu:
the first machine learning model comprising a reinforcement learning model ((P. 1, Abstract) “there is an essential need for algorithms that, within a given search space, can find near-optimal hyperparameters for the best achievable privacy-utility tradeoffs efficiently. We formulate this problem into a general optimization framework for establishing a desirable privacy-utility tradeoff, and systematically study three cost-effective algorithms for being used in the proposed framework: evolutionary, Bayesian, and reinforcement learning.”),
the reinforcement learning model being configured to receive as inputs at least the first subset of the multiple parameters extracted from the target machine learning model, and to generate as an output the predicted first learning rate ((P. 3, Sec. 3.4) “this classical problem can also be dealt with by reinforcement learning. … We start by sampling a random collection of hyperparameters used to train the DPSGD model and obtain the reward to fit the regression network. We then proceed to extract the estimated reward of the entire search space for our hyperparameter tuning. The best-performing hyperparameters are obtained from this estimation.”
(P. 1, Sec. 1, Last Paragraph) “we specifically focus on two important hyperparameters: noise multiplier
σ
(i.e., the standard deviation of the Gaussian noise) and learning rate
η
. We optimize for these two parameters specifically”
Reinforcement learning uses hyperparameters used to train a DPSGD model (‘target machine learning model’) to generate a reward that contains the best-performing learning rate for the DPSGD model, therefore the hyperparameters of the DPSGD model are a ‘first subset of multiple parameters extracted from a target machine learning model’ used to estimate (output) a learning rate for the DPSGD model.);
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the method of Fu with the reinforcement learning method disclosed by Priyanshu to use reinforcement learning to estimate learning rates. By using reinforcement learning to estimate learning rates, an agent can explore learning rates to avoid over-estimating a learning rate for a target model, thereby dynamically adapting a learning rate and accelerating model convergence.
Regarding Claims 2, 11 and 20, the combined method of Fu / Priyanshu teaches:
The method according to claim 1, wherein choosing, based on the first learning rate and the second learning rate, the learning rate having the minimum loss value in the first learning rate and the second learning rate comprises: comparing the first learning rate and the second learning rate respectively … to choose a learning rate … (Fu discloses (P. 2, Sec. 2.1, ¶1)“training a deep model with n parameters can be formulated as the problem of minimizing a function l … Usually, l is optimized by iteratively adjusting
w
t
(the weight vector at time step t) using gradient information”
Fu discloses (P. 2, Sec. 2.1, ¶2) “an LSTM optimizer g with its own set of parameters φ, is used to minimize the loss of optimizee l”
To train the optimizee model, the loss function l for the model is minimized. LSTM optimizer g is used to predict (propose) learning rates to minimize the loss function l. Therefore, when the LSTM optimizer has its two LSTMs with different hidden states predict learning rates (see P. 3, Sec. 2.2, ¶3), the two learning rates (first and second learning rates) are compared to determine (choose) which learning rate minimizes the loss function l.).
However, Fu does not teach choosing a learning rate having a minimum difference from a target benchmark learning rate, which is taught by Priyanshu:
comparing the first learning rate and the second learning rate respectively with a target benchmark learning rate to choose a learning rate having a minimum difference from the target benchmark learning rate ((P. 1, Sec. 1, Last Paragraph) “we study three cost-effective algorithms: evolutionary, Bayesian, and reinforcement learning and compare the results with the grid search base-line”
(P. 2, Sec. 3.1, ¶1) “Let
H
=
{
h
1
,
…
,
h
N
}
denotes the set of 𝑁 hyperparameters that are used during training on 𝐷𝑡𝑟𝑎𝑖𝑛 and have impact on both validation loss (val_loss) and privacy loss … To provide a general but customizable framework, we define reward as a weighted linear combination of val_loss and
ϵ
”
(P. 1, Sec. 1, Last Paragraph) “we specifically focus on two important hyperparameters: noise multiplier
σ
(i.e., the standard deviation of the Gaussian noise) and learning rate
η
. We optimize for these two parameters specifically”
Priyanshu discloses best-performing hyperparameters are obtained from rewards, “We then proceed to extract the estimated reward of the entire search space for our hyperparameter tuning. The best-performing hyperparameters are obtained from this estimation” (P. 3, Sec. 3.4).
Priyanshu discloses “Although finding the best reward is our goal, we also evaluate the computational time required by each algorithm to achieve the maximum reward attained by Grid Search. The time consumed is calculated based on the time taken for an optimization algorithm to achieve a reward equal to or greater than the baseline reward. Here, baseline reward refers to the highest reward achieved by the Grid Search algorithm” (P. 3, Sec. 4.3, ¶1).
Optimization algorithms (evolutionary, Bayesian, and reinforcement learning) achieve a reward equal or greater to a baseline reward (the highest award achieved by performing Grid Search). The baseline reward is a “target benchmark learning rate” since hyperparameters (learning rate) are extracted from rewards and the baseline reward is used to compare optimization algorithms. Each reward calculated using a respective optimization algorithm contains a learning rate. To find the best or equal rewards, a minimum difference between an optimization algorithm reward and the baseline reward must be calculated (therefore comparing first and second learning rates with a target benchmark learning rate to choose a learning rate having a minimum difference from the target benchmark learning rate).);
and wherein the target benchmark learning rate has been determined by using a grid search method (A baseline reward is the highest reward achieved by the Grid Search algorithm, see (P. 3, Sec. 4.3, ¶1).
(P. 1, Abstract) “Our experiments, for hyperparameter tuning in DPSGD conducted on MNIST and CIFAR-10 datasets, show that these three algorithms significantly outperform the widely used grid search baseline”
(P. 1, Sec. 2, ¶1) “The most widely used methods for hyperparameter tuning in deep learning are manual search, random search, and grid search [16]. … grid search is utilized to provide a sufficient exploration within a restricted search space”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the technique disclosed by Priyanshu to compare learning rates against a baseline learning rate obtained through grid search. By comparing learning rates against a baseline learning rate obtained through grid search, learning rates achieved by hyperparameter optimization algorithms can be compared against widely used grid search, thereby providing a clear comparison to determine if optimization algorithms can approach or outperform an established baseline.
Regarding Claims 3 and 12, the combined method of Fu / Priyanshu teaches:
the method according to claim 1, further comprising: predicting, by the first machine learning model, the first learning rate associated with the target machine learning model based on the state information and the loss value of the target machine learning model in the multiple parameters (Fu discloses (P. 3, Sec. 2.2, ¶2) “To solve the above problems, our proposed one is trained to predict the learning rates. … we adopt the following rule:
w
t
+
1
=
w
t
-
g
t
h
∇
l
w
t
;
ϕ
⋅
∇
l
w
t
, where h(·), the input to the LSTM optimizer, is defined as the state description vector of the optimizee gradients at iteration t, and
g
t
h
∇
l
w
t
;
ϕ
=
α
t
.”
The LSTM optimizer
g
t
h
∇
l
w
t
;
ϕ
=
α
t
is used to predict learning rates. To compute a gradient
∇
l
using weights
w
t
, loss values must be calculated for the optimizee model, therefore the optimizer uses loss values of the optimizee model to predict a learning rate. The LSTM optimizer is comprised of two LSTMs with different hidden states
h
k
⋅
, see (P. 3, Sec. 2.2, ¶3). Since the LSTM optimizer
g
t
·
requires a hidden state h(·) as input to predict a learning rate, each LSTM with a different hidden state
h
k
⋅
therefore predicts its own learning rate (and therefore a first LSTM (‘first machine learning model’) predicts a first learning rate for the optimizee model (‘target machine learning model’) using a hidden state (‘state information’) and loss values of the optimizee model.).
Regarding Claims 4 and 13, the combined method of Fu / Priyanshu teaches:
The method according to claim 1, further comprising: predicting, by the second machine learning model, the second learning rate associated with the first machine learning model based on the weight and the loss value of the target machine learning model in the multiple parameters (Fu discloses (P. 3, Sec. 2.2, ¶2) “To solve the above problems, our proposed one is trained to predict the learning rates. … we adopt the following rule:
w
t
+
1
=
w
t
-
g
t
h
∇
l
w
t
;
ϕ
⋅
∇
l
w
t
, where h(·), the input to the LSTM optimizer, is defined as the state description vector of the optimizee gradients at iteration t, and
g
t
h
∇
l
w
t
;
ϕ
=
α
t
.”
The LSTM optimizer
g
t
h
∇
l
w
t
;
ϕ
=
α
t
is used to predict learning rates. To compute a gradient
∇
l
using weights
w
t
, loss values must be calculated for the optimizee model, therefore the optimizer uses loss values and weights of the optimizee model to predict a learning rate. The LSTM optimizer is comprised of two LSTMs with different hidden states
h
k
⋅
, see (P. 3, Sec. 2.2, ¶3). Since the LSTM optimizer
g
t
·
requires a hidden state h(·) as input to predict a learning rate, each LSTM with a different hidden state
h
k
⋅
therefore predicts its own learning rate (and therefore a second LSTM (‘second machine learning model’) predicts a second learning rate for the optimizee model (‘target machine learning model’) using weights and loss values of the optimizee model.
The second learning rate is associated with a first LSTM (‘first machine learning model’) since a learning rate predicted by the first LSTM and the second learning rate are used to minimize a loss function l for the optimizee model (see P. 2, Sec. 2.1, ¶1-2).).
Regarding Claims 6 and 15, the combined method of Fu / Priyanshu teaches:
The method according to claim 1, further comprising: adjusting the first machine learning model based on that the second learning rate is determined as the learning rate having the minimum loss value, wherein the adjustment comprises: reducing a difference between the first learning rate generated by the first machine learning model and a target benchmark learning rate by adjusting the first machine learning model by means of one or more methods of a Q learning method and/or a strategy gradient method (Priyanshu discloses “Our aim for the following experiments remains to optimize the reward given by Equation (1)” (P. 2, Sec. 3.1, Last Paragraph).
Priyanshu discloses “Although finding the best reward is our goal, we also evaluate the computational time required by each algorithm to achieve the maximum reward attained by Grid Search. The time consumed is calculated based on the time taken for an optimization algorithm to achieve a reward equal to or greater than the baseline reward. Here, baseline reward refers to the highest reward achieved by the Grid Search algorithm” (P. 3, Sec. 4.3, ¶1).
Priyanshu discloses best-performing hyperparameters are obtained from rewards, see (P. 3, Sec. 3.4). Priyanshu discloses Figure 2 on P. 4 depicting the time taken by reinforcement learning to achieve a reward greater than or equal to a baseline reward.
Priyanshu discloses “this classical problem can also be dealt with by reinforcement learning. In our application of this method, we begin by initializing a regression network capable of estimating the reward output of training on a particular set of hyperparameters” (P. 3, Sec. 3.4).
A reinforcement learning model that is capable of estimating a reward of training on a particular set of hyperparameters must use a policy to generate those rewards. The policy used by reinforcement learning to generate rewards must be updated to optimize rewards, therefore a gradient that updates a policy (a strategy gradient) is implied.
The baseline reward is a “target benchmark learning rate” since hyperparameters (learning rate) are extracted from rewards and the baseline reward is used to compare optimization algorithms. The baseline reward is a ‘second learning rate’ since the baseline reward is a learning rate obtained by Grid Search that minimizes validation loss (see P. 2, Sec. 3.1, ¶1).
The reward calculated by reinforcement learning contains a learning rate. Over time, a reward is optimized by reinforcement learning to achieve a reward greater than or equal to a baseline reward, therefore, a reward (learning rate) is optimized over time to be closer to the baseline reward (and therefore reducing a difference between a learning rate and a target benchmark learning rate by adjusting the first machine learning model by means of a strategy gradient method).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the technique disclosed by Priyanshu to optimize a learning rate obtained from reinforcement learning to be close to a baseline learning rate. By optimizing a learning rate obtained from reinforcement learning to be close to a baseline learning rate, it can be ensured that a learning rate obtained by reinforcement learning is close to a reliable baseline, thereby ensuring that a model using the learning rate achieves consistent and stable performance.
Regarding Claims 7 and 16, the combined method of Fu / Priyanshu teaches:
The method according to claim 6, further comprising: in response to reduction of the difference between the first learning rate generated by the first machine learning model and the target benchmark learning rate, obtaining a reward value associated with the first learning rate (Priyanshu discloses best-performing hyperparameters are obtained from rewards, see (P. 3, Sec. 3.4).
The reward calculated by reinforcement learning contains a learning rate. Over time, a reward is optimized by reinforcement learning to achieve a reward greater than or equal to a baseline reward (see Figure 2 on P. 4), therefore, a reward (learning rate) is optimized over time to be closer to the baseline reward (and therefore reducing the difference between the first learning rate generated by the first machine learning model and the target benchmark learning rate). );
and adjusting the first machine learning model based on the reward value (A reinforcement learning model that can estimate a reward of training on a particular set of hyperparameters (see Priyanshu, P. 3, Sec. 3.4), must use a policy to generate those rewards. The policy used by reinforcement learning to generate rewards must be updated to optimize rewards, therefore the rewards generated by reinforcement learning are used to adjust the first machine learning model based on the reward value.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the technique disclosed by Priyanshu to optimize a learning rate using reinforcement learning. By optimizing a learning rate using reinforcement learning, reinforcement learning can use rewards to receive feedback from its environment, thereby maximizing a long-term reward and resulting in an optimal learning rate that can be used to accelerate model convergence.
Regarding Claim 10, the rejection of claim 1 is incorporated. The difference in scope being:
An electronic device, comprising (Fu discloses (P. 1, Abstract) “we train a neural optimizer”
A computer (‘electronic device’) is implied by training a neural optimizer.):
at least one processor (A processor is implied by using a computer.);
and memory, the memory being coupled to the at least one processor and storing instructions, wherein the instructions, when executed by the at least one processor, cause the electronic device to perform actions comprising (A computer that trains a neural optimizer implies a memory (coupled to a processor) that stores executable instructions.).
Regarding Claim 19, the rejection of claim 1 is incorporated. The difference in scope being:
A computer program product comprising a non-transitory computer-readable storage medium having machine-executable instructions stored therein, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising (Fu discloses (P. 1, Abstract) “we train a neural optimizer”
Training a neural optimizer implies the use of a computer, which further implies a computer program product comprising a non-transitory computer-readable storage medium having stored machine-executable instructions.).
The following are the references relied upon in the rejections below:
Triplet (US 20210174246 A1)
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Fu / Priyanshu / Triplet.
Regarding Claims 5 and 14, the combined method of Fu / Priyanshu teaches:
the method according to claim 1, further comprising: adjusting the second machine learning model based on that the first learning rate is determined as the learning rate having the minimum loss value, wherein the adjustment comprises: reducing a difference between the second learning rate generated by the second machine learning model and a target benchmark learning rate (Priyanshu discloses rewards are optimized, “start searching for the optimal hyperparameters in
H
… we consider
H
=
{
σ
,
η
}
,
where
σ
denotes the noise multiplier and
η
denotes the learning rate; in DPSGD. Our aim for the following experiments remains to optimize the reward given by Equation (1)” (P. 2, Sec. 3.1, Last Paragraph).
Priyanshu discloses “Although finding the best reward is our goal, we also evaluate the computational time required by each algorithm to achieve the maximum reward attained by Grid Search. The time consumed is calculated based on the time taken for an optimization algorithm to achieve a reward equal to or greater than the baseline reward. Here, baseline reward refers to the highest reward achieved by the Grid Search algorithm” (P. 3, Sec. 4.3, ¶1).
Priyanshu discloses best-performing hyperparameters are obtained from rewards, see (P. 3, Sec. 3.4).
Priyanshu discloses Figure 2 on P. 4 (reproduced below) depicting the time taken by each optimization algorithm to achieve a reward greater than or equal to a baseline reward.
PNG
media_image2.png
806
996
media_image2.png
Greyscale
The baseline reward is a “target benchmark learning rate” since hyperparameters (learning rate) are extracted from rewards and the baseline reward is used to compare optimization algorithms. Each reward calculated using a respective optimization algorithm contains a learning rate. Over time, a reward is optimized by each optimization algorithm to achieve a reward greater than or equal to a baseline reward, therefore, a reward (learning rate) is optimized over time to be closer to the baseline reward (and therefore reducing a difference between a learning rate and a target benchmark learning rate).)
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the technique disclosed by Priyanshu to optimize a learning rate to be close to a baseline learning rate. By optimizing a learning rate to be close to a baseline learning rate, it can be ensured that a learning rate obtained by an optimization algorithm is close to a reliable baseline, thereby ensuring that a model using the learning rate achieves consistent and stable performance.
However, the combination does not teach adjusting hyperparameters by adjusting sample data and quantity of the sample data, which is taught by Triplet:
[adjusting hyperparameters] generated by the second machine learning model … by adjusting sample data of the second machine learning model for predicting the [hyperparameters] and a quantity of the sample data (Triplet discloses “use Reinforcement Learning (RL) to tune hyperparameters of one or more ML techniques and to cause the processing device to train a ML model using the one or more ML techniques in which the respective hyperparameters were tuned in the RL” [0015].
Triplet discloses “the action of “tuning” hyperparameters may include adjusting the hyperparameters so as to strengthen, augment, or enhance the hyperparameters. A goal for example is to tune the hyperparameters so to as approach optimized values or to improve upon previous values by using the reward function. By learning to tune or strengthen these hyperparameters, the systems and methods of the present disclosure are able to significantly reduce the training time compared with conventional systems, minimize amount of data for training” [0029].
Triplet discloses “The “rewards” of the RL-based system 70 may rely on: a) maximizing the accuracy, precision, and/or recall; b) minimizing the amount of data required for training” [0064].
Triplet discloses “The RL-base system 70 also improves sample efficiency and reduces amount of data required, thereby accelerating deployments of ML models in production” [0065].
A reinforcement learning model is used to find optimal hyperparameters for a ML technique that is used to train a model. The reinforcement learning model uses a reward function to tune hyperparameters. The rewards of the reinforcement learning rely on minimizing the amount of data required for training, therefore hyperparameters are tuned (adjusted) by adjusting a quantity of sample data (and therefore adjusting sample data).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the technique disclosed by Triplet to tune hyperparameters based on minimizing the amount of data required for training. By tuning hyperparameters based on minimizing the amount of data required for training, deployments of machine learning models can be accelerated and less computational resources are used to process training data.
The following are the references relied upon in the rejections below:
Wen, Long, et al. "Convolutional neural network with automatic learning rate scheduler for fault classification." IEEE Transactions on Instrumentation and Measurement 70 (2021): 1-12.
Claims 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Fu / Priyanshu / Wen.
Regarding Claims 8 and 17, the combined method of Fu / Priyanshu teaches: The method according to claim 1, however the combination does not teach adjusting a first or second machine learning model in each iteration of multiple iteration adjustments of a target machine learning model, which is taught by Wen:
further comprising: adjusting one of the first machine learning model or the second machine learning model in each iteration of multiple iteration adjustments of the target machine learning model ((P. 3, Sec. II-B, Last Paragraph) “an AutoLR is designed for the CNN-based fault classification. AutoLR uses the historical CNN training information to construct an RL agent to adjust the learning rate of CNN model at each step, which can make the full use of the training information to guide the control on learning rate automatically”
(P. 5, Sec. IV-C, ¶1) “The main working flow of the AutoLR-CNN is as follows. At each step t, first, generate the state st of the MainCNN or GameCNN network. Then, the actor network in DDPG predicts its action … Third, the agent will update the learning rate according to the action using (8) and (9), and then conduct one mini-batch training step. Finally, the agent will receive its reward rt+1, and arrive at the next time step t +1.”
(P. 5, Sec. IV-C, Last Paragraph) “AutoLR-CNN applies the LSTM network to learn the feature of the past M loss on mini-batch training and then trains the agent to control the learning rate of CNN”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the technique disclosed by Wen to update how a machine learning model predicts a learning rate in each iteration. By updating how a machine learning model predicts a learning rate in each iteration, the machine learning model can be re-trained on new training data to update how it predicts learning rates for a target model, thereby ensuring that the most optimal learning rates are predicted at each iteration for the target model.
The following are the references relied upon in the rejections below:
Xu (CN 113625336 A)
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Fu / Priyanshu / Xu.
Regarding Claims 9 and 18, the combined method of Fu / Priyanshu teaches:
The method according to claim 1, further comprising: predicting, respectively by the first machine learning model and the second machine learning model based on the multiple parameters extracted from the target machine learning model, … [learning rates] that have minimum loss values and are for the target machine learning model (Fu discloses (P. 2, Sec. 2.1, ¶1)“training a deep model with n parameters can be formulated as the problem of minimizing a function l”
Fu discloses (P. 2, Sec. 2.1, ¶2) “an LSTM optimizer g with its own set of parameters φ, is used to minimize the loss of optimizee l”
Fu discloses (P. 3, Sec. 2.2, ¶2) “we adopt the following rule:
w
t
+
1
=
w
t
-
g
t
h
∇
l
w
t
;
ϕ
⋅
∇
l
w
t
, where h(·), the input to the LSTM optimizer, is defined as the state description vector of the optimizee gradients at iteration t, and
g
t
h
∇
l
w
t
;
ϕ
=
α
t
.”
To train the optimizee model (‘target machine learning model’), the loss function l for the model is minimized. The LSTM optimizer is used to predict learning rates to minimize the loss function l. Therefore, when the LSTM optimizer has its two LSTMs with different hidden states (‘multiple parameters extracted from the target machine learning model’) predict learning rates (see P. 3, Sec. 2.2, ¶3), the two learning rates (first and second learning rates) are compared to determine (choose) which learning rate minimizes the loss function l.).
However, the combination does not teach predicting a momentum, weight attenuation, a loss rate, and a batch size that have minimum loss values, which is taught by Xu:
predicting … based on the multiple parameters extracted from the target machine learning model, a momentum, weight attenuation, a loss rate, and a batch size that have minimum loss values and are for the target machine learning model (Xu discloses “Further, the hyper-parameter comprises training times, batch sample number, learning rate, discarding rate, weight attenuation coefficient, optimizer momentum, convolution kernel size. … the weight attenuation coefficient adjusting range is [0, 1e-4]” (P. 3, ¶4-5).
Xu discloses batch size, “The number of batches of samples is too small to reduce the effective capacity, and the number of samples is selected according to the capacity of the hardware” (P. 4, Last Paragraph).
Xu discloses “loss rate representation, less discarding parameter means the lifting of the model parameter number, the adaptability of the parameter is improved, model capacity is improved, but the model effective tolerance is not necessarily increased, the adjusting range is generally [0.1, 0.5]. the weight attenuation coefficient can effectively limit the amplitude of the parameter change; it has a certain regular action; the adjusting range is generally [0, 1e-4]. optimizer momentum is used for accelerating training, avoiding local optimal solution” (P. 5, ¶1).
Xu discloses “each established neural network will have the best hyper-parameter with it, such as convolution kernel size, learning rate, batch size, loss function part hyper-parameter discard rate and so on each of hyper-parameter value of the setting. The set of optimal combinations of the best hyper-parameter generally capable of obtaining relatively small loss values and the inversion model obtained by training will be better. Of course, for each neural network, there is no direct method of the best hyper-parameter the determined combination, which is obtained by repeatedly testing … first, pre-processing the data; secondly, through training, verifying the obtained loss value change curve … so as to repeatedly adjust the parameter training of the network” (P. 7, ¶3).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Fu / Priyanshu with the hyperparameter optimization technique disclosed by Xu to optimize a momentum, weight attenuation, a loss rate, and a batch size for a machine learning model. By optimizing a momentum, weight attenuation, a loss rate, and a batch size for a machine learning model, the settings for these hyperparameters can be adjusted during parameter training, thereby resulting in optimal set of hyperparameters that minimizes loss and improves model generalization.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Xu et al. (US 20210034973 A1) teaches using a control network (trained by reinforcement learning) to predict learning rates to train a trainee machine learning model.
Wu et al. (“Selecting and Composing Learning Rate Policies for Deep Neural Networks”) teaches iteratively training a DNN by selecting an optimal learning rate from a set of learning rates predicted by different learning rate policies.
Meier et al. (“Online Learning of a Memory for Learning Rates”) teaches having local models predict a learning rate and then averaging each of the predicted learning rates to generate a single learning rate to update a target machine learning model.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PEDRO J MORALES/Examiner, Art Unit 2124
/Kevin W Figueroa/Primary Examiner, Art Unit 2124