DETAILED ACTION
This action is responsive to the application filed on 04/18/2024. Claims 2-21 are pending in the case. Claims 2, 13, and 14 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement filed 04/18/2024 fails to comply with 37 CFR 1.98(a)(2), which requires a legible copy of each cited foreign patent document; each non-patent literature publication or that portion which caused it to be listed; and all other information or that portion which caused it to be listed. It has been placed in the application file, but the information referred to therein has not been considered.
The patent document numbered 2016/0328644 has been stricken through and not considered because the date provided in the citation does not match the date of the document. All the foreign patent documents and non-patent literature documents have been stricken through and not considered because no copy has been provided. All other references are being considered by the examiner.
Claim Objections
Claims 2, 8, and 20-21 are objected to because of the following informalities:
Claim 2, line 4, two instances of “the one or more computers” should read “the one or more first computers” to maintain consistency among claim limitations.
Claim 8, line 5, “training iteration;” should read “training iteration.” as it appears to be a typographical error.
Claim 20, line 6, “input setting,” should read “input setting.” as it appears to be a typographical error.
Claim 21, line 5, “network output” should read “network output.” as it appears to be a typographical error.
Appropriate correction is required.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 2-21 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 2-20 of U.S. Patent No. 12008445. Although the claims at issue are not identical, they are not patentably distinct from each other because the instant claims are obvious over the corresponding claims as follows:
Instant Application
Patent No. 12008445
Claim 2
Claim 2
A system for determining an optimized setting for one or more process
parameters, the system comprising:
A system for determining an optimized setting for one or more process parameters, the system comprising:
one or more first computers and one or more first storage devices storing instructions
that when executed by the one or more computers cause the one or more computers to implement:
one or more first computers and one or more first storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement:
a first neural network having a plurality of network parameters and being configured to:
a first neural network that corresponds to a plurality of worker computing units, the first neural network having a plurality of network parameters and being configured to:
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting,
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters, the one or more process parameters comprising one or more architecture settings of a machine learning model being trained using a machine learning training process, and (ii) a measure of a performance of the machine learning training process with the input setting,
process the sequence of network inputs to generate a respective network output
for each network input that defines an updated setting for the one or more process parameters, and
process the sequence of network inputs in accordance with the network parameters to generate a respective network output for each network input that defines an updated setting for the one or more process parameters, and
provide the updated setting to one of a plurality of worker computing units to
evaluate performance of the machine learning training process in training the machine learning model with the updated setting;
the plurality of worker computing units;
provide the updated setting to one of the plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting; wherein the first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration;
a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices, the subsystem being configured to:
a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices, the subsystem being configured to:
determine, for each of a plurality of candidate settings for the one or more process
parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising:
determine, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising:
processing, using the first neural network in accordance with first values
of the network parameters, a current network input to obtain a current network output that
defines an updated setting of the one or more process parameters, and
processing, using the first recurrent neural network in accordance with first values of the network parameters, the current network input to obtain a current network output that defines an updated setting of the one or more process parameters,
obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first
neural network; and
obtaining a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first recurrent neural network,
and generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output;
select a candidate setting from the plurality of candidate settings as an optimized
setting for the one or more process parameters using the measures of the performance for the
candidate settings.
select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings
and the plurality of worker computing units, wherein each worker computing unit is configured to: receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process, wherein the one or more process parameters are not learned as part of the machine learning training process; execute the machine learning training process with one or more process parameters specified by the input setting; and measure the performance of the machine learning training process with the one or more process parameters specified by the input setting, wherein obtaining a measure of the performance of the machine learning training process with the updated setting defined by the current network output comprises: providing the updated setting defined by the current network output generated by the first neural network to one of the plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process.
Page 16, Lines 17-23 and Page 16, Line 31 – Page 17, Line 2 of Patent No. 12008445 make it clear the subsystem comprises one or more second computers and one or more second storage devices.
Instant Application
Patent No. 12008445
Claim 3
Claim 3
wherein each worker computing unit operates
asynchronously from each other worker computing unit
wherein each worker computing unit operates
asynchronously from each other worker computing unit
Instant Application
Patent No. 12008445
Claim 4
Claim 4
wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises:
generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance;
processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input; and
providing the updated settings defined by the plurality of initial network outputs to
respective worker computing units in the plurality of worker computing units.
wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises:
generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance;
processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input; and
providing the updated settings defined by the initial network outputs to
respective worker computing units in the plurality of worker computing units.
Instant Application
Patent No. 12008445
Claim 5
Claim 5
wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
Instant Application
Patent No. 12008445
Claim 6
Claim 6
wherein the first neural network is a differentiable neural computer (DNC).
wherein the first neural network is a differentiable neural computer (DNC).
Instant Application
Patent No. 12008445
Claim 7
Claim 7
wherein the first neural network is a long short-term memory (LSTM) neural network.
wherein the first neural network is a long short-term memory (LSTM) neural network.
Instant Application
Patent No. 12008445
Claim 8
Claim 1
wherein first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration;
…wherein the first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration;
Instant Application
Patent No. 12008445
Claim 9
Claim 8
wherein the training function for each iteration is sampled from a training distribution.
wherein the training function for each iteration is sampled from a training distribution.
Instant Application
Patent No. 12008445
Claim 10
Claims 9 and 10
wherein the loss function is a summed loss function or an expected posterior improvement loss function.
Claim 9: wherein the loss function is a summed loss function.
Claim 10: wherein the loss function is an expected posterior improvement loss function.
Instant Application
Patent No. 12008445
Claim 11
Claim 11
wherein the loss function is an observed improvement loss function
wherein the loss function is an observed improvement loss function
Instant Application
Patent No. 12008445
Claim 12`
Claim 12
wherein the subsystem is further configured to:
train the machine learning model using the machine learning training process with the
optimized setting for the process parameters; and
output the trained machine learning model.
wherein the subsystem is further configured to:
train the machine learning model using the machine learning training process with the
optimized setting for the process parameters; and
output the trained machine learning model.
Instant Application
Patent No. 12008445
Claim 13
Claim 13
One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement a system for determining an optimized setting for one or more process parameters, the system comprising:
One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement a system for determining an optimized setting for one or more process parameters, the system comprising:
a first neural network, the first neural network having a plurality of network parameters and being configured to:
a first neural network that corresponds to a plurality of worker computing units, the first neural network having a plurality of network parameters and being configured to:
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting,
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters, the one or more process parameters comprising one or more architecture settings of a machine learning model being trained using a machine learning training process, and (ii) a measure of a performance of a machine learning training process with the input setting, and
process the sequence of network inputs to generate a respective network output
for each network input that defines an updated setting for the one or more process parameters,
process the sequence of network inputs in accordance with the network parameters to generate a respective network output for each network input that defines an updated setting for the one or more process parameters, and
provide the updated setting to one of a plurality of worker computing units to
evaluate performance of the machine learning training process in training the machine learning model with the updated setting;
the plurality of worker computing units;
provide the updated setting to one of the plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting; wherein the first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration;
a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices, the subsystem being configured to:
a subsystem configured to:
determine, for each of a plurality of candidate settings for the one or more process
parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising:
determine, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing the following:
processing, using the first neural network in accordance with first values
of the network parameters, a current network input to obtain a current network output that
defines an updated setting of the one or more process parameters, and
processing, using the first recurrent neural network in accordance with first values of the network parameters, the current network input to obtain a current network output that defines an updated setting of the one or more process parameters,
obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first
neural network; and
obtaining a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first recurrent neural network,
and generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output;
select a candidate setting from the plurality of candidate settings as an optimized
setting for the one or more process parameters using the measures of the performance for the
candidate settings.
select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings
and the plurality of worker computing units, wherein each worker computing unit is configured to: receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process, wherein the one or more process parameters are not learned as part of the machine learning training process; execute the machine learning training process with the one or more process parameters specified by the input setting; and measure the performance of the machine learning training process with the one or more process parameters specified by the input setting, wherein obtaining a measure of the performance of the machine learning training process with the updated setting defined by the current network output comprises: providing the updated setting defined by the current network output generated by the first neural network to one of the plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process.
Instant Application
Patent No. 12008445
Claim 14
Claim 14
A method of determining an optimized setting for one or more process parameters of a machine learning training process, the method comprising:
A system for determining an optimized setting for one or more process parameters, the system comprising:
determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of a machine learning training process for training a machine learning model with the candidate setting,
determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the one or more process parameters comprising one or more architecture settings of a machine learning model being trained using a machine learning training process,
wherein the determining comprises, for each time step of a plurality of time steps, performing the following:
wherein the determining comprises, for each time step of a plurality of time steps, performing the following:
providing, as input to a first neural network, a current network input comprising (i) a current setting of the one or more process parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting;
providing, as input to a first neural network that corresponds to a plurality of worker computing units, a current network input comprising (i) a current setting of the one or more process parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting;
processing, using the first neural network, the current network input to obtain a current network output that defines an updated setting of the one or more process parameters,
processing, using the first neural network in accordance with first values of the network parameters, the current network input to obtain a current network output that defines an updated setting of the one or more process parameters, and
providing the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting, and
provide the updated setting to one of the plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting,
obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first neural network, comprising providing the updated setting defined by the current network output generated by the first neural network to one of a plurality of worker computing units and,
obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first neural network, comprising providing the updated setting defined by the current network output generated by the first neural network to one of a plurality of worker computing units and,
obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process;
obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process,
wherein the first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration,
wherein each worker computing unit is configured to: receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process,
wherein the one or more process parameters are not learned as part of the machine learning training process;
execute the machine learning training process with one or more hyper parameters specified by the input setting; and
measure the performance of the machine learning training process with the one or more process parameters specified by the input setting, generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output;
and selecting a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings.
and selecting a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings.
Instant Application
Patent No. 12008445
Claim 15
Claim 15
The method of claim 14, wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises: generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance; processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input; and providing the updated settings defined by the initial network outputs to respective worker computing units in the plurality of worker computing units.
The method of claim 14, wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises: generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance; processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input; and providing the updated settings defined by the initial network outputs to respective worker computing units in the plurality of worker computing units.
Instant Application
Patent No. 12008445
Claim 16
Claim 16
The method of claim 15, wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
The method of claim 15, wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
Instant Application
Patent No. 12008445
Claim 17
Claims 17, 18, and 19
wherein first neural network has been trained on a loss function, wherein the loss function is a summed loss function, an expected posterior improvement loss function, or an observed improvement loss function.
Claim 17: wherein the loss function is a summed loss function.
Claim 18: wherein the loss function is an expected posterior improvement loss function
Claim 19: wherein the loss function is an observed improvement loss function
Instant Application
Patent No. 12008445
Claim 18
Claim 20
training the machine learning model using the machine learning training process with the optimized setting for the process parameters; and outputting the trained machine learning model
training the machine learning model using the machine learning training process with the optimized setting for the process parameters; and outputting the trained machine learning model
Instant Application
Patent No. 12008445
Claim 19
Claim 14
where in the first neural network corresponds to the plurality of worker computing units
…the first neural network that corresponds to the plurality of worker computing units…
Instant Application
Patent No. 12008445
Claim 20
Claim 14
wherein each worker computing unit is configured to: receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process; and measure the performance of the machine learning training process with the one or more process parameters specified by the input setting,
…wherein each worker computing unit is configured to: receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process…; and measure the performance of the machine learning training process with the one or more process parameters specified by the input setting,
Instant Application
Patent No. 12008445
Claim 21
Claim 14
generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output
…generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output…
Claims 2-21 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 5-9, 11-17, 21-22, 24-27 of U.S. Patent No. 11354594. Although the claims at issue are not identical, they are not patentably distinct from each other because the instant claims are obvious over the corresponding claims as follows:
Instant Application
Patent No. 11354594
Claim 2
Claim 1
A system for determining an optimized setting for one or more process
parameters, the system comprising:
A system for determining an optimized setting for one or more hyper-parameters of a machine learning process of a machine learning model, the system comprising:
one or more first computers and one or more first storage devices storing instructions
that when executed by the one or more computers cause the one or more computers to implement:
one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement:
a first neural network having a plurality of network parameters and being configured to:
a first recurrent neural network that corresponds to a plurality of worker computing units, the first recurrent neural network having a plurality of network parameters and being configured to:
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting,
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more hyper-parameters and (ii) a measure of a performance of the machine learning training process with the input setting,
process the sequence of network inputs to generate a respective network output
for each network input that defines an updated setting for the one or more process parameters, and
process the sequence of network inputs in accordance with the network parameters to generate a respective network output for each network input that defines an updated setting for the one or more hyper-parameters, and
provide the updated setting to one of a plurality of worker computing units to
evaluate performance of the machine learning training process in training the machine learning model with the updated setting;
the plurality of worker computing units;
Provide the updated setting to one of the plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting;
a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices, the subsystem being configured to:
a subsystem configured to:
determine, for each of a plurality of candidate settings for the one or more process
parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising:
determine, for each of a plurality of candidate settings for the one or more hyper-parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising:
providing, as input to the first recurrent neural network, a current network input comprising (i) a current setting of the one or more hyper-parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting;
processing, using the first neural network in accordance with first values
of the network parameters, a current network input to obtain a current network output that
defines an updated setting of the one or more process parameters, and
processing, using the first recurrent neural network in accordance with first values of the network parameters, the current network input to obtain a current network output that defines an updated setting of the one or more hyper-parameters
obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first
neural network; and
obtaining a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first recurrent neural network, and
generating a new network input to be provided as input to the first recurrent neural network at the next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output; and
select a candidate setting from the plurality of candidate settings as an optimized
setting for the one or more process parameters using the measures of the performance for the
candidate settings.
select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more hyper-parameters using the measures of the performance for the candidate settings; and
the plurality of worker computing units, wherein each worker computing unit is configured to:
receive, from the first recurrent neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more hyper- parameters of the machine learning training process, wherein the one or more hyper-parameters are not learned as part of the machine learning training process;
execute the machine learning training process with the one or more hyper parameters specified by the input setting; and
measure the performance of the machine learning training process with the one or more hyper-parameters specified by the input setting, wherein obtaining a measure of the performance of the machine learning training process with the updated setting defined by the current network output comprises:
providing the updated setting defined by the current network output generated by the first recurrent neural network to one of the plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process, and
wherein the first values of the plurality of network parameters are obtained by training the first recurrent neural network to, at each iteration of the training, optimize a differentiable training function by minimizing a loss associated with the differentiable training function
Page 16, Lines 11-16 and Lines 25-30 of Patent No. 11354594 make it clear the subsystem comprises one or more second computers and one or more second storage devices.
Instant Application
Patent No. 11354594
Claim 3
Claim 5
wherein each worker computing unit operates
asynchronously from each other worker computing unit.
wherein each worker computing unit operates asynchronously from each other worker computing unit
Instant Application
Patent No. 11354594
Claim 4
Claim 6
wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises:
generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance;
processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input; and
providing the updated settings defined by the plurality of initial network outputs to
respective worker computing units in the plurality of worker computing units.
wherein determining, for each of a plurality of candidate settings for the one or more hyper-parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises:
generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance;
processing each of the plurality of initial network inputs using the first recurrent neural network to generate a respective initial network output for each initial network input; and
providing the updated settings defined by the initial network outputs to
respective worker computing units in the plurality of worker computing units.
Instant Application
Patent No. 11354594
Claim 5
Claim 7
wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
Instant Application
Patent No. 11354594
Claim 6
Claim 8
wherein the first neural network is a differentiable neural computer (DNC).
wherein the first recurrent neural network is a differentiable neural computer (DNC).
Instant Application
Patent No. 11354594
Claim 7
Claim 9
wherein the first neural network is a long short-term memory (LSTM) neural network.
wherein the first recurrent neural network is a long short-term memory (LSTM) neural network.
Instant Application
Patent No. 11354594
Claim 9
Claim 11
wherein the training function for each iteration is sampled from a training distribution.
wherein the differentiable training function for each iteration is sampled from a training distribution.
Instant Application
Patent No. 11354594
Claim 10
Claims 12 and 13
wherein the loss function is a summed loss function or an expected posterior improvement loss function.
Claim 12: wherein the loss is a summed loss function
Claim 13: wherein the loss is an expected posterior improvement loss function
Instant Application
Patent No. 11354594
Claim 11
Claim 14
wherein the loss function is an observed improvement loss function.
wherein the loss is an observed improvement loss function.
Instant Application
Patent No. 11354594
Claim 12
Claim 15
wherein the subsystem is further configured to: train the machine learning model using the machine learning training process with the optimized setting for the process parameters; and output the trained machine learning model.
wherein the subsystem is further configured to: train the machine learning model using the machine learning training process with the optimized setting for the hyper-parameters; and output the trained machine learning model.
Instant Application
Patent No. 11354594
Claim 13
Claim 16
One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement a system for determining an optimized setting for one or more process parameters, the system comprising:
One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement a system for determining an optimized setting for one or more hyper- parameters of a machine learning training process, the system comprising:
a first neural network having a plurality of network parameters and being configured to:
a first recurrent neural network that corresponds to a plurality of worker computing units, the first recurrent neural network having a plurality of network parameters and being configured to:
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting,
receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more hyper-parameters and (ii) a measure of a performance of the machine learning training process with the input setting, and
process the sequence of network inputs to generate a respective network output
for each network input that defines an updated setting for the one or more process parameters, and
process the sequence of network inputs in accordance with the network parameters to generate a respective network output for each network input that defines an updated setting for the one or more hyper-parameters, and
provide the updated setting to one of a plurality of worker computing units to
evaluate performance of the machine learning training process in training the machine learning model with the updated setting;
the plurality of worker computing units;
Provide the updated setting to one of the plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting;
a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices, the subsystem being configured to:
a subsystem configured to:
determine, for each of a plurality of candidate settings for the one or more process
parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising:
determine, for each of a plurality of candidate settings for the one or more hyper-parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing the following:
providing, as input to the first recurrent neural network, a current network input comprising (i) a current setting of the one or more hyper-parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting;
processing, using the first neural network in accordance with first values
of the network parameters, a current network input to obtain a current network output that
defines an updated setting of the one or more process parameters,
processing, using the first recurrent neural network in accordance with first values of the network parameters, the current network input to obtain a current network output that defines an updated setting of the one or more hyper-parameters
obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first
neural network; and
obtaining a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first recurrent neural network, and
generating a new network input to be provided as input to the first recurrent neural network at the next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output; and
select a candidate setting from the plurality of candidate settings as an optimized
setting for the one or more process parameters using the measures of the performance for the
candidate settings.
select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more hyper-parameters using the measures of the performance for the candidate settings; and
the plurality of worker computing units, wherein each worker computing unit is configured to:
receive, from the first recurrent neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more hyper- parameters of the machine learning training process, wherein the one or more hyper-parameters are not learned as part of the machine learning training process;
execute the machine learning training process with the one or more hyper parameters specified by the input setting; and
measure the performance of the machine learning training process with the one or more hyper-parameters specified by the input setting, wherein obtaining a measure of the performance of the machine learning training process with the updated setting defined by the current network output comprises:
providing the updated setting defined by the current network output generated by the first recurrent neural network to one of the plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process, and
wherein the first values of the plurality of network parameters are obtained by training the first recurrent neural network to, at each iteration of the training, optimize a differentiable training function by minimizing a loss associated with the differentiable training function
Instant Application
Patent No. 11354594
Claim 14
Claim 17
A method of determining an optimized setting for one or more process parameters of a machine learning training process, the method comprising:
A method of determining an optimized setting for one or more process parameters of a machine learning training process, the method comprising:
determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of a machine learning training process for training a machine learning model with the candidate setting,
determining, for each of a plurality of candidate settings for the one or more hyper- parameters, a respective measure of the performance of the machine learning training process with the candidate setting,
wherein the determining comprises, for each time step of a plurality of time steps, performing the following:
wherein the determining comprises, for each time step of a plurality of time steps, performing the following:
providing, as input to a first neural network, a current network input comprising (i) a current setting of the one or more process parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting;
providing, as input to a first recurrent neural network that corresponds to a plurality of worker computing units, a current network input comprising (i) a current setting of the one or more hyper-parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting;
processing, using the first neural network, the current network input to obtain a current network output that defines an updated setting of the one or more process parameters,
processing, using the first recurrent neural network in accordance with first values of the network parameters, the current network input to obtain a current network output that defines an updated setting of the one or more hyper-parameters, and
providing the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting, and
provide the updated setting to one of the plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting,
obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first neural network, comprising providing the updated setting defined by the current network output generated by the first neural network to one of a plurality of worker computing units and,
obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first recurrent neural network, comprising providing the updated setting defined by the current network output generated by the first recurrent neural network to one of a plurality of worker computing units and,
obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process;
obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process,
wherein each worker computing unit is configured to:
receive, from the first recurrent neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more hyper-parameters of the machine learning training process, wherein the one or more hyper- parameters are not learned as part of the machine learning training process;
execute the machine learning training process with the one or more hyper parameters specified by the input setting; and measure the performance of the machine learning training process with the one or more hyper-parameters specified by the input setting,
generating a new network input to be provided as input to the first recurrent neural network at the next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output; and
and selecting a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings.
selecting a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings,
and wherein the first values of the plurality of network parameters are obtained by training the first recurrent neural network to, at each iteration of the training, optimize a differentiable training function by minimizing a loss associated with the differentiable training function
Instant Application
Patent No. 11354594
Claim 15
Claim 21
wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises: generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance; processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input; and providing the updated settings defined by the initial network outputs to respective worker computing units in the plurality of worker computing units.
wherein determining, for each of a plurality of candidate settings for the one or more hyper-parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises: generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance; processing each of the plurality of initial network inputs using the first recurrent neural network to generate a respective initial network output for each initial network input; and providing the updated settings defined by the initial network outputs to respective worker computing units in the plurality of worker computing units.
Instant Application
Patent No. 11354594
Claim 16
Claim 22
wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values.
Instant Application
Patent No. 11354594
Claim 17
Claims 24-26
wherein first neural network has been trained on a loss function, wherein the loss function is a summed loss function, an expected posterior improvement loss function, or an observed improvement loss function.
24: wherein the loss is a summed loss function
25: wherein the loss is an expected posterior improvement loss function
26: wherein the loss is an observed improvement loss function
Instant Application
Patent No. 11354594
Claim 19
Claim 17
wherein the first neural network corresponds to the plurality of worker computing units.
…a first recurrent neural network that corresponds to a plurality of worker computing units…
Instant Application
Patent No. 11354594
Claim 20
Claim 17
wherein each worker computing unit is configured to: receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process; and measure the performance of the machine learning training process with the one or more process parameters specified by the input setting,
wherein each worker computing unit is configured to: receive, from the first recurrent neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more hyper-parameters of the machine learning training process, wherein the one or more hyper- parameters are not learned as part of the machine learning training process; execute the machine learning training process with the one or more hyper parameters specified by the input setting; and measure the performance of the machine learning training process with the one or more hyper-parameters specified by the input setting,
Instant Application
Patent No. 11354594
Claim 21
Claim 17
generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output
…generating a new network input to be provided as input to the first recurrent neural network at the next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output;
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 2-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 2:
Step 1 Statutory Category: Claim 2 is directed to a system, which falls under one of the four statutory categories.
Step 2A Prong 1 Judicial exception: Claim 2 recites, in part, “process the sequence of network inputs to generate a respective network output for each network input that defines an updated setting for the one or more process parameters”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “determine, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising: processing, … in accordance with first values of the network parameters, a current network input to obtain a current network output that defines an updated setting of the one or more process parameters”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “A system … comprising: one or more first computers and one or more first storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement…”, “the plurality of worker computing units”, and “a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “a first neural network having a plurality of network parameters”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “provide the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “…using the first neural network…”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first neural network”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g).
Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “A system … comprising: one or more first computers and one or more first storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement…”, “the plurality of worker computing units”, and “a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional elements: “a first neural network having a plurality of network parameters” and “…using the first neural network…” amount to generally linking the use of the judicial exception to a particular technological environment or field of use. Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional elements “receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting”, “provide the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting”, and “obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first neural network” amount to adding insignificant extra-solution activity to the judicial exception and are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible.
Regarding claim 3, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein each worker computing unit operates asynchronously from each other worker computing unit”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 4, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting further comprises: generating a plurality of initial network inputs, each initial network input comprising (i) a placeholder setting and (ii) a placeholder measure of performance; processing each of the plurality of initial network inputs using the first neural network to generate a respective initial network output for each initial network input”. This limitation recites mental processes in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception.
Further, the claim recites: “providing the updated settings defined by the plurality of initial network outputs to respective worker computing units in the plurality of worker computing units”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible.
Regarding claim 5, the rejection of claim 4 is incorporated, and further, the claim recites: “wherein each network input further includes a binary variable that indicates whether or not the network input includes placeholder values, and wherein the binary variable in each initial network input indicates that the initial network input includes placeholder values”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process, see MPEP §2106.05(f) and generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process or generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 6, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein the first neural network is a differentiable neural computer (DNC)”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 7, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein the first neural network is a long short-term memory (LSTM) neural network”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 8, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration”. This limitation covers the recitation of a mathematical formula or equation, as directed to “a claim that recites a numerical formula or equation will be considered as falling within the "mathematical concepts" grouping. In addition, there are instances where a formula or equation is written in text format that should also be considered as falling within this grouping”. See MPEP § 2106.04(a)(2)(I)(B).
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 9, the rejection of claim 8 is incorporated, and further, the claim recites: “wherein the training function for each iteration is sampled from a training distribution”. This limitation recites mathematical concepts in addition to those identified the rejection of the parent claim. Thus, the claim recites a judicial exception.
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 10, the rejection of claim 8 is incorporated, and further, the claim recites: “wherein the loss function is a summed loss function or an expected posterior improvement loss function”. This limitation is a continuation of the “wherein first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration” limitation identified as an abstract idea in the parent claim.
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 11, the rejection of claim 8 is incorporated, and further, the claim recites: “wherein the loss function is an observed improvement loss function”. This limitation is a continuation of the “wherein first neural network has been trained on a loss function that, for each training iteration of one or more training iterations, depends on values of a training function for a plurality of queries, each of the plurality of queries being defined by a respective network output generated by first neural network at a respective time step of a plurality of time steps of the training iteration” limitation identified as an abstract idea in the parent claim.
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 12, the rejection of claim 2 is incorporated, and further, the claim recites: “train the machine learning model using the machine learning training process with the optimized setting for the process parameters”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 13:
Step 1 Statutory Category: Claim 13 is directed to an article of manufacture, which falls under one of the four statutory categories.
Step 2A Prong 1 Judicial exception: Claim 13 recites, in part, “process the sequence of network inputs to generate a respective network output for each network input that defines an updated setting for the one or more process parameters”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “determine, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of the machine learning training process with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing operations comprising: processing, … in accordance with first values of the network parameters, a current network input to obtain a current network output that defines an updated setting of the one or more process parameters”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement a system for determining an optimized setting for one or more process parameters”, “the plurality of worker computing units”, and “a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “a first neural network, the first neural network having a plurality of network parameters”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “provide the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “…using the first neural network…”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first neural network”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g).
Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to implement a system for determining an optimized setting for one or more process parameters”, “the plurality of worker computing units”, and “a subsystem for executing the first neural network, the subsystem comprising one or more second computers and one or more second storage devices” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional elements: “a first neural network, the first neural network having a plurality of network parameters” and “…using the first neural network…” amount to generally linking the use of the judicial exception to a particular technological environment or field of use. Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional elements “receive a sequence of network inputs, each network input in the sequence comprising (i) a respective input setting that specifies the one or more process parameters of a machine learning training process for training a machine learning model, and (ii) a measure of a performance of the machine learning training process with the input setting”, “provide the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting”, and “obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first neural network” amount to adding insignificant extra-solution activity to the judicial exception and are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible.
Regarding claim 14:
Step 1 Statutory Category: Claim 14 is directed to a method, which falls under one of the four statutory categories.
Step 2A Prong 1 Judicial exception: Claim 14 recites, in part, “determining, for each of a plurality of candidate settings for the one or more process parameters, a respective measure of the performance of a machine learning training process for training a machine learning model with the candidate setting, wherein the determining comprises, for each time step of a plurality of time steps, performing the following”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “processing, …, the current network input to obtain a current network output that defines an updated setting of the one or more process parameters”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Further, the claim recites: “selecting a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case a judgment. See MPEP § 2106.04(a)(2)(III).
Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “a machine learning training process”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “providing, as input to a first neural network, a current network input comprising (i) a current setting of the one or more process parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the claim recites: “…using the first neural network…”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “providing the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting” and “obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first neural network, comprising providing the updated setting defined by the current network output generated by the first neural network to one of a plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process”. These limitations are additional elements that amount to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g).
Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a machine learning training process” and “…using the first neural network…” generally link the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Further, the additional elements: “providing, as input to a first neural network, a current network input comprising (i) a current setting of the one or more process parameters of the machine learning training process and (ii) a measure of performance of the machine learning training process in training the machine learning model with the current setting”, “providing the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting”, and “obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first neural network, comprising providing the updated setting defined by the current network output generated by the first neural network to one of a plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process” amount to adding insignificant extra-solution activity to the judicial exception and are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible.
Regarding claim 15, the rejection of claim 14 is incorporated, and further, claim 15 is substantially similar to claim 4 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 16, the rejection of claim 15 is incorporated, and further, claim 16 is substantially similar to claim 5 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 17, the rejection of claim 14 is incorporated, and further, the claim recites: “wherein first neural network has been trained on a loss function, wherein the loss function is a summed loss function, an expected posterior improvement loss function, or an observed improvement loss function”. This limitation covers the recitation of a mathematical formula or equation, as directed to “a claim that recites a numerical formula or equation will be considered as falling within the "mathematical concepts" grouping. In addition, there are instances where a formula or equation is written in text format that should also be considered as falling within this grouping”. See MPEP § 2106.04(a)(2)(I)(B).
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Regarding claim 18, the rejection of claim 14 is incorporated, and further, claim 18 is substantially similar to claim 12 respectively, and is rejected in the same manner and reasoning applying.
Regarding claim 19, the rejection of claim 14 is incorporated, and further, the claim recites: “where in the first neural network corresponds to the plurality of worker computing units”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible.
Regarding claim 20, the rejection of claim 14 is incorporated, and further, the claim recites: “measure the performance of the machine learning training process with the one or more process parameters specified by the input setting”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III).
Further, the claim recites: “receive, from the first neural network that corresponds to the plurality of worker computing units, a respective input setting that specifies one or more process parameters of the machine learning training process”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). The claim is not patent eligible.
Regarding claim 21, the rejection of claim 14 is incorporated, and further, the claim recites: “generating a new network input to be provided as input to the first neural network at a next time step, the new network input comprising (i) the updated setting defined by the current network output and (ii) the measure of the performance of the machine learning training process with the updated setting defined by the current network output”. This limitation covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III).
The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible.
Conclusion
Claims 2-20 have been rejected under 35 U.S.C. 101 and Double Patenting only. A complete prior art search was performed for these claims; however, no prior art was uncovered that discloses or fairly suggests the following claimed features:
After detailed search the cited arts, neither alone nor in combination, teach the claimed subject matter of claims 2, 13, and 14:
In claims 2 and 13:
…provide the updated setting to one of a plurality of worker computing units to evaluate performance of the machine learning training process in training the machine learning model with the updated setting…
…obtaining, from one of the plurality of worker computing units, a measure of the performance of the machine learning training process in training the machine learning model with the updated setting defined by the current network output generated by the first neural network; and
select a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings
In claim 14:
…providing the updated setting to one of a plurality of worker computing units to
evaluate performance of the machine learning training process in training the machine learning model with the updated setting, and
obtaining a measure of the performance of the machine learning training process in training the machine learning model with an updated setting defined by the current network output generated by the first neural network, comprising providing the updated setting defined by the current network output generated by the first neural network to one of a plurality of worker computing units and, obtaining, from the one of the plurality of worker computing units, the measure of the performance of the machine learning training process; and
selecting a candidate setting from the plurality of candidate settings as an optimized setting for the one or more process parameters using the measures of the performance for the candidate settings
The closest prior art of record includes:
Gibiansky et al., U.S. Patent Application Publication No. 20160110657, teaches a machine learning selection and parameter optimization process involving a determination and selection of one or more candidate machine learning methods. Gibiansky includes an architecture with a server and client devices, however the method does not include sending updated settings to client devices and obtaining a measure of performance for that setting from the client device as required by the claims.
Faivishevsky et al., U.S. Patent Application Publication No. 20180240010, teaches optimization of machine learning training including a computing device to train a machine learning network that is configured to configuration parameters; the device inputs the configuration parameters to a neural network to generate a representation and inputs that to another neural network. However, Faivishevsky does not teach a plurality of worker computing units that receive settings and output a measure of performance for that setting as required by the claims.
Zoph et al., U.S. Patent Application Publication No. 20190251439, teaches using a controller neural network to generate architecture configurations for a child neural network, evaluating performance of the configurations and using the performance metrics to adjust the controller neural network. Zoph also discloses a system including clients and servers but does not teach sending updated settings to client devices and obtaining a measure of performance for that setting from the client device as required by the claims.
Lin et al., U.S Patent Application Publication No. 20160328644, teaches training a neural network and makes cursory mention of a neural network and a plurality of processing units, Lin does not explicitly teach a neural network feeding its output values to a plurality of computing/processing units and obtaining a measure of performance for that setting from the client device as required by the claims.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOLLY CLARKE SIPPEL whose telephone number is (571)272-3270. The examiner can normally be reached Monday - Friday, 7:30 a.m. - 4:30 p.m. ET..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571)272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.C.S./ Examiner, Art Unit 2122
/KAKALI CHAKI/ Supervisory Patent Examiner, Art Unit 2122