DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 1 is objected to because of the following informalities:
Claim 1 recites “outputting the control agent is trained to use a state signal and a weight value a control signal optimizing the target function” in last paragraph of d) that have grammar errors. Should be “the control agent is trained to use the state signals and the weight values to output a control signal optimizing the target function”.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 13-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim 13 does not fall within at least one of the four categories of patent eligible subject matter because the claim is directed to a signal - claim 13 recites “A computer program product, comprising a computer readable hardware storage device having computer readable program code stored therein”, the computer readable hardware storage device could be transitory.
Claim 14 does not fall within at least one of the four categories of patent eligible subject matter because the claim is directed to a signal - claim 14 recites “A computer-readable storage medium having a computer program product”, the computer readable storage medium could be transitory.
[0008] in specification recites non-transitory computer readable medium in a non-exclusive example list. By broadest interpretation, “computer readable hardware storage device” and “computer-readable storage medium” could be transitory.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-4, 7 and 10-14 are rejected under 35 U.S.C. 103 as being unpatentable over HENTSCHEL US 20200064788 A1 in view of Mohan US 20190210460 A1.
Regarding claim 1, HENTSCHEL teaches a computer-implemented method for controlling a machine by a learning-based control agent (Fig. 3 [0054] control of technical system by trained control device), comprising:
a) providing a performance evaluator and using a control signal to determine a performance for controlling the machine by the control signal (Fig.3 [0054] – [0057] trained model BM+RM use control actions SA1, …, SAN and current state SZ as inputs to predict the subsequent state PZ and the deviations to the PZ, PD1, …, PDN that are caused by the control actions, and the deviations are added to PZ to predict the modified subsequent states MZ1, …, MZN that are used by the optimization mode to evaluate a predefined performance),
b) providing an action evaluator and using the control signal to determine a deviation from a predefined control sequence (Fig.3 [0054] – [0057] trained model BM+RM use control actions SA1, …, SAN and current state SZ as inputs to predict the subsequent state PZ and the deviations to the subsequent state PZ, PD1, …, PDN that are caused by the control actions),
d) feeding a multiplicity of state signals into the control agent (Fig.3 the modified subsequent states signals are input into the optimization module), wherein
feeding a respectively resulting output signal from the control agent as a control signal into the performance evaluator and into the action evaluator (Fig.3 [0054] – [0057] the control actions output from the optimization module SA1, …, SAN are fed back into the trained model BM+RM to predict predefined performance of subsequent state and the deviations to the subsequent state), and
the control agent is trained to use the state signals to output a control signal optimizing the target function (Fig. 3 [0058] an optimized control action OSA is output by the trained control device using a state signal SZ and predefined performance criteria i.e. “the target function”), and
e) in order to control the machine
feeding an operating state signal from the machine into the trained control agent (Fig. 3 [0055] [0058] current system state is fed into the trained control device), and
supplying a resulting output signal from the trained control agent to the machine (Fig. 3 [0058] an optimized control action OSA is output by the trained control device to the technical system).
HENTSCHEL does not explicitly further teach:
weighting a multiplicity of weight values for the performance with respect to the deviation are generated;
feeding the weight values into the control agent;
weighting a performance respectively determined by the performance evaluator with respect to a deviation respectively determined by the action evaluator by a target function according to the respective weight value; and
the target function is optimized using the weight values.
Mohan explicitly teaches in an analogous art:
weighting a multiplicity of weight values for the performance with respect to the deviation are generated; feeding the weight values into the control; weighting a performance respectively determined by the performance evaluator with respect to a deviation respectively determined by the action evaluator by a target function according to the respective weight value; and the target function is optimized using the weight values ([0040] weight value for the fuel consumption i.e. “the performance” and the weight value for the corresponding deviation of the average engine speed i.e. “with respect to the deviation” are provided for the cost function, [0044] [0050] the weight values are adapted in real time).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified HENTSCHEL to incorporate the teachings of Mohan, because they all directed to technical system control, to make the method wherein weighting a multiplicity of weight values for the performance with respect to the deviation are generated; feeding the weight values into the control agent; weighting a performance respectively determined by the performance evaluator with respect to a deviation respectively determined by the action evaluator by a target function according to the respective weight value; and the target function is optimized using the weight values. One of ordinary skill in the art would have been motivated to do this modification so as to tailor the objective function to each specific level of the operating parameter, as Mohan teaches in [0107].
Regarding claim 3, HENTSCHEL further teaches a respective performance is determined by the performance evaluator and/or a respective deviation is determined by the action evaluator on the basis of a respective state signal (Fig.3 [0054] – [0057] the performance and the deviations to the subsequent state are based on the predicted state).
Regarding claim 4, HENTSCHEL further teaches a performance value is respectively read in for a multiplicity of state signals and control signals and quantifies a performance resulting from application of a respective control signal to a state of the machine specified by a respective state signal, in that the performance evaluator is trained to reproduce an associated performance value on the basis of a state signal and a control signal (Fig.3 [0054] – [0057] trained model BM+RM use control actions SA1, …, SAN and current state SZ as inputs to predict the subsequent state PZ and the deviations to the PZ, PD1, …, PDN that are caused by the control actions, and the deviations are added to PZ to predict the modified subsequent states MZ1, …, MZN that are used by the optimization mode to evaluate a predefined performance, and the step is repeat until the performance criteria is met, i.e. “reproduce an associated performance value”).
Regarding claim 7, HENTSCHEL further teaches the deviation is determined by the action evaluator by a variational autoencoder, by an autoencoder, by generative adversarial networks and/or by a comparison, a state-signal-dependent comparison (Fig.3 [0054] – [0057] trained model BM+RM use control actions SA1, …, SAN and current state SZ as inputs to predict the subsequent state PZ and the deviations to the subsequent state PZ, PD1, …, PDN that are caused by the control actions, i.e. by “a state-signal-dependent comparison”), with predefined control signals.
Regarding claim 10, HENTSCHEL further teaches the control agent, the performance evaluator and/or the action evaluator comprise an artificial neural network, a recurrent neural network, a convolutional neural network, a multilayer perceptron, a Bayesian neural network, an autoencoder, a variational autoencoder, a deep learning architecture, a support vector machine, a data-driven trainable regression model, a k-nearest neighbor classifier, a physical model and/or a decision tree (neural network, a support vector machine, a decision tree).
Regarding claim 11, HENTSCHEL further teaches the machine is a robot, a motor, a manufacturing plant, a factory, an energy supply device, a gas turbine, a wind turbine ([0003] wind turbine), a steam turbine, a milling machine or another device or another installation.
Regarding claim 12, HENTSCHEL further teaches a controller for controlling a machine (Fig. 3 [0054] CTL), configured to carry out a method as claimed in claim 1.
Regarding claim 13, HENTSCHEL further teaches a computer program product, comprising a computer readable hardware storage device having computer readable program code stored therein, said program code executable by a processor of a computer system (Fig. 3 [0054] CTL) to implement a method as claimed in claim 1.
Regarding claim 14, HENTSCHEL further teaches a computer-readable storage medium having a computer program product (Fig. 3 [0054] CTL) as claimed in claim 13.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over HENTSCHEL in view of Mohan as applied to claims 1, 3-4, 7 and 10-14 above, further in view of Fujimoto “Off-Policy Deep Reinforcement Learning without Exploration” 2019.
Regarding claim 5, HENTSCHEL further teaches the performance evaluator is trained to determine a performance accumulated over a future period of time (Fig.3 [0054] – [0057] trained model BM+RM use control actions SA1, …, SAN and current state SZ as inputs to predict the subsequent state PZ and the deviations to the PZ, PD1, …, PDN that are caused by the control actions, and the deviations are added to PZ to predict the modified subsequent states MZ1, …, MZN that are used by the optimization mode to evaluate a predefined performance, i.e. “performance accumulated over a future period of time”).
Neither HENTSCHEL nor Mohan explicitly further teaches the training is done by a Q-learning method and/or another Q-function-based reinforcement learning method.
Fujimoto explicitly teaches in an analogous art that the training is done by a Q-learning method and/or another Q-function-based reinforcement learning method (page 1 right column paragraph 2 from the bottom, Batch-Constrained deep Q-learning).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified HENTSCHEL and Mohan to incorporate the teachings of Fujimoto, because they all directed to technical system control, to make the method wherein the training is done by a Q-learning method and/or another Q-function-based reinforcement learning method. One of ordinary skill in the art would have been motivated to do this modification so that agents are trained to maximize reward, as Fujimoto teaches in page 1 right column paragraph 2 from the bottom.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over HENTSCHEL in view of Mohan as applied to claims 1, 3-4, 7 and 10-14 above, further in view of YAO CN 109742773 B.
Regarding claim 8, neither HENTSCHEL nor Mohan explicitly further teaches the weight values are generated in a randomized manner.
YAO explicitly teaches in an analogous art that the weight values are generated in a randomized manner (page 6 paragraph 6 from the bottom, setting initial weight randomly).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified HENTSCHEL and Mohan to incorporate the teachings of Fujimoto, because they all directed to technical system control, to make the method wherein the weight values are generated in a randomized manner. One of ordinary skill in the art would have been motivated to do this modification so as to provide initial weight, as YAO teaches in page 6 paragraph 6 from the bottom.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over HENTSCHEL in view of Mohan as applied to claims 1, 3-4, 7 and 10-14 above, further in view of So US 20200355392 A1.
Regarding claim 8, neither HENTSCHEL nor Mohan explicitly further teaches a gradient-based optimization method, a stochastic optimization method, particle swarm optimization and/or a genetic optimization method is/are used to train the control agent, the performance evaluator and/or the action evaluator.
So explicitly teaches in an analogous art that a gradient-based optimization method, a stochastic optimization method, particle swarm optimization and/or a genetic optimization method is/are used to train the control agent, the performance evaluator and/or the action evaluator ([0020] particle swarm optimization and/or a genetic algorithm).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified HENTSCHEL and Mohan to incorporate the teachings of Fujimoto, because they all directed to technical system control, to make the method wherein a gradient-based optimization method, a stochastic optimization method, particle swarm optimization and/or a genetic optimization method is/are used to train the control agent, the performance evaluator and/or the action evaluator. One of ordinary skill in the art would have been motivated to do this modification so as to optimize the applicable control signals with relatively low computational effort, as So teaches in [0020].
Allowable Subject Matter
Claims 2 and 6 are objected to as being dependent upon rejected base claims, but would be allowable if rewritten to overcome the objections set forth in this Office action and in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim 2, claim 2 depends on claim 1, HENTSCHEL and Mohan together teach the claim limitations of claim 1. Mohan further teaches the weight is adapted in real-time ([0044] [0050]); SCHWAIBOLD US 20230095821 A1 teaches the weight value is gradually changed when controlling the machine ([0019] weight performance factors more recent in the progression over time more strongly). However, HENTSCHEL, Mohan and SCHWAIBOLD do not teach or suggest individually or in combination:
the weight value is gradually changed when controlling the machine in such a manner that the performance is increasingly given a higher weighting with respect to the deviation.
Regarding claim 6, claim 6 depends on claim 1, HENTSCHEL and Mohan together teach the claim limitations of claim 1. Starke US 12138543 B1 teaches use a control signal to reproduce the control signal following information reduction, wherein a reproduction error is determined (Fig. 2C character control variables are input to autoencoder for reconstruction, the error between the output and the input character control variables are used to determine the reconstruction loss). However, HENTSCHEL, Mohan and Starke do not teach or suggest individually or in combination:
the action evaluator is trained to use a state signal and a control signal to reproduce the control signal following information reduction, wherein a reproduction error is determined, and
in that the deviation is determined on the basis of the reproduction error.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Menner US 20230038215 A1 teaches performance objective including cost function defining deviation of states, deviation of the control inputs, and tuning weights of a cost function.
QIAO CN 102611118 B teaches weights of safety index and target tracking deviation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Tang whose telephone number is (571)272-7437. The examiner can normally be reached M-F 7:30-4 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamini Shah can be reached on (571)272-2279. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.T./ Examiner, Art Unit 2115
/KAMINI S SHAH/ Supervisory Patent Examiner, Art Unit 2115