DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is made non-final.
The preliminary amendment of claims 1-8 and 10-13 filed on 09/12/2024 has been reviewed and considered by this office action.
Priority
Acknowledgment is made of applicant's claim for foreign priority under 35 U.S.C. 119 (a)-(d) based on Application No. CN202210886335.2 filed on 07/26/2022. Copies of certified papers required by 37 CFR 1.55 have been received.
Information Disclosure Statement
The information disclosure statement filed on 12/31/2024 has been reviewed and considered by this office action.
Drawings
The drawings filed on 09/12/2024 have been reviewed and are considered acceptable.
Specification
The specification filed on 09/12/2024 has been reviewed and is considered acceptable.
Applicant is reminded of the proper content of an abstract of the disclosure.
A patent abstract is a concise statement of the technical disclosure of the patent and should include that which is new in the art to which the invention pertains. The abstract should not refer to purported merits or speculative applications of the invention and should not compare the invention with the prior art.
If the patent is of a basic nature, the entire technical disclosure may be new in the art, and the abstract should be directed to the entire disclosure. If the patent is in the nature of an improvement in an old apparatus, process, product, or composition, the abstract should include the technical disclosure of the improvement. The abstract should also mention by way of example any preferred modifications or alternatives.
Where applicable, the abstract should include the following: (1) if a machine or apparatus, its organization and operation; (2) if an article, its method of making; (3) if a chemical compound, its identity and use; (4) if a mixture, its ingredients; (5) if a process, the steps.
Extensive mechanical and design details of an apparatus should not be included in the abstract. The abstract should be in narrative form and generally limited to a single paragraph within the range of 50 to 150 words in length.
See MPEP § 608.01(b) for guidelines for the preparation of patent abstracts.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 5-7, and 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Ma et al. (CN 114217524 A), in view of Zhu et al. (CN 114048903 A), Gao et al. (CN 112529727 A) (Note: machine translations are used for mapping, attached to this action), and in view of Yu (Yu, Liang, et al. “Optimal operation of a hydrogen-based building multi-energy system based on deep reinforcement learning.” arXiv preprint arXiv:2109.10754 (2021).)
Regarding claim 1, Ma teaches a method for power grid real-time dispatch optimization, the method comprising:
acquiring power grid model parameters and power grid operation data ([0086]: “Based on the current power grid operating conditions, add random faults in the Gird2OP power grid simulation environment to simulate the actual operating conditions. After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0089]: “105 real power grid scenario data are used as the input of the agent”);
obtaining, according to the power grid model parameters and the power grid operation data, a power grid real-time dispatch adjustment strategy through a preset reinforcement learning and training model for power grid real-time dispatch ([0089]: “the obtained power grid dispatching agent is applied to real power grid scenario data and can output the corresponding action strategy of power grid dispatching in real time, so as to maximize the consumption of new energy under the premise of stable operation of the new power grid”; [0092]: “ In the context of power grid environments with load changes, parameter disturbances, and random faults, this invention proposes an adaptive scheduling decision-making method based on reinforcement learning. By sensing changes in the power grid environment in real time and adaptively adjusting the scheduling strategy according to the sensing, the active power output and voltage values of thermal power units are output in real time”);
wherein the preset reinforcement learning and training model for the power grid real- time dispatch comprises an agent and a reinforcement learning and training environment ([0089]: “a power grid dispatching agent based on the IL-SAC algorithm is constructed”; [0091]: “Based on this agent, it can interact with the real-time operating environment of the power grid and provide adaptive control decisions in sub-seconds”);
wherein the obtaining the power grid real-time dispatch adjustment strategy through the preset reinforcement learning and training model for the power grid real-time dispatch, comprises: repeating interaction operations for a preset number of times ([0191]: “the total number of training steps was set to approximately 5000 steps, meaning that the performance of the two agents was compared after about 5000 training steps”);
wherein the interaction operations comprise that: the reinforcement learning and training environment obtains a state space through a preset power flow simulation function according to the power grid model parameters and the power grid operation data ([0086]: “After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0019]: “All these variables are system observation state quantities that can be directly observed or called through the Grid2Op power grid system simulation model”), obtains a reward feedback through a preset reward feedback function according to the state space ([0106]: “The 4-dimensional tuple (S,A,P,R) is used to describe the power grid system, where S represents the state set of the power grid system, A represents the action set of the power grid system, P: S×A×S→[0,1] represents the state transition probability, and R: S×A→R represents the reward mechanism”), and transmits the state space and the reward feedback to the agent ([0088]: “In the expert action space, the optimal action is greedily searched based on the greedy optimization criterion of step (3-1). The corresponding power grid scenario state and action are combined to form an action-state pair (a,s), that is, a better action label is found for each state”);
the agent obtains an action strategy according to the state space and the reward feedback and transmits the action strategy to the reinforcement learning and training environment ([0091]: “the proposed IL-SAC algorithm agent was applied to the IEEE 118-node novel power grid system in the Grid2Op environment. Based on this agent, it can interact with the real-time operating environment of the power grid and provide adaptive control decisions in sub-seconds”; [0092]: “the agent is continuously trained in real time to ensure that the agent can obtain the maximum reward value within a decision cycle”); and
the reinforcement learning and training environment verifies the action strategy according to an action space, and updates the power grid operation data by executing the verified action strategy ([0182]: “in the action space after discretization by equation (5), the optimal action is greedily searched based on a greedy algorithm. The optimal greedy index is to maximize the new energy consumption rate index in equation (8) while ensuring that the maximum rho on each transmission line does not exceed 100%. After performing the greedy algorithm, we obtain a simulated expert action space”); and
taking the action strategy executed when the reward feedback is the highest as the power grid real-time dispatch adjustment strategy ([0155]: “solve for the policy that maximizes the cumulative reward value of the MDP model in Step 1”);
wherein the state space of the reinforcement learning and training environment comprises ([0019]: “All these variables are system observation state quantities that can be directly observed or called through the Grid2Op power grid system simulation model”) an active power output of generating units ([0019]: “the active power output… at the j-th generator node”), a reactive power output of the generating units ([0019]: “reactive power output… at the j-th generator node”), a voltage magnitude of the generating units ([0019]: “voltage at the j-th generator node”), a load active power ([0019]: “the active power demand… at the k-th load node”), a load reactive power ([0019]: “reactive power demand… at the k-th load node”), a load voltage magnitude ([0019]: “voltage at the k-th load node”), ([0019]: “Fi represents the on/off state of the i-th power transmission line”), a line loading rate ([0019]: “rhoi represents the load rate on the i-th power transmission line”), ([0019]: “the predicted upper limit of active power output at the m-th renewable energy generator node at the next time step,” where the upper bound describes a feasible or “legal” range for actions),([0040]: “the maximum output of the new energy unit m at the current time step”), a maximum active power output of the renewable energy generating units at a next time step ([0019]: “the predicted upper limit of active power output at the m-th renewable energy generator node at the next time step”), a load at a next time step ([0019]: “the active power demand… at the k-th load node”)
wherein the reward feedback function is a weighted sum ([0152-0154]: “the overall reward function rt∈R at time t is as follows: rt=c1r1+c2r2+c3r3+c4r4+c51r5+c6r61(18) Wherein, ci(i=1,2,..,6) represents the coefficients of each reward function”) of a generation cost of the generating units ([0138]: “Set a negative bonus r4 based on the unit operating costs”), ([0134]: “Set a negative reward r3 based on the over-limit status of the balancing unit's power,” which functions as a reserve capacity usage cost by penalizing the use of the balancing unit's power output beyond its limits), a line loading rate ([0128]: “Set the reward function r<sub>1</sub> based on the transmission line over-limit situation”) and a degree of node voltage exceeding the limit ([0145]: “Set negative reward r6 based on the voltage over-limit situation of unit nodes and load nodes”);
wherein weight coefficients of the generation cost of the generating units([0152]: “the range of reward function r1 is (-1, 1), the range of r1 is [0, 1], and the range of r3, r4, r5, and r6 is (-1, 0)”).
Ma does not explicitly teach “wherein the state space of the reinforcement learning and training environment comprises… a charging and discharging power of an energy storage battery, a power grid loss, a startup-shutdown state of the generating units, … and a power flow convergence flag.”
Zhu further teaches wherein the state space of the reinforcement learning and training environment comprises a power grid loss ([0043]: “rNER32 is the network loss optimization reward”), a startup-shutdown state of the generating units ([0182]: “generating unit state”), and a power flow convergence flag ([0182]: “Simulate the power grid environment using a simulator, and return the reward value r of the previous step's state-action, the round end flag 'done', and the current step's observation space”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to adapt the method of Ma to incorporate the teachings of Zhu so as to include the state space of the reinforcement learning and training environment comprising a power grid loss, a startup-shutdown state of the generating units, and a power flow convergence flag. Doing so would allow the state space representation to include power grid loss, generating unit state, and a power flow convergence flag with the aim of improving real-time operation adjustment strategies (Zhu, [0203]: the model is trained to “quickly provide a power grid safe operation strategy according to the real-time state of the power grid”).
Ma and Zhu do not explicitly teach “wherein the state space of the reinforcement learning and training environment comprises… a charging and discharging power of an energy storage battery.”
Gao further teaches wherein the state space of the reinforcement learning and training environment comprises a charging and discharging power of an energy storage battery ([0012]: “Pb(t) represents the charging or discharging power of the energy storage system at each time t”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to adapt the method of Ma in view of Zhu to incorporate the teachings of Gao so as to include the state space of the reinforcement learning and training environment comprising a charging and discharging power of an energy storage battery. Doing so would allow the state space representation to include energy storage behavior with the aim of improving renewable energy absorption and reducing microgrid operating costs (Gao, [0031]: “energy storage systems can improve the microgrid's ability to absorb distributed renewable energy sources and can also reduce the operating costs of the microgrid in its daily operation by working with the microgrid's energy management strategies”).
Ma, Zhu, and Gao do not explicitly teach “wherein the reward feedback function is a weighted sum of… a carbon emission cost of the generating units, a loss cost of the energy storage battery.”
Yu further teaches wherein the reward feedback function is a weighted sum of a carbon emission cost of the generating units, a loss cost of the energy storage battery (Page 5, Section II: “The operational cost of the HBMES consists of six parts… carbon emission cost C2,t, BESS depreciation cost C3,t”);;
wherein weight coefficients of the carbon emission cost of the generating units, the loss cost of the energy storage battery are negative (Page 5, Section II: “Let vt and τt be the buying and selling prices of electricity, respectively… C1,t = vtPg,t if Pg,t ≥ 0; Otherwise, C1,t = τtPg,t… C2,t = μcμe,tPg,tΔt,”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to adapt the method of Ma in view of Zhu and Gao to incorporate the teachings of Yu so as to include the reward feedback function being a weighted sum of a carbon emission cost of the generating units and a loss cost of the energy storage battery. Doing so would allow carbon emissions and energy storage loss to be accounted for in a reward function with the aim of reducing operational cost (Yu, Page 2, Section I: “we intend to minimize the expected operational cost of an HBMES by intelligently scheduling thermal loads and various ESSs, including hydrogen, thermal, and electric ESSs. However, several challenges are involved in achieving the above aim. Firstly, there are many uncertain parameters, e.g., renewable generation output, electric load, electricity price, outdoor temperature, and carbon emission rate”).
Regarding claim 2, Ma in view of Zhu, Gao, and Yu teaches the method for power grid real-time dispatch optimization of claim 1.
Ma further teaches wherein the method further comprises acquiring equipment failure information of a power grid, and updating the power grid model parameters according to the equipment failure information ([0181]: “Based on the current power grid operating conditions, add random faults in the Gird2OP power grid simulation environment to simulate the actual operating conditions. After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0190]: “The random fault rules designed in the simulation process of this invention are as follows: In each time step, a transmission line outage probability of 1% is designed, that is, the probability of each transmission line failing at time t is 1%”).
Regarding claim 3, Ma in view of Zhu, Gao, and Yu teaches the method for power grid real-time dispatch optimization of claim 1.
Ma further teaches wherein the action space comprises respective action variables and action constraints of thermal power units ([0092]: “the active power output and voltage values of thermal power units are output in real time”), PV-type renewable energy generating units ([0019]: “active power output at the m-th renewable energy generator node”; [0187]: “The operable actions provided by Grid2Op in this new power grid system are the active power output and unit voltage values of the units”), PQ-type renewable energy generating units ([0112]: “reactive power Output” corresponds to PQ-type)
wherein the action variable of the thermal power units comprises an active power adjustment amount and a terminal voltage adjustment amount ([0023]: “X represents the number of controllable generating units in the power grid system; represents the active power output regulation value at the x-th generating unit node; and represents the voltage regulation value at the x-th generating unit node”);
the action variable of the PV-type renewable energy generating units comprises an active power adjustment amount and a terminal voltage adjustment amount ([0023]: “represent the active power output, reactive power output, and voltage at the j-th generator node, respectively”);
the action variable of the PQ-type renewable energy generating units comprises an active power adjustment amount and a reactive power adjustment amount ([0023]: “represent the active power output, reactive power output… at the j-th generator node, respectively”);
the action constraint of the thermal power units comprises a power output constraint of the generating units ([0040]: “the maximum output of the new energy unit m at the current time step”; [0051]: “the upper and lower limits of the unit's reactive power output”), ([0056]: “the upper and lower limits of the voltage of each generator node and load node”)
the action constraint of the PV-type renewable energy generating units comprises a terminal voltage constraint of the renewable energy generating units ([0056]: “the upper and lower limits of the voltage of each generator node and load node”) and a maximum allowable power output constraint of PV-type renewable energy ([0040]: “the maximum output of the new energy unit m at the current time step”; [0051]: “the upper and lower limits of the unit's reactive power output”);
the action constraint of the PQ-type renewable energy generating units comprises a maximum allowable power output constraint of PQ-type renewable energy ([0040]: “the maximum output of the new energy unit m at the current time step”) and a reactive power constraint of the generating units ([0051]: “the upper and lower limits of the unit's reactive power output”);
Ma does not explicitly teach “wherein the action space comprises an energy storage battery,” “the action variable of the energy storage battery comprises an active power adjustment amount,” or “the action constraint of the energy storage battery comprises a battery charging and discharging constraint and a battery capacity constraint.”
Gao further teaches wherein the action space comprises respective action variables and action constraints of an energy storage battery ([0077]: “the energy storage battery”);
the action variable of the energy storage battery comprises an active power adjustment amount ([0012]: “Pb(t) represents the charging or discharging power of the energy storage system at each time t”; [0027]: “The actions of the energy storage system include charging, discharging, and no action”);
the action constraint of the energy storage battery comprises a battery charging and discharging constraint and a battery capacity constraint ([0080-0082]: “For the established energy storage model, in order to ensure its normal operation, its charging and discharging power Pb(t) and state of charge (SOC) are constrained:Pbmin≤Pb(t)≤Pbmax(2)SOCmin≤SOC(t)≤SOCmax(3)”).
Ma and Gao do not explicitly teach “the action constraint of the thermal power units comprises a power output ramping constraint of the generating units and a startup-shutdown constraint of the generating units.”
Zhu further teaches the action constraint of the thermal power units comprises a power output ramping constraint of the generating units ([0049]: “Unit ramping constraint: The active power output adjustment value of any thermal power unit must be less than the ramping rate”) and a startup-shutdown constraint of the generating units ([0050]: “Unit start-up and shutdown constraints: The shutdown rule for thermal power units is that the active power output of the unit must be adjusted to the lower limit of the output before the unit is shut down, and then adjusted to 0”);
Regarding claim 5, Ma teaches a system for power grid real-time dispatch optimization, the system comprising:
wherein the processor is configured to: acquire power grid model parameters and power grid operation data ([0086]: “Based on the current power grid operating conditions, add random faults in the Gird2OP power grid simulation environment to simulate the actual operating conditions. After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0089]: “105 real power grid scenario data are used as the input of the agent”); and
obtain a power grid real-time dispatch adjustment strategy through a preset reinforcement learning and training model for power grid real-time dispatch according to the power grid model parameters and the power grid operation data ([0089]: “the obtained power grid dispatching agent is applied to real power grid scenario data and can output the corresponding action strategy of power grid dispatching in real time, so as to maximize the consumption of new energy under the premise of stable operation of the new power grid”; [0092]: “ In the context of power grid environments with load changes, parameter disturbances, and random faults, this invention proposes an adaptive scheduling decision-making method based on reinforcement learning. By sensing changes in the power grid environment in real time and adaptively adjusting the scheduling strategy according to the sensing, the active power output and voltage values of thermal power units are output in real time”);
wherein the preset reinforcement learning and training model for the power grid real- time dispatch comprises an agent and a reinforcement learning and training environment ([0089]: “a power grid dispatching agent based on the IL-SAC algorithm is constructed”; [0091]: “Based on this agent, it can interact with the real-time operating environment of the power grid and provide adaptive control decisions in sub-seconds”);
wherein the processor is further configured to repeat interaction operations for a preset number of times ([0191]: “the total number of training steps was set to approximately 5000 steps, meaning that the performance of the two agents was compared after about 5000 training steps”);
wherein the interaction operations comprise that: the reinforcement learning and training environment obtains a state space through a preset power flow simulation function according to the power grid model parameters and the power grid operation data ([0086]: “After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0019]: “All these variables are system observation state quantities that can be directly observed or called through the Grid2Op power grid system simulation model”), obtains a reward feedback through a preset reward feedback function according to the state space ([0106]: “The 4-dimensional tuple (S,A,P,R) is used to describe the power grid system, where S represents the state set of the power grid system, A represents the action set of the power grid system, P: S×A×S→[0,1] represents the state transition probability, and R: S×A→R represents the reward mechanism”), and transmits the state space and the reward feedback to the agent ([0088]: “In the expert action space, the optimal action is greedily searched based on the greedy optimization criterion of step (3-1). The corresponding power grid scenario state and action are combined to form an action-state pair (a,s), that is, a better action label is found for each state”);
the agent obtains an action strategy according to the state space and the reward feedback and transmits the action strategy to the reinforcement learning and training environment ([0091]: “the proposed IL-SAC algorithm agent was applied to the IEEE 118-node novel power grid system in the Grid2Op environment. Based on this agent, it can interact with the real-time operating environment of the power grid and provide adaptive control decisions in sub-seconds”; [0092]: “the agent is continuously trained in real time to ensure that the agent can obtain the maximum reward value within a decision cycle”); and
the reinforcement learning and training environment verifies the action strategy according to an action space, and updates the power grid operation data by executing the verified action strategy ([0182]: “in the action space after discretization by equation (5), the optimal action is greedily searched based on a greedy algorithm. The optimal greedy index is to maximize the new energy consumption rate index in equation (8) while ensuring that the maximum rho on each transmission line does not exceed 100%. After performing the greedy algorithm, we obtain a simulated expert action space”); and
taking the action strategy executed when the reward feedback is the highest as the power grid real-time dispatch adjustment strategy ([0155]: “solve for the policy that maximizes the cumulative reward value of the MDP model in Step 1”);
wherein the state space of the reinforcement learning and training environment comprises ([0019]: “All these variables are system observation state quantities that can be directly observed or called through the Grid2Op power grid system simulation model”) an active power output of generating units ([0019]: “the active power output… at the j-th generator node”), a reactive power output of the generating units ([0019]: “reactive power output… at the j-th generator node”), a voltage magnitude of the generating units ([0019]: “voltage at the j-th generator node”), a load active power ([0019]: “the active power demand… at the k-th load node”), a load reactive power ([0019]: “reactive power demand… at the k-th load node”), a load voltage magnitude ([0019]: “voltage at the k-th load node”), ([0019]: “Fi represents the on/off state of the i-th power transmission line”), a line loading rate ([0019]: “rhoi represents the load rate on the i-th power transmission line”), ([0019]: “the predicted upper limit of active power output at the m-th renewable energy generator node at the next time step,” where the upper bound describes a feasible or “legal” range for actions),([0040]: “the maximum output of the new energy unit m at the current time step”), a maximum active power output of the renewable energy generating units at a next time step ([0019]: “the predicted upper limit of active power output at the m-th renewable energy generator node at the next time step”), a load at a next time step ([0019]: “the active power demand… at the k-th load node”)
wherein the reward feedback function is a weighted sum ([0152-0154]: “the overall reward function rt∈R at time t is as follows: rt=c1r1+c2r2+c3r3+c4r4+c51r5+c6r61(18) Wherein, ci(i=1,2,..,6) represents the coefficients of each reward function”) of a generation cost of the generating units ([0138]: “Set a negative bonus r4 based on the unit operating costs”), ([0134]: “Set a negative reward r3 based on the over-limit status of the balancing unit's power,” which functions as a reserve capacity usage cost by penalizing the use of the balancing unit's power output beyond its limits), a line loading rate ([0128]: “Set the reward function r<sub>1</sub> based on the transmission line over-limit situation”) and a degree of node voltage exceeding the limit ([0145]: “Set negative reward r6 based on the voltage over-limit situation of unit nodes and load nodes”);
wherein weight coefficients of the generation cost of the generating units([0152]: “the range of reward function r1 is (-1, 1), the range of r1 is [0, 1], and the range of r3, r4, r5, and r6 is (-1, 0)”).
While Ma teaches a program interface for a simulation environment ([0086]: “in the simulation environment, the corresponding observation state space is obtained by calling the program interface”), Ma does not explicitly teach “a processor” and “a memory configured to store an instruction executable on the processor.” Also, Ma does not explicitly teach “wherein the state space of the reinforcement learning and training environment comprises… a charging and discharging power of an energy storage battery, a power grid loss, a startup-shutdown state of the generating units, … and a power flow convergence flag.”
Zhu further teaches a processor ([0205]: “These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus”); and
a memory configured to store an instruction executable on the processor ([0206]: “These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner”),
wherein the state space of the reinforcement learning and training environment comprises a power grid loss ([0043]: “rNER32 is the network loss optimization reward”), a startup-shutdown state of the generating units ([0182]: “generating unit state”), and a power flow convergence flag ([0182]: “Simulate the power grid environment using a simulator, and return the reward value r of the previous step's state-action, the round end flag 'done', and the current step's observation space”).
The reasons to combine Zhu into Ma are the same as articulated in the rejection of claim 1 above.
Ma and Zhu do not explicitly teach “wherein the state space of the reinforcement learning and training environment comprises… a charging and discharging power of an energy storage battery.”
Gao further teaches wherein the state space of the reinforcement learning and training environment comprises a charging and discharging power of an energy storage battery ([0012]: “Pb(t) represents the charging or discharging power of the energy storage system at each time t”).
The reasons to combine Gao into Ma and Zhu are the same as articulated in the rejection of claim 1 above.
Ma, Zhu, and Gao do not explicitly teach “wherein the reward feedback function is a weighted sum of… a carbon emission cost of the generating units, a loss cost of the energy storage battery.”
Yu further teaches wherein the reward feedback function is a weighted sum of a carbon emission cost of the generating units, a loss cost of the energy storage battery (Page 5, Section II: “The operational cost of the HBMES consists of six parts… carbon emission cost C2,t, BESS depreciation cost C3,t”);;
wherein weight coefficients of the carbon emission cost of the generating units, the loss cost of the energy storage battery are negative (Page 5, Section II: “Let vt and τt be the buying and selling prices of electricity, respectively… C1,t = vtPg,t if Pg,t ≥ 0; Otherwise, C1,t = τtPg,t… C2,t = μcμe,tPg,tΔt,”).
The reasons to combine Yu into Ma, Zhu, and Gao are the same as articulated in the rejection of claim 1 above.
Regarding claim 6, Ma in view of Zhu, Gao, and Yu teaches the system for power grid real-time dispatch optimization of claim 5.
Ma further teaches wherein the processor is further configured to: acquire equipment failure information of a power grid, and updating the power grid model parameters according to the equipment failure information ([0181]: “Based on the current power grid operating conditions, add random faults in the Gird2OP power grid simulation environment to simulate the actual operating conditions. After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0190]: “The random fault rules designed in the simulation process of this invention are as follows: In each time step, a transmission line outage probability of 1% is designed, that is, the probability of each transmission line failing at time t is 1%”).
Regarding claim 7, Ma in view of Zhu, Gao, and Yu teaches the system for power grid real-time dispatch optimization of claim 5.
Ma further teaches wherein the action space comprises respective action variables and action constraints of thermal power units ([0092]: “the active power output and voltage values of thermal power units are output in real time”), PV-type renewable energy generating units ([0019]: “active power output at the m-th renewable energy generator node”; [0187]: “The operable actions provided by Grid2Op in this new power grid system are the active power output and unit voltage values of the units”), PQ-type renewable energy generating units ([0112]: “reactive power Output” corresponds to PQ-type)
wherein the action variable of the thermal power units comprises an active power adjustment amount and a terminal voltage adjustment amount ([0023]: “X represents the number of controllable generating units in the power grid system; represents the active power output regulation value at the x-th generating unit node; and represents the voltage regulation value at the x-th generating unit node”);
the action variable of the PV-type renewable energy generating units comprises an active power adjustment amount and a terminal voltage adjustment amount ([0023]: “represent the active power output, reactive power output… at the j-th generator node, respectively”);
the action variable of the PQ-type renewable energy generating units comprises an active power adjustment amount and a reactive power adjustment amount ([0023]: “represent the active power output, reactive power output… at the j-th generator node, respectively”);
the action constraint of the thermal power units comprises a power output constraint of the generating units ([0040]: “the maximum output of the new energy unit m at the current time step”; [0051]: “the upper and lower limits of the unit's reactive power output”), ([0056]: “the upper and lower limits of the voltage of each generator node and load node”)
the action constraint of the PV-type renewable energy generating units comprises a terminal voltage constraint of the renewable energy generating units ([0056]: “the upper and lower limits of the voltage of each generator node and load node”) and a maximum allowable power output constraint of PV-type renewable energy ([0040]: “the maximum output of the new energy unit m at the current time step”; [0051]: “the upper and lower limits of the unit's reactive power output”);
the action constraint of the PQ-type renewable energy generating units comprises a maximum allowable power output constraint of PQ-type renewable energy ([0040]: “the maximum output of the new energy unit m at the current time step”) and a reactive power constraint of the generating units ([0051]: “the upper and lower limits of the unit's reactive power output”);
Ma does not explicitly teach “wherein the action space comprises an energy storage battery,” “the action variable of the energy storage battery comprises an active power adjustment amount,” or “the action constraint of the energy storage battery comprises a battery charging and discharging constraint and a battery capacity constraint.”
Gao further teaches wherein the action space comprises respective action variables and action constraints of an energy storage battery ([0077]: “the energy storage battery”);
the action variable of the energy storage battery comprises an active power adjustment amount ([0012]: “Pb(t) represents the charging or discharging power of the energy storage system at each time t”; [0027]: “The actions of the energy storage system include charging, discharging, and no action”);
the action constraint of the energy storage battery comprises a battery charging and discharging constraint and a battery capacity constraint ([0080-0082]: “For the established energy storage model, in order to ensure its normal operation, its charging and discharging power Pb(t) and state of charge (SOC) are constrained:Pbmin≤Pb(t)≤Pbmax(2)SOCmin≤SOC(t)≤SOCmax(3)”).
Ma and Gao do not explicitly teach “the action constraint of the thermal power units comprises a power output ramping constraint of the generating units and a startup-shutdown constraint of the generating units.”
Zhu further teaches the action constraint of the thermal power units comprises a power output ramping constraint of the generating units ([0049]: “Unit ramping constraint: The active power output adjustment value of any thermal power unit must be less than the ramping rate”) and a startup-shutdown constraint of the generating units ([0050]: “Unit start-up and shutdown constraints: The shutdown rule for thermal power units is that the active power output of the unit must be adjusted to the lower limit of the output before the unit is shut down, and then adjusted to 0”);
Regarding claim 10, Ma teaches
wherein the method comprises: acquiring power grid model parameters and power grid operation data ([0086]: “Based on the current power grid operating conditions, add random faults in the Gird2OP power grid simulation environment to simulate the actual operating conditions. After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0089]: “105 real power grid scenario data are used as the input of the agent”);
obtaining, according to the power grid model parameters and the power grid operation data, a power grid real-time dispatch adjustment strategy through a preset reinforcement learning and training model for power grid real-time dispatch ([0089]: “the obtained power grid dispatching agent is applied to real power grid scenario data and can output the corresponding action strategy of power grid dispatching in real time, so as to maximize the consumption of new energy under the premise of stable operation of the new power grid”; [0092]: “ In the context of power grid environments with load changes, parameter disturbances, and random faults, this invention proposes an adaptive scheduling decision-making method based on reinforcement learning. By sensing changes in the power grid environment in real time and adaptively adjusting the scheduling strategy according to the sensing, the active power output and voltage values of thermal power units are output in real time”);
wherein the preset reinforcement learning and training model for the power grid real- time dispatch comprises an agent and a reinforcement learning and training environment ([0089]: “a power grid dispatching agent based on the IL-SAC algorithm is constructed”; [0091]: “Based on this agent, it can interact with the real-time operating environment of the power grid and provide adaptive control decisions in sub-seconds”);
wherein the obtaining the power grid real-time dispatch adjustment strategy through the preset reinforcement learning and training model for the power grid real-time dispatch, comprises: repeating interaction operations for a preset number of times ([0191]: “the total number of training steps was set to approximately 5000 steps, meaning that the performance of the two agents was compared after about 5000 training steps”);
wherein the interaction operations comprise that: the reinforcement learning and training environment obtains a state space through a preset power flow simulation function according to the power grid model parameters and the power grid operation data ([0086]: “After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0019]: “All these variables are system observation state quantities that can be directly observed or called through the Grid2Op power grid system simulation model”), obtains a reward feedback through a preset reward feedback function according to the state space ([0106]: “The 4-dimensional tuple (S,A,P,R) is used to describe the power grid system, where S represents the state set of the power grid system, A represents the action set of the power grid system, P: S×A×S→[0,1] represents the state transition probability, and R: S×A→R represents the reward mechanism”), and transmits the state space and the reward feedback to the agent ([0088]: “In the expert action space, the optimal action is greedily searched based on the greedy optimization criterion of step (3-1). The corresponding power grid scenario state and action are combined to form an action-state pair (a,s), that is, a better action label is found for each state”);
the agent obtains an action strategy according to the state space and the reward feedback and transmits the action strategy to the reinforcement learning and training environment ([0091]: “the proposed IL-SAC algorithm agent was applied to the IEEE 118-node novel power grid system in the Grid2Op environment. Based on this agent, it can interact with the real-time operating environment of the power grid and provide adaptive control decisions in sub-seconds”; [0092]: “the agent is continuously trained in real time to ensure that the agent can obtain the maximum reward value within a decision cycle”); and
the reinforcement learning and training environment verifies the action strategy according to an action space, and updates the power grid operation data by executing the verified action strategy ([0182]: “in the action space after discretization by equation (5), the optimal action is greedily searched based on a greedy algorithm. The optimal greedy index is to maximize the new energy consumption rate index in equation (8) while ensuring that the maximum rho on each transmission line does not exceed 100%. After performing the greedy algorithm, we obtain a simulated expert action space”); and
taking the action strategy executed when the reward feedback is the highest as the power grid real-time dispatch adjustment strategy ([0155]: “solve for the policy that maximizes the cumulative reward value of the MDP model in Step 1”);
wherein the state space of the reinforcement learning and training environment comprises ([0019]: “All these variables are system observation state quantities that can be directly observed or called through the Grid2Op power grid system simulation model”) an active power output of generating units ([0019]: “the active power output… at the j-th generator node”), a reactive power output of the generating units ([0019]: “reactive power output… at the j-th generator node”), a voltage magnitude of the generating units ([0019]: “voltage at the j-th generator node”), a load active power ([0019]: “the active power demand… at the k-th load node”), a load reactive power ([0019]: “reactive power demand… at the k-th load node”), a load voltage magnitude ([0019]: “voltage at the k-th load node”), ([0019]: “Fi represents the on/off state of the i-th power transmission line”), a line loading rate ([0019]: “rhoi represents the load rate on the i-th power transmission line”), ([0019]: “the predicted upper limit of active power output at the m-th renewable energy generator node at the next time step,” where the upper bound describes a feasible or “legal” range for actions),([0040]: “the maximum output of the new energy unit m at the current time step”), a maximum active power output of the renewable energy generating units at a next time step ([0019]: “the predicted upper limit of active power output at the m-th renewable energy generator node at the next time step”), a load at a next time step ([0019]: “the active power demand… at the k-th load node”)
wherein the reward feedback function is a weighted sum ([0152-0154]: “the overall reward function rt∈R at time t is as follows: rt=c1r1+c2r2+c3r3+c4r4+c51r5+c6r61(18) Wherein, ci(i=1,2,..,6) represents the coefficients of each reward function”) of a generation cost of the generating units ([0138]: “Set a negative bonus r4 based on the unit operating costs”), ([0134]: “Set a negative reward r3 based on the over-limit status of the balancing unit's power,” which functions as a reserve capacity usage cost by penalizing the use of the balancing unit's power output beyond its limits), a line loading rate ([0128]: “Set the reward function r<sub>1</sub> based on the transmission line over-limit situation”) and a degree of node voltage exceeding the limit ([0145]: “Set negative reward r6 based on the voltage over-limit situation of unit nodes and load nodes”);
wherein weight coefficients of the generation cost of the generating units([0152]: “the range of reward function r1 is (-1, 1), the range of r1 is [0, 1], and the range of r3, r4, r5, and r6 is (-1, 0)”).
While Ma teaches a program interface for a simulation environment ([0086]: “in the simulation environment, the corresponding observation state space is obtained by calling the program interface”), Ma does not explicitly teach “a non-transitory computer-readable storage medium, storing computer programs.” Also, Ma does not explicitly teach “wherein the state space of the reinforcement learning and training environment comprises… a charging and discharging power of an energy storage battery, a power grid loss, a startup-shutdown state of the generating units, … and a power flow convergence flag.”
Zhu further teaches a non-transitory computer-readable storage medium, storing computer programs ([0206]: “These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner”);
wherein the state space of the reinforcement learning and training environment comprises a power grid loss ([0043]: “rNER32 is the network loss optimization reward”), a startup-shutdown state of the generating units ([0182]: “generating unit state”), and a power flow convergence flag ([0182]: “Simulate the power grid environment using a simulator, and return the reward value r of the previous step's state-action, the round end flag 'done', and the current step's observation space”).
The reasons to combine Zhu into Ma are the same as articulated in the rejection of claim 1 above.
Ma and Zhu do not explicitly teach “wherein the state space of the reinforcement learning and training environment comprises… a charging and discharging power of an energy storage battery.”
Gao further teaches wherein the state space of the reinforcement learning and training environment comprises a charging and discharging power of an energy storage battery ([0012]: “Pb(t) represents the charging or discharging power of the energy storage system at each time t”).
The reasons to combine Gao into Ma and Zhu are the same as articulated in the rejection of claim 1 above.
Ma, Zhu, and Gao do not explicitly teach “wherein the reward feedback function is a weighted sum of… a carbon emission cost of the generating units, a loss cost of the energy storage battery.”
Yu further teaches wherein the reward feedback function is a weighted sum of a carbon emission cost of the generating units, a loss cost of the energy storage battery (Page 5, Section II: “The operational cost of the HBMES consists of six parts… carbon emission cost C2,t, BESS depreciation cost C3,t”);;
wherein weight coefficients of the carbon emission cost of the generating units, the loss cost of the energy storage battery are negative (Page 5, Section II: “Let vt and τt be the buying and selling prices of electricity, respectively… C1,t = vtPg,t if Pg,t ≥ 0; Otherwise, C1,t = τtPg,t… C2,t = μcμe,tPg,tΔt,”).
The reasons to combine Yu into Ma, Zhu, and Gao are the same as articulated in the rejection of claim 1 above.
Regarding claim 11, Ma in view of Zhu, Gao, and Yu teaches the non-transitory computer-readable storage medium of claim 10.
Ma further teaches wherein the method further comprises: acquiring equipment failure information of a power grid, and updating the power grid model parameters according to the equipment failure information ([0181]: “Based on the current power grid operating conditions, add random faults in the Gird2OP power grid simulation environment to simulate the actual operating conditions. After power flow calculation is performed in the simulation environment, the corresponding observation state space is obtained by calling the program interface”; [0190]: “The random fault rules designed in the simulation process of this invention are as follows: In each time step, a transmission line outage probability of 1% is designed, that is, the probability of each transmission line failing at time t is 1%”).
Regarding claim 12, Ma in view of Zhu, Gao, and Yu teaches the non-transitory computer-readable storage medium of claim 10.
Ma further teaches wherein the action space comprises respective action variables and action constraints of thermal power units ([0092]: “the active power output and voltage values of thermal power units are output in real time”), PV-type renewable energy generating units ([0019]: “active power output at the m-th renewable energy generator node”; [0187]: “The operable actions provided by Grid2Op in this new power grid system are the active power output and unit voltage values of the units”), PQ-type renewable energy generating units ([0112]: “reactive power Output” corresponds to PQ-type)
wherein the action variable of the thermal power units comprises an active power adjustment amount and a terminal voltage adjustment amount ([0023]: “X represents the number of controllable generating units in the power grid system; represents the active power output regulation value at the x-th generating unit node; and represents the voltage regulation value at the x-th generating unit node”);
the action variable of the PV-type renewable energy generating units comprises an active power adjustment amount and a terminal voltage adjustment amount ([0023]: “represent the active power output, reactive power output… at the j-th generator node, respectively”);
the action variable of the PQ-type renewable energy generating units comprises an active power adjustment amount and a reactive power adjustment amount ([0023]: “represent the active power output, reactive power output… at the j-th generator node, respectively”);
the action constraint of the thermal power units comprises a power output constraint of the generating units ([0040]: “the maximum output of the new energy unit m at the current time step”; [0051]: “the upper and lower limits of the unit's reactive power output”), ([0056]: “the upper and lower limits of the voltage of each generator node and load node”)
the action constraint of the PV-type renewable energy generating units comprises a terminal voltage constraint of the renewable energy generating units ([0056]: “the upper and lower limits of the voltage of each generator node and load node”) and a maximum allowable power output constraint of PV-type renewable energy ([0040]: “the maximum output of the new energy unit m at the current time step”; [0051]: “the upper and lower limits of the unit's reactive power output”);
the action constraint of the PQ-type renewable energy generating units comprises a maximum allowable power output constraint of PQ-type renewable energy ([0040]: “the maximum output of the new energy unit m at the current time step”) and a reactive power constraint of the generating units ([0051]: “the upper and lower limits of the unit's reactive power output”);
Ma does not explicitly teach “wherein the action space comprises an energy storage battery,” “the action variable of the energy storage battery comprises an active power adjustment amount,” or “the action constraint of the energy storage battery comprises a battery charging and discharging constraint and a battery capacity constraint.”
Gao further teaches wherein the action space comprises respective action variables and action constraints of an energy storage battery ([0077]: “the energy storage battery”);
the action variable of the energy storage battery comprises an active power adjustment amount ([0012]: “Pb(t) represents the charging or discharging power of the energy storage system at each time t”; [0027]: “The actions of the energy storage system include charging, discharging, and no action”);
the action constraint of the energy storage battery comprises a battery charging and discharging constraint and a battery capacity constraint ([0080-0082]: “For the established energy storage model, in order to ensure its normal operation, its charging and discharging power Pb(t) and state of charge (SOC) are constrained:Pbmin≤Pb(t)≤Pbmax(2)SOCmin≤SOC(t)≤SOCmax(3)”).
Ma and Gao do not explicitly teach “the action constraint of the thermal power units comprises a power output ramping constraint of the generating units and a startup-shutdown constraint of the generating units.”
Zhu further teaches the action constraint of the thermal power units comprises a power output ramping constraint of the generating units ([0049]: “Unit ramping constraint: The active power output adjustment value of any thermal power unit must be less than the ramping rate”) and a startup-shutdown constraint of the generating units ([0050]: “Unit start-up and shutdown constraints: The shutdown rule for thermal power units is that the active power output of the unit must be adjusted to the lower limit of the output before the unit is shut down, and then adjusted to 0”);
Claims 4, 8, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Ma et al. (CN 114217524 A), in view of Zhu et al. (CN 114048903 A), in view of Gao et al. (CN 112529727 A), in view of Yu, and in view of Chen et al. (CN 105046395 A) (Note: a machine translation is used for mapping, attached to this action).
Regarding claim 4, Ma in view of Zhu, Gao, and Yu teaches the method for power grid real-time dispatch optimization of claim 1.
While Gao teaches day-ahead forecast state input ([0036]: “Obtain the day-ahead forecast data of the microgrid, including the microgrid state information required in (2.1), and initialize the battery state SOC”), Ma, Zhu, Gao, and Yu do not explicitly teach “wherein the state space further comprises a reference value of a day-ahead planned active power output of the generating units.”
Chen further teaches wherein the state space further comprises a reference value of a day-ahead planned active power output of the generating units ([0098]: “day-ahead plan is published once every day at midnight, including the start-up and shutdown plans and output curves of each conventional unit for the next day”; [0102]: “the scheduling results of the day-ahead plan are also used as input data”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to adapt the method of Ma in view of Zhu, Gao, and Yu to incorporate the teachings of Chen so as to include the state space further comprising a reference value of a day-ahead planned active power output of the generating units. Doing so would allow forecasted impacts on the power grid to be considered in the optimization with the aim of reducing the impact of uncertainty (Chen, [0007]: “This invention can reduce the impact of the uncertainty of new energy sources on the power grid and improve the power grid's ability to absorb new energy sources. On the one hand, this invention is based on ultra-short-term power forecast data with higher prediction accuracy to compile intraday rolling plans, which together with the day-ahead plan constitute a multi-time-scale scheduling mode to reduce the uncertainty of new energy sources step by step. On the other hand, it adopts a robust scheduling method to theoretically ensure that the system has the ability to absorb the uncertainty of new energy sources”).
Regarding claim 8, Ma in view of Zhu, Gao, and Yu teaches the system for power grid real-time dispatch optimization of claim 5.
While Gao teaches day-ahead forecast state input ([0036]: “Obtain the day-ahead forecast data of the microgrid, including the microgrid state information required in (2.1), and initialize the battery state SOC”), Ma, Zhu, Gao, and Yu do not explicitly teach “wherein the state space further comprises a reference value of a day-ahead planned active power output of the generating units.”
Chen further teaches wherein the state space further comprises a reference value of a day-ahead planned active power output of the generating units ([0098]: “day-ahead plan is published once every day at midnight, including the start-up and shutdown plans and output curves of each conventional unit for the next day”; [0102]: “the scheduling results of the day-ahead plan are also used as input data”).
The reasons to combine Chen into Ma, Zhu, Gao, and Yu are the same as articulated in the rejection of claim 4 above.
Regarding claim 13, Ma in view of Zhu, Gao, and Yu teaches the non-transitory computer-readable storage medium of claim 10.
While Gao teaches day-ahead forecast state input ([0036]: “Obtain the day-ahead forecast data of the microgrid, including the microgrid state information required in (2.1), and initialize the battery state SOC”), Ma, Zhu, Gao, and Yu do not explicitly teach “wherein the state space further comprises a reference value of a day-ahead planned active power output of the generating units.”
Chen further teaches wherein the state space further comprises a reference value of a day-ahead planned active power output of the generating units ([0098]: “day-ahead plan is published once every day at midnight, including the start-up and shutdown plans and output curves of each conventional unit for the next day”; [0102]: “the scheduling results of the day-ahead plan are also used as input data”).
The reasons to combine Chen into Ma, Zhu, Gao, and Yu are the same as articulated in the rejection of claim 4 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 2021/0367424 A1: Uses reinforcement learning for multi-objective power flow control, including regulating voltage profiles, line flows, and transmission losses
US 2020/0403408 A1: Scheduling startup and shutdown of generating units and batteries in a power system microgrid based on local operating measurements
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Magdalena Kossek whose telephone number is (571)272-5603. The examiner can normally be reached Mon-Fri 8:00-5:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Robert Fennema can be reached at (571)272-2748. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.I.K./Examiner, Art Unit 2117
/ROBERT E FENNEMA/Supervisory Patent Examiner, Art Unit 2117