Control-Barrier-Function Shielded Conservative Distributional Reinforcement Learning for Anti-Misoperation AGC of Thermal Power Units
Mujie Zhang1, Yajun Wu1, Dongsheng Li1, Yankai Zhu2,*, Zhengyan Zhao2
1 Central China Branch of State Grid Corporation of China, Wuhan, China
2 State Key Laboratory for Alternate Electrical Power System with Renewable Energy Sources, School of Control and Computer Engineering, North China Electric Power University, Beijing, China
* Corresponding Author: Yankai Zhu. Email:
(This article belongs to the Special Issue: Digital Twin and AI-Enabled Engineering Applications for Modern Power Energy Systems)
Energy Engineering https://doi.org/10.32604/ee.2026.087017
Received 09 June 2026; Accepted 28 August 2026; Published online 07 September 2026
Abstract
High shares of variable renewable generation increasingly require thermal units to provide deeper, faster, and more frequent automatic generation control (AGC) while operating close to low-load, steam-pressure, ramp-rate, and actuator limits. Under these conditions, an area-level regulation request that is valid for grid balancing can become unsafe at the plant interface because of nonlinear boiler–turbine dynamics, delayed measurements, sensor bias, and compound disturbances. This study therefore formulates AGC from the plant side: an area-level dispatcher allocates regulation among thermal generation, hydro generation, battery storage, and flexible load, while the proposed controller modifies only the command assigned to a 600 MW reheat thermal unit before it enters the plant distributed control system. A control-barrier-function (CBF) shielded conservative distributional reinforcement learning method is developed to prevent excessive ramping, steam-pressure excursions, valve over-actuation, low-load instability, and errors caused by delayed or biased measurements. Every transition used for training, validation, and testing is generated numerically by the declared behavior policy and benchmark generator; no measured AGC record or plant-historian data enter the reported results. Asynchronous signals are causally aligned on a 4 s decision grid, and a dedicated distributional safety-cost critic estimates the 0.95 conditional value-at-risk (CVaR) of discounted constraint cost. The nonlinear simulator is additionally cross-checked against the published plant-derived Bell–Åström model at seven operating points. Normalized power-transient NRMSE is 0.021–0.079 with correlation
, while pressure-transient NRMSE is 0.084–0.254 with
–0.970. Ten independent training seeds and 3500 held-out numerical episodes show that the controller reduces cumulative absolute ACE from 218 to 165 MW min and improves the frequency nadir from
to
Hz relative to PI AGC. No hard violation occurs in 525,000 feasible test decisions, giving a one-sided 95% binomial upper bound of 0.00057% on the per-decision violation probability. The safety statement remains conditional on the declared plant model, uncertainty set, robust control-invariant subset, and bounded prediction errors. End-to-end CPU execution requires 6.8 ms on average and 9.6 ms at the 95th percentile.
Keywords
Automatic generation control; thermal power unit; safe reinforcement learning; control barrier function; conservative offline learning; conditional value-at-risk; anti-misoperation