Return
Shielding Twin Delayed Deep Deterministic Policy Gradient Algorithm for Optimal Operational Control
DOI:10.1109/tii.2026.3728036.png)
Abstract
En 中文
This article addresses the optimal operational control problem in complex industrial processes, focusing on the automatic determination of setpoints to improve system performance. The challenges in solving this problem include ensuring operational safety, addressing unknown nonlinear dynamics, and mitigating strong disturbances. To address these challenges, we propose a shielding twin delayed deep deterministic policy gradient reinforcement learning algorithm. First, the optimal setpoint decision policy is approximated by a deep actor network guided by a critic network, which determines the setpoints and optimizes the defined performance index. Second, the boundaries of the safe action set are computed through an approximate model, and a shielding mechanism is designed to guarantee the safety of system operations. Finally, experiments are conducted on a hardware-in-the-loop digital-twin system to verify the effectiveness of the proposed method.
Keywords:
Complex industrial processes
hardware-in-the-loop
optimal operational control (OOC)
shielding twin delayed deep deterministic policy gradient (STD3)
setpoint decision
Journal
IF:
9.9
Papers:
8.6K
Citations:
6.0W
Organization
Cited Papers
No cited papers available

