Return
Safe policy optimization with stretchable penalties
DOI:10.1016/j.neucom.2026.134968.png)
Abstract
En 中文
In safe reinforcement learning (RL), ensuring cost-constraint satisfaction while minimizing the sacrifice of reward acquisition presents a significant challenge. This paper proposes an efficient safe RL algorithm called stretchable penalty-based safe policy optimization (S2P2O). S2P2O constructs a non-monotonic Swish-shaped penalty with a sign-changing gradient and a tunable stretched optimal penalty point, enabling a smooth transition from reward promotion to safety correction. A Kullback–Leibler (KL) divergence stretching mechanism rescales the reward and penalty gradients based on the current reward and constraint status once the KL threshold is exceeded, enabling continued policy optimization and improved sample efficiency by fully utilizing collected trajectories. The theoretical error bound between the optimal values of the Swish-shaped objective and the original constrained objective is analyzed. Comprehensive experiments are conducted to benchmark S2P2O against several state-of-the-art safe RL algorithms. The results demonstrate that S2P2O exhibits enhanced cost constraint satisfaction, superior reward acquisition capacity, and accelerated cost convergence rates.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

