arrow
Return

Safe policy optimization with stretchable penalties

delete2026-08-31
delete0
PRE
AI
N
Ning Pang
L
Longyang Huang
B
Botao Dong
张卫东 (Weidong Zhang) *
DOI:10.1016/j.neucom.2026.134968delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In safe reinforcement learning (RL), ensuring cost-constraint satisfaction while minimizing the sacrifice of reward acquisition presents a significant challenge. This paper proposes an efficient safe RL algorithm called stretchable penalty-based safe policy optimization (S2P2O). S2P2O constructs a non-monotonic Swish-shaped penalty with a sign-changing gradient and a tunable stretched optimal penalty point, enabling a smooth transition from reward promotion to safety correction. A Kullback–Leibler (KL) divergence stretching mechanism rescales the reward and penalty gradients based on the current reward and constraint status once the KL threshold is exceeded, enabling continued policy optimization and improved sample efficiency by fully utilizing collected trajectories. The theoretical error bound between the optimal values of the Swish-shaped objective and the original constrained objective is analyzed. Comprehensive experiments are conducted to benchmark S2P2O against several state-of-the-art safe RL algorithms. The results demonstrate that S2P2O exhibits enhanced cost constraint satisfaction, superior reward acquisition capacity, and accelerated cost convergence rates.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

S
shanghai jiao tong university
Scholars:
15.6W
Papers: 11.6W
Citations: 159
T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137