arrow
返回

CVaR-Constrained Policy Optimization for Safe Reinforcement Learning

delete2025-01-01
delete2
PRE
AI
Q
Qiyuan Zhang
S
Shu Leng
X
Xiaoteng Ma
Q
Qihan Liu
王
王学谦 (Xueqian Wang)
梁
梁斌 (Bin Liang)
刘
刘昱 (Yu Liu) *
J
Jun Yang *
DOI:10.1109/TNNLS.2023.3331304delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Current constrained reinforcement learning (RL) methods guarantee constraint satisfaction only in expectation, which is inadequate for safety-critical decision problems. Since a constraint satisfied in expectation remains a high probability of exceeding the cost threshold, solving constrained RL problems with high probabilities of satisfaction is critical for RL safety. In this work, we consider the safety criterion as a constraint on the conditional value-at-risk (CVaR) of cumulative costs, and propose the CVaR-constrained policy optimization algorithm (CVaR-CPO) to maximize the expected return while ensuring agents pay attention to the upper tail of constraint costs. According to the bound on the CVaR-related performance between two policies, we first reformulate the CVaR-constrained problem in augmented state space using the state extension procedure and the trust-region method. CVaR-CPO then derives the optimal update policy by applying the Lagrangian method to the constrained optimization problem. In addition, CVaR-CPO utilizes the distribution of constraint costs to provide an efficient quantile-based estimation of the CVaR-related value function. We conduct experiments on constrained control tasks to show that the proposed method can produce behaviors that satisfy safety constraints, and achieve comparable performance to most safe RL (SRL) methods.
Keyword:
Costs
Safety
Optimization
Measurement
Tail
Reactive power
Random variables
Artificial intelligence (AI) safety
conditional value-at-risk (CVaR)
constrained policy optimization
safe reinforcement learning (SRL)

期刊

IEEE Transactions on Neural Networks and Learning Systems 封面图
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
论文数:
7.6K
被引数:
7.2W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
引用论文

引用论文

Miscibility and interactions in a mixture of poly(ethylene oxide) and an aromatic poly(ether amide)
err1998-03-01
err0
PREAI
errA. Etxeberria; S. Guezala; J.J. Iruin; J.G. de la Campa; J. de Abajo
err分享
err收藏
Optimization of conditional value-at-risk条件风险价值的优化
err2000-01-01
err0
PREAI
errR. Tyrrell Rockafellar; Stanislav Uryasev
err分享
err收藏
An electrophoretic e‐paper device with stretchable, washable, and rewritable functions
err2022-04-29
err0
PREAI
errZhiguang Qiu; Simu Zhu; Hao Lu; Yifan Gu; Ziyi Wu; Gaofan Zhang; Shaozhi Deng; Bo‐Ru Yang
err分享
err收藏
Robust Reinforcement Learning: A Review of Foundations and Recent Advances
err2022-03-19
err60
errOAAI
errMoos, Janosch; Hansel, Kay; Abdulsamad, Hany; Stark, Svenja; Clever, Debora; Peters, Jan
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容