arrow
返回

Safe Value Functions

delete2023-05-01
delete4
delete
OA
AI
P
Pierre-François Massiani *
S
Steve Heim
F
Friedrich Solowjow
S
Sebastian Trimpe
DOI:10.1109/TAC.2022.3200948delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Safety constraints and optimality are important but sometimes conflicting criteria for controllers. Although these criteria are often solved separately with different tools to maintain formal guarantees, it is also common practice in reinforcement learning (RL) to simply modify reward functions by penalizing failures, with the penalty treated as a mere heuristic. We rigorously examine the relationship of both safety and optimality to penalties, and formalize sufficient conditions for safe value functions (SVFs): value functions that are both optimal for a given task, and enforce safety constraints. We reveal this structure by examining when rewards preserve viability under optimal control, and show that there always exists a finite penalty that induces an SVF. This penalty is not unique, but upper-unbounded: larger penalties do not harm optimality. Although it is often not possible to compute the minimum required penalty, we reveal clear structure of how the penalty, rewards, discount factor, and dynamics interact. This insight suggests practical, theory-guided heuristics to design reward functions for control problems where safety is important.
Keyword:
Safety
Task analysis
Trajectory
Dynamical systems
Reinforcement learning
Optimal control
Kernel
reinforcement learning (RL)
safety
strong duality
value functions
viability

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

R
RWTH Aachen University
学者数:
3.5W
论文数: 2.6W
被引数: 3.6W
M
Max Planck Society
学者数:
8.2W
论文数: 7.7W
被引数: 3.3W
引用论文

引用论文

err分享
err收藏
Formation and properties of radiation-induced defects and radiolysis products in lithium orthosilicate
err1991-12-01
err0
PREAI
errJ.E. Tiliks; G.K. Kizane; A.A. Supe; A.A. Abramenkovs; J.J. Tiliks; V.G. Vasiljev
err分享
err收藏
Learning agile and dynamic motor skills for legged robots
err2019-01-30
err795
errOAAI
errHwangbo, Jemin; Lee, Joonho; Dosovitskiy, Alexey; Bellicoso, Dario; Tsounis, Vassilios; Koltun, Vladlen; Hutter, Marco
err分享
err收藏
学者 查看更多内容