arrow
返回

Safe Reinforcement Learning via Shielding

delete2018-04-29
delete0
delete
OA
AI
DOI:10.1609/aaai.v32i1.11797delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Reinforcement learning algorithms discover policies that maximize reward, but do not necessarily guarantee safety during learning or execution phases. We introduce a new approach to learn optimal policies while enforcing properties expressed in temporal logic. To this end, given the temporal logic specification that is to be obeyed by the learning system, we propose to synthesize a reactive system called a shield. The shield monitors the actions from the learner and corrects them only if the chosen action causes a violation of the specification. We discuss which requirements a shield must meet to preserve the convergence guarantees of the learner. Finally, we demonstrate the versatility of our approach on several challenging reinforcement learning scenarios.

期刊

暂无期刊信息

机构

暂无机构信息
引用论文

引用论文

暂无论文信息